Chinese AI lab MiniMax launched H3 (also referred to as Hailuo 3.0), an omni-modal generation model that takes text, images, video, and audio as input and produces up to 15-second clips at 2K resolution with native, synchronized stereo audio — meaning the model generates sound and dialogue as part of the same generation pass rather than requiring a separate audio pipeline stitched on afterward. It also supports editing existing footage and transferring motion from a reference clip onto new content based on natural-language instructions. On the independent Artificial Analysis leaderboard, H3 ranked #1 in video editing and placed top-3 in both text-to-video and image-to-video, and MiniMax priced it aggressively at $0.13 per second of output (about $1.95 for a 15-second clip), undercutting several closed competitors. Access at launch is through the Hailuo app, MiniMax's own hub, its Open Platform API, and third-party hosts like OpenRouter, with open weights promised to follow under a MiniMax Community License permitting commercial use for organizations under $20M in revenue. While this is a generative-media model rather than a coding or agent tool, it's relevant to builders for a specific reason: an API-accessible, competitively priced, natively-audio video model is a building block, not just a novelty — it lowers the cost of adding programmatic video generation (product demos, localized marketing variants, procedurally generated content, synthetic training data for vision models) directly into an application via API calls rather than through a manual creative tool. Teams evaluating video-generation APIs for product features now have a meaningfully cheaper, higher-ranked option to benchmark against the more established closed providers, and the promised open weights would additionally make self-hosted, cost-controlled deployment possible once released.