HiDream.ai has announced the launch of HiDream-O1-Video-1.0, also known as HiDream V1, a native omnimodal video generation model intended to strengthen physical consistency and creative intent understanding. The system supports text, image, and video inputs and can generate high-fidelity 1080p videos lasting five to twenty seconds with natively synchronized audio. Its debut positions the company among the leading global competitors in AI video generation.
Benchmark Recognition and Market Context
HiDream V1 made its first appearance on two independent international leaderboards, ranking fourth globally on the Artificial Analysis Image to Video Leaderboard With Audio and eighth on the Arena.ai Image-to-Video leaderboard. The results reflect strong performance across visual quality, prompt adherence, motion quality, narrative coherence, and audiovisual coordination. In a competitive market where Chinese models such as Seedance and MiniMax H3 already hold strong positions, this debut further expands China's presence among leading video generation systems.
A Unified Framework for Intent and Physics
Chief Technology Officer Yao Ting stated that future video models will not be defined solely by higher resolution or longer duration. Instead, they must understand a creator's intent and how objects, actions, and sounds interact in the real world. HiDream V1 was designed from the outset to represent text, video, and audio within a unified framework, supporting coherent performances and sound aligned with visuals.
Planning, Generation, and Physical Reasoning
HiDream V1 incorporates multimodal intent understanding and planning before generation begins, structuring user requests across elements such as shot duration, setting, character state, movement, composition, camera motion, dialogue, and ambient sound. The company uses a three stage technical framework involving global planning, joint generation, and alignment through multimodal reward signals. The model also factors physical rules such as gravity, inertia, collisions, deformation, materials, lighting, and spatial continuity into generation, helping object movement and environmental responses reflect real world behavior.
Content-Adaptive Duration and Native Audio
Most current video generation models require users to specify a fixed duration in advance, which can force content into a predetermined time window. HiDream V1 instead incorporates duration into its narrative planning process and can determine an appropriate video length between five and twenty seconds. This capability is designed to make duration serve the narrative, producing more natural pacing and more complete sequences.
Joint Modeling of Text, Video, and Audio
Conventional AI video systems often generate visuals first and add sound afterward, leading to mismatches between lip movements, physical actions, and ambient sound. HiDream V1 uses a native omnimodal architecture to jointly model text, video, and audio signals within a single generation process. Visual motion influences sound timing, dialogue and sound effects contribute to emotional tone, and text conditions guide generation throughout.
Completing the Omnimodal Portfolio
HiDream V1 is a core component of HiDream's native omnimodal world model strategy, which is built on a unified transformer foundation. The portfolio now includes image, video, world, and embodied model families covering spatial understanding, temporal modeling, 3D interaction, and embodied action. HiDream's image model previously ranked second globally and first among Chinese models on the Artificial Analysis text-to-image leaderboard.
Series C+ Financing and Availability
Alongside the product launch, HiDream announced the completion of its Series C+ financing round backed by New Micro Capital, Jiaozi Capital, and ICBC Capital. The proceeds will support native omnimodal model research and development, product iteration, AI computing infrastructure, and industry application expansion. HiDream V1 is currently in internal testing and is expected to become available in the near future through the company's hiharness.ai platform.
The launch of HiDream V1 signals a shift in AI video generation from visual sharpness and smooth motion toward deeper instruction understanding, sustained coherence, and synchronized audio. By integrating text, video, and audio within a unified architecture, HiDream aims to model how a dynamic world operates over time rather than simply reproducing appearances. With new financing and a coordinated model portfolio, the company is positioning itself to compete globally in the next generation of foundation models.