Annual Sale!Get 30% offClaim Now→
Alibaba's 15-billion parameter video model — joint text-to-video, image-to-video, and synchronized audio in a single unified transformer. Topped every leaderboard at launch.
A unified multimodal transformer purpose-built for joint video and audio generation — no separate models, no post-processing.
A 40-layer single-stream transformer with 32 shared layers. Text, video, and audio tokens are processed in one sequence — no cross-attention bottleneck.
Generates synchronized dialogue, ambient sound, and Foley alongside video frames in a single forward pass. No post-production dubbing pipeline required.
DMD-2 distillation reduces denoising from 50+ steps to just 8 without classifier-free guidance — accelerated by an in-house compiler runtime.
Native support for English, Mandarin, Cantonese, Japanese, Korean, German, and French — with industry-leading low Word Error Rate for digital humans.
5–15 second clips at full 1080p in 5 aspect ratios — suitable for social media, advertising, and cinematic production.
Backed by published architecture details and benchmark transparency. Designed for both creators and researchers who want to understand what's under the hood.
Blind-tested Elo ratings from thousands of human-evaluated comparisons. Happy Horse leads both text-to-video and image-to-video categories.
| # | Model | Developer | Elo |
|---|---|---|---|
| 1 | HappyHorse 1.0 | Alibaba | 1332 |
| 2 | Seedance 2.0 720p | ByteDance | 1273 |
| 3 | SkyReels V4 | Skywork AI | 1245 |
| 4 | PixVerse V6 | PixVerse | 1241 |
| 5 | Kling 3.0 1080p | KlingAI | 1241 |
Source: Artificial Analysis Video Arena, April 2026.
From idea to finished video in under five minutes. No model setup, no GPU required.
Choose Text-to-Video to describe a scene, or Image-to-Video to animate a still frame you upload.
Pick 720p or 1080p, drag the slider for any length from 3 to 15 seconds, then choose an aspect ratio for text-to-video.
Submit the job and your video lands in your library in 2–5 minutes — credits are automatically refunded if generation fails.
Real videos generated by Happy Horse 1.0 at 1080p — physically grounded motion, sharp detail, and smooth temporal consistency.
Generate your first 1080p video in minutes. Pay only for what you use — 60 credits per second at 720p, 100 at 1080p.
Open the GeneratorHave a different question and can't find the answer you're looking for? Reach out to our support team by sending us an email and we'll get back to you as soon as we can.
Happy Horse 1.0 is a 15-billion parameter AI video generation model from Alibaba's Model Studio. It jointly produces video and synchronized audio from text or image prompts, and ranked #1 on the Artificial Analysis Video Arena with an Elo score of 1332 for text-to-video and 1391 for image-to-video.
On the Artificial Analysis blind-tested leaderboard, Happy Horse 1.0 (1332 Elo) outperforms Seedance 2.0 (1273), SkyReels V4 (1245), PixVerse V6 (1241), and Kling 3.0 (1241). It leads in both text-to-video and image-to-video categories.
Two modes — Text-to-Video, where you describe the scene in any supported language and pick an aspect ratio, and Image-to-Video, where you upload a first-frame image and an optional prompt. The aspect ratio is auto-detected from the image in I2V mode.
Resolution: 720p or 1080p. Duration: any integer from 3 to 15 seconds, default 5 seconds. Aspect ratios for text-to-video: 16:9, 9:16, 1:1, 4:3, and 3:4.
Pricing is linear by duration. 720p costs 60 credits per second; 1080p costs 100 credits per second. A 5-second 720p clip is 300 credits; a 10-second 1080p clip is 1000 credits.
Most jobs finish in 2–5 minutes depending on duration and resolution. The page polls every 30 seconds while your video is rendering, and you can leave and check back via the My Videos page at any time.
Credits are automatically refunded for any job that fails, gets cancelled, or otherwise doesn't return a video. You can also retry a failed prompt at a discounted credit cost.
No. Watermarks are disabled by default for every video generated through this site, at every resolution and duration.