Annual Sale!Get 30% offClaim Now→

#1 on Artificial Analysis Video Arena

QwenHappy Horse 1.0

Alibaba's 15-billion parameter video model — joint text-to-video, image-to-video, and synchronized audio in a single unified transformer. Topped every leaderboard at launch.

1332
T2V Elo
1391
I2V Elo
15B
Parameters
1080p
Max Resolution

What makes Happy Horse different

A unified multimodal transformer purpose-built for joint video and audio generation — no separate models, no post-processing.

Unified Transformer

A 40-layer single-stream transformer with 32 shared layers. Text, video, and audio tokens are processed in one sequence — no cross-attention bottleneck.

Joint Video + Audio

Generates synchronized dialogue, ambient sound, and Foley alongside video frames in a single forward pass. No post-production dubbing pipeline required.

8-Step Inference

DMD-2 distillation reduces denoising from 50+ steps to just 8 without classifier-free guidance — accelerated by an in-house compiler runtime.

Multilingual Lip-Sync

Native support for English, Mandarin, Cantonese, Japanese, Korean, German, and French — with industry-leading low Word Error Rate for digital humans.

1080p Cinematic Output

5–15 second clips at full 1080p in 5 aspect ratios — suitable for social media, advertising, and cinematic production.

Open Research Spirit

Backed by published architecture details and benchmark transparency. Designed for both creators and researchers who want to understand what's under the hood.

Top of the leaderboard

Blind-tested Elo ratings from thousands of human-evaluated comparisons. Happy Horse leads both text-to-video and image-to-video categories.

#ModelDeveloperElo
1HappyHorse 1.0Alibaba1332
2Seedance 2.0 720pByteDance1273
3SkyReels V4Skywork AI1245
4PixVerse V6PixVerse1241
5Kling 3.0 1080pKlingAI1241

Source: Artificial Analysis Video Arena, April 2026.

Generate in three steps

From idea to finished video in under five minutes. No model setup, no GPU required.

1

Pick your starting point

Choose Text-to-Video to describe a scene, or Image-to-Video to animate a still frame you upload.

2

Configure resolution & duration

Pick 720p or 1080p, drag the slider for any length from 3 to 15 seconds, then choose an aspect ratio for text-to-video.

3

Generate and download

Submit the job and your video lands in your library in 2–5 minutes — credits are automatically refunded if generation fails.

Sample outputs

Real videos generated by Happy Horse 1.0 at 1080p — physically grounded motion, sharp detail, and smooth temporal consistency.

Ready to ride Happy Horse?

Generate your first 1080p video in minutes. Pay only for what you use — 60 credits per second at 720p, 100 at 1080p.

Open the Generator

Happy Horse — Frequently Asked Questions

Have a different question and can't find the answer you're looking for? Reach out to our support team by sending us an email and we'll get back to you as soon as we can.

What is Happy Horse 1.0?

Happy Horse 1.0 is a 15-billion parameter AI video generation model from Alibaba's Model Studio. It jointly produces video and synchronized audio from text or image prompts, and ranked #1 on the Artificial Analysis Video Arena with an Elo score of 1332 for text-to-video and 1391 for image-to-video.

How does Happy Horse compare to Seedance, Kling, and Sora?

On the Artificial Analysis blind-tested leaderboard, Happy Horse 1.0 (1332 Elo) outperforms Seedance 2.0 (1273), SkyReels V4 (1245), PixVerse V6 (1241), and Kling 3.0 (1241). It leads in both text-to-video and image-to-video categories.

What input modes does Happy Horse support?

Two modes — Text-to-Video, where you describe the scene in any supported language and pick an aspect ratio, and Image-to-Video, where you upload a first-frame image and an optional prompt. The aspect ratio is auto-detected from the image in I2V mode.

What resolutions and durations are available?

Resolution: 720p or 1080p. Duration: any integer from 3 to 15 seconds, default 5 seconds. Aspect ratios for text-to-video: 16:9, 9:16, 1:1, 4:3, and 3:4.

How are credits calculated?

Pricing is linear by duration. 720p costs 60 credits per second; 1080p costs 100 credits per second. A 5-second 720p clip is 300 credits; a 10-second 1080p clip is 1000 credits.

How long does generation take?

Most jobs finish in 2–5 minutes depending on duration and resolution. The page polls every 30 seconds while your video is rendering, and you can leave and check back via the My Videos page at any time.

What happens if a generation fails?

Credits are automatically refunded for any job that fails, gets cancelled, or otherwise doesn't return a video. You can also retry a failed prompt at a discounted credit cost.

Is the output watermarked?

No. Watermarks are disabled by default for every video generated through this site, at every resolution and duration.