About Video Arena
Video Arena is a platform for blind, human-preference benchmarks of AI video generation models.
AI video quality is difficult to reduce to a single automated score. A clip can be sharp but physically implausible, faithful to the prompt but unpleasant to watch, or technically clean while missing the idea that makes a scene work. Video Arena measures a direct outcome: when viewers see two model outputs made from the same input, which one do they prefer?
What Video Arena measures
Every benchmark defines its own content distribution, generation rules, and ranking scope. Visitors compare anonymous outputs side by side. Model names, providers, and price information stay hidden until after the vote, reducing brand and expectation bias.
The result is a relative preference ranking for the models, scenarios, audience, and rules in that benchmark. It is not a claim that one model is universally best for every video task.
MemeBench (梗Bench)
MemeBench is the first Video Arena benchmark. It tests image-to-video models on operator-curated Chinese internet meme scenarios: short, expressive scenes where motion, timing, prompt following, character consistency, and entertainment value all matter.
Both models in a matchup receive the same source image and motion prompt. A viewer watches both anonymous clips and selects A, B, a tie, or both bad. The identities are revealed only after that choice.
What makes the benchmark auditable
- Versioned evaluation sets: source images and prompts are curated and tracked rather than submitted by visitors.
- Same-input comparisons: each pair shares the same image and prompt, reducing content as a confounding factor.
- Blind evaluation: model identity remains hidden until the decision is recorded.
- Exact attribution: each vote records the precise model outputs that were presented.
- Uncertainty shown: the leaderboard reports vote counts and 95% confidence intervals alongside its Bradley-Terry score.
- Limits stated: the methodology explains what the rankings do and do not establish.
For viewers and model teams
Anyone can vote without creating an account, uploading media, or writing a prompt. Model teams and researchers can use the leaderboard as evidence about human preference on the published MemeBench task, then pair it with technical, safety, and domain-specific evaluation before making broader decisions.