MemeBench methodology
How Video Arena compares AI video models and turns blind human choices into a leaderboard.
Published August 16, 2026 · Updated August 17, 2026
1. What MemeBench measures
MemeBench (梗Bench) measures which AI-generated video a viewer prefers when two image-to-video models receive the same source image and the same motion prompt. Its current content distribution emphasizes short Chinese internet meme scenarios.
The leaderboard estimates relative model preference under the current MemeBench content, audience, serving policy, and voting rules.
The benchmark is designed around delivered viewing experience. Prompt following, motion coherence, physical plausibility, subject preservation, pacing, artifacts, and entertainment value can all influence a vote. It is not a pixel-only codec test or a universal measure of every video-generation use case.
2. Dataset and generation rules
- Versioned inputs. Source images and Chinese motion prompts belong to a curated, versioned MemeBench dataset. Retired items no longer enter new matchups.
- No visitor prompts or uploads. Operators review and import the benchmark inputs. Public visitors only evaluate pre-generated clips.
- Same-input pairs. Both sides of a matchup use the same source image and prompt, which reduces content differences within that comparison.
- Generation and publication are separate. A provider response must pass storage, media compatibility, and activation checks before it can be shown.
- Cross-organization comparison. Served pairs use models from different makers, preserving evaluator attention for cross-provider evidence.
- Native delivery characteristics remain relevant. Public clips share a browser-playability contract—MP4, H.264, yuv420p, silent, fast-start—but model-native dimensions, frame rate, and duration remain part of delivered quality.
3. Blind pairwise voting
- The server selects an eligible dataset item and an allowed pair of model outputs.
- The A/B position is randomized deterministically for that anonymous session and round.
- The viewer sees the common source image and prompt, then watches both clips without model identities.
- The viewer chooses A, B, Tie, or Both bad.
- Only after the decision are the model names and details revealed.
A vote records the exact two generation IDs that were displayed. The server recalculates the eligible pair when accepting the vote so a client cannot substitute a different model. One presentation round can produce one accepted vote, while a visitor may evaluate the same benchmark item again in a later round.
High-risk activity can be retained for audit while being marked ignored_for_ranking. Ranking uses pre-reveal votes that are not excluded by this abuse review. Public voting does not require an account or an IP-based one-person-one-vote rule.
4. Regularized Bradley-Terry ranking
For two models i and j, the Bradley-Terry model represents the probability that i is preferred as:
P(i beats j) = sigmoid(θᵢ − θⱼ)
An A win contributes 1, a B win contributes 0, and Tie or Both bad contributes 0.5 to the current fit. The raw four-way outcome remains available for analysis even though the two neutral outcomes have the same ranking contribution.
The fit uses ridge regularization around a prior centered at 1200 Elo. This shrinks small-sample estimates toward the center and reduces extreme rankings caused by a few early outcomes. The fitted values are transformed to an Elo-style scale for readability; they are not updated with a traditional chess Elo formula after each vote.
The public leaderboard includes:
- relative rank and Elo-scaled Bradley-Terry score;
- an approximate 95% confidence interval;
- the number of eligible comparison votes;
- first-party public model pricing where a comparable price is available.
Price is descriptive and never enters the ranking score. Image-to-video and prompt-only models are ranked separately because they solve different tasks.
5. How to interpret the results
A higher score means that a model was more often preferred, after accounting for opponent strength, within the observed comparison graph. Confidence intervals communicate sampling uncertainty; overlapping intervals are a warning against treating close positions as a definitive ordering.
Important limitations
- Visitors are a self-selected audience and may not represent every geography, profession, or use case.
- MemeBench focuses on short meme-style scenes. Results should not be generalized directly to advertising, cinema, medicine, industrial video, or long-form generation.
- Model coverage is sparse rather than a complete model-by-prompt grid. Uneven context coverage can influence a simple global Bradley-Terry fit.
- Anonymous sessions can contribute multiple evaluations across different items and rounds. Current intervals are approximate and are not evaluator-clustered bootstrap intervals.
- Serving weights affect which model pairs collect evidence. The current published score does not apply inverse-propensity reweighting.
Frequently asked questions
What does MemeBench measure?
Relative human preference between AI-generated videos made from the same image and motion prompt on the current MemeBench scenarios.
How are the models ranked?
Eligible blind votes are fit with a regularized Bradley-Terry model and displayed on an Elo-style scale with vote counts and 95% confidence intervals.
Are identities hidden during voting?
Yes. Names, organizations, endpoints, and prices are revealed only after the viewer chooses an outcome.
Can the public upload images or prompts?
No. MemeBench uses operator-curated, pre-generated evaluation material so every matchup follows the benchmark rules.