AI video generation and blind human-preference evaluation.
A practical explanation of blind pairwise AI video evaluation: same-input comparisons, anonymous voting, randomized presentation, and preference rankings.
Evaluate AI video generators across prompt following, temporal consistency, motion, subject preservation, reliability, cost, and human preference.
MemeBench methodology: dataset curation, blind pairwise voting, model sampling, regularized Bradley-Terry rankings, confidence intervals, and limits.