Open Figure 1 at full size: ZoomFigure 1An item in PlaylistEval.
Two 30-minute segments of a 100-hour Documentary playlist, paired by embedding similarity, each supply one of the question’s two entities. The question names neither entity and identifies each only by its surroundings, so its evidence must be found across the collection and confirmed by what is seen and heard. From the gold answer, four wrong answers are derived by injecting visual errors of graded severity (rating 4 to 1), invisible in the transcript. Any two answers form a pair, so the rating gap sets how hard each pair is.
01The paper in brief
Abstract
Video-language models are increasingly used as judges of video understanding, both for evaluating model outputs and for training reward models. Whether their judgments remain reliable when the evidence is buried in day-long videos has yet to be established. Existing benchmarks cannot answer this. Their videos are typically only a few minutes long, many answer pairs can be separated from the transcript alone, and collecting human judgments does not scale to ultra-long videos. We introduce PlaylistEval, an agentic framework that builds video-language judge benchmarks over 100-hour playlist collection without human annotation. It automatically generates questions with paired answers whose differences are controlled by causal degradation, so that every pair demands retrieval across the collection. The resulting benchmark contains 630 pairs across seven domains spanning both static and dynamic knowledge, and on a stratified subset of 152 pairs it agrees with human judgments 93.0% of the time (IAA 0.781). Evaluating 17 omnimodal and multimodal models from eight families reveals that frontier judges reach only 75.4% pairwise accuracy, while open-source judge models perform far behind. We further show that both retrieval and final judgment depend on using multiple modalities, and that judge accuracy degrades as the playlist set grows. We release our pipeline, benchmark, and evaluation code at playlisteval.github.io.
02How the benchmark is built
Pipeline
Open Figure 2 at full size: ZoomFigure 2Overview of PlaylistEval.A From indexed and paired playlist segments, B Phase I generates a QA and C verifies it through three gates, D Phase II generates graded distractors and verifies them likewise, and E the feedback loop returns every rejection to its generator. F Surviving items pass a difficulty gate and G form the preference pairs for judge meta-evaluation.
03Main results
Judges fall far short of humans
Pairwise accuracy (%) of 17 judges on the 630 pairs, per content domain and overall, with domain-retrieved frames versus uniform frame sampling. Retrieved frames come from the 30-second chunks most similar to the question; uniform frames are spread evenly over the ~100-hour domain. Chance is 50%.
Input: video frames,
transcript,
audio track.
Gain is retrieved minus uniform, in points, as reported in the paper from unrounded accuracies, so it can differ by 0.1 from the rounded columns.
Kimi-K2.6 is 1T-A32B and Qwen-3.8-Max is 2.4T-A95B.
The retriever is Qwen3-VL-Embedding-8B over video and transcript.
01
Even the strongest judges fall far short of humans
The best judge, Gemini-3.7-Flash, reaches only 75.4% with retrieved evidence, against 93.0% human agreement. The other hosted judges cluster below it, from Qwen-3.8-Max (73.9%), GPT-5.6-Terra (72.5%) and Kimi-K2.6 (71.4%) down to Gemini-3.5-Flash-Lite (59.5%).
02
Smaller and open-weight judges approach chance
Only Gemma-4-26B-A4B (59.7%) rises clearly above the 50% floor; the other open-weight general-purpose models fall between 44 and 57%. Accuracy broadly tracks scale, with the largest member of each family ahead of its smaller variants.
03
Short-video reward models do not transfer
InternLM-XComposer-2.5-Reward and VideoJudge-3B/7B, all trained for video preference tasks, score 47.8–52.7%. Retrieval leaves them unchanged or slightly worse (−1.6 points for InternLM-XComposer-2.5-Reward).
04
Retrieval helps, but only judges that can use it
Retrieved frames beat uniform sampling for nearly every capable judge, by up to +10.5 points for Gemma-4-26B-A4B and +4.4 for Gemini-3.7-Flash. The right evidence is necessary but not sufficient: it raises accuracy only when the judge can reason over it.
04How to cite
BibTeX
@article{islam2026playlisteval,
title={PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?},
author={Shayekh Bin Islam and Hwanjun Song},
year={2026},
eprint={2609.34314},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.34314},
journal={arXiv preprint arXiv:2609.34314},
}
05Funding & credits
Acknowledgements
This work was supported by Institute of Information & communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (MSIT) (No. RS-2024-00445087 & RS-2025-25464461) and National IT Industry Promotion Agency (NIPA) grant funded by the Korea government (MSIT) (No. RS-2026-25621604).
The design of this page is adapted from the World Tracing project page by World Labs. Model logos are trademarks of their respective owners.