PlaylistEval

Preprint

PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?

Shayekh Bin Islam Hwanjun Song (corresponding author)

KAIST Corresponding author

hours of playlist video
700
domains, ~100 h each
7
preference pairs
630
judges evaluated
17
Open Figure 1 at full size: An example PlaylistEval item. Top: two 30-minute segments from a 100-hour Documentary playlist, one about the Pacific Northwest showing Portland's White Stag sign, one about Washington showing Tacoma's clock tower. Middle: a two-segment question that describes both landmarks only by their surroundings, and the gold answer with excerpts from each segment. Bottom: four graded wrong answers, rated 4 to 1, each injecting more visual errors, with causal-record tags. A final strip gives, for the pairs 3 vs 2 and 2 vs 1, the verdicts of GPT, Gemini, Qwen and the human annotators; all pick the higher-rated answer except GPT on 2 vs 1.
Figure 1 An item in PlaylistEval. Two 30-minute segments of a 100-hour Documentary playlist, paired by embedding similarity, each supply one of the question’s two entities. The question names neither entity and identifies each only by its surroundings, so its evidence must be found across the collection and confirmed by what is seen and heard. From the gold answer, four wrong answers are derived by injecting visual errors of graded severity (rating 4 to 1), invisible in the transcript. Any two answers form a pair, so the rating gap sets how hard each pair is.

01The paper in brief

Abstract

Video-language models are increasingly used as judges of video understanding, both for evaluating model outputs and for training reward models. Whether their judgments remain reliable when the evidence is buried in day-long videos has yet to be established. Existing benchmarks cannot answer this. Their videos are typically only a few minutes long, many answer pairs can be separated from the transcript alone, and collecting human judgments does not scale to ultra-long videos. We introduce PlaylistEval, an agentic framework that builds video-language judge benchmarks over 100-hour playlist collection without human annotation. It automatically generates questions with paired answers whose differences are controlled by causal degradation, so that every pair demands retrieval across the collection. The resulting benchmark contains 630 pairs across seven domains spanning both static and dynamic knowledge, and on a stratified subset of 152 pairs it agrees with human judgments 93.0% of the time (IAA 0.781). Evaluating 17 omnimodal and multimodal models from eight families reveals that frontier judges reach only 75.4% pairwise accuracy, while open-source judge models perform far behind. We further show that both retrieval and final judgment depend on using multiple modalities, and that judge accuracy degrades as the playlist set grows. We release our pipeline, benchmark, and evaluation code at playlisteval.github.io.

02How the benchmark is built

Pipeline

Open Figure 2 at full size: The PlaylistEval pipeline. A: playlists from 7 domains are split into segments of up to 30 minutes, and similar segments from the same playlist are paired by embedding similarity. Phase 1: B synthesizes a question and answer, which pass structural, video-necessity and visual-evidence gates (C). Phase 2: D synthesizes graded distractors, which pass structural, parametric-detectability and grounding gates. E: a feedback loop returns every rejection reason to its generator. F: a difficulty gate drops pairs that two small VLMs both answer correctly. G: judge meta-evaluation over 630 preference pairs with pairwise accuracy.
Figure 2 Overview of PlaylistEval. A From indexed and paired playlist segments, B Phase I generates a QA and C verifies it through three gates, D Phase II generates graded distractors and verifies them likewise, and E the feedback loop returns every rejection to its generator. F Surviving items pass a difficulty gate and G form the preference pairs for judge meta-evaluation.

03Main results

Judges fall far short of humans

Pairwise accuracy (%) of 17 judges on the 630 pairs, per content domain and overall, with domain-retrieved frames versus uniform frame sampling. Retrieved frames come from the 30-second chunks most similar to the question; uniform frames are spread evenly over the ~100-hour domain. Chance is 50%.

Judge
Input
Accuracy chart
Retrieved
Uniform
Gain

Input: video frames, transcript, audio track. Gain is retrieved minus uniform, in points, as reported in the paper from unrounded accuracies, so it can differ by 0.1 from the rounded columns. Kimi-K2.6 is 1T-A32B and Qwen-3.8-Max is 2.4T-A95B. The retriever is Qwen3-VL-Embedding-8B over video and transcript.

01

Even the strongest judges fall far short of humans

The best judge, Gemini-3.7-Flash, reaches only 75.4% with retrieved evidence, against 93.0% human agreement. The other hosted judges cluster below it, from Qwen-3.8-Max (73.9%), GPT-5.6-Terra (72.5%) and Kimi-K2.6 (71.4%) down to Gemini-3.5-Flash-Lite (59.5%).

02

Smaller and open-weight judges approach chance

Only Gemma-4-26B-A4B (59.7%) rises clearly above the 50% floor; the other open-weight general-purpose models fall between 44 and 57%. Accuracy broadly tracks scale, with the largest member of each family ahead of its smaller variants.

03

Short-video reward models do not transfer

InternLM-XComposer-2.5-Reward and VideoJudge-3B/7B, all trained for video preference tasks, score 47.8–52.7%. Retrieval leaves them unchanged or slightly worse (−1.6 points for InternLM-XComposer-2.5-Reward).

04

Retrieval helps, but only judges that can use it

Retrieved frames beat uniform sampling for nearly every capable judge, by up to +10.5 points for Gemma-4-26B-A4B and +4.4 for Gemini-3.7-Flash. The right evidence is necessary but not sufficient: it raises accuracy only when the judge can reason over it.

04How to cite

BibTeX

@article{islam2026playlisteval,
      title={PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?}, 
      author={Shayekh Bin Islam and Hwanjun Song},
      year={2026},
      eprint={2609.34314},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.34314}, 
      journal={arXiv preprint arXiv:2609.34314}, 
}

05Funding & credits

Acknowledgements

This work was supported by Institute of Information & communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (MSIT) (No. RS-2024-00445087 & RS-2025-25464461) and National IT Industry Promotion Agency (NIPA) grant funded by the Korea government (MSIT) (No. RS-2026-25621604).

The design of this page is adapted from the World Tracing project page by World Labs. Model logos are trademarks of their respective owners.