Video-Index: A Curated Meta-Benchmark for Video Understanding
Abstract: A video benchmark should reward the capability it claims to measure, yet models can exploit answer options, question text, or partial visual evidence. We introduce the attack pyramid, five levels of shortcut attacks with increasing access to each item, and audit 115 video benchmarks with it. On 35 benchmarks, attackers that never see a frame approach full-video accuracy. On 51 benchmarks with temporal probes, shuffled frames keep a median 96% of full-video accuracy. Near-duplicate questions make up at least half the items in 63 benchmarks. We screen 505,518 question-answer pairs from 112 of them into an audited pool. Agents turn evaluation requests into specifications, and a deterministic selector with a red-team gate composes reproducible benchmarks. We release Video-Index, the 210 hardest verified items under these attacks in each of four capability groups, 840 items from 76 sources. With the same fixed input, Claude Opus 5 outscores every open-source model by over 37 percentage points, and agent tools add about 20 more, yet all systems leave room to improve efficiency and accuracy. Blog: https://www.enxinsong.com/blog/video-index/ GitHub: https://github.com/Espere-1119-Song/Video-Index Hugging Face: https://huggingface.co/datasets/Video-Index/Video-Index
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is the paper about?
This paper studies how well video-understanding tests measure what they claim to measure.
A video benchmark is like an exam for an AI system. It might ask questions such as:
- What happened first?
- How many times did someone perform an action?
- Where was an object located?
- What happened during a long lecture or sports match?
The researchers worry that some AI models may get the right answer without really watching or understanding the video. For example, a model might guess from the answer choices or from the wording of the question.
To investigate this problem, the authors created a testing system called the attack pyramid and used it to examine 115 video benchmarks. They also created a new, harder benchmark called Video-Index.
2. What questions did the researchers ask?
The researchers mainly wanted to find out:
- Can models answer video questions without seeing the video?
- Can they use clues in the answer choices or question wording?
- Can they use answers from other questions to guess new answers?
- Can they answer correctly from only one picture or a written description of the video?
- Does the order of video frames really matter?
- Which existing video benchmarks are genuinely difficult and reliable?
- Can the researchers build a better benchmark using questions that resist these shortcuts?
In simple terms, they asked:
Is the AI truly understanding the video, or is it finding an easier way to guess the answer?
3. How did they do the research?
Examining many existing benchmarks
The researchers first looked at hundreds of video benchmarks created during the previous five years. They selected 115 English multiple-choice benchmarks that did not require audio.
Together, these benchmarks contained more than half a million questions about over 100,000 videos.
They then tested each benchmark using five kinds of “attack.” Here, an attack means a shortcut that tries to answer questions without doing the full task.
The attack pyramid
The attack pyramid has five levels. Each level gives the attacker a little more information:
| Level | What the attacker can see | What it tests |
|---|---|---|
| Options | Only the answer choices | Whether the choices themselves give away the answer |
| Text | The question and answer choices | Whether the wording reveals the answer |
| Question pool | Other questions and answers from the test | Whether similar earlier questions can be copied or used |
| Frame | One video frame or captions | Whether a single image is enough |
| Order | Shuffled or shortened video frames | Whether the AI really needs the correct order and the whole video |
An analogy is a school exam:
- At the first level, a student sees only the multiple-choice answers.
- At the second level, the student sees the question too.
- At later levels, the student gets a picture or parts of the lesson.
- Only the full test allows the student to watch the entire lesson in the correct order.
If a shortcut method performs almost as well as a model that sees the full video, the benchmark is considered broken at that level.
Creating Video-Index
After testing the old benchmarks, the researchers removed questions that were too easy to solve using shortcuts.
They also removed questions that were almost duplicates of one another. This is important because 200 questions that all ask nearly the same thing do not provide as much information as 200 very different questions.
The researchers used computer programs and AI agents to:
- Find difficult questions.
- Remove questions answerable without enough visual information.
- Group questions by skills, such as perception, timing, spatial understanding, and reasoning.
- Avoid choosing too many questions about the same video.
- Have another system check that the questions and answers were correct.
- Test the final collection again for shortcuts.
The result was Video-Index, containing 840 difficult and checked questions from 76 different sources. There were 210 questions in each of four broad skill groups.
4. What did they find?
Many tests can be solved without watching the video
One of the most important findings was that some benchmarks could be solved surprisingly well without seeing any video frames.
- On 35 benchmarks, attackers that never saw a frame came close to the accuracy of the full-video system.
- On 51 benchmarks, shuffling the video frames still preserved a median of 96% of the original accuracy.
- This suggests that many supposedly temporal questions did not truly require understanding the order of events.
For example, if an AI gets nearly the same score from shuffled frames as from correctly ordered frames, the test may not really be testing “what happened first.”
Question wording can reveal answers
The researchers found that question text sometimes gave away the answer. A model could read the question and answer choices without watching the video, yet still perform well.
This problem was more common in newer benchmarks. The authors suggest that questions written or improved by LLMs may accidentally contain clues that make the correct answer predictable.
Many questions were too similar
At least half of the questions in 63 benchmarks were near-duplicates of other questions.
This means that a benchmark may appear large while actually testing only a small number of different ideas. In one analysis, a sample of 200 questions had a median of only about 32 genuinely different items.
This is similar to studying for an exam using 200 practice questions that repeat the same pattern. The number 200 sounds impressive, but the student may only have learned 32 different types of problems.
Different skills fail in different ways
The researchers found that different kinds of video understanding had different weaknesses:
- Reasoning and knowledge questions were often answerable from the text or captions.
- Perception questions were often solvable from a single frame.
- Spatial and physical questions were affected by answer-choice clues or frame changes.
- Temporal questions sometimes did not really depend on the correct order of events.
This means that benchmark designers should not use exactly the same checks for every skill.
More frames and higher image quality can help
The study also tested whether models needed:
- More detailed images, or
- More frames from the video.
Higher resolution helped with small details, such as text on slides or tiny objects. More frames helped when an important event happened briefly.
For long videos, most benchmarks became harder when the model was given too few frames. A model might miss the key moment if it samples the video too sparsely.
Video-Index was much harder
The researchers tested several open-source models and commercial AI systems on Video-Index.
The results showed a large difference:
- The best fixed-input commercial model scored about 56.8%.
- Most open-source models scored between about 9% and 19%.
- When AI agents were allowed to use tools to search and inspect the video, the best system reached about 79.3%.
- Human volunteers scored about 55% in the reported test.
Tools improved performance by about 20 percentage points for one model. These tools helped agents search through frames, zoom into details, and look for important moments.
However, even the best systems still made many mistakes. This shows that difficult video understanding remains an unsolved problem.
5. Why are these findings important?
The paper shows that a high benchmark score does not always prove that an AI has the ability being tested.
A model might score well because it:
- Learns common answer patterns.
- Uses clues in the question.
- Recognizes repeated questions.
- Looks at only one useful frame.
- Ignores the order of events.
- Guesses from language instead of understanding the video.
The attack pyramid gives researchers a way to discover these weaknesses. Instead of simply saying that a benchmark is good or bad, it shows which shortcut causes problems.
The new Video-Index benchmark is designed to be harder to cheat. It keeps questions that require more genuine inspection of the video and removes many questions that can be answered using text, answer choices, captions, or limited visual information.
Conclusion: What could this research change?
This research could lead to better and fairer tests for video AI.
In the future, benchmark designers may:
- Test whether questions can be answered without the video.
- Check whether frame order really matters.
- Remove repeated or nearly identical questions.
- Include more questions that require finding evidence across time.
- Report which shortcuts a benchmark resists.
- Give models tools for searching videos while measuring whether they use those tools efficiently.
The main lesson is simple:
A good video test should reward an AI for understanding the video, not for finding clues that make the video unnecessary.
The authors also point out that no test is perfect forever. As AI systems improve, they may discover new shortcuts. Therefore, video benchmarks will need to be regularly checked and updated.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- The audit covers only 115 English, multiple-choice, video-only benchmarks and excludes audio-dependent tasks, leaving the validity of the attack pyramid for open-ended, multilingual, multimodal, and audio-visual evaluation unresolved.
- The benchmark census identifies 605 video benchmarks, but only 115 are audited and 112 contribute items to the screened pool; the reasons why the excluded benchmarks were not analyzed and how this selection affects generalizability remain unclear.
- The attack pyramid tests a fixed set of five attacker levels, so it does not establish robustness against stronger shortcuts such as external web retrieval, OCR, speech recognition, video metadata, memorized training examples, timestamp search, or multimodal agents with iterative reasoning.
- The paper does not quantify how much attacker performance changes when attackers are given larger compute budgets, multiple attempts, chain-of-thought reasoning, tool use, or access to benchmark-specific instructions.
- The “strongest tested shortcut” is model-dependent, but the attacker set is small and changes across levels; it remains unknown whether substantially stronger or differently trained attackers would identify additional broken items.
- The reference accuracy relies primarily on Gemini 3.5 Flash at fixed frame budgets, making breaking-level assignments dependent on one model’s capabilities, sampling strategy, and long-context limitations.
- The study does not determine whether benchmark failures reflect genuine dataset shortcuts or weaknesses of the selected reference and attacker models; for example, an attacker performing near reference accuracy may still be using different capabilities.
- The 5% breaking tolerance is justified using an approximate sampling margin, but the paper does not provide systematic confidence intervals, corrections for multiple comparisons across 115 benchmarks and several attack levels, or power analyses for individual break classifications.
- Only 300 items are sampled per benchmark, and some benchmarks contain fewer items; the stability of exploitability estimates and breaking levels under alternative samples is not fully established.
- The results do not separate item-level shortcut prevalence from benchmark-level accuracy effects, leaving unclear whether a small number of highly predictable questions drives many of the reported breaks.
- The near-duplicate analysis depends on SigLIP2 and BGE-large embeddings and fixed similarity procedures; the sensitivity of duplicate counts and effective-size estimates to embedding models, thresholds, and semantic definitions of duplication is unresolved.
- The claim that newer benchmarks break earlier is observational and may be confounded by benchmark domain, item-generation methods, dataset size, model familiarity, release-year differences, or the greater prevalence of synthetic questions.
- The paper suggests that LLM-generated questions contribute to text shortcuts but does not experimentally compare human-authored and machine-generated items under matched domains, videos, difficulty, and annotation procedures.
- Cross-benchmark question and video overlap is measured, but the extent to which model pretraining or instruction-tuning contamination explains the observed performance is not directly tested.
- The screening and selection pipeline is strongly model-relative: Qwen3-VL models remove items and Claude Opus 5 verifies and labels candidates, so the released benchmark may preferentially retain items that are difficult for these models rather than intrinsically valid measures of video understanding.
- Video-Index excludes the screening model from some ranking statistics, but it does not establish whether the selected items are robust against models with architectures, training data, or inference strategies unlike Qwen3-VL and Claude Opus.
- The selection objective prioritizes resistance to measured attacks and high attacker-relative difficulty, but it does not demonstrate that the resulting items have stronger construct validity, better human interpretability, or greater correlation with real-world video understanding.
- The 840-item Video-Index is drawn from existing benchmarks, so it may inherit their annotation errors, source-specific biases, duplicated concepts, and narrow coverage rather than constitute an independently validated benchmark.
- Answer verification is performed by Claude Opus 5 against video frames, but the paper does not report comprehensive human adjudication, inter-annotator agreement, or error rates for gold answers and explanations.
- The reported human baseline is based on volunteers, but participant expertise, demographics, video familiarity, time limits, tool access, and sample size are insufficiently characterized to support a reliable human-performance comparison.
- The paper does not evaluate whether humans and models fail on the same items or for the same reasons, limiting interpretation of the benchmark’s difficulty and the meaning of model–human gaps.
- The multiple-choice format remains vulnerable to guessing, option-position effects, distractor quality, and elimination strategies even after option randomization; the relationship between Video-Index scores and performance on open-ended responses is not established.
- The audit largely treats temporal understanding as resistance to shuffled frames or reduced temporal coverage, but these perturbations do not test causal interpretation, event duration, temporal localization, anticipation, streaming latency, or memory over very long intervals.
- Audio, speech, subtitles, and sound-based temporal evidence are excluded, leaving unresolved whether the reported shortcut patterns persist in realistic audiovisual video understanding.
- Long-video conclusions are limited by frame budgets and the reference model’s context capacity; the study does not determine the frame or token budgets needed for reliable understanding of hour-scale videos or continuous streaming inputs.
- The frame, caption, and order attacks use particular preprocessing pipelines, so the results may change substantially with adaptive frame selection, high-quality video summarization, OCR, speech transcripts, or learned temporal retrieval.
- The paper reports that agent tools improve accuracy, but it does not isolate which tools, prompting strategies, evidence-selection policies, or amounts of computation cause the improvement.
- Accuracy, runtime, image counts, and operations are reported, but there is no standardized cost–accuracy–latency analysis across proprietary and open systems, making efficiency comparisons difficult to reproduce.
- Proprietary model versions and tool behavior may change over time, creating reproducibility and temporal-validity concerns for the reported rankings.
- The benchmark’s lockfile improves composition reproducibility, but the full pipeline still depends on nondeterministic proprietary models, undocumented model updates, and potentially unavailable APIs.
- The evaluation does not test robustness to adversarially authored questions, deliberately misleading options, ambiguous wording, corrupted videos, distribution shifts, or domain changes.
- The four broad capability groups and 18 fine-grained categories are assigned by labeling agents and majority rules; the reliability, validity, and possible category bias of this taxonomy are not independently validated.
- Error attribution uses model-generated reasoning and judgments confirmed by two independent judgments, but it remains unclear whether these attributions reflect actual causal failure modes rather than plausible post hoc explanations.
- The paper does not investigate how benchmark scores change after models are trained on the released Video-Index items or after attackers are explicitly optimized against the attack pyramid, leaving the benchmark’s resistance to adaptive overfitting unknown.
- It remains unresolved whether repeated auditing and filtering will cause an arms race that removes difficult but valid items, producing increasingly narrow benchmarks optimized for known attacks.
- The relationship between exploitability profiles and downstream model capabilities—such as robotics, video search, accessibility, surveillance, or embodied planning—is not measured, so the practical significance of surviving the audit remains uncertain.
- The study does not provide a principled method for choosing the appropriate attack threshold or threat model for different capability claims, despite showing that breaking levels vary substantially across capability groups.
- The claim that a positive residual “rules out only the shortcuts tested” is conceptually acknowledged but not operationalized into a broader validity framework that quantifies the remaining uncertainty about what a benchmark score certifies.
Practical Applications
Immediate Applications
The paper’s methods can be used now because they rely largely on existing video benchmarks, multimodal models, embedding models, and reproducible evaluation scripts rather than requiring new sensors or hardware.
- Industry — Auditing video-AI benchmarks before product deployment. Organizations developing video-LLMs can apply the five-level attack pyramid to determine whether a benchmark score reflects actual video understanding or merely exploits answer options, question wording, duplicate items, captions, single frames, or shuffled footage. A model card or vendor evaluation report could include the benchmark’s breaking level, residual score above the strongest shortcut, and results under ordered and shuffled video. Dependencies: Access to benchmark items and videos, a defined reference model, an agreed tolerance such as the paper’s 5%, and sufficient compute for repeated attacker evaluations.
- Industry — Quality assurance for video analytics products. Healthcare monitoring, security surveillance, sports analytics, manufacturing inspection, and retail systems can test whether models depend on static visual cues rather than events over time. For example, a safety-monitoring system can be evaluated on original, shuffled, truncated, and sparsely sampled footage to determine whether it truly detects temporal sequences such as falls, unsafe machine operation, or intrusion events. Dependencies: Representative domain videos, privacy-compliant data handling, and task-specific perturbations. Shuffling is informative for temporal claims but is not by itself a complete test of real-world reliability.
- Software and MLOps — Automated benchmark health checks.
The screening pipeline can become a continuous-integration tool for evaluation datasets. Each new benchmark release could automatically run option-only, text-only, retrieval, single-frame, caption, shuffle, and truncated-video baselines. A release could fail if attackers approach the full-video score, if near-duplicates exceed a specified threshold, or if answer-position distributions are unbalanced.
Potential product: A
video-benchmark-lintpackage integrated with model-evaluation platforms and dataset repositories. Dependencies: Stable attack implementations, versioned model checkpoints, reproducible preprocessing, and safeguards against overfitting the benchmark exclusively to the tested attacks. - Academia — More reliable reporting of video-model performance. Researchers can report ordinary accuracy together with blind-reader accuracy, video gain, exploitability at each pyramid level, effective dataset size, duplicate rate, frame-budget curves, and performance under temporal perturbation. This would distinguish genuine temporal or spatial reasoning from answer priors and static-frame recognition. Dependencies: Community agreement on reporting standards and access to comparable reference models. Results remain relative to the tested attacker set and cannot certify resistance to all shortcuts.
- Academia — Dataset deduplication and effective-sample-size analysis. The paper’s joint visual-text similarity and Vendi-score procedure can identify near-duplicate questions, repeated videos, and templated items. Dataset curators can replace nominal item counts with effective sample sizes, weight duplicate clusters appropriately, and prevent shared content across training and test sets. Dependencies: Suitable embedding models, threshold calibration, and manual review of false duplicate matches. Similarity metrics may miss semantic duplicates or incorrectly merge visually similar but meaningfully different events.
- Industry and academia — Cost-aware video inference. The frame-budget experiments support adaptive inference workflows that increase resolution for small visual details and increase frame density for brief events or long videos. A system might first inspect a sparse sample, then request additional frames only when confidence is low or when the task requires temporal evidence. Potential tools: Adaptive frame samplers, evidence-selection agents, timestamp search modules, and resolution-versus-frame-rate schedulers. Dependencies: Reliable confidence estimates, latency and compute budgets, and the model’s ability to use additional frames effectively. More frames do not automatically improve accuracy.
- Policy and procurement — Independent evaluation requirements for high-risk video systems. Regulators or public-sector buyers can require suppliers to disclose whether a system’s performance survives blind, single-frame, shuffled-order, and partial-video tests. This is relevant to automated surveillance, traffic safety, workplace monitoring, medical video analysis, and public-service video triage. Dependencies: Sector-specific risk thresholds, access to auditable evaluation data, privacy protections, and procedures for testing proprietary systems without exposing sensitive material.
- Daily life and education — More trustworthy video search and summarization. Users of lecture, meeting, tutorial, and home-video assistants could benefit from systems that identify timestamps and cite the frames supporting an answer rather than relying only on captions or a few sampled images. The paper’s findings suggest that caption-only systems may answer many questions without processing the underlying video. Dependencies: Accurate timestamp grounding, adequate sampling of long videos, accessibility of the source video, and human review for consequential claims.
- Benchmark and model selection — Using Video-Index as a stress test. Organizations can use the released 840-item Video-Index to compare models on perception, temporal, spatial/physical, and reasoning/knowledge capabilities. The benchmark is particularly suitable for regression testing after model updates or for selecting models for pilot deployments. Dependencies: The dataset’s licensing and source availability, domain relevance, language coverage, and awareness that the benchmark is intentionally difficult and may not predict performance on every operational distribution.
Long-Term Applications
These applications require further research, broader data, domain validation, or integration into production systems.
- Healthcare — Clinically validated temporal-video assessment. The attack pyramid could support evaluation of models analyzing endoscopy, radiology procedures, rehabilitation, emergency-room footage, or remote patient monitoring. A clinically meaningful benchmark would require evidence that the answer depends on the relevant sequence of events, not merely on a frame, caption, or textual prior. Potential product: A regulated clinical video model-evaluation suite with timestamped evidence, uncertainty estimates, and clinician-verified explanations. Dependencies: Expert annotation, patient consent, de-identification, prospective validation, demographic and institutional diversity, and regulatory approval. Benchmark robustness alone does not establish clinical safety.
- Robotics — Testing perception and planning from egocentric video. Robot policies could be evaluated on whether they understand action order, object state changes, physical interactions, and long-horizon task progress. Shuffled and truncated probes could expose systems that recognize objects but cannot infer actionable sequences. Potential product: A simulation-and-real-world evaluation harness for household, warehouse, agricultural, and service robots, with video evidence linked to decisions. Dependencies: Coupling between visual understanding and motor control, real-time latency, sensor synchronization, safety constraints, and transfer from benchmark videos to embodied environments.
- Industrial safety and energy — Early-warning systems for sequential events. Robust temporal evaluation could improve monitoring of power plants, factories, construction sites, mines, and transport infrastructure. Models could be trained and tested to recognize precursor-event chains rather than isolated frames—for example, a leak followed by pressure change or a worker entering a hazardous zone followed by equipment activation. Dependencies: Rare-event data, calibrated false-alarm rates, domain shift across sites and camera configurations, reliable timestamps, and human-in-the-loop response procedures.
- Education — Validating lecture and instructional-video assistants. The pipeline could generate curricula-specific tests for whether an assistant follows demonstrations, tracks multi-step procedures, remembers earlier explanations, and answers questions using the correct segment of a lecture. This is more demanding than checking whether a model can infer an answer from the transcript. Potential product: A learning platform that returns answers with timestamped evidence and distinguishes transcript-based recall from genuine visual or procedural understanding. Dependencies: Alignment with learning objectives, teacher validation, accessibility requirements, multilingual support, and protection against using benchmark questions as hidden training data.
- Policy — Standardized certification for multimodal AI claims. Government agencies and standards bodies could develop capability-specific certification schemes. A “temporal reasoning” claim might require resistance to options-only, text-only, single-frame, shuffled-order, and partial-coverage attacks, while a “video knowledge” claim might use a different threat model. Dependencies: Internationally consistent protocols, model-access agreements, regularly refreshed test sets, audit independence, and mechanisms for handling proprietary or continuously updated models.
- AI safety and governance — Continuous, adversarial benchmark maintenance. The paper’s pool, watcher, specification agent, deterministic selector, red-team gate, and lockfile could evolve into a continuously refreshed evaluation service. New items would be selected from diverse sources, screened against current models, verified by humans or trusted judges, and released with immutable provenance. Potential product: A versioned “live” benchmark registry with reproducible snapshots, attack profiles, source caps, and automatic regression dashboards. Dependencies: Preventing contamination, controlling model-generated benchmark artifacts, preserving test-set secrecy where appropriate, and avoiding a feedback loop in which models are optimized directly against the audit.
- Research infrastructure — Evidence-grounded, open-ended video evaluation. Video-Index currently uses multiple-choice questions, but its principles could be extended to require answers accompanied by timestamps, frame regions, event order, and confidence. This would reduce the risk that a model selects the correct option for the wrong reason. Dependencies: High-quality temporal and spatial evidence annotations, scalable judging, robust scoring of partially correct explanations, and safeguards against plausible but unsupported rationales.
- Long-video personal assistants — Memory over meetings, home videos, and wearable footage. An assistant could search hours of footage, locate relevant evidence, summarize events, and answer questions while adaptively increasing its frame budget. The paper’s finding that many long-video benchmarks improve with larger frame budgets indicates that production systems may need evidence retrieval rather than uniform sparse sampling. Dependencies: Efficient indexing, local or privacy-preserving processing, accurate event segmentation, user consent, storage costs, and protection against sensitive-person identification.
- Finance, insurance, and legal workflows — Auditable video evidence extraction. Robust temporal models could assist with claims inspection, incident reconstruction, compliance review, and legal discovery by identifying relevant timestamps and presenting the sequence supporting a conclusion. The attack pyramid would help test whether conclusions depend on the actual footage rather than descriptions or dataset priors. Dependencies: Chain-of-custody requirements, admissibility standards, explainability, human authorization, privacy law, and extremely low tolerance for unsupported conclusions. These systems should assist rather than autonomously determine liability or coverage.
- Model and hardware co-design — Optimizing evidence selection rather than raw frame volume. The observed differences among agents suggest a research direction in which models learn when to tile, crop, zoom, decode densely, or stop inspecting. Future systems could optimize accuracy, latency, energy consumption, and the relevance of retrieved evidence jointly. Dependencies: Reliable utility signals for selected frames, standardized compute and energy measurements, hardware support for adaptive decoding, and evaluation that rewards both correctness and efficient evidence gathering.
Glossary
- Adversarial filtering: Selecting or removing data by testing whether models can exploit it through unwanted shortcuts. “After selection, a red-team gate applies adversarial filtering \citep{le2020adversarial} to check the composition for residual shortcuts.”
- Adversarial attack: A deliberately designed method for causing or measuring model failure by exploiting weaknesses. “The attack pyramid organizes five attacker sets of growing access.”
- Answer-position entropy: A measure of uncertainty or distributional imbalance among answer-choice positions. “Appendix~\ref{app:t0} details the option level, with answer-position entropy, option counts against declared chance, and permutation sensitivity.”
- Benchmark exploitability: The extent to which a benchmark can be solved using a tested shortcut rather than the intended capability. “The exploitability of a benchmark at a level is the largest margin over chance that its attackers reach.”
- Blind baseline: A model evaluation that removes the principal input modality to measure performance based on non-visual information alone. “Blind baselines \citep{chen2024we} give a mean text-only accuracy of 6\% and a video gain of 11\% over ten models, with video above blind for all ten.”
- Bounded observer: An evaluator or attacker restricted to specified information and computational resources. “Following \citet{finzi2026entropy}, each level represents an observer with bounded access and compute, as developed in the appendix.”
- Coreset selection: Choosing a small, representative subset from a larger dataset while preserving important coverage or diversity. “Coverage-centric coreset selection \citep{zheng2022coverage} and QDIT \citep{bukharin2024data} balance coverage, quality, and diversity.”
- Cosine similarity: A measure of the angular similarity between two vectors, commonly used for comparing embeddings. “A copy attacker copies the answer of the nearest earlier question when their BGE-large \citep{xiao2024c} cosine similarity exceeds 0.9.”
- Cumulative profile: A sequence describing how a quantity accumulates across ordered evaluation levels. “We write for the exploitability at level , and the cumulative profile $\varepsilon_{\mathrm{opt},\varepsilon_{\mathrm{text},\varepsilon_{\mathrm{pool},\varepsilon_{\mathrm{frame},\varepsilon_{\mathrm{order}$ is nondecreasing.”
- Data contamination: The presence of evaluation examples or related information in a model’s training data. “Appendix~\ref{app:related} places the pyramid in the literature on benchmark validity, blind baselines, dataset artifacts, evidence grounding and evidence-seeking agents, data contamination, and benchmark composition and selection.”
- Deterministic selector: A selection procedure that produces the same output when given the same inputs and configuration. “Agents turn evaluation requests into specifications, and a deterministic selector with a red-team gate composes reproducible benchmarks.”
- Eigenvalue: A scalar associated with a matrix that describes how the matrix transforms a corresponding eigenvector. “where is the th eigenvalue of .”
- Embedding: A numerical vector representation of an item, such as text or an image, designed to preserve semantic information. “We then drop near-duplicate questions across sources and restrict repeated videos during selection.”
- Evidence grounding: Connecting an answer to the specific information or observations that support it. “Evidence-grounded benchmarks score whether a model finds the evidence behind its answer.”
- Evidence-seeking agent: An agent that actively searches for relevant observations before producing an answer. “Evidence-seeking systems choose frames before answering.”
- Frame budget: The maximum number of video frames provided to a model for processing. “We also test whether long-video questions need evidence that sparse sampling misses, with frame budgets from 32 to 1024.”
- Frame perturbation: A modification to the selection or ordering of video frames used to test temporal dependence. “Much of the score survives temporal perturbation.”
- Ground truth: The correct label or answer against which a model’s prediction is evaluated. “No attacker sees the ground truth of the item under evaluation.”
- Hypothesis-only baseline: A model or evaluation that predicts an answer using only the hypothesis or question, without the evidence normally required. “Hypothesis-only baselines expose annotation artifacts in natural language inference \citep{gururangan2018annotation, poliak2018hypothesis}.”
- Item-bootstrap analysis: A resampling-based method for estimating statistical reliability from individual evaluation items. “Item-bootstrap signal-to-noise analysis \citep{heineman2026signal} gives a noise floor of 2.0\% and a signal-to-noise ratio of 10.8, against 3.7 for the median source.”
- Lockfile: A file that records exact versions, inputs, parameters, or identifiers needed to reproduce a computation or dataset release. “A registrar randomizes option order, reports residual exploitability, and logs the pool snapshot, query, selector version, and seed in a lockfile whose hash names the release.”
- Long-context ability: A model’s capacity to process and reason over large amounts of sequential input. “Since the observed gains also depend on the reference model's long-context ability, other benchmarks may still place high demands on processing events across long time spans.”
- Meta-benchmark: A benchmark constructed from or used to evaluate other benchmarks or collections of evaluation items. “We release Video-Index, a meta-benchmark of the 840 hardest verified items of the pool, and evaluate open-source video models and proprietary systems on it.”
- Near-duplicate: An item that is highly similar to another item but may differ slightly in wording or presentation. “Within each benchmark, we keep one item per near-duplicate group, which removes 1.1\% of sampled items, refill to where possible, and screen the replacements.”
- Option-position permutation: A rearrangement of the positions assigned to multiple-choice answer options. “Across 70 complete runs, eight permutations of the option order give Qwen3-VL-8B \citep{bai2025qwen3} a median accuracy range of 5\%, significant on 15 benchmarks.”
- Predictive auditing: Auditing a dataset by measuring how well available information predicts labels or answers. “Following the predictive auditing of \citet{brown2025benchmark}, an online logistic regression on BGE-large embeddings predicts answer letters out of sample over five rotating splits.”
- Residual exploitability: The performance advantage remaining after accounting for the strongest tested shortcut. “A registrar randomizes option order, reports residual exploitability, and logs the pool snapshot, query, selector version, and seed in a lockfile whose hash names the release.”
- Red-team gate: A validation stage in which adversarial tests are used to identify and remove vulnerable items. “It tests selected items with Qwen3-VL-8B \citep{bai2025qwen3} at 32 frames and the caption attacker, then reruns all five levels.”
- Reference accuracy: The performance of a designated model under the complete evaluation protocol, used as a comparison standard. “Let the reference accuracy be the full-protocol accuracy of a reference model outside the attacker sets.”
- Signal-to-noise ratio: The strength of meaningful variation relative to random or measurement-related variation. “Item-bootstrap signal-to-noise analysis \citep{heineman2026signal} gives a noise floor of 2.0\% and a signal-to-noise ratio of 10.8, against 3.7 for the median source.”
- Shortcut attack: A strategy that solves evaluation items using unintended cues rather than the target capability. “We argue that a certificate is worth the difficulty of obtaining it without the capability, measured for every benchmark under one protocol.”
- Shortcut resistance: The degree to which a benchmark prevents models from succeeding through unintended cues. “Progress requires accuracy, shortcut resistance, and efficient evidence gathering.”
- Temporal coverage: The extent to which sampled video frames represent the full time span of a video. “We test whether accuracy depends on frame order and on coverage of the full clip.”
- Temporal perturbation: An intentional alteration of the temporal ordering or sampling of video frames. “Among the 51 benchmarks with both probes at this level, shuffling retains a median 96\% of the same model's full-video accuracy and one tenth retains 78\%.”
- Threat model: A formal description of the capabilities, information, and access available to an adversary or evaluator. “A benchmark's threat model specifies the levels its capability claim must resist.”
- Vendi score: A diversity measure estimating the effective number of distinct items represented in a dataset. “The Vendi score counts how many distinct items carry the observed diversity, and with we call it the effective size $n_{\mathrm{eff}$.”
- Video dependence: The performance improvement attributable to access to video compared with text and options alone. “Video dependence is reference accuracy minus the accuracy of the stronger blind reader, and a larger marks stronger resistance to text shortcuts.”
- Visual grounding: Linking a model’s interpretation or answer to visual evidence in an image or video. “These benchmarks check the evidence directly, while Video-Index keeps the multiple-choice format and removes the items that attackers answer without it, and the two checks complement each other.”