研究与学习 5.0 · 必读 2026-08-06 · 论文

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

论文用 trace-grounded parametric profiling 检查视频语言模型的事件计数能力,在 bouncing-ball contactsvisual blinksstate transitions 三类可控任务中生成 2190 个带 executable event trace 的视频结果显示模型在低频低事件数区域更容易成功,但在高计数高频场景只有 0.2% final counts 正确;提高采样率会提升最终分数,却不能显著恢复真实事件序列

打开原文回到归档

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

Source: https://arxiv.org/abs/2608.06361
PDF: https://arxiv.org/pdf/2608.06361v1
Content fetched: 2026-08-09T04:20:50.800307+00:00
Grounding: OpenCLI arXiv metadata and abstract

Metadata

  • Author(s): Sarvesh Baskar, Zikui Cai, Shayan Shabihi, Anirudh Satheesh, Muhammad R. Islam, Udari Madhushani Sehwag, Tom Goldstein, Furong Huang
  • Original date: 2026-08-06
  • Platform: arxiv
  • AAIF quality score: 5
  • Primary category: cs.AI
  • Categories: cs.AI

中文摘要

论文用 trace-grounded parametric profiling 检查视频语言模型的事件计数能力,在 bouncing-ball contactsvisual blinksstate transitions 三类可控任务中生成 2190 个带 executable event trace 的视频结果显示模型在低频低事件数区域更容易成功,但在高计数高频场景只有 0.2% final counts 正确;提高采样率会提升最终分数,却不能显著恢复真实事件序列

English Summary

This paper introduces trace-grounded parametric profiling for video event counting, using 2,190 controlled videos with executable event traces across wall contacts, blinks, and categorical state transitions. It finds staged temporal failures: high-count/high-frequency cases reach only 0.2% correct final counts, and increased sampling can raise final accuracy without faithful event recovery.

Intake Rationale

VLM evaluation needs timestamp-level event traces because final counts can improve without faithful temporal understanding.

Source Metadata

Source Abstract

Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, making failure modes hard to isolate. While existing programmatic benchmarks offer better control, they score only the final answer rather than auditing reported events against executable ground truth. To bridge this gap, we introduce trace-grounded parametric profiling for event counting in three controlled video tasks: bouncing-ball wall contacts, visual blinks, and categorical state transitions. Across 2,190 videos, we vary event count N and frequency F while holding rendering fixed. Each video includes an executable event trace for capability-surface estimation and timestamp-level evaluation. Our results reveal a staged temporal failure. At an 80% reliability threshold, Gemini 3.6 Flash reliably counts persistent state transitions up to 12 events at 0.5 and 1.0 Hz, yet demonstrates no reliable positive-count region for transient blinking events. Thus, event representation dictates whether a model initially accesses evidence -- a limitation that compounds as count and frequency increase. In the high-count, high-frequency regime, only 0.2% of final counts are correct and the model recovers just 18.1% of true events. To test if visual access is the primary bottleneck, we increase sampling rate. Although this boosts Bounce Ball accuracy from 19.6% to 29.3%, the reported sequence agrees with ground truth only 3.7% of the time. Extra frames can therefore inflate final scores without producing faithful event recovery. Different prompting strategies yield similarly limited gains, and real-world video evaluations show the same concentration of success at low event counts. Ultimately, trace-grounded profiling shifts video evaluation from aggregate accuracy metrics to a detailed diagnostic of where temporal reasoning fails.

<!-- aaif-entry-id: 7e003753 -->