Pith. sign in

Paper Citation Record · LEDGER

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

As of 10 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2608.06361.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06361 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:24:55.008240Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d7c84f27-de22-4830-958c-c31c17eba640 · outbound

This paper cites L.; Panda, S.; Meghwani, H.; Singh, J.; Dua, K.; Li, P.; Sheng, T.; Ravi, S.; and Roth, D.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping L.; Panda, S.; Meghwani, H.; Singh, J.; Dua, K.; Li, P.; Sheng, T.; Ravi, S.; and Roth, D

Reference 1

Resolution
verified exact
raw_fallback, observed 2026-08-07T04:24:56.000894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:52.144137Z digest=sha256:145bd0984d20df79c0cfc41013b179d043b353eef93fd54300cdad5c467535ac

Observation 07944c51-e15a-4a6c-a568-eb97ff54b32d · outbound

This paper cites Qwen3-VL Technical Report.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:52.231517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:52.231517Z digest=sha256:923165cf574d85c0499dd19816f77464618b948ea2b36d01891682803e9e735d

Observation 8c0f946c-1274-42ff-8e08-c1fa16448b78 · outbound

This paper cites MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:52.386597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:52.386597Z digest=sha256:54291cb1799f1225223f088d00b5d4c4645e5cc21c3017a9cc6fe94b53449ed9

Observation 599d86f9-2b43-4dd5-94d9-a12d01fc3ef8 · outbound

This paper cites M.; Kota, T.; He, J.; Eyzaguirre, C.; Durante, Z.; Li, M.; Wu, J.; and Fei-Fei, L.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping M.; Kota, T.; He, J.; Eyzaguirre, C.; Durante, Z.; Li, M.; Wu, J.; and Fei-Fei, L

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.691178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:52.488003Z digest=sha256:6cc88d93a77e5a4fa1dbb33903fef69850a227e23c55ffde73c4ae4419667098

Observation 671a9d24-71dd-4040-a285-22661e74e1c7 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:52.643870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:52.643870Z digest=sha256:d417cf8c299115f5990a016bb055cb699ccc7dab79daaf6e76517310302de514

Observation 0805b9db-5acd-4cf9-8d63-735f9a1322c1 · outbound

This paper cites P.; Li, W.; and Gong, S.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping P.; Li, W.; and Gong, S

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.672317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:52.734869Z digest=sha256:3fc542e85b5c959697b7d156204d3e76c3a249019f1c29b7a82bc8472d8fafaf

Observation 07139ee8-aa73-416a-984a-5340b765adc4 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.655945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:52.830178Z digest=sha256:fc520dad050c71a05445783fa24bde70620c4d4233a6c11561df148473ce75fe

Observation cfcddce4-ae72-4b94-b290-f290aef5a431 · outbound

This paper cites Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:52.914333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:52.914333Z digest=sha256:e64329fbc49b516c3c95e0787dd2967332de6d2894415bc411afcf3ae0040a32

Observation 823347c9-3cd6-4b20-86b5-a49d03ba5008 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.639322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:53.012924Z digest=sha256:4157f06d48155c41d28a7b4320c9975a1488f87fafddde60df20346ce9ede81c

Observation 92ed64fa-c936-432c-a326-d779fce80713 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.620228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:53.214323Z digest=sha256:99768a3276e7010e74c0ffde94a6d50990e8a25b7cf186316fffb6a1e2de21fb

Observation 2f4fecf4-dedb-4797-aab3-cfc5ecd2ec38 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:53.296769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:53.296769Z digest=sha256:19fefdba3dfb8bc689ea7aeca28c9c5955e2287e534aec1c43c9e09c18cf4260

Observation 30ee07bf-454e-4f99-823a-6742c436921c · outbound

This paper cites W.; Li, L.; Yang, Z.; Wang, L.; and Cheng, Y.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping W.; Li, L.; Yang, Z.; Wang, L.; and Cheng, Y

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.602422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:53.418701Z digest=sha256:241fe647985e57ef09ef15129f28aabe0db06cd5375a6c4a705a1512019a6165

Observation 5a226575-a1be-4601-b9d0-1b1f8ab7ca45 · outbound

This paper cites L.; and Girshick, R.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping L.; and Girshick, R

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.578561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:53.726452Z digest=sha256:c1255a1541830ec393a752f10d19b0f5884b2ee67c3b4c6064e061151232fedb

Observation 81537453-421d-4dd5-add1-b1ca663a9126 · outbound

This paper cites VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:53.867176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:53.867176Z digest=sha256:85f2259c4a3d71c8c5cfdade87ac35b224890a235c5487e43c47003553c92087

Observation d91c9d61-57eb-42aa-9c96-71709166d49e · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.025621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.025621Z digest=sha256:69ae0e4a1c496b3fae8104b3b70b7f03ea009dc8268a735dd59fb93820fb1fa2

Observation 7677c8b8-5d86-44cb-9938-e5aa4cda0546 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.559694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.245785Z digest=sha256:c336451b404662c2618a2bdc61f3f1b10807abdf4580de926de66bdad3ecbd3e

Observation 5d1fd752-2a2d-4fd6-905c-cd9935f0481b · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.540311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.384918Z digest=sha256:491568fa3d232fef0e1fe143bc38ea8ae463d7f1d2e36c9ead6805b747c34901

Observation 6ebfd26c-e9d8-4278-9bdb-640d56553c47 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.511970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.511970Z digest=sha256:580d5609df3e3cc300e0e971610860a0e14b913fe72f4a0bbacf35e50e967bfe

Observation 17e824f4-ed54-4ee8-8f78-5cd733c5cd47 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.524376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.664409Z digest=sha256:7c0fa9fb490e573ba3ab897dd5e5bbfb3306017dc0682a6fad12615cc6c18589

Observation db54f330-9e3e-4d4c-ac75-3c597c212482 · outbound

This paper cites The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.749022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.749022Z digest=sha256:aa25b3282f7c614dbcbdbf4cc64f75c22f575f9d92d100217440a4e2fc91f1fa

Observation 62aba644-7176-4f22-b1ed-11b45250e50a · outbound

This paper cites Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.755039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.755039Z digest=sha256:c2d41223b0905548c4b478c22884b45583365c00bec14e6d739f439b8adbad32

Observation bb9418fb-fd26-484e-91b6-02a8fea9c261 · outbound

This paper cites OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.759751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.759751Z digest=sha256:38111629bf422e031fe3b45e0d501a89cdb9d6bc3b8d7866c90c3af8141dba3d

Observation bf8a1780-e20e-43be-a311-3ad12ba35cb3 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.764865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.764865Z digest=sha256:8c41a9f6192cdf9c946a518cc8fabce1933c6e2faebe8a65ae6308d3f174bac1

Observation 8e51ee67-2932-4e6c-a1ff-4b3e0843eb3e · outbound

This paper cites Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.770078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.770078Z digest=sha256:f900dc0a97d24b5085b34a0def3f09a8befd28dd66fb06e94874ed2085de8741

Observation 37f66cc9-eb61-476a-bf83-4e07eab41068 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.777262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.777262Z digest=sha256:68247e610b1efa051a4ead327c3b7b30586747e5f0a5f627d3a16b17bb8ae5cb

Observation 9a10bc43-c406-4344-ad18-d58920138353 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.505970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.788332Z digest=sha256:42c21a6125eb191286d7bd21161697c87170e974ac95a55e8042d20e0d0901ae

Observation abc2a3de-42ec-41a4-9a5c-bffa94bcc04d · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.793516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.793516Z digest=sha256:ab2774899fcba2926a64c0cd13ab21392b79b62e869aa648f673d29b7dd85511

Observation cb2382fa-fb81-4374-8239-61807de95740 · outbound

This paper cites J.; Huang, Y.; Liu, Z.; Qu, P.; He, J.; Chen, J.; Yuan, Y.-J.; Han, J.; Xu, H.; Li, H.; Sachan, M.; and Liang, X.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping J.; Huang, Y.; Liu, Z.; Qu, P.; He, J.; Chen, J.; Yuan, Y.-J.; Han, J.; Xu, H.; Li, H.; Sachan, M.; and Liang, X

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.799125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.799125Z digest=sha256:e367ec9ed0ac45064f6054b77bda2ae8809225bfa3f79cce719699190bc17eb1

Observation 19cbbeb2-51d4-4e30-8d45-ca82d8efd493 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.809548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.809548Z digest=sha256:827576c43fa14b169593fafeb99fbd9be1df5a70eec3023907b2e56930576bfa

Observation 5f0f0fda-46cf-4be4-b234-0ebd0af11645 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.815382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.815382Z digest=sha256:a979f6601e1c589d7f6d1f99988c3a55ab903e79aabd7b2115c7689b0a7b3230

Observation 5a27b7c3-fa34-427b-bced-52db97aefa1a · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.474953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.821197Z digest=sha256:7ef4b80fa6c4d346433576a33d38aeb09856699e3d8c3c2a804da7d16f1c9920

Observation 1b8f7891-dae9-4ca9-8a64-7008f05e4b0a · outbound

This paper cites T emp C ompass: Do Video LLM s Really Understand Videos?.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping T emp C ompass: Do Video LLM s Really Understand Videos?

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.453882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.826419Z digest=sha256:8f1c5cc65991172dcb19ed42a5c718cb8bc15fa1e4c08742fc243ced750202f7

Observation fd09c947-e957-4f7e-a9c0-a0ae21b9e44c · outbound

This paper cites Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.435710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.831264Z digest=sha256:8ef8afa907ab1ed33515f256801a01fb31158c63b8be012f58b17c55290690c8

Observation 7f2dcae5-0272-4da8-9aad-5be1c79680d3 · outbound

This paper cites and Kota, Taran and He, Jimming and Eyzaguirre, Cristobal and Durante, Zane and Li, Manling and Wu, Jiajun and Fei-Fei, Li , booktitle =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping and Kota, Taran and He, Jimming and Eyzaguirre, Cristobal and Durante, Zane and Li, Manling and Wu, Jiajun and Fei-Fei, Li , booktitle =

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.418861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.836116Z digest=sha256:85c15869617e00b614ec634e8b0fec83696e655ae013b55ffb0f4849239430de

Observation c3b7d9a1-f192-4f65-9048-b2c86a063eeb · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding , volume =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding , volume =

Reference 44

Resolution
verified exact
doi, observed 2026-08-07T04:24:55.067902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.842447Z digest=sha256:7e1124c0ecb803a9bb6015b665c0e47bf9e7fa4393a12613b908fa1a260f83da

Observation 8f5df2b8-86b4-45ec-be73-207a60f414fc · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.401075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.848233Z digest=sha256:f41d78b77f35f0694503fa7a62d68dfce29cf4f5af6238739e0135c442eb097e

Observation ace52e79-d713-41aa-9ddc-050d36c33230 · outbound

This paper cites International Conference on Learning Representations , year =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping International Conference on Learning Representations , year =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.381945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.853320Z digest=sha256:0dcf46ce0175c530ea616dc028d01961d6ff37b38dc76a159388410f359d8917

Observation 15810289-8cee-4a45-9f3b-74ca14ececa6 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.363864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.858091Z digest=sha256:1bee7ccf9cddb229240f9553077c6ae005229f099afd3c0c11b6eea6ec03a183

Observation 6db93d44-64d1-43b9-b14c-0093100b1754 · outbound

This paper cites PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.863248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.863248Z digest=sha256:c27e72c04f83b8f8d0a21c39f5e9f3a186ac489b9e25b6124c6eb4112aaaefb5

Observation 4ccd3904-b946-4a70-a5d7-9cfbd058001f · outbound

This paper cites arXiv preprint arXiv:2512.05091 , year=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping arXiv preprint arXiv:2512.05091 , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.868214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.868214Z digest=sha256:632deb984669b80db89e145952ce483e7862258e40c16bc114fe29259ea84e39

Observation 679bf50f-250d-406c-b904-95b71e93cf43 · outbound

This paper cites arXiv preprint arXiv:2505.23359 , year=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping arXiv preprint arXiv:2505.23359 , year=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.873262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.873262Z digest=sha256:833fe1f8dffa6416c652ee83ab87b5e8c1e5e244074416f07fb0766ba45656c8

Observation b2341d0e-c233-400c-8f1b-1e587130ecf7 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.346770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.878051Z digest=sha256:c42d8421a42e4d103313a3f4fcd9bbc4e99ccc52b2e9b8a4128c51f7eea0ea64

Observation fa2ed78d-82ba-4cda-bd07-c3d25ac51509 · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs , volume =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs , volume =

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.330599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.883261Z digest=sha256:7e8d80f8c41aeeaa00641b3884727a6af7e687c288a25a3e2d3f1beb1409e3aa

Observation 5788ee92-2d80-4450-8b1d-78067fbe5342 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.310306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.887607Z digest=sha256:29d451b1cd8e5255d0cc8aec548ead31f863ee0136fa1dfc75b241353e7b1d94

Observation aa320b76-db6a-4cf0-82d2-21f43de77622 · outbound

This paper cites TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T04:24:55.795302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.891710Z digest=sha256:c27174c389676d853422f6370613d13598b72fb32da98b67d14b0a734294bf2f

Observation d5e61dbc-7926-43b7-8d8c-e253a4975770 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.895820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.895820Z digest=sha256:82789127f7e212825c93a79272832d07b9b12d76208e58cbeeaa5df9c38b020d

Observation 4dce49a0-883c-435e-8dcb-31d07b9cface · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.900531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.900531Z digest=sha256:4c7531ea45c5260c66254487ffc17a92b230e33208270a09ad5977bdde2fae5a

Observation e964b1e2-40c0-4100-bca9-cc1c452ce07f · outbound

This paper cites 2026 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2026 , eprint=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.904905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.904905Z digest=sha256:40361ae26a3b1ea643a5cb5c29d4177591655bd0713364bd9117e89a3deb3ba8

Observation 5e58298e-02f4-4315-aaad-bb7d062ea2b8 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.909336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.909336Z digest=sha256:10b626723130a84f762937b55f1e1983e9a084100c04df499c8aa6ffb1bc5278

Observation 22ba40ce-b387-4a60-892b-8dcfff417bf8 · outbound

This paper cites Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.242939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.913877Z digest=sha256:ff4632c802685df6cc2e1e3392c0324c1826ce121a12d8c49e6a4d02fc9265a1

Observation 4290319c-0200-434a-b1a4-ef745eae8798 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.223694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.918039Z digest=sha256:512a5c166eaabe3ee53e87dc96b82d52a6ea85214b30b340aeafde85df4d98f1

Observation fe6093ce-4615-4fca-a0e7-e9bed56f58aa · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.922152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.922152Z digest=sha256:c123f5fc9401666a93fb8968302a41d1d31cf2e150cf305f3bdcbf63b854b972

Observation 95aa17f3-0196-46c6-b638-0ca70215d7d3 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.926274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.926274Z digest=sha256:8858e980feb77103446f8aa8e2a70645caca491accc13bfb47e7d3b3622ad269

Observation 42bea15d-cbf8-483c-b6a9-f9a4b53578d6 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.179831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.931250Z digest=sha256:3af35a017151aef35d7d3e01b77839134b21d9ebf1a33a04380bf57abfd4363e

Observation 15e4ae1a-fc2e-4373-9667-8bfd1a58c922 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.935746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.935746Z digest=sha256:f3dc77fe2cc0f9065556585fa75119cf0c5ed3150141384fa92cf3556dcd6888

Observation 9aeaa774-301f-46f9-ab75-6440462d76a2 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.939825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.939825Z digest=sha256:44f458a8dc08a3854f8c19133fc9f34756a8883de19c86551b0e8ebda6af00e8

Observation a6797bd3-346a-4119-82f5-e3ecdddf8424 · outbound

This paper cites Proceedings of the 42nd International Conference on Machine Learning (ICML 2025) , year =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Proceedings of the 42nd International Conference on Machine Learning (ICML 2025) , year =

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.133964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.944352Z digest=sha256:d770d6a7bcb502ba729d94be3b820abb1c06b248bb66fd339236a96d5580a866

Observation 8e7c17d2-d4e7-448a-9acf-e130fe5d8cea · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.114458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.949338Z digest=sha256:ac8aed261d24ce9be0f1dae088ade7df743ccb178b38ef36e54044eab80a6eee

Observation 779460ca-fa69-4b33-b59a-c3a0e5d7bffa · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.955160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.955160Z digest=sha256:70e996cc008ebda49bdf1e4622a1de18397a1587393acaef484b7b3234d95f6d

Observation 2283cd70-da8b-488a-a33d-4d1eaa6eec6e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Advances in Neural Information Processing Systems , volume=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.960617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.960617Z digest=sha256:71f38e75994c90c6733656516743b1ab26ecb4d4871e35f7bb59f616d369557d

Observation bcea1e89-b1a5-452d-9e38-2ce30b87d4b6 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.965931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.965931Z digest=sha256:eac3b969293f00705e73f8cea50c3c9eef9330a26066e6bb192223ba6ee1eb2a

Observation 95794f40-93bf-49ef-8cf6-3c44bd445f98 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.970881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.970881Z digest=sha256:7b1827ef557df9372769e1be215b3ad921b87d7f1c1f57fa1b16287d4bc3263b

Observation cbe1aa48-295d-4d78-82f9-45ef25c57e8e · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.054314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.975540Z digest=sha256:6255cae640b0ff45779c6086046c7c01400f4554656c0f3729e1cba9fd6cdaec

Observation e059e240-3680-4a48-8baa-dbd830e98b72 · outbound

This paper cites VideoVista: A Versatile Benchmark for Video Understanding and Reasoning.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.981358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.981358Z digest=sha256:bbccb64d39b0dff1b820df1c042f626c84ef103894131887bd7482a9ff906f7b

Observation 44b4a782-ef62-4217-b57a-179974c28124 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.985958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.985958Z digest=sha256:c4fc7ef9dff664526aa182f9f558cab01af4da17617edfb475fb352c4046049c

Observation 4e648f65-6112-46d9-9f6a-e5484246083f · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.991413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.991413Z digest=sha256:486a54d2617b3df800313d90aaa8102ef8e950a2147a22ecf8aaca0ab41a603c

Observation 6472183e-2c3b-4b5c-9936-59eaf987834c · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.036757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.996689Z digest=sha256:77b94393da8e8f27aeb8d0bc78f463c3c1b6d076a2b0fa59b28994dc1ffcf77c

Observation cd04825c-b05d-4ba5-9fad-add7deb39880 · outbound

This paper cites Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:55.001360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:55.001360Z digest=sha256:15763174371fca4d2c81399cd58ff717777e0eaaa8f67d125fc8cddea05ddcc2

Observation 529b2d0c-38a0-4bd4-8414-23b60bbb4008 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings , year=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings , year=

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.018883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:55.008240Z digest=sha256:0aa49c453c821e1619ab4450f93c8bbf7383e3bb9c9d35a6beb12539b21db1ae

Pith citing papers

No inbound Pith citation observations are available.