Pith. sign in

Paper Citation Record · LEDGER

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding

As of 7 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 1 inbound Pith citation observation for arXiv:2506.06275.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06275 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:00:57.046150Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T04:37:46.090777Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T06:06:24.603613Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact2
  • verified fuzzy26
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3dbbe518-562b-43e6-b03c-8de5cd7c6d21 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.837684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.837684Z digest=sha256:c9db96942806950abe39c6eefea2f88c76377e9dd9f2fddb183c62eea39d0479

Observation 95538f57-de00-4044-aec6-42e5ecb61d05 · outbound

This paper cites Infinibench: A comprehensive benchmark for large multimodal models in very long video understanding, 2024.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Infinibench: A comprehensive benchmark for large multimodal models in very long video understanding, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.841252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.841252Z digest=sha256:d9e0fe449cf6df3e59d80e060055f4cdff4311ddcb85a15ba3a3679f5f99d13a

Observation 7fb0bc7b-3ca1-46ea-92d6-e076dd9f3974 · outbound

This paper cites Qwen2.5-VL Technical Report.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.843888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.843888Z digest=sha256:7705e44d5d52a03ec110cc4bffaafe3ffb1df797251d3a6b96527b649eb109f1

Observation a3a76b66-8636-4821-b3f8-77c24e7b0085 · outbound

This paper cites Memory consolidation enables long-context video understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Memory consolidation enables long-context video understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.847080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.847080Z digest=sha256:0cc2fb17c03986a54c0c77a147b756d1fb047891f7a4e82b3d92bbe02b1ef1db

Observation 179b2d20-9325-4ff1-b447-3b8fc0cc03b0 · outbound

This paper cites LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.849677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.849677Z digest=sha256:22c366ef9ee42f7a244c5044088a02cd083df7020055799fd318f16d348d61eb

Observation b2c254c6-ec9c-4f92-9412-132a96c5f277 · outbound

This paper cites Hadzic, Taran Kota, Jimming He, Cristobal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Fei-Fei Li.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Hadzic, Taran Kota, Jimming He, Cristobal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Fei-Fei Li

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.592958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:56.852611Z digest=sha256:720731a3358c0484e89c24057254e1f0ea2a884e39f5d89fd309ccf5922880c6

Observation d426b64f-a31c-4848-bfb2-48061f7ec977 · outbound

This paper cites CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.855381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.855381Z digest=sha256:7b7b06c04fa0ac26fbec85590e3d91e42d7439b84eb7c8e3d6683760a13995d8

Observation b4aafe70-e76b-4ea2-978e-0adce15a8fe8 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.858110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.858110Z digest=sha256:c443ef2f1c3659b35b7a4dd1bea7ac734f7a9e939302e6bc1b721e2fdb4182e6

Observation b8268873-e78d-4ac4-b87e-286d9f98e358 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.860531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.860531Z digest=sha256:26add0f69b4bc7dadc110adda32b897a0fa96373d07efe1fe1532730dd51e1fe

Observation e8a93478-79eb-43dc-9b5d-4b9eed8e11e0 · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.863411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.863411Z digest=sha256:1f14eab2c47b18b44405ca85d139109a56af876001215bf698f0921b8a0e3e1c

Observation 616ae0b2-6f27-4c8e-ad54-e29f750b57f8 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis, 2024.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.581669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:56.866203Z digest=sha256:71bee85ebdf4735c0f8d2c1fd93cc0770d4a63e33c0786afc0fd8f733676dff1

Observation 88b87ae7-8177-4abd-ba3e-c8ff682564b9 · outbound

This paper cites Video-mmmu: Evaluating knowledge acquisition from multi-discipline professional videos.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Video-mmmu: Evaluating knowledge acquisition from multi-discipline professional videos

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.574381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:56.871356Z digest=sha256:8b246b9920cf54711108301f8304a9ffb7e00805dd1c0b56edc4f47a888533b5

Observation 5ce38d0b-2dba-4664-bd79-57751542e4b5 · outbound

This paper cites Movienet: A holistic dataset for movie understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Movienet: A holistic dataset for movie understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.566263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:56.876297Z digest=sha256:c55d2ded0d16066afa2a6400ae80964420c6e6ad8a4f9efeee45843e23942da5

Observation a35c169a-c774-4c9d-98fa-faf3a21aa74c · outbound

This paper cites Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.557448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:56.878674Z digest=sha256:a3b3a1002bdb0619a25853ab0842b7cb5d1dfc6dfa81db231000a0070c077ee2

Observation 23e8bf43-a19a-48ba-be62-8e24581ad37f · outbound

This paper cites Needle in a haystack - pressure testing LLMs, 2024.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Needle in a haystack - pressure testing LLMs, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.548694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:56.883819Z digest=sha256:3e3c8a17a18e54d5947a8062be4a3b4381d4d62f104d641596a66dcb0447713b

Observation 43e4fb47-36f0-4e12-9846-706375296cd4 · outbound

This paper cites One thousand and one pairs: A “novel” challenge for long-context language models.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding One thousand and one pairs: A “novel” challenge for long-context language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.886154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.886154Z digest=sha256:0616c7ba663cc06578b82cfc42b05d7cd7ca44364f87bf9d49aac0ef456982e5

Observation 7f4ee9c5-c60f-4997-8830-9cd8754bbeba · outbound

This paper cites TVQA: Localized, compositional video question answering.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding TVQA: Localized, compositional video question answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.888635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.888635Z digest=sha256:874dce783f0ed98873177418f656c761e1083ebe0aba6ba833bfefc8b9e6579b

Observation 51d57200-a121-49dd-aee8-b6e92f6f5435 · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474, 2020

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.891037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.891037Z digest=sha256:8f7e79b80af884d387b26fd87fbc3593dc61b91bea4523e566fd147aa97bb6f4

Observation 75aff597-cc40-43a7-83aa-f1702c5a4ade · outbound

This paper cites Merlot reserve: Neural script knowledge through vision and language and sound.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Merlot reserve: Neural script knowledge through vision and language and sound

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.536206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:56.893414Z digest=sha256:d1228eb8ec4b75a0c2b777a0fbb391767a4aeb748fd1baa2d3b1c3b990fabd50

Observation 61a95390-c468-4b95-992e-c6f6636ede24 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.895831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.895831Z digest=sha256:7ed4eac1510519eb1eba98f6cd3813e489c8d7f0db18732473326c603b177ea7

Observation d0161f92-a0c5-4a8c-a35f-243b89c77a80 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding VideoChat: Chat-Centric Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.898389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.898389Z digest=sha256:e90aa4c01626930bcdee3ba379db4e1ff10398d0aa467eb4063c57e91ae6de6c

Observation 4eb44370-18c3-4b9d-9de5-3e902697d56d · outbound

This paper cites Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.901023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.901023Z digest=sha256:41d90e3309553e115aec87332cfb74dcae45c06bf548ccb4c4c600e45a51b47a

Observation 23218241-3d3b-4161-b4b2-224ebe739eb4 · outbound

This paper cites Contrastive decoding: Open-ended text generation as optimization.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Contrastive decoding: Open-ended text generation as optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.903725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.903725Z digest=sha256:da17f988619aa36d184a6588a2a495732330ad0787c8289287f86ad4b5c3bc92

Observation 4985de5e-2abd-4ad8-9b8b-15a208ca95b9 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models, 2023.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Llama-vid: An image is worth 2 tokens in large language models, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.906287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.906287Z digest=sha256:658059dd9fabcb3d15ab2a0aeddca4e3faa120ad1073a4be6c65c92c7216ebf1

Observation 10def0f5-5245-457a-b658-355be0530ddf · outbound

This paper cites World model on million-length video and language with blockwise ringattention.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding World model on million-length video and language with blockwise ringattention

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.524542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:56.908532Z digest=sha256:f2878b67569872f3969993eabaf59f4930914bdd2ecd8951fc15a62fff7dd83f

Observation 967fb098-2b18-4567-a4d8-4143df8e4060 · outbound

This paper cites Visual instruction tuning, 2023.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Visual instruction tuning, 2023

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.910861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.910861Z digest=sha256:3820e08700ecb6a8d136f18de8d0bbecea8da52f600e2899769a55c63d0fb970

Observation 3d98c38e-8bda-4777-b07f-04ba65877ee6 · outbound

This paper cites Is your video language model a reliable judge? InThe Thirteenth International Conference on Learning Representations, 2025.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Is your video language model a reliable judge? InThe Thirteenth International Conference on Learning Representations, 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.513146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:56.913190Z digest=sha256:0a1119bb2df737dfd95529fbeaba1e63bc00930779e9aa8e3a6ffb69f8cc64b3

Observation 76eae7de-0c9e-4fe0-901d-087e9b3766f9 · outbound

This paper cites Nvila: Efficient frontier visual language models, 2024.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Nvila: Efficient frontier visual language models, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.505461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:56.915429Z digest=sha256:f3163dda55f485a57c29c521f636c896b300b547470dbf78d980a1340044be7b

Observation dd4ac394-dfec-4cd4-8db4-7047bd0bcaf3 · outbound

This paper cites Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:00:57.241715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:56.917881Z digest=sha256:4892577b9585b3e71eb694b7da5d9a76626e2f9df399423327a3c4bb32dd0293

Observation e7ea8119-c001-4313-b8b4-2dd6293465b7 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.920387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.920387Z digest=sha256:c4633cbef0db5fae0f309d5056a5fe8c4ded5f576301345aad4e606c524fb1fb

Observation 30c21deb-f754-42cd-8196-a1b5eeebe7c5 · outbound

This paper cites Valley: Video assistant with large language model enhanced ability, 2023.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Valley: Video assistant with large language model enhanced ability, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.497326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:56.923163Z digest=sha256:134f3a999b93fb786067e3cfc6572bf87d7447f165a846b2dfca32aac3de2449

Observation 58b872c7-8679-4cdd-bba8-3cc549cdbb46 · outbound

This paper cites VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.925414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.925414Z digest=sha256:a3a1aa1adc02ded2c8ab9dfc52a19c37b0ed24d6503407f163beb8300ac88a88

Observation bc071e67-7a13-4624-bd0d-15159caaee83 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.928000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.928000Z digest=sha256:f87352122050c49e899b2fc84baf20242253a5aa2e019413877d1ad09a7fda79

Observation d3ebb7c6-b8f5-4554-81cc-2595dd0a4c2b · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.930356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.930356Z digest=sha256:b4696c147979767d5e3b53a9a183a90733c9df5037b9dbdfa9d99990a111fe81

Observation 98d7f9db-20c5-445b-8e17-87402094ef20 · outbound

This paper cites Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.933234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.933234Z digest=sha256:1103836a54116d3c5bfb4f1023f894aee0eedb4535a5742a746b30b8d22ae915

Observation f37af5f2-89b8-4bcd-af6f-c0557e44a17c · outbound

This paper cites Neptune: The Long Orbit to Benchmarking Long Video Understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Neptune: The Long Orbit to Benchmarking Long Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.935935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.935935Z digest=sha256:d5c4199b54d6255f78539e9353bc3a01bc85a755975326fc13079130ade548a0

Observation 65d1812c-b78e-4778-84de-5d824fb31caf · outbound

This paper cites GPT-4o System Card.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding GPT-4o System Card

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.938513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.938513Z digest=sha256:68b90375ebecf243f57afcd77a14f683822248e122ab4b040a7b40685237d516

Observation 88b23c33-f7a5-4752-8eda-c3b0699c32fc · outbound

This paper cites Movie plot analysis via turning point identification.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Movie plot analysis via turning point identification

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.941552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.941552Z digest=sha256:7e4cfd3cf028a538446deea68871302dbc77bf263997bcb1fb13572224edd889

Observation b5976af9-30bf-423d-aa0e-455f78386ac3 · outbound

This paper cites Screenplay summariza- tion using latent narrative structure.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Screenplay summariza- tion using latent narrative structure

Reference 40

Resolution
verified exact
doi, observed 2026-08-07T06:00:57.084836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:56.944168Z digest=sha256:603bf7e04b106bc658d3d6d542cfafdb08ae87cdafafe81c5f0a2df3cdcbc5ef

Observation 5900ee82-62fd-4757-bae9-cfd7a231a0fc · outbound

This paper cites EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.946680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.946680Z digest=sha256:5bcaa8597076f27925fd6759460f2a154d711530fef9a891cb93d9913eccbaeb

Observation 886a64cd-0adb-4b7c-a86c-86f03b35f197 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Learning transferable visual models from natural language supervision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.949537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.949537Z digest=sha256:e73b53b1d8968ebd4a2df1f410a8b4c8be1ca3a3422af0361067a9b34c570a82

Observation 28df8268-ac92-432a-820f-81a3a296aedd · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Robust speech recognition via large-scale weak supervision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.951890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.951890Z digest=sha256:b16707ff54ce402e5661cb47e35470ba5959459987d7351bcf975b645a6363ee

Observation 13e2c0fb-30ac-409e-9f52-7726c30ad6ae · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.954513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.954513Z digest=sha256:ee7221f1fada28937cd1e583ab5057e71e3b98de25f1eb6ea849f27847061916

Observation 6f61abf0-2a63-4875-8b84-cfeeb0d97834 · outbound

This paper cites $\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding $\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.957099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.957099Z digest=sha256:a20032bbca546d447698343a4bd6ae9a4b24682edcdd5e32ef6318a6e19ad803

Observation ae2d1c1f-8cc0-49d6-834a-5a54be98c14a · outbound

This paper cites Trusting your evidence: Hallucinate less with context-aware decoding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Trusting your evidence: Hallucinate less with context-aware decoding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.479032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:56.959860Z digest=sha256:aeddb51f22da067ce8ef2700b9c9486e8a4f8a75583a71ad5c6f84848a104312

Observation 835245b2-76a1-49c6-86be-dea589a25624 · outbound

This paper cites It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.964951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.964951Z digest=sha256:3bd8d0b76274b837a225d2a7bc736190fa47761fffaac240a9e2cdd10f175a8e

Observation 7e5fa39a-d49d-44f5-9ae8-4c8effec6e48 · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.967687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.967687Z digest=sha256:039c5e33694f50b658b96800d06021a7877b84bb4505eba1299ab951f8b3bc65

Observation a72b4ca9-7698-4895-9ef0-3b6547706f43 · outbound

This paper cites MovieChat+: Question-aware Sparse Memory for Long Video Question Answering.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding MovieChat+: Question-aware Sparse Memory for Long Video Question Answering

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.970276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.970276Z digest=sha256:3d49f4e6b5de80bac234598d1dd645c97383a36159cb5508df7eb06b34feaf66

Observation 075f24ad-0cd7-47cc-b7f8-d969bae98396 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.973042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.973042Z digest=sha256:1c01481a3cc73a5feeacbb35807f18549b8b520aec7cb8ef1444b8be5b30a3cf

Observation a695d49b-f432-4f8e-bcdb-a69c8560da44 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.975636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.975636Z digest=sha256:af6a1be27517e1426cd0cf16bde60093ea18b2cf5fc4b75fdff6b5a3a4304014

Observation 08be5ddc-2573-4a46-9640-3a4072847048 · outbound

This paper cites AdaCAD: Adaptively decoding to balance conflicts between contextual and parametric knowledge.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding AdaCAD: Adaptively decoding to balance conflicts between contextual and parametric knowledge

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.471151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:56.978154Z digest=sha256:b661e931dd5fbcb5f6447d2b947a79e694c86e1d6cdc0f0ce189e3acbbc8ce06

Observation 7f3d84b0-1d89-45bf-afde-9d63081f84c0 · outbound

This paper cites Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.980537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.980537Z digest=sha256:cb675640397b4e208ddea09d6e6a5bc8cd29eb60e8971f8f9d4d4225604cfa53

Observation 0bfbd83a-e957-481a-a63d-885e263d8b0d · outbound

This paper cites Lvbench: An extreme long video understanding benchmark, 2024.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Lvbench: An extreme long video understanding benchmark, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.983169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.983169Z digest=sha256:5b01a2a44922b2f67d9025f7dae8c1a38964cc2c309130491bda7e2ad60aefda

Observation 9fb1d83c-2578-4f51-8ecc-0ca55ff4e8ab · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Videoagent: Long-form video understanding with large language model as agent

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.985963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.985963Z digest=sha256:55ca62536544a47335b48f3e9c94b6eb6d34ee2440f1bc8f15a6783055301147

Observation 462885d5-2bf7-4164-9e8f-e35caa3bab70 · outbound

This paper cites Videollamb: Long-context video understanding with recurrent memory bridges, 2024.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Videollamb: Long-context video understanding with recurrent memory bridges, 2024

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.455993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:56.988510Z digest=sha256:e54bcdc357428621154f5040050774ec8d2ce8453e0446b6230028bae7df6eb6

Observation bafa5580-7d0f-462f-a507-2c1989dff989 · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.990818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.990818Z digest=sha256:8cf378c0b666ff632a8742709f2434c7bcb6f60b14adb2c19e131420d9ae409c

Observation ccb18b42-d2ee-4817-86cf-18b7b39e0c50 · outbound

This paper cites Tenenbaum, and Chuang Gan.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Tenenbaum, and Chuang Gan

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.448759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:56.993378Z digest=sha256:faeb3f3570c1d9a7047d9fa158594d64f9f21d601130eb0679c860e5ed466da1

Observation 6fc27423-4722-45b9-9f75-2e30d3e446ed · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.995632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.995632Z digest=sha256:5f7c4c6f27ca1f5493b8392d7eea43a07a1be0182eee9ece966d7748a8ee5fd5

Observation 6760223f-58d2-4ec1-bbd7-148a5b92981f · outbound

This paper cites Next-qa: Next phase of question- answering to explaining temporal actions.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Next-qa: Next phase of question- answering to explaining temporal actions

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.997977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.997977Z digest=sha256:85316b4036991a1728c174c43cf7242281fbd236087b1a6ecf0997afd37f6a32

Observation 0dbb8263-fde6-49b7-8c90-ba9d9797e36a · outbound

This paper cites Qwen2.5-Omni Technical Report.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Qwen2.5-Omni Technical Report

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:57.000343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:57.000343Z digest=sha256:fa07421d02c9e3bcb9b470fdf2e71df368691a8426d855f69675c4e4e69b1e5e

Observation 64d2832e-edb9-4cf8-be2f-ef3b14498fc6 · outbound

This paper cites Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:57.002924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:57.002924Z digest=sha256:0ac00afb8ccc90fbef1d189405ac34fe422d6870711ef43eb5aa88b99d1e607e

Observation a62d46c6-1036-497c-96f4-38cdbb35b660 · outbound

This paper cites Just ask: Learning to answer questions from millions of narrated videos.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Just ask: Learning to answer questions from millions of narrated videos

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.430391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:57.005333Z digest=sha256:b114d4b124378dddecf0e1f516ad6710b26ec71ec9a2e75fac93057e2df8a00b

Observation 9d7c74ba-3881-466f-8a23-3ed41b43d464 · outbound

This paper cites Justice or prejudice? quantifying biases in LLM-as-a-judge.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Justice or prejudice? quantifying biases in LLM-as-a-judge

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.423281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:57.007608Z digest=sha256:775027a34fc515da42de448edbb9a5e45c9e24be02744723827bcc3d81e10dd5

Observation c3633974-9125-46c9-90b8-b5523148329a · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.415781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:57.009993Z digest=sha256:7ac2a0037d734f19a9b021714c0b31e9cc9b3c506a1a1e4db7c71a65e972791c

Observation 50f3b593-8d14-4364-9b52-56505b331cd0 · outbound

This paper cites Merlot: Multimodal neural script knowledge models.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Merlot: Multimodal neural script knowledge models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.408407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:57.012364Z digest=sha256:e5f05637e8a5ea7d2457575bd6dd7e53e11533f389673576a30891946f96323d

Observation 1944a34a-1968-42dc-abc7-646a0bf944a0 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:57.014587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:57.014587Z digest=sha256:bcd13b3e54043742c7c83838b24a0a06090346d1175841d60d341e75de43b622

Observation 10bd0a87-1047-4102-b7ed-356e09ae2a06 · outbound

This paper cites A simple llm framework for long-range video question-answering.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding A simple llm framework for long-range video question-answering

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.401439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:57.017356Z digest=sha256:68f3737c3f16acd68f9cd844255e211f2d0a6c107c7c9bbc07ce38ffe7b82403

Observation 1e4a762f-7acb-4b53-9d18-3bfc005b705b · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:57.019680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:57.019680Z digest=sha256:a27fb9a4f2a19a3df523f0cc6c7f7a184e9e62a26e5b464eaf596cbd12a6a7c6

Observation 21a9f31e-4cd3-4917-a883-0863f50c32fb · outbound

This paper cites LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:57.022279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:57.022279Z digest=sha256:2dee1f265c4c62af2cfe75826f6b523b66297c35f25532a50c8118f3cb85dccf

Observation 3b014608-1e4e-4c3d-bae1-f2fe67a62cb5 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:57.024929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:57.024929Z digest=sha256:0d9f7b262d78da145e3d9d28f800c4e366505905c4474247dc517af6f065eb34

Observation a42027df-1fc4-417e-b038-935d905e835b · outbound

This paper cites Needle in a video haystack: A scalable synthetic evaluator for video MLLMs.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Needle in a video haystack: A scalable synthetic evaluator for video MLLMs

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.393887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:57.027321Z digest=sha256:38c43ee8b0e6b5b8a6b4b06ce24f6478c778eb3add54dc4577e7da0ceee7a167

Observation 1bede31b-70d3-40c3-9001-27aa1b135876 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:57.029729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:57.029729Z digest=sha256:d715bafed5ce7d12797c2d606334ff1601b402e3315977c69bf0961a2fd3c76f

Observation ef4cd2f1-7031-42bb-99f4-518c23f15f8a · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:57.032403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:57.032403Z digest=sha256:1b09ccdc658b179dffc69909fbea1888bb3c14605b3c550047ffc1167a557b0b

Observation b500fd35-3dca-4ee8-9c01-3246711c998a · outbound

This paper cites The two claims should differ by minimal edits, meaning they should be as similar as possible while maintaining contrast.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding The two claims should differ by minimal edits, meaning they should be as similar as possible while maintaining contrast

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.386454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:57.035514Z digest=sha256:612f09770708188f24aaa64c114bd0c74f5a74b6c918ceb0c63d155475239e45

Observation 9ae08d2f-83f7-4ef2-b045-6e00a89a15c4 · outbound

This paper cites Examples for Reasoning Granularity.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Examples for Reasoning Granularity

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.378904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:57.037997Z digest=sha256:6b38d798198a6a2383ebf5418d6a005db950fe3ce1bf3280f68283889b1f0130

Observation 18497632-1550-42c2-84ff-5226ef626da0 · outbound

This paper cites Other" and suggest a new category. Note:The categorization is based on both claims (fact and fib). Check the examples provided in the “Examples for Comprehension Dimensions.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Other" and suggest a new category. Note:The categorization is based on both claims (fact and fib). Check the examples provided in the “Examples for Comprehension Dimensions

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.371793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:57.040534Z digest=sha256:1b95650f9684e2ec3a069d178e73d57c86694e97789a60265cdbfc2cedeca9d8

Observation dda476c8-6b53-4b78-8c16-5a594fce1e60 · outbound

This paper cites Pay attention to details and context in the movie, as some claims may be subtle or require careful reasoning.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Pay attention to details and context in the movie, as some claims may be subtle or require careful reasoning

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.364350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:57.043281Z digest=sha256:03c190be1ec524e95c8a2e25fef263a74a4182e23da97a3953f6d6cfecdc7b1a

Observation d9a24bfa-ba3a-4c3b-8cc8-b974e1502508 · outbound

This paper cites Start Classifying Claims.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Start Classifying Claims

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:00:57.356777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:00:57.046150Z digest=sha256:3be2dc706332c91123c2d6fb6189ac146c0a6d9823d897ff91b10b3b08f84744

Observation e2eb4797-bc4b-4baa-8094-4ec82f17dc74 · outbound

This paper cites doi: 10.18653/v1/2023.emnlp-main.308.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding doi: 10.18653/v1/2023.emnlp-main.308

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.881019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.881019Z digest=sha256:35eb35eed4ba76bcff5c6c9f4e04c5d891e7c373b61fe5cdea7fba189e4a8348

Observation 8dda0679-3ce2-4ba9-9350-a42df4b54a9f · outbound

This paper cites doi: 10.18653/v1/2024.naacl-short.69.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding doi: 10.18653/v1/2024.naacl-short.69

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.962238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.962238Z digest=sha256:80a82d00c6934e0f44da26d9a660cf64c6bd61f34f5bec07dd1981c81432dd48

Observation 83866043-2fe4-421c-baad-5889f5fab71d · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.873858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.873858Z digest=sha256:957d6e2123e17fd775f3cdb88d33f8999f42b715245b28c0b39e994f5d1675db

Pith citing papers

Observation 0247282f-46b0-4ec8-a7af-5cfc4b9fa650 · inbound

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding cites this paper.

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:24.607724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T04:37:46.090777Z digest=sha256:fdb11d9960cca9b8389b9aca7e89e79bd536c6f722c83d8891f99fdb326354ca