Pith. sign in

Paper Citation Record · LEDGER

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

As of 14 August 2026, this Paper Citation Record lists 100 of 127 outbound references and 3 inbound Pith citation observations for arXiv:2511.20272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.20272 v2

Coverage vector

measured 100 of 127 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:23:07.705686Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T16:37:06.384435Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:38:39.619743Z

Reference resolution

100 of 127 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1c2785ed-4cd7-488a-8e38-8df42133fc3d · outbound

This paper cites GPT-4 Technical Report.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T20:22:59.834670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:22:59.834670Z digest=sha256:ee6d7fc4fb7eddff46248cc0dbe009d53fcc7134f1bcd6ae837e2b27f7602564

Observation 0e4113ed-1a88-4f11-bed2-153b4d47d61b · outbound

This paper cites On seeing stuff: The perception of materials by humans and machines.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs On seeing stuff: The perception of materials by humans and machines

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T20:22:59.937363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:22:59.937363Z digest=sha256:c86f1a310aa9c87c0806d2e6b131a384065769aa1b31d45a43578dd87cd0dc7c

Observation 40d649dd-1e16-41bd-ad37-01207a4f42b5 · outbound

This paper cites Vqa: Visual question answering.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Vqa: Visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.093535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.093535Z digest=sha256:bcb483da382f2f1fd1c23a5251cb91a6161ded1edccd1e2223a4c1e32d8224d4

Observation d19b8ecb-31a9-4742-8a73-f796e7c8d198 · outbound

This paper cites Eye-contact, distance and affiliation.Sociometry, pages 289–304, 1965.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Eye-contact, distance and affiliation.Sociometry, pages 289–304, 1965

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.331240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.331240Z digest=sha256:359a8c6aec328c597d46ab816f2722e1c0ebbbcd3f27fe6977f2f6502f9eaa51

Observation 14fda1d8-9815-4d28-9671-ecacd20eb630 · outbound

This paper cites Qwen2.5-VL Technical Report.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.406527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.406527Z digest=sha256:7c78b03dc7e04b84d08659c364569159b58e4e1bd617747850fd468a1516cada

Observation 1dae183b-ab4f-4652-9b30-8dcee8737c6d · outbound

This paper cites Representing the existence and the lo- cation of hidden objects: Object permanence in 6-and 8- month-old infants.Cognition, 23(1):21–41, 1986.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Representing the existence and the lo- cation of hidden objects: Object permanence in 6-and 8- month-old infants.Cognition, 23(1):21–41, 1986

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.484396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.484396Z digest=sha256:0df53b2c2b8743fcf1366c911b414bc3239c92c5755a3cfc4a085a2d85ea199b

Observation c53a5db6-45f7-4cc0-8170-3321c28fdfa7 · outbound

This paper cites theory of mind.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs theory of mind

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.572585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.572585Z digest=sha256:934d71c03aef6734a48ca783c02ab3da8475bb52f4682483ac10c8de5a89ff4a

Observation 858ff559-25c1-4076-bb04-fd8184e2384b · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.688451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.688451Z digest=sha256:028e065a9238834816320b9ddecbb7d3a572dc0a6731f833e61f081e67b18bdd

Observation 98860f4a-b3b8-4ce3-91d0-263f87216e4c · outbound

This paper cites IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.771656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.771656Z digest=sha256:eb5169a449c8228a32469b5d4392249481ed694070cb65c3eba6650dcd87d26e

Observation 0d473e7b-2837-4ab2-b10e-58c29a3aeabb · outbound

This paper cites Routledge, 1995.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Routledge, 1995

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.854703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.854703Z digest=sha256:9bf9ee9369f423fbdbbd9a45432c78964a44de73d86b49a175a1a4568698dff4

Observation 0c586aa5-8c9d-4252-b923-e4201d3f038e · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Activitynet: A large-scale video benchmark for human activity understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.937721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.937721Z digest=sha256:512ddadb77aba90b61111f78c265678db03590ed89d3f1a131a637e2f83165c3

Observation 4715a1fd-4758-4444-81f4-330e1b44d5c2 · outbound

This paper cites Rextime: A benchmark suite for reasoning-across-time in videos.Advances in Neural In- formation Processing Systems, 37:28662–28673, 2024.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Rextime: A benchmark suite for reasoning-across-time in videos.Advances in Neural In- formation Processing Systems, 37:28662–28673, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:00.971832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:00.971832Z digest=sha256:07bec8d86c2c6eb3657a5aec32e3ee683cfb3b1cfc44c83b68ade4df44e0434b

Observation f61da4e4-5d2f-474c-8887-1cd84e193873 · outbound

This paper cites Are we on the right way for evaluating large vision-language models?Advances in Neural Infor- mation Processing Systems, 37:27056–27087, 2024.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Are we on the right way for evaluating large vision-language models?Advances in Neural Infor- mation Processing Systems, 37:27056–27087, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.064028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.064028Z digest=sha256:6507ced029947de94ad34bd27ea0db5a0dc13508f239df52d8913d0776ae00db

Observation 96d07928-9fd3-4a1b-a2cd-89e6c7c79f24 · outbound

This paper cites Video-holmes: Can mllm think like holmes for complex video reasoning?, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Video-holmes: Can mllm think like holmes for complex video reasoning?, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.175032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.175032Z digest=sha256:ccda7dd56f788f6f4ce943e9340f3eff545d72fa5a4d27ae6df215f35f1e3878

Observation ed3b046b-4305-4e84-b01a-dee2201bcd7b · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.290978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.290978Z digest=sha256:b75b174e53d56c3e36c7163bf6fcf3a99b0c2a0a416fca30bf74de3303c19a19

Observation cbe8be57-de9b-4f62-8e83-9a4890fa0284 · outbound

This paper cites Le, Sergey Levine, and Yi Ma.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Le, Sergey Levine, and Yi Ma

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.403982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.403982Z digest=sha256:f4a33506b3c9625c31fe6925c51fbaa6b5a22d6c3b6a3530907cab4b22cadd4c

Observation f2bc5c7d-9545-4df6-9dde-7ab60c9cf8a1 · outbound

This paper cites Whatever next? predictive brains, situated agents, and the future of cognitive science.Behavioral and brain sciences, 36(3):181–204, 2013.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Whatever next? predictive brains, situated agents, and the future of cognitive science.Behavioral and brain sciences, 36(3):181–204, 2013

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.472386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.472386Z digest=sha256:9e47f9e81052590c3194b02c8801b92d82a0376c55c8cc144630e8f110089603

Observation 7e7f228a-2550-4553-9916-365f8ead4e66 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.531512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.531512Z digest=sha256:716e0a7803fe8b4195b275a40b8655d64be4a169b39f66a3125ceb44a0df406f

Observation 5223b74d-761a-450a-9ae5-af8fed1c7b9a · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Instructblip: Towards general- purpose vision-language models with instruction tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.579499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.579499Z digest=sha256:615afe6da55140deaf0bd9729eafd86e5ecf103733c5ae642395445d27535d1c

Observation ca6d9e86-e844-4883-b7c6-7f82d5ec0538 · outbound

This paper cites MIT press, 1987.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs MIT press, 1987

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.643970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.643970Z digest=sha256:5bae0bd853068fd5dc13dd8ee05457c4f80638c58213908038a2f8bbe3af654d

Observation 35bb7093-aff4-45a3-ba58-47c1837078ba · outbound

This paper cites The llama 3 herd of models.arXiv e-prints, pages arXiv–2407,.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs The llama 3 herd of models.arXiv e-prints, pages arXiv–2407,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.727373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.727373Z digest=sha256:5a3110a6f9a415c2a60f02437a3ebe22d6413e9bbc3b0c46f14f731863dd4685

Observation 049c5852-561a-4b6b-9c0a-9ad9ce89d05c · outbound

This paper cites An argument for basic emotions.Cognition & emotion, 6(3-4):169–200, 1992.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs An argument for basic emotions.Cognition & emotion, 6(3-4):169–200, 1992

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.812885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.812885Z digest=sha256:8048961ba977cedcb132ee24f01811a873989e17910f8e8ea3fa11b2c0dfa4a7

Observation 15443549-ad36-4cf8-a9df-f10b57c5db37 · outbound

This paper cites Unreal engine.https : / / www.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Unreal engine.https : / / www

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.869913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.869913Z digest=sha256:5548f62cc171bbb6ff47194721536a3ff0d30f20f9aa748985ea4249a577dfe9

Observation 0141789d-b487-475d-867e-941fac8706f2 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:01.923190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:01.923190Z digest=sha256:a1b1a33bea0f3dda0c2df252f9c62b165603d0c5fb622ebbfbbe0d9891fd850e

Observation df25d2c5-5169-468f-a084-b69b1fc46b41 · outbound

This paper cites Material perception.Annual review of vision science, 3:365–388, 2017.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Material perception.Annual review of vision science, 3:365–388, 2017

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.003342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.003342Z digest=sha256:a078ec0a4cd8d86509f30d7b939325148b855d84e08141d7d3ae7cbf6267ec57

Observation 6163625b-6a14-42e8-8f7b-afdba3cc4faf · outbound

This paper cites The free-energy principle: a unified brain the- ory?Nature reviews neuroscience, 11(2):127–138, 2010.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs The free-energy principle: a unified brain the- ory?Nature reviews neuroscience, 11(2):127–138, 2010

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.033703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.033703Z digest=sha256:f0a8f31cd5697ce78a399aee360b8d114e6d4f3fc017dbbd405c75bb25987b08

Observation 9948e519-a3c9-4ab7-a4cc-a60d18381849 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.091017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.091017Z digest=sha256:abac314b624e891a332d4c16dc40e852be635b2cefa86ab7c3c8f21b23701049

Observation 2f5062b3-88b0-4b28-bca1-026698eb8739 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Blink: Multimodal large language models can see but not perceive

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.138277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.138277Z digest=sha256:0fdecab66514d56ece102a2867484dd75fd9b2331b6680acf7b8e233211f6ba1

Observation a30fc0d5-5e85-45b8-9ccc-84d95dcfeaf0 · outbound

This paper cites Atomistic Control in Molecular Beam Epitaxy Growth of Intrinsic Magnetic Topological Insulator MnBi2Te4.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Atomistic Control in Molecular Beam Epitaxy Growth of Intrinsic Magnetic Topological Insulator MnBi2Te4

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.182852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.182852Z digest=sha256:db180c1f3adb5f8fb33f8167783f2c72b823b00e5ee145486266ba95164b329d

Observation b3b05772-dd96-409b-b6e7-3cbca0d93443 · outbound

This paper cites Houghton Mifflin, 1979.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Houghton Mifflin, 1979

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.249905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.249905Z digest=sha256:13a924901697c273cdc67486bd814db194a9a0070732a53deee9cbe5865b9876

Observation 37ef0134-653e-4edd-bef5-ff46717a1752 · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs The” something something” video database for learning and evaluating visual common sense

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.321899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.321899Z digest=sha256:59dd021c7d06d4bbb284b19fc5d365381adc203640f5753e623e6087dca0856a

Observation f3e67a5c-82eb-49a4-b17f-4932b2dacfb9 · outbound

This paper cites The Llama 3 Herd of Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs The Llama 3 Herd of Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.394429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.394429Z digest=sha256:8dcb2e9b03cb11b69f20afebb3f2875868992374099816db626f020f722d5852

Observation 0cba16a0-eb5e-4666-88cf-e2a04c2e0cf4 · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language halluci- nation and visual illusion in large vision-language models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Hallusionbench: an advanced diagnostic suite for entangled language halluci- nation and visual illusion in large vision-language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.480865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.480865Z digest=sha256:db8aed9d04f1a9390353a5810d1dc3b600c06c0669d4ab00cd3966aabe736fbe

Observation 6411cd8f-eded-43a2-b78b-3cae95e7e5d4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.549143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.549143Z digest=sha256:abe95759529cfcf97ab9eeed77a85b16d9230b87ec2373e5ec795ffa7aa6f5b8

Observation dc1af7d2-65d2-48c4-8ed2-51d6db7532a2 · outbound

This paper cites Doubleday, 1966.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Doubleday, 1966

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.624952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.624952Z digest=sha256:f95cc0743da51d3ffd5a0e99bfda1d831046caf04d8a0836278e6bc61473e762

Observation 10724494-018f-430d-935e-7415a1293efd · outbound

This paper cites Glm-4.1 v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Glm-4.1 v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.726794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.726794Z digest=sha256:5b0ffe20c0ce903ac79b1206f89e76b85c5bd69a905a67d0cc0264a64b8e944c

Observation f56e9cb6-14ad-429a-82cd-1483b68685bf · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.778666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.778666Z digest=sha256:f527d38acaed864799fc2c6a2d874bb62e87f0547fcbe49dc63c2cb02eebf136

Observation 5f8a72b2-eb6f-47ee-9d55-dd7ae9dce488 · outbound

This paper cites GPT-4o System Card.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs GPT-4o System Card

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:02.877204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:02.877204Z digest=sha256:d44431e6f244ed09ae7ec4398c0345db0ffeeed532d82db94f463fccc40c01a3

Observation 678d0199-f37c-47c2-98c6-88ad754fba18 · outbound

This paper cites an unresolved cited work.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.044818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.044818Z digest=sha256:561a04a2448d90d7a75088d41e573bb498accaa83fea8284558f50df7d6f7c91

Observation da0fb4f2-cd52-41d2-a5f9-f584920f3438 · outbound

This paper cites From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.179935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.179935Z digest=sha256:f99698a10ea85d1b004fcb6e2af5407b2d3386256ff40dd6e2a67b20671203e0

Observation f9851049-ad81-41e2-94d3-c00edd3d4bf3 · outbound

This paper cites Towards Social AI: A Survey on Understanding Social Interactions.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Towards Social AI: A Survey on Understanding Social Interactions

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.378338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.378338Z digest=sha256:cdf959c2bf91427cb1d1fb4fa57e8f04b17faf797f423778b5fa1f7cd3bf91af

Observation d511dc7e-7895-4907-9c2b-c3b73d49c0dd · outbound

This paper cites What is More Likely to Happen Next? Video-and-Language Future Event Prediction.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs What is More Likely to Happen Next? Video-and-Language Future Event Prediction

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.468092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.468092Z digest=sha256:5cad7f65c06589832ab2a37671dc6281e1dc36202b828aa6ab0fc8cd35f912ae

Observation 6f908d3f-9d90-465a-9c6f-a5905c14acff · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Detecting mo- ments and highlights in videos via natural language queries

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.619120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.619120Z digest=sha256:0f78926b6d14eb07276aab79e47e9677f15d46816f24a3c230002bf952a17996

Observation 224dc2ff-e5db-4249-bfec-2fe6ff741b00 · outbound

This paper cites Anisotropic flow, flow fluctuation and flow decorrelation in relativistic heavy-ion collisions: the roles of sub-nucleon structure and shear viscosity.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Anisotropic flow, flow fluctuation and flow decorrelation in relativistic heavy-ion collisions: the roles of sub-nucleon structure and shear viscosity

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.810455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.810455Z digest=sha256:a2da162ddc444cb6a088195cdb73e8a1866f11b8a1b332b2f8e9c1dd055823cd

Observation fd73e304-7c6a-4557-b36c-1e76ed112956 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.031544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.031544Z digest=sha256:6cd051f68bd277bf60ca8c5507ab50b5e78ad39d146ea529ec9756568f86d5b6

Observation cf5d0930-c3d0-443f-a63d-92be2ecef23d · outbound

This paper cites From represen- tation to reasoning: Towards both evidence and common- sense reasoning for video question-answering.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs From represen- tation to reasoning: Towards both evidence and common- sense reasoning for video question-answering

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.223426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.223426Z digest=sha256:9b9b894b4a24c5624dfe3b1cef6e625985b228164056857baddbbe80ac20b1cf

Observation f72b758d-fcdf-49d2-b4ba-33bbc058f01c · outbound

This paper cites Mvbench: A comprehensive multi-modal video under- standing benchmark.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Mvbench: A comprehensive multi-modal video under- standing benchmark

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.327558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.327558Z digest=sha256:15ae8ba1169c5413279ff64953782a18f8fdd66d1e85550b1170d49c7d6e5113

Observation 4fc421a3-3b21-4ed9-bdc0-1e7dec310b92 · outbound

This paper cites Pope: A simple method to hallucination eval- uation in visual question answering.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Pope: A simple method to hallucination eval- uation in visual question answering

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.422362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.422362Z digest=sha256:78052f1ef633ee496466b25d84eae8fa5e67c3bbe628bda5330963be42ca3d21

Observation 278caf4c-6ebd-4563-beb2-ea12b21b56cc · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.532245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.532245Z digest=sha256:8df7c8a035ab05c1a0ab1fabc616faf60dda3428d1293c8ca7bdc4b0dfe6f26c

Observation f1a887c0-7353-4b9f-bf56-c24768bc63e2 · outbound

This paper cites Core Knowledge Deficits in Multi-Modal Language Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Core Knowledge Deficits in Multi-Modal Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.626496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.626496Z digest=sha256:542ab9c2cb153a5085ae8aeb18dab869c70a5d9f62bb983ae4eb1ac73c398975

Observation 9ea4dfa3-11d9-4158-9b50-b17e8074ccdc · outbound

This paper cites Self-Rewarding Vision-Language Model via Reasoning Decomposition.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.725008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.725008Z digest=sha256:c3eaaa5245f966377be58f42820ba87f9a0b6080b0c3a06587fc1930ec308de5

Observation a7619fe7-b755-48d1-89b7-3d73dba6fe7b · outbound

This paper cites Explainable Multimodal Emotion Recognition.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Explainable Multimodal Emotion Recognition

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.826214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.826214Z digest=sha256:6aaf67766515a31500c2a2837b18e18bafdec365c925fb94f5e088ac4f2017b3

Observation b4d6e639-687f-48bd-b459-b9d41f707997 · outbound

This paper cites Improved Visual-Spatial Reasoning via R1-Zero-Like Training.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.926420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.926420Z digest=sha256:02d06166efea277b689a42e0923035b4840651348d857b874c5425539062cbb9

Observation f150adbc-f5dc-46ac-81f6-0a1fa312d5fb · outbound

This paper cites Microsoft coco: Common objects in context.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Microsoft coco: Common objects in context

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.034079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.034079Z digest=sha256:8a8cf2497d2e5554e4b8282ed8c4f6721fb019d114ff9963a7d9adaf528dc6bc

Observation 9038ad2b-7b01-4e72-b0ee-82aab3cd54fa · outbound

This paper cites SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.200466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.200466Z digest=sha256:14774d1009c8e153516abbbcf3269a1e31e64b49bf259b8047da5584653d7321

Observation c6d62e94-c635-4f87-a6b4-aa202eed3c8b · outbound

This paper cites Generative Physical AI in Vision: A Survey.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Generative Physical AI in Vision: A Survey

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.304254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.304254Z digest=sha256:858da0c950f949214aa33c93808818609f26bb306a31849f90c4aa4258460a73

Observation 4725a73b-cd00-4335-bd38-b554a4a47c85 · outbound

This paper cites Visual instruction tuning.Advances in neural infor- mation processing systems, 36:34892–34916, 2023.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Visual instruction tuning.Advances in neural infor- mation processing systems, 36:34892–34916, 2023

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.407372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.407372Z digest=sha256:a35fb8e85770ed20b05e18ec300b8f8bc993513bcceb4ddc5ac75634486f7219

Observation e9aebeb5-ba66-4fb6-b7fc-a4e731d52021 · outbound

This paper cites Improved baselines with visual instruction tuning.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Improved baselines with visual instruction tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.487338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.487338Z digest=sha256:08e02b87d6f0b4286888b2613e680c55b022823375b93f3841b4f8b372a777dd

Observation 4b6cbdf4-27dd-42ad-b025-e4b7e614b806 · outbound

This paper cites Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.553933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.553933Z digest=sha256:7e9d7c01e5b2bb86ff08bb0b59ab624dffbf7917c4f6f745770885349f1ee9a6

Observation 2c559219-0d9f-435a-a826-b1d462cbcf51 · outbound

This paper cites Mmbench: Is your multi- modal model an all-around player? InEuropean conference on computer vision, pages 216–233.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Mmbench: Is your multi- modal model an all-around player? InEuropean conference on computer vision, pages 216–233

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.605866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.605866Z digest=sha256:508312ab53a5a9787bf57a447a9d3811abbfed50a453e0d064c003ec14f5026f

Observation 75e2c42f-5191-4b9b-bbce-3ebab6794364 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs TempCompass: Do Video LLMs Really Understand Videos?

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.644159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.644159Z digest=sha256:79c55dd986105cbf319a57ffe1bd933af005b74ae292e72b9166b8cbc834a423

Observation 0eb26b62-1cfd-4c03-b18b-baa9752989bb · outbound

This paper cites Videoreasonbench: Can mllms perform vision-centric complex video reasoning?arXiv preprint arXiv:2505.23359, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Videoreasonbench: Can mllms perform vision-centric complex video reasoning?arXiv preprint arXiv:2505.23359, 2025

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.701446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.701446Z digest=sha256:624a281a7642f50e2e20900651a794a351405c9b78198cef0b882b5d9650812d

Observation 98b90b1a-9759-4ffa-bb92-81e248525fb2 · outbound

This paper cites When thinking drifts: Evidential grounding for robust video reasoning.arXiv preprint arXiv:2510.06077, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs When thinking drifts: Evidential grounding for robust video reasoning.arXiv preprint arXiv:2510.06077, 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.749238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.749238Z digest=sha256:1e9f05977f98d87f43cfd74394541eed67e25f378d3f98a163db2daaf66c1013

Observation cf187c95-adac-4098-b1bc-88bddfc558b9 · outbound

This paper cites MIT press, 2010.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs MIT press, 2010

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.792172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.792172Z digest=sha256:a91a62909b50857dd3b707ffd094a0e3686eddbfe5c723630467541cf0af38e6

Observation a349a015-7c69-4313-89ee-8323b3926575 · outbound

This paper cites Ba- sic books, 1988.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Ba- sic books, 1988

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.833758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.833758Z digest=sha256:7784bd21f2e98a426a9b1576bc6e4e38de15843304ecd24fc80a9bd1a1746464

Observation e1d5a2e2-5c32-4bba-a7c1-b626cb290a91 · outbound

This paper cites Clarendon Press, 1978.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Clarendon Press, 1978

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.895137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.895137Z digest=sha256:5d498d63725869fbd8c87cc70385583a9dc0a7d26310790a97e5d447742caf97

Observation 2e884371-8bcb-4f6f-b24f-2d478ec58b54 · outbound

This paper cites On visual knowledge.Frontiers of Information Technology & Electronic Engineering, 20(8):1021–1025,.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs On visual knowledge.Frontiers of Information Technology & Electronic Engineering, 20(8):1021–1025,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.945973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.945973Z digest=sha256:0c0a3436b66665f36bba34f4decc15b544f12d8896e6c006ba42148b2d9cfcf0

Observation 8baf9970-6bcf-4ac3-a585-5cac0dc0bd69 · outbound

This paper cites Basic Books, 1954.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Basic Books, 1954

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.986585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.986585Z digest=sha256:08929321990027a7edd1c48445e3bd78b745eef3cb34da416cbb82fd1df80a4b

Observation 814d2dfb-5841-4200-b9ad-4dbed6af28e2 · outbound

This paper cites Free Press,.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Free Press,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.041280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.041280Z digest=sha256:a1f4d141cd289081a4c8e5a6d2bb25b98590a6c68606755c02acc6071ca34b8b

Observation 383f3a06-06dc-41c2-888f-4915c0f5158e · outbound

This paper cites Does the chimpanzee have a theory of mind?Behavioral and brain sciences, 1 (4):515–526, 1978.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Does the chimpanzee have a theory of mind?Behavioral and brain sciences, 1 (4):515–526, 1978

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.169383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.169383Z digest=sha256:7eff565b9b1587362a3b45b137dfd3157c276f260ae91568922bbde71bcf82e6

Observation 8085e3b1-b330-4261-bfe7-ff04d5b2faa8 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Learn- ing transferable visual models from natural language super- vision

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.206756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.206756Z digest=sha256:4a946817d70faa8916b7ce5be0662e60d81de5f71c7c8521039b34458336cf99

Observation 76b23818-ceb9-458e-a77e-a1b1a0fec841 · outbound

This paper cites Robust speech recognition via large-scale weak supervision, 2022.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Robust speech recognition via large-scale weak supervision, 2022

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.241697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.241697Z digest=sha256:518d2e946269439d448502b725de36a68accbf92c3ac96713fcd7bfbaa85402d

Observation 8e03999a-7e75-490b-8664-0a14bed4c345 · outbound

This paper cites Reka flash 3, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Reka flash 3, 2025

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.292109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.292109Z digest=sha256:e547171fe371a7c94e06bc0eb9fe1a0a2b9f7fc9a8ac3dd3d877099c1d124d90

Observation 37092672-20ea-423e-a20d-b2de8430b06b · outbound

This paper cites IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.339842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.339842Z digest=sha256:25daca84af8f8a086b893cf7f7380e4ad629875e492d5ac587d66c9867af5b57

Observation 0835e414-529d-4a0f-bb76-38dd27af434c · outbound

This paper cites Wellness Insti- tute, Inc., 2001.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Wellness Insti- tute, Inc., 2001

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.400203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.400203Z digest=sha256:3b92a7be8afceaf5293b2320fc3bc9bf3401e1a08035ac266661fd9b59dc9d18

Observation f2da25a9-8ede-48f8-89bd-94ad7b8478ef · outbound

This paper cites Lawrence Erlbaum Associates, 1977.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Lawrence Erlbaum Associates, 1977

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.439459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.439459Z digest=sha256:ebbb3a493c0bb6b7a82feb68c224c25f10055a461ead22bf42b91197fc39ca8c

Observation 655df557-4eac-465e-882e-bdf5827dc90f · outbound

This paper cites Core dimensions of human material perception.Proceedings of the National Academy of Sci- ences, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Core dimensions of human material perception.Proceedings of the National Academy of Sci- ences, 2025

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.482435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.482435Z digest=sha256:db7b9f206d0d8ec73cad6ff96950bec27792ff61336f230abb01a46a8c364934

Observation 85483ddc-b72c-445c-a6be-a44e7720ef2f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.535503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.535503Z digest=sha256:4dab5bfd342bc23229db249063d0f8d0014c2c56d27cd244ecf171632163c789

Observation dd66fe63-c1d4-4f88-802a-62595ad9b606 · outbound

This paper cites Towards vqa models that can read.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Towards vqa models that can read

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.578820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.578820Z digest=sha256:63ec4a0fd9d545346bcba2b9bca22a27593fbd3b3062653ce6c0037fd27d4028

Observation 4babefde-5f86-4720-bc87-6e121df70ef2 · outbound

This paper cites A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.635646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.635646Z digest=sha256:bb620cdefa8d72f22159f90ba76e160920768fade048e453b6eb9edebe9e9a99

Observation bd8084a8-b7ec-405c-a36d-98be1261629c · outbound

This paper cites Origins of knowledge.Psychologi- cal review, 99(4):605, 1992.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Origins of knowledge.Psychologi- cal review, 99(4):605, 1992

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.685266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.685266Z digest=sha256:453ae7776e4765e2c22aceebc540838ae81279a9355559492e262b8ea1c2f368

Observation c5496cb1-740a-4094-994e-73864e367016 · outbound

This paper cites Mimo-vl technical report, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Mimo-vl technical report, 2025

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.722348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.722348Z digest=sha256:fefa90f30722648a43c623ed8787fe7ff3bb8e5f4feb616b130fcfe43e059c8b

Observation 595fa52c-9b86-407c-876f-a754121fa28b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.763176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.763176Z digest=sha256:69b156dea5fe555d3ee2829916eae27cebbafb59544b1edeb0d354aca0cf61e7

Observation 19aaa2ea-9a29-43ce-92bb-673ffdf86ce9 · outbound

This paper cites Qwen2.5: A party of foundation models,.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Qwen2.5: A party of foundation models,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.811188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.811188Z digest=sha256:1691f13f0636c0ed731ca28436095569f66c08b28ac80b2560cf25fa60740878

Observation 300bd0a9-c482-4590-9ca2-ee6940c64633 · outbound

This paper cites Trl: Trans- former reinforcement learning.https : / / github.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Trl: Trans- former reinforcement learning.https : / / github

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.869311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.869311Z digest=sha256:4d984623b7b4cbfde3829ae26ae4034ab780ee9b6bb36c93fe88d913a72ebb8e

Observation a632ed77-af1f-40e5-b70b-9ade9057808b · outbound

This paper cites Make Your Training Flexible: Towards Deployment-Efficient Video Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Make Your Training Flexible: Towards Deployment-Efficient Video Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.938062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.938062Z digest=sha256:fe42de9e81e95aba80b85807f57ca7cfdb17700c73e14bde9c76b29e43419596

Observation 7e6f7a3f-13d5-4433-b15e-9a32d1ba58a9 · outbound

This paper cites VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.982542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.982542Z digest=sha256:fe32e976b4797dc59717ede78faddf5e9e758345d7689220184b18db04aa2bfd

Observation 04382e29-c0ac-4578-8504-e3b22ab54b6d · outbound

This paper cites Videorft: Incentivizing video reasoning capabil- ity in mllms via reinforced fine-tuning.arXiv preprint arXiv:2505.12434, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Videorft: Incentivizing video reasoning capabil- ity in mllms via reinforced fine-tuning.arXiv preprint arXiv:2505.12434, 2025

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.036858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.036858Z digest=sha256:9a5cc28a54d6a3772af195ff1429727fc7a26891b6d10bb4fd484ecd72a9c1bb

Observation 7976c6c5-d3fa-427d-883a-80ef94acb2d8 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.078012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.078012Z digest=sha256:96215f474dc94199616626558f00364a3d7df95c546c38138d59b9cf0abf7667

Observation 5e3a3b12-e1d4-4821-9fa0-158de14183ab · outbound

This paper cites Visual knowl- edge in the big model era: Retrospect and prospect.Fron- tiers of Information Technology & Electronic Engineering, 26(1):1–19, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Visual knowl- edge in the big model era: Retrospect and prospect.Fron- tiers of Information Technology & Electronic Engineering, 26(1):1–19, 2025

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.123948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.123948Z digest=sha256:8258fff138b3c40a7d4d5656b17edde9eea59d3109cf8842351bad94137f8220

Observation 446e61eb-466a-431c-b28f-f2d23c2e59e7 · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding, 2024.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Internvideo2: Scaling foundation models for multimodal video understanding, 2024

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.176103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.176103Z digest=sha256:2a9c73d6db6006fd23c32b4d0d1e5958aa4c5a59ec98cff6b6f5eb2d893493aa

Observation cf47c66d-3022-43d2-8fab-2778e12eba31 · outbound

This paper cites Social-iq 2.0 challenge: Benchmarking multimodal social understanding.https : / / github.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Social-iq 2.0 challenge: Benchmarking multimodal social understanding.https : / / github

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.201464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.201464Z digest=sha256:fc745a821142460769fa54b43fd3c8e25d9faea51cc91ba806adfea35fe917a5

Observation 4b7087a9-7e98-47b4-92b3-56b13ebd9674 · outbound

This paper cites Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception.Cognition, 13(1):103–128, 1983.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception.Cognition, 13(1):103–128, 1983

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.239523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.239523Z digest=sha256:dac950297a945ce5e1a26b9f449e1e9ea80d235d4ff11032028278c9f48ccd09

Observation 905837be-2b9a-4bfc-bfc4-713a1d332ae9 · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.271934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.271934Z digest=sha256:41d36244289b753bdd22640c3e7388e50d524b5ac346f5f10c52bc3df0b557de

Observation 7cceb949-8946-4862-ab0c-74e261ec11a1 · outbound

This paper cites Visionary-r1: Mitigating shortcuts in vi- sual reasoning with reinforcement learning.arXiv preprint arXiv:2505.14677, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Visionary-r1: Mitigating shortcuts in vi- sual reasoning with reinforcement learning.arXiv preprint arXiv:2505.14677, 2025

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.319389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.319389Z digest=sha256:b0571942748cd99d5288d1ecfb0225566663c5971bc295324bd08f3794ebbe9c

Observation 42a0454b-2425-4809-989e-d2fa46412140 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Next-qa: Next phase of question-answering to explaining temporal actions

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.349331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.349331Z digest=sha256:d2bbc7b86ce9ccae194e14a236574facb451aa064aafcc62f6d93712ef875314

Observation 7330fcfe-4bed-43db-9130-ee96a05e3ba3 · outbound

This paper cites Advancing multi- modal reasoning capabilities of multimodal large language models via visual perception reward.arXiv preprint arXiv:2506.07218, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Advancing multi- modal reasoning capabilities of multimodal large language models via visual perception reward.arXiv preprint arXiv:2506.07218, 2025

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.392482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.392482Z digest=sha256:df9d8b956b6440b72fcea86775126c619e61d572c242f0134652aec0c1330f4b

Observation 36820249-8bf6-4b53-b020-81b06c3cd21c · outbound

This paper cites Expvid: A benchmark for experi- ment video understanding & reasoning.arXiv preprint arXiv:2510.11606, 2025.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Expvid: A benchmark for experi- ment video understanding & reasoning.arXiv preprint arXiv:2510.11606, 2025

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.464600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.464600Z digest=sha256:c57e71e7879718756c37af58fa2dc1fd6effe57d6fc19857bca9fdf2dda57ea3

Observation c280d29b-db4e-4c29-98d3-f92a41a34247 · outbound

This paper cites Qwen3 Technical Report.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Qwen3 Technical Report

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.540368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.540368Z digest=sha256:71518325386c2cba17b7730b87c60226026a7388cabbb6f91c9f85fa35006377

Observation a88715d9-15f4-484d-b8e2-edd3ccdb777f · outbound

This paper cites Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:07.705686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:07.705686Z digest=sha256:3d8ca72a6005d1ef9aeb0be32b2fae3de1575b2f6900ff22f1f9371e233d06d7

Pith citing papers

Observation 243b57be-37fc-49e0-b526-59420390fa47 · inbound

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction cites this paper.

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:16:00.259742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T02:46:50.373450Z digest=sha256:4dabfb0e239a553495cb3a0fec24d63175ebcb9c43310e4e2d9be349f5f6692d

Observation 9e3662f1-eca9-4aa7-bc4c-f4783f099257 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

Reference 289

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:16:00.259742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:c229252d2708bf6bb2f444682aa2afffb411ffd5308a32a4ac1f5a727c570bf8

Observation d83de7e4-9f52-4e18-acfc-1e5e5fef7c06 · inbound

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning cites this paper.

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:16:00.259742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-03T16:37:06.384435Z digest=sha256:fefcb2f2586a3c19ca7ec05a800757f4a4f1eebb3c22c93fb107c8e8d018ef82