Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T20:23:07.705686Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 127 outbound references and 3 inbound Pith citation observations for arXiv:2511.20272.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T20:23:07.705686Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-03T16:37:06.384435Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T16:38:39.619743Z
100 of 127 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1c2785ed-4cd7-488a-8e38-8df42133fc3d · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e4113ed-1a88-4f11-bed2-153b4d47d61b · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs On seeing stuff: The perception of materials by humans and machines
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40d649dd-1e16-41bd-ad37-01207a4f42b5 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Vqa: Visual question answering
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d19b8ecb-31a9-4742-8a73-f796e7c8d198 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Eye-contact, distance and affiliation.Sociometry, pages 289–304, 1965
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14fda1d8-9815-4d28-9671-ecacd20eb630 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Qwen2.5-VL Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dae183b-ab4f-4652-9b30-8dcee8737c6d · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Representing the existence and the lo- cation of hidden objects: Object permanence in 6-and 8- month-old infants.Cognition, 23(1):21–41, 1986
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c53a5db6-45f7-4cc0-8170-3321c28fdfa7 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs theory of mind
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 858ff559-25c1-4076-bb04-fd8184e2384b · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98860f4a-b3b8-4ce3-91d0-263f87216e4c · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d473e7b-2837-4ab2-b10e-58c29a3aeabb · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Routledge, 1995
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c586aa5-8c9d-4252-b923-e4201d3f038e · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Activitynet: A large-scale video benchmark for human activity understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4715a1fd-4758-4444-81f4-330e1b44d5c2 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Rextime: A benchmark suite for reasoning-across-time in videos.Advances in Neural In- formation Processing Systems, 37:28662–28673, 2024
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f61da4e4-5d2f-474c-8887-1cd84e193873 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Are we on the right way for evaluating large vision-language models?Advances in Neural Infor- mation Processing Systems, 37:27056–27087, 2024
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96d07928-9fd3-4a1b-a2cd-89e6c7c79f24 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Video-holmes: Can mllm think like holmes for complex video reasoning?, 2025
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed3b046b-4305-4e84-b01a-dee2201bcd7b · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbe8be57-de9b-4f62-8e83-9a4890fa0284 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Le, Sergey Levine, and Yi Ma
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2bc5c7d-9545-4df6-9dde-7ab60c9cf8a1 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Whatever next? predictive brains, situated agents, and the future of cognitive science.Behavioral and brain sciences, 36(3):181–204, 2013
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e7f228a-2550-4553-9916-365f8ead4e66 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5223b74d-761a-450a-9ae5-af8fed1c7b9a · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Instructblip: Towards general- purpose vision-language models with instruction tuning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca6d9e86-e844-4883-b7c6-7f82d5ec0538 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs MIT press, 1987
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35bb7093-aff4-45a3-ba58-47c1837078ba · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs The llama 3 herd of models.arXiv e-prints, pages arXiv–2407,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 049c5852-561a-4b6b-9c0a-9ad9ce89d05c · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs An argument for basic emotions.Cognition & emotion, 6(3-4):169–200, 1992
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15443549-ad36-4cf8-a9df-f10b57c5db37 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Unreal engine.https : / / www
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0141789d-b487-475d-867e-941fac8706f2 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df25d2c5-5169-468f-a084-b69b1fc46b41 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Material perception.Annual review of vision science, 3:365–388, 2017
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6163625b-6a14-42e8-8f7b-afdba3cc4faf · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs The free-energy principle: a unified brain the- ory?Nature reviews neuroscience, 11(2):127–138, 2010
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9948e519-a3c9-4ab7-a4cc-a60d18381849 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f5062b3-88b0-4b28-bca1-026698eb8739 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Blink: Multimodal large language models can see but not perceive
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a30fc0d5-5e85-45b8-9ccc-84d95dcfeaf0 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Atomistic Control in Molecular Beam Epitaxy Growth of Intrinsic Magnetic Topological Insulator MnBi2Te4
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3b05772-dd96-409b-b6e7-3cbca0d93443 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Houghton Mifflin, 1979
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37ef0134-653e-4edd-bef5-ff46717a1752 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs The” something something” video database for learning and evaluating visual common sense
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3e67a5c-82eb-49a4-b17f-4932b2dacfb9 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs The Llama 3 Herd of Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cba16a0-eb5e-4666-88cf-e2a04c2e0cf4 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Hallusionbench: an advanced diagnostic suite for entangled language halluci- nation and visual illusion in large vision-language models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6411cd8f-eded-43a2-b78b-3cae95e7e5d4 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc1af7d2-65d2-48c4-8ed2-51d6db7532a2 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Doubleday, 1966
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10724494-018f-430d-935e-7415a1293efd · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Glm-4.1 v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f56e9cb6-14ad-429a-82cd-1483b68685bf · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f8a72b2-eb6f-47ee-9d55-dd7ae9dce488 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs GPT-4o System Card
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 678d0199-f37c-47c2-98c6-88ad754fba18 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da0fb4f2-cd52-41d2-a5f9-f584920f3438 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9851049-ad81-41e2-94d3-c00edd3d4bf3 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Towards Social AI: A Survey on Understanding Social Interactions
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d511dc7e-7895-4907-9c2b-c3b73d49c0dd · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs What is More Likely to Happen Next? Video-and-Language Future Event Prediction
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f908d3f-9d90-465a-9c6f-a5905c14acff · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Detecting mo- ments and highlights in videos via natural language queries
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 224dc2ff-e5db-4249-bfec-2fe6ff741b00 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Anisotropic flow, flow fluctuation and flow decorrelation in relativistic heavy-ion collisions: the roles of sub-nucleon structure and shear viscosity
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd73e304-7c6a-4557-b36c-1e76ed112956 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs LLaVA-OneVision: Easy Visual Task Transfer
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf5d0930-c3d0-443f-a63d-92be2ecef23d · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs From represen- tation to reasoning: Towards both evidence and common- sense reasoning for video question-answering
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f72b758d-fcdf-49d2-b4ba-33bbc058f01c · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Mvbench: A comprehensive multi-modal video under- standing benchmark
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fc421a3-3b21-4ed9-bdc0-1e7dec310b92 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Pope: A simple method to hallucination eval- uation in visual question answering
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 278caf4c-6ebd-4563-beb2-ea12b21b56cc · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1a887c0-7353-4b9f-bf56-c24768bc63e2 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Core Knowledge Deficits in Multi-Modal Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ea4dfa3-11d9-4158-9b50-b17e8074ccdc · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7619fe7-b755-48d1-89b7-3d73dba6fe7b · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Explainable Multimodal Emotion Recognition
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4d6e639-687f-48bd-b459-b9d41f707997 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Improved Visual-Spatial Reasoning via R1-Zero-Like Training
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f150adbc-f5dc-46ac-81f6-0a1fa312d5fb · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Microsoft coco: Common objects in context
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9038ad2b-7b01-4e72-b0ee-82aab3cd54fa · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6d62e94-c635-4f87-a6b4-aa202eed3c8b · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Generative Physical AI in Vision: A Survey
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4725a73b-cd00-4335-bd38-b554a4a47c85 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Visual instruction tuning.Advances in neural infor- mation processing systems, 36:34892–34916, 2023
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9aebeb5-ba66-4fb6-b7fc-a4e731d52021 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Improved baselines with visual instruction tuning
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b6cbdf4-27dd-42ad-b025-e4b7e614b806 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c559219-0d9f-435a-a826-b1d462cbcf51 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Mmbench: Is your multi- modal model an all-around player? InEuropean conference on computer vision, pages 216–233
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75e2c42f-5191-4b9b-bbce-3ebab6794364 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs TempCompass: Do Video LLMs Really Understand Videos?
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eb26b62-1cfd-4c03-b18b-baa9752989bb · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Videoreasonbench: Can mllms perform vision-centric complex video reasoning?arXiv preprint arXiv:2505.23359, 2025
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b90b1a-9759-4ffa-bb92-81e248525fb2 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs When thinking drifts: Evidential grounding for robust video reasoning.arXiv preprint arXiv:2510.06077, 2025
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf187c95-adac-4098-b1bc-88bddfc558b9 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs MIT press, 2010
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a349a015-7c69-4313-89ee-8323b3926575 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Ba- sic books, 1988
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1d5a2e2-5c32-4bba-a7c1-b626cb290a91 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Clarendon Press, 1978
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e884371-8bcb-4f6f-b24f-2d478ec58b54 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs On visual knowledge.Frontiers of Information Technology & Electronic Engineering, 20(8):1021–1025,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8baf9970-6bcf-4ac3-a585-5cac0dc0bd69 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Basic Books, 1954
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 814d2dfb-5841-4200-b9ad-4dbed6af28e2 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Free Press,
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 383f3a06-06dc-41c2-888f-4915c0f5158e · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Does the chimpanzee have a theory of mind?Behavioral and brain sciences, 1 (4):515–526, 1978
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8085e3b1-b330-4261-bfe7-ff04d5b2faa8 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Learn- ing transferable visual models from natural language super- vision
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76b23818-ceb9-458e-a77e-a1b1a0fec841 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Robust speech recognition via large-scale weak supervision, 2022
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e03999a-7e75-490b-8664-0a14bed4c345 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Reka flash 3, 2025
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37092672-20ea-423e-a20d-b2de8430b06b · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0835e414-529d-4a0f-bb76-38dd27af434c · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Wellness Insti- tute, Inc., 2001
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2da25a9-8ede-48f8-89bd-94ad7b8478ef · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Lawrence Erlbaum Associates, 1977
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 655df557-4eac-465e-882e-bdf5827dc90f · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Core dimensions of human material perception.Proceedings of the National Academy of Sci- ences, 2025
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85483ddc-b72c-445c-a6be-a44e7720ef2f · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd66fe63-c1d4-4f88-802a-62595ad9b606 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Towards vqa models that can read
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4babefde-5f86-4720-bc87-6e121df70ef2 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd8084a8-b7ec-405c-a36d-98be1261629c · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Origins of knowledge.Psychologi- cal review, 99(4):605, 1992
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5496cb1-740a-4094-994e-73864e367016 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Mimo-vl technical report, 2025
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 595fa52c-9b86-407c-876f-a754121fa28b · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19aaa2ea-9a29-43ce-92bb-673ffdf86ce9 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Qwen2.5: A party of foundation models,
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 300bd0a9-c482-4590-9ca2-ee6940c64633 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Trl: Trans- former reinforcement learning.https : / / github
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a632ed77-af1f-40e5-b70b-9ade9057808b · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Make Your Training Flexible: Towards Deployment-Efficient Video Models
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e6f7a3f-13d5-4433-b15e-9a32d1ba58a9 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04382e29-c0ac-4578-8504-e3b22ab54b6d · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Videorft: Incentivizing video reasoning capabil- ity in mllms via reinforced fine-tuning.arXiv preprint arXiv:2505.12434, 2025
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7976c6c5-d3fa-427d-883a-80ef94acb2d8 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e3a3b12-e1d4-4821-9fa0-158de14183ab · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Visual knowl- edge in the big model era: Retrospect and prospect.Fron- tiers of Information Technology & Electronic Engineering, 26(1):1–19, 2025
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 446e61eb-466a-431c-b28f-f2d23c2e59e7 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Internvideo2: Scaling foundation models for multimodal video understanding, 2024
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf47c66d-3022-43d2-8fab-2778e12eba31 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Social-iq 2.0 challenge: Benchmarking multimodal social understanding.https : / / github
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b7087a9-7e98-47b4-92b3-56b13ebd9674 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception.Cognition, 13(1):103–128, 1983
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 905837be-2b9a-4bfc-bfc4-713a1d332ae9 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs STAR: A Benchmark for Situated Reasoning in Real-World Videos
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cceb949-8946-4862-ab0c-74e261ec11a1 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Visionary-r1: Mitigating shortcuts in vi- sual reasoning with reinforcement learning.arXiv preprint arXiv:2505.14677, 2025
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42a0454b-2425-4809-989e-d2fa46412140 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Next-qa: Next phase of question-answering to explaining temporal actions
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7330fcfe-4bed-43db-9130-ee96a05e3ba3 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Advancing multi- modal reasoning capabilities of multimodal large language models via visual perception reward.arXiv preprint arXiv:2506.07218, 2025
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36820249-8bf6-4b53-b020-81b06c3cd21c · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Expvid: A benchmark for experi- ment video understanding & reasoning.arXiv preprint arXiv:2510.11606, 2025
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c280d29b-db4e-4c29-98d3-f92a41a34247 · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Qwen3 Technical Report
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a88715d9-15f4-484d-b8e2-edd3ccdb777f · outbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 243b57be-37fc-49e0-b526-59420390fa47 · inbound
Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9e3662f1-eca9-4aa7-bc4c-f4783f099257 · inbound
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
Reference 289
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d83de7e4-9f52-4e18-acfc-1e5e5fef7c06 · inbound
Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.