Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:46:57.758923Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2411.18932.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:46:57.758923Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T00:57:34.846303Z
A source-named dated measurement, never combined with another source.
Source: cited_works
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fbd6dcb9-32c0-4513-8a6b-478096d4f08b · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges Gemini: A Family of Highly Capable Multimodal Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f32acab7-fbfc-4213-9c5e-28ad4f82cc59 · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a9f01a7-7a3c-41f2-8bb9-17cebb34d4e3 · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4fc68b4-c8c6-4700-bd7b-c29d0f693ac9 · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49fb90f3-54ad-4db7-8250-17a21b4aba36 · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d070049-fa1a-48e2-a934-e1df4690b0d6 · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d209b181-8732-499b-a9de-7b53329468d9 · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges GOAT-Bench: Safety Insights to Large Multimodal Models through Meme-Based Social Abuse
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b924bf87-87a2-4a6b-b74e-83f5a23273b0 · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges MMBench: Is Your Multi-modal Model an All-around Player?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 955733a6-807e-42ac-936f-e99038201a4a · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges In The Twelfth International Conference on Learning Representa- tions, ICLR 2024, Vienna, Austria, May 7-11,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 87e19035-6054-4176-bd9a-45336f31ebc5 · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc32126e-5e31-460d-bcf4-00178deafe7f · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges GPT-4 Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc0d99ad-6f8c-4971-bb71-a8cc4c6dbb0a · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b17e664-339d-4a19-8be3-f4970aac4f89 · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a95fe574-63cc-4bd6-90ec-2cbc93579048 · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eee2850f-023b-436c-8635-04cff51f94d5 · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e021e21-b4e8-43be-adcb-a2c27d2206d4 · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38a07c97-623a-48c9-8ba1-2d5a914b9922 · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e030f8ab-9b07-4d6d-b1c8-c95bb822ae2b · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges In Forty-first In- ternational Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 43356f96-68b8-4e93-b330-fcf878ca1778 · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39653d53-7604-40d4-ba12-59324df9df46 · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges In Proceedings of the 2017 CHI conference on human factors in computing systems, pages 3620–3631
Reference 2017
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dd6e4ed9-b048-44aa-b8c3-805d2f50c2a1 · outbound
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges Pixtral 12B
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be7d77e5-0c73-44fd-8f8b-f0ddc2e1903f · inbound
RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.