Pith. sign in

Paper Citation Record · LEDGER

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges

As of 17 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2411.18932.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18932 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:46:57.758923Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T00:57:34.846303Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fbd6dcb9-32c0-4513-8a6b-478096d4f08b · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges Gemini: A Family of Highly Capable Multimodal Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:57.685260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:57.685260Z digest=sha256:b63c1e7192fc63d3edd8caa118aa4bf0b873615328856bb828647bc81fd5a2d3

Observation f32acab7-fbfc-4213-9c5e-28ad4f82cc59 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:57.689807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:57.689807Z digest=sha256:7d4bb6181748a2b8561639ccc725cd6fc8728915c3c900d47e0a6b153987225a

Observation 3a9f01a7-7a3c-41f2-8bb9-17cebb34d4e3 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:57.697851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:57.697851Z digest=sha256:0bfed40c95e56ed7690c8de01d18b18ccb36f72c8b4151dcaffe943d6ff74c14

Observation a4fc68b4-c8c6-4700-bd7b-c29d0f693ac9 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:57.701819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:57.701819Z digest=sha256:1ca79fa30f890f893ed1e90caa7bb500275efa0c267733c7babfc16d2995d777

Observation 49fb90f3-54ad-4db7-8250-17a21b4aba36 · outbound

This paper cites Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:57.705977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:57.705977Z digest=sha256:d8b12a974ecae5c3264189c30b69da59075a0096aec7bcf5bf62f1db085dbd4e

Observation 1d070049-fa1a-48e2-a934-e1df4690b0d6 · outbound

This paper cites MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:57.710361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:57.710361Z digest=sha256:d0b7071307e3d52cb772bf1b8bb9407e3eb25f7b9198575897b02f8d8304447f

Observation d209b181-8732-499b-a9de-7b53329468d9 · outbound

This paper cites GOAT-Bench: Safety Insights to Large Multimodal Models through Meme-Based Social Abuse.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges GOAT-Bench: Safety Insights to Large Multimodal Models through Meme-Based Social Abuse

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:57.714538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:57.714538Z digest=sha256:8e085d154ca60c298a11ca31f196f4cb474ded2d41c113a75d7e7fff19caef24

Observation b924bf87-87a2-4a6b-b74e-83f5a23273b0 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges MMBench: Is Your Multi-modal Model an All-around Player?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:57.718508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:57.718508Z digest=sha256:a6f24c4ad6ed3fe90cf61a0a7fd5fc8a7a758cca7313764ea39ceb57c6211b4c

Observation 955733a6-807e-42ac-936f-e99038201a4a · outbound

This paper cites In The Twelfth International Conference on Learning Representa- tions, ICLR 2024, Vienna, Austria, May 7-11,.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges In The Twelfth International Conference on Learning Representa- tions, ICLR 2024, Vienna, Austria, May 7-11,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:57.986588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:57.722117Z digest=sha256:d8efcb6ca6d410d58fb4d85c90bb3cad5c94bd1d8d4f79d62462e061ede9fdb1

Observation 87e19035-6054-4176-bd9a-45336f31ebc5 · outbound

This paper cites VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:57.725830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:57.725830Z digest=sha256:9bc4f76b034a74c68ff4d8c2ec65a131fb7e5f70983843071e4d72b2e0a16941

Observation fc32126e-5e31-460d-bcf4-00178deafe7f · outbound

This paper cites GPT-4 Technical Report.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges GPT-4 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:57.729742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:57.729742Z digest=sha256:b4c553bc65299c503a500e1313607ca0eb2c5bd912d3a49ba14b42c4f80989aa

Observation cc0d99ad-6f8c-4971-bb71-a8cc4c6dbb0a · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:57.733101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:57.733101Z digest=sha256:b653763f04e695b248390ccb706fa42e8c42a5e8a5c92de1f35f4e184c18e229

Observation 5b17e664-339d-4a19-8be3-f4970aac4f89 · outbound

This paper cites ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:57.736829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:57.736829Z digest=sha256:3e3c03051ad00e16de6d053501c1df3bd2a4f70428a81ba2b81bfe6346a450f7

Observation a95fe574-63cc-4bd6-90ec-2cbc93579048 · outbound

This paper cites Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:57.740670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:57.740670Z digest=sha256:09970f148c6f7b6a2860909ecfe495fdbaf90dc8f7920a5415a8677ec772e60e

Observation eee2850f-023b-436c-8635-04cff51f94d5 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:57.744094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:57.744094Z digest=sha256:fefe49977d838d11dc4fc2da9ce18fbd86531621cfe982e2c7387339a0d2df9c

Observation 3e021e21-b4e8-43be-adcb-a2c27d2206d4 · outbound

This paper cites Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:57.747574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:57.747574Z digest=sha256:bb523e5f9cd045f1d0cf765c230c604c4f297488b3326791ee7157a56169832a

Observation 38a07c97-623a-48c9-8ba1-2d5a914b9922 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:57.751164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:57.751164Z digest=sha256:a77042a4f8f00a0bdf05780f8310201c3c2344c28c9f842a9db22660ed388f6e

Observation e030f8ab-9b07-4d6d-b1c8-c95bb822ae2b · outbound

This paper cites In Forty-first In- ternational Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges In Forty-first In- ternational Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:57.975322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:57.755063Z digest=sha256:d2361000e01a112b12bc459452ac73d92fefa18ef96f7e8b98abe3ab903fdbdb

Observation 43356f96-68b8-4e93-b330-fcf878ca1778 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:57.758923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:57.758923Z digest=sha256:34510573ac2ee3a573bcfcdc22c0c2ce7476142f2b03105d6f9129fdc0e1ed04

Observation 39653d53-7604-40d4-ba12-59324df9df46 · outbound

This paper cites In Proceedings of the 2017 CHI conference on human factors in computing systems, pages 3620–3631.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges In Proceedings of the 2017 CHI conference on human factors in computing systems, pages 3620–3631

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:46:58.000316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:46:57.693717Z digest=sha256:c987d625ac556e329e66b2c3af06c36455a67d446a7c40cf870370a1968beef9

Observation dd6e4ed9-b048-44aa-b8c3-805d2f50c2a1 · outbound

This paper cites Pixtral 12B.

ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges Pixtral 12B

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:57.681325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:57.681325Z digest=sha256:dfcdbe7d23e76c1f1ef9bca95db3fd9d901c0aeef877f536cccb936cab5285aa

Pith citing papers

Observation be7d77e5-0c73-44fd-8f8b-f0ddc2e1903f · inbound

RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models cites this paper.

RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T00:57:34.846303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:57:34.846303Z digest=sha256:15c753ea2efea1fc53daf8b1ee2aba1c2de0fb9ff55e11966969f387e03ca70b