Pith. sign in

Paper Citation Record · LEDGER

BabyVision: Visual Reasoning Beyond Language

As of 22 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 31 inbound Pith citation observations for arXiv:2601.06521.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.06521 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T11:25:37.168330Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:15:39.766849Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T09:37:00.824668Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2de732fb-c481-4680-b05a-50591f5ae38a · outbound

This paper cites Accessed: 2025-01-09.

BabyVision: Visual Reasoning Beyond Language Accessed: 2025-01-09

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.254842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.254842Z digest=sha256:c388731321bd14ae8b2695820fb59a734e67b0afb0f02ff4750b5d74f50a71df

Observation bfb5f875-3479-4275-8e03-d259b1939353 · outbound

This paper cites Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, and Feng Zhao.

BabyVision: Visual Reasoning Beyond Language Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, and Feng Zhao

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.324358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.324358Z digest=sha256:01792969d417ff34f8e439f128a957392bd7aa14bbf5c069d2042fb7b0ec871f

Observation 801ef178-dbe7-43fd-9ae7-febbc2b9b640 · outbound

This paper cites Accessed: 2025-01-09.

BabyVision: Visual Reasoning Beyond Language Accessed: 2025-01-09

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.583244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.583244Z digest=sha256:bd81764ed5cd2be394331de150c4c5ec65d5160d907a179dde2abfd1cfc49822

Observation 4af5246c-d98b-474a-b118-4a42e8a91ecb · outbound

This paper cites MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation.

BabyVision: Visual Reasoning Beyond Language MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.655028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.655028Z digest=sha256:40fdf8d7b1b60c3e324f2b2dd660ee6c34988275d2633b0e687ecc1c533544ed

Observation d0e1e774-3d1e-47c6-846e-c86c5eabbe7f · outbound

This paper cites Accessed: 2025-01-09.

BabyVision: Visual Reasoning Beyond Language Accessed: 2025-01-09

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.738689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.738689Z digest=sha256:83d72f8905a8cfe132f84807cd51bfcde22dc937e9bf9032ce3a4c01aae29958

Observation d644983c-2a3e-4d98-8be0-1c4e66419cdc · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

BabyVision: Visual Reasoning Beyond Language HybridFlow: A Flexible and Efficient RLHF Framework

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.780815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.780815Z digest=sha256:0f12dc67c767e118c38b90cb34487a3f2e8136f324dbeaa41e5dd76efbc3fa86

Observation bdaee791-7db0-4c97-bf96-ee1f340191ed · outbound

This paper cites Core Team, Zihao Yue, Zhenru Lin, Yifan Song, Weikun Wang, Shuhuai Ren, Shuhao Gu, Shicheng Li, et al.

BabyVision: Visual Reasoning Beyond Language Core Team, Zihao Yue, Zhenru Lin, Yifan Song, Weikun Wang, Shuhuai Ren, Shuhao Gu, Shicheng Li, et al

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.862991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.862991Z digest=sha256:2a683429739cd70ad7e634a1f4b325482d738c394d0c88889febf80a250eea42

Observation fc4343a4-c7c2-4d2a-a26e-0760b32da638 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

BabyVision: Visual Reasoning Beyond Language Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:37.014871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:37.014871Z digest=sha256:cf810e8c32483a99570f05881975051f2e8ae6801b666bd59752499cdd896aa4

Observation f743877f-9b8b-4328-956c-c8e6acb7a2e5 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

BabyVision: Visual Reasoning Beyond Language InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:37.086374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:37.086374Z digest=sha256:1ff53b6d20d8846359f0c15484d55044bb0a523dbf6c441101ca33c152288b5e

Observation e54c4f02-70dd-455d-bed4-3f1f0daa5138 · outbound

This paper cites DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?.

BabyVision: Visual Reasoning Beyond Language DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:37.168330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:37.168330Z digest=sha256:4e8aab6990a91541fb07592add868f20751c7a8d7b651e4719fcc7b924b3a85e

Observation 241372c7-3494-44a1-a7be-4c5019d3e86e · outbound

This paper cites Humanity's Last Exam.

BabyVision: Visual Reasoning Beyond Language Humanity's Last Exam

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.947259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.947259Z digest=sha256:0889299dee36760a3bb31008ae701459cd93182e6db5bdaf0e785d9c908c7e66

Observation 4882c10a-7f71-48fa-8989-421a225fdd51 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

BabyVision: Visual Reasoning Beyond Language MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.510939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.510939Z digest=sha256:ae7c5676ebcc612c7f0fef6414be408aeaf0b084550cb81a9313f81ba159e536

Observation 768ad5cf-84db-4e70-a90b-7928f392fe03 · outbound

This paper cites Qwen3-VL Technical Report.

BabyVision: Visual Reasoning Beyond Language Qwen3-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.210179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.210179Z digest=sha256:78854ed1017d8d57665c532cefdc80244b24f20adbc5ba94dbb14b864c645421

Pith citing papers

Observation ecbae974-0770-40b4-b6da-45b8cb86e6d2 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence BabyVision: Visual Reasoning Beyond Language

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:b7e3c0f34625d3f8a410f9f3f61836e4686525a657f31301f9472c01b8a1ec12

Observation 3be3dc69-02b2-43c8-ba62-ad5c117f3a99 · inbound

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors cites this paper.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors BabyVision: Visual Reasoning Beyond Language

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:8f31eba55e95a0630c5252b95ba2f6b43037761bb20be438a1e2190536702388

Observation 9846bd0a-5d1f-42ee-ae57-920cd691042d · inbound

EXAONE 4.5 Technical Report cites this paper.

EXAONE 4.5 Technical Report BabyVision: Visual Reasoning Beyond Language

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T17:47:34.692414Z digest=sha256:0d505c7bc2d124ab3bc4b0175cc1256b2f264c6a03cc00c516e097c71182c10f

Observation 37bffe5c-46df-406f-b6b1-b7718587d055 · inbound

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm cites this paper.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm BabyVision: Visual Reasoning Beyond Language

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T23:56:02.856878Z digest=sha256:ecb3877f9b480bb2ad990e60f6b9e005d6ba10d7e5b3fda6ce433c5e10bb8185

Observation 0eb02fec-f321-4728-bf03-db7c1a25d8e8 · inbound

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm cites this paper.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm BabyVision: Visual Reasoning Beyond Language

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:8ffbda7a5bd7c70f32c7eeac6f66e4fd45b451a4d9c19ba37221a9051dc2009c

Observation 5cbe58c9-2535-4f0f-b992-fe8167643885 · inbound

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression cites this paper.

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression BabyVision: Visual Reasoning Beyond Language

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-09T14:36:29.666730Z digest=sha256:6cc30016c8b8fd380060826f2cd3826201d390e26362605f943cb54038adf0b5

Observation f8c77675-e600-4f98-9826-bd26e24b018d · inbound

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression cites this paper.

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression BabyVision: Visual Reasoning Beyond Language

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T04:52:09.685243Z digest=sha256:d2e5a44a0d3d9fc5fb6532907cf4ad58a39a7a6e8e8c4c01692e5067e5d7215e

Observation 51fcbd8a-30ee-41a5-9c1b-2a117f04ae4d · inbound

Do multimodal models imagine electric sheep? cites this paper.

Do multimodal models imagine electric sheep? BabyVision: Visual Reasoning Beyond Language

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T03:24:01.933339Z digest=sha256:704a37ab47b2e48a36b3ff5bfa9da5f12499404329de507224fad64a77ad50ef

Observation e69986e0-e556-419c-b975-eceda920f3c1 · inbound

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space cites this paper.

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space BabyVision: Visual Reasoning Beyond Language

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T04:30:54.053958Z digest=sha256:bcb4d6f795a8ca97f18653a7c7dee91945da6639fa4816c5563f0883a5687192

Observation 17c6eedd-5c6b-498e-9742-17acee3eae00 · inbound

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space cites this paper.

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space BabyVision: Visual Reasoning Beyond Language

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T22:58:04.574536Z digest=sha256:154ca127e04e56980f40aacc0b76134499101ccf7d2a5600a6cf5ea1ef5258a2

Observation 2eecc1b5-7e55-4014-b21c-e4c0d30ef38b · inbound

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture cites this paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture BabyVision: Visual Reasoning Beyond Language

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:823269b8b4dec3759c6a8fed07e52053ac63e6064a8d736280502c8566be8673

Observation b2c596a6-d638-446c-a9c5-1e75c5c40fd1 · inbound

VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following cites this paper.

VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following BabyVision: Visual Reasoning Beyond Language

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T20:14:34.739062Z digest=sha256:26bc82bb0dcec91ba06bb1d31a10b1e2bc140017d2fb0f19eb9a5c6c087d1ad1

Observation 5851d4f6-769c-4b6f-92e9-0d896e34dc6a · inbound

Step-wise Rubric Rewards for LLM Reasoning cites this paper.

Step-wise Rubric Rewards for LLM Reasoning BabyVision: Visual Reasoning Beyond Language

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:0e5cebad523cf65cb896a87e10f8e939b58e069836f685a84e1961e98880392f

Observation 2127c17d-1a86-413a-9736-bdbbf64a8041 · inbound

EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data cites this paper.

EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data BabyVision: Visual Reasoning Beyond Language

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T12:12:45.251924Z digest=sha256:74d63400f2365e2fd3a04e82208c0a32d314c8a0bb000cf2f2d9b2ee6d1a45b8

Observation 04dbbbd5-7c41-42c4-8c80-707d31f8d328 · inbound

Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding cites this paper.

Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding BabyVision: Visual Reasoning Beyond Language

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:25:20.511933Z digest=sha256:e8ff4c2ec9443765348e9b0e5a39d978fd73bd791304882826e55d8bbac714fb

Observation 3ebebce8-606e-4713-9819-be895d60852c · inbound

DeepLatent: Think with Images via Parallel Latent Visual Reasoning cites this paper.

DeepLatent: Think with Images via Parallel Latent Visual Reasoning BabyVision: Visual Reasoning Beyond Language

Reference 108

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T18:44:39.545911Z digest=sha256:fe18f810575b3abdcfa9b192acacdf8858a3399341ed5320c6422b17dc79b902

Observation 187c584c-1f2d-406b-94cd-aa8f92e5307c · inbound

ATLAS: Agentic Test-time Learning-to-Allocate Scaling cites this paper.

ATLAS: Agentic Test-time Learning-to-Allocate Scaling BabyVision: Visual Reasoning Beyond Language

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T15:27:28.290178Z digest=sha256:f147f4d36bd034d288a98438f373c7325a3525a6b922c781ee50b823de348a3a

Observation abdb3657-0c8e-497f-a05f-882bf39451c9 · inbound

LEVANTE-bench: Multi-Scale Comparison of VLMs to Children Using Cognitive Tasks (or, "Is Your VLM Smarter Than a 5th Grader?") cites this paper.

LEVANTE-bench: Multi-Scale Comparison of VLMs to Children Using Cognitive Tasks (or, "Is Your VLM Smarter Than a 5th Grader?") BabyVision: Visual Reasoning Beyond Language

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T06:39:08.246174Z digest=sha256:6765cc68e753be3cdcd128ff0b14ed3c36c8f12c490c829fd22bb9db301aa8d4

Observation 381440d4-1561-4548-a217-0477d06c4de3 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients BabyVision: Visual Reasoning Beyond Language

Reference 131

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:1e3c570e6ec8f59fd22cae5a3bdca07851f37dbebb7a2b97bf4ecdd931f3c68e

Observation 27ad2c73-da42-4e3e-a21a-860c8bceef53 · inbound

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning cites this paper.

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning BabyVision: Visual Reasoning Beyond Language

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-25T21:23:10.051805Z digest=sha256:6a5b54da5c421ea98692704afeec436f915d83c147e80885b13d1dc4c8d0f9de

Observation dba8cbd6-253b-42d7-a467-25bf97e808a5 · inbound

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity cites this paper.

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity BabyVision: Visual Reasoning Beyond Language

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-02T18:57:46.841456Z digest=sha256:d9da20efb6e59fbca28251592cf96be217c66a50b2f3aa4b04d4c652067d89ab

Observation 71f17e32-ee9b-4509-b9e9-2e4ec44e263e · inbound

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models cites this paper.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models BabyVision: Visual Reasoning Beyond Language

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.825956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:47fc274690e05ba508ff3e4ecc384a5e018bac21c60e53cd5c31e8bd4eac9542

Observation 5f823fbd-a930-48f6-a4e3-720357328807 · inbound

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning cites this paper.

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning BabyVision: Visual Reasoning Beyond Language

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T05:37:19.632869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T05:37:19.632869Z digest=sha256:b745b0f6dffb4d0a71be560c6ca8699f57f5ab3a73abb27543b669de091379d3

Observation 17b1b511-f44e-4f07-b535-a58fa7242f42 · inbound

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning cites this paper.

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning BabyVision: Visual Reasoning Beyond Language

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T07:01:19.711542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:01:19.711542Z digest=sha256:670cd06ab06941f67639a01dfd15f694f337a7b6b7587179f849d81618267292

Observation c4384486-1fbb-4686-9187-89c420c32b05 · inbound

An Exam for Active Observers cites this paper.

An Exam for Active Observers BabyVision: Visual Reasoning Beyond Language

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T21:12:03.005752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:12:03.005752Z digest=sha256:997d381664217fc3a7d647d7365ba917eb15f27217c98ee19eebf4e7daeb7775

Observation 0986a07a-a4f6-43b1-a070-8b88c1d6301c · inbound

Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests cites this paper.

Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests BabyVision: Visual Reasoning Beyond Language

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:32:27.377351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:32:27.377351Z digest=sha256:88721130e139e06128f48c781f3261938d547dd4cc9f980dcc8dc1af52c312bd

Observation 747b163b-6bef-4176-97fe-c820f1d48fc4 · inbound

RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection cites this paper.

RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection BabyVision: Visual Reasoning Beyond Language

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-02T11:50:20.183933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:50:20.183933Z digest=sha256:f481e918a0a14f15573268f60d39a8b42dead98a19ecff70640de17c8740fded

Observation d51a5525-1455-4600-a670-736d17b93e09 · inbound

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models cites this paper.

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models BabyVision: Visual Reasoning Beyond Language

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T05:01:28.838169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T05:01:28.838169Z digest=sha256:899c43ff5789f868312974cb32a1281c8570c4cd9efd85df4eee3c68c17f09a6

Observation 52f1f3d1-c9a0-418c-9ea2-43d2dd02a6fd · inbound

Beacon: Knowing When and How to Perform Agentic Visual Reasoning cites this paper.

Beacon: Knowing When and How to Perform Agentic Visual Reasoning BabyVision: Visual Reasoning Beyond Language

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T02:45:28.580613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:45:28.580613Z digest=sha256:a78cd2c07f098a45253f5951ffd3168c06d171dd369fee0533c4933410427e62

Observation 16cf6ca3-e2f6-4e97-8a37-0cf641434036 · inbound

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making cites this paper.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making BabyVision: Visual Reasoning Beyond Language

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.121965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.121965Z digest=sha256:4f4106c6d5c77a337d0ff0cfb2f9e49be3f435529debd085b4595cd852c1c587

Observation 4069ec4b-08de-4012-854e-4aafb89148dd · inbound

Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams cites this paper.

Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams BabyVision: Visual Reasoning Beyond Language

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:39.766849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:39.766849Z digest=sha256:ed69979a1b199923c8dc29ebdb95918f684f6d4d098b3a3ce3b8fd465eafe9ec