Pith. sign in

Paper Citation Record · LEDGER

The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2504.10462.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.10462 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:49:27.833950Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:33:28.015082Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a3172c1e-ea2e-428d-96a7-4e2d469dff55 · inbound

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training cites this paper.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.833950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.833950Z digest=sha256:a5b0949087f75e5f0a9e79beb0f50caed11f64d83fa1d44fb4dead3d2fc732da

Observation 97348b64-4114-4889-9a2a-b871dc941f75 · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.804522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:e000f2117be65e8958ab1b13cf049c2a3c4a8222018b7c2c9ebc49ff34e5c0d2

Observation cb0a40c4-3a26-4ea0-bf95-5f2fffdf4283 · inbound

VGR: Visual Grounded Reasoning cites this paper.

VGR: Visual Grounded Reasoning The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:12:14.474028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T09:11:00.295700Z digest=sha256:c1b9605f3656427f3b7108fb8b32462d955fff5525b617b98852e51eeb2c97e2

Observation 8e3d7bf0-8038-4b4e-9817-32ed004715e8 · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.867758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.867758Z digest=sha256:1d7bb5605de6f59c214e75b9714068232beee351195ac872609894a2723f75de

Observation 42056ed7-d647-444e-86bf-c2139b686188 · inbound

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models cites this paper.

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:19:28.552429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-14T20:13:18.813131Z digest=sha256:95ee2ecdb6a5af97dabad3c53feb3f8d9d291f8de4ed62afa1218ab37e2c6061

Observation 9299fd8b-3ac9-4be1-b289-3704a39184a0 · inbound

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning cites this paper.

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:18:05.174455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-20T06:16:47.650748Z digest=sha256:95bc86982bf259fc760e3886b1c3fcc290751367365bcc04e22c986dde3eb001

Observation d3dcd580-27df-4fee-9729-3b3de8ca34c9 · inbound

From Pixels to Words -- Towards Native One-Vision Models at Scale cites this paper.

From Pixels to Words -- Towards Native One-Vision Models at Scale The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:33:28.016525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T13:29:53.006064Z digest=sha256:08eb3f557dc33b7c412723ac0c70aa06b58bbb8f988cfcd454b2bc46df891070

Observation 458afbc7-108a-4f0e-9017-244ed6609077 · inbound

MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization cites this paper.

MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:32:21.339189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:32:21.339189Z digest=sha256:b504d4e9aa698d29c320b9fbae6ef073d5b5b34aced012de4e496999ea4c60ad