Pith. sign in

Paper Citation Record · LEDGER

Scaling Native Multimodal Pre-Training From Scratch

As of 7 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 1 inbound Pith citation observation for arXiv:2607.22043.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.22043 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:05:51.822157Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:55:28.165189Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T11:55:28.599019Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 856c3335-7f7d-4384-b7a3-a456a6083ad4 · outbound

This paper cites B Experiment Results This section details the comprehensive per-benchmark results supporting the analyses presented in Section 4, organized into three primary categories.

Scaling Native Multimodal Pre-Training From Scratch B Experiment Results This section details the comprehensive per-benchmark results supporting the analyses presented in Section 4, organized into three primary categories

Reference 1

Resolution
malformed identifier
no resolver link, observed 2026-08-01T06:05:51.822157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:51.822157Z digest=sha256:3c291b2c9a846efd1c374795154993153ebfd951c1a16d99dbd88d740a21b5e3

Observation ead061e3-8709-4d91-b02a-fdde4282e2bf · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

Scaling Native Multimodal Pre-Training From Scratch SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:50.162416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:50.162416Z digest=sha256:735f48d7b1d4dfdcdccd4f90a2868afb34ede1147b149889f2eef5a427bfa6db

Observation e85c9bb8-a03c-4297-808f-ef5efe7d45ef · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

Scaling Native Multimodal Pre-Training From Scratch TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:50.526418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:50.526418Z digest=sha256:32db40290d7b2a12e3c65b5dc5f1e960b92c7531cc9e00b85791eedb4ce7205e

Observation bbd9c77e-1b29-4b8a-90f5-239837ed3207 · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

Scaling Native Multimodal Pre-Training From Scratch SocialIQA: Commonsense Reasoning about Social Interactions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:50.965502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:50.965502Z digest=sha256:c7264370cdbe08bd765d3c1559bcf60bbdbe915c658419196923383a2ac89bc5

Observation 84f1ed92-845b-42d2-b20c-d7589845fe0d · outbound

This paper cites CountQA: How Well Do MLLMs Count in the Wild?.

Scaling Native Multimodal Pre-Training From Scratch CountQA: How Well Do MLLMs Count in the Wild?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:51.176081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:51.176081Z digest=sha256:41b8778a59203a68311a4c95b2ca2c91cd1c3b4bc9e23048d240f6c4545c81db

Observation 614566c2-0175-4d9a-aa16-8704f6bdfed6 · outbound

This paper cites Beyond Language Modeling: An Exploration of Multimodal Pretraining.

Scaling Native Multimodal Pre-Training From Scratch Beyond Language Modeling: An Exploration of Multimodal Pretraining

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:51.300505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:51.300505Z digest=sha256:12efcfc73726b003ac0ea9bd9a75e982997fba4660e6cea826077922e5dd5d56

Observation be91083e-1d09-4308-a721-39023b8eb3ee · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Scaling Native Multimodal Pre-Training From Scratch SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:51.388385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:51.388385Z digest=sha256:d06a37a80c81cb3e874f7f64c9bf176a241b8ed31c2e5893e1506935f58835d2

Observation 37482601-c0a9-4281-8423-757addebef24 · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

Scaling Native Multimodal Pre-Training From Scratch LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:51.483271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:51.483271Z digest=sha256:f02d731232d73b7df4dadeb07bc09aa066afb4667f3cf6999ca0af9654ea515a

Observation 0cbdd42e-40b8-47ee-b422-847ec4b20f7f · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Scaling Native Multimodal Pre-Training From Scratch HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:51.578252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:51.578252Z digest=sha256:d8b46fd0a1977d107ed9ce351408be7d14d9bf1a5e538cc0fda54c14ba17a17b

Observation 0299e798-3028-48b6-9547-ec7c6453b70c · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

Scaling Native Multimodal Pre-Training From Scratch AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:51.646144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:51.646144Z digest=sha256:9fc224ef5775cc5713d0c9c3460b755c3c5892d9ddbae534d77755110668ae21

Observation 4e62a012-efbd-40a0-969e-fc653e72f039 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Scaling Native Multimodal Pre-Training From Scratch Kimi K2.5: Visual Agentic Intelligence

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:50.744535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:50.744535Z digest=sha256:1ac3f30bafe9e89e17a297b60f0fd2c6065757f400d43f767b332220a6707039

Observation 6c65647e-5163-42cb-9a4d-fc7daf5f8f56 · outbound

This paper cites Scaling Laws for Neural Language Models.

Scaling Native Multimodal Pre-Training From Scratch Scaling Laws for Neural Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:50.633139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:50.633139Z digest=sha256:c56628855d0b89e5ca98ef0af4e20ca0188a0fb3c43e718a1c60aed4c1f2103f

Observation c06469c9-7efd-4bf6-bd39-c9338a0f66c5 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Scaling Native Multimodal Pre-Training From Scratch Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:51.077650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:51.077650Z digest=sha256:7754c72cc4643490cd2e561f2e1eb1ad375aa528e263d34dee5225c69e538811

Observation eca071a0-9487-4488-8b9e-a93547fca486 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Scaling Native Multimodal Pre-Training From Scratch Training Compute-Optimal Large Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:50.323293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:50.323293Z digest=sha256:130a2c4a1a4b0ae48d74fec36fc655bbc4cfe2eb56e81814b2d0f981e92469fb

Observation 72283a80-6ac1-414d-989d-5a71c8d3f60f · outbound

This paper cites Emu3.5: Native Multimodal Models are World Learners.

Scaling Native Multimodal Pre-Training From Scratch Emu3.5: Native Multimodal Models are World Learners

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:50.098846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:50.098846Z digest=sha256:a6377f7e90f62c69b0b437bb9ef32d430e5f0bfba8aaf9c93a2fe17e4462629e

Observation 0024aedb-602f-46ae-9c12-c734a14323da · outbound

This paper cites OmniS- patial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models.arXiv preprint arXiv:2506.03135,.

Scaling Native Multimodal Pre-Training From Scratch OmniS- patial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models.arXiv preprint arXiv:2506.03135,

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:50.428860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:50.428860Z digest=sha256:0e90eedc1bebdeaa80cb822e3f5876f44800b696df2628377db688b06d83c515

Observation 4d8c3356-65c3-4221-bd68-e3cb80d0dd06 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Scaling Native Multimodal Pre-Training From Scratch Measuring Massive Multitask Language Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:50.263534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:50.263534Z digest=sha256:365d396957bc450c1a8f25c2656d0eef39fd0efba3b699dbb5b81dd9c1ab611d

Observation 9cd0fced-9622-4c04-9a67-8736f94141b9 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Scaling Native Multimodal Pre-Training From Scratch Training Verifiers to Solve Math Word Problems

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:50.023614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:50.023614Z digest=sha256:b3cd6c17edd10dee8c650e11834bea33c02aee02934e0bc6366e5fa24d7e5640

Observation f40d8a51-4782-428a-9029-7ee9b258ee8e · outbound

This paper cites DeepSeek-V3 Technical Report.

Scaling Native Multimodal Pre-Training From Scratch DeepSeek-V3 Technical Report

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:50.854597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:50.854597Z digest=sha256:52556480075b09ee5b14d1fb2ef54ccc3c822378273232c73fb7a5372e0a2966

Observation d57737de-fe33-415e-896e-4766186d75d6 · outbound

This paper cites We detail the specific architecture configurations and training settings in Table.

Scaling Native Multimodal Pre-Training From Scratch We detail the specific architecture configurations and training settings in Table

Reference 4096

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:51.724178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:51.724178Z digest=sha256:075aabc50dfcbcaa37e2eadf8c257040bb118b2544eb3124386a211557af64e4

Pith citing papers

Observation 41f0d0e1-821b-4096-8db3-3eb8253c53e9 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Scaling Native Multimodal Pre-Training From Scratch

Reference 137

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:55:28.602138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:55:28.165189Z digest=sha256:52970d466d8a567235e0f85fc99b77e4efde4320b260928ce62242da61bdbe3e