Pith. sign in

Paper Citation Record · LEDGER

Scalable Vision Language Model Training via High Quality Data Curation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2501.05952.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05952 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:12:30.679583Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.276573Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 221fb2f8-0767-4ecc-b0ff-add9d8e1fe3e · inbound

Scaling Pre-training to One Hundred Billion Data for Vision Language Models cites this paper.

Scaling Pre-training to One Hundred Billion Data for Vision Language Models Scalable Vision Language Model Training via High Quality Data Curation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T12:12:30.679583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:12:30.679583Z digest=sha256:2ee7569b4cd7a3cd5a93359400ecf1f0a40f4c345a79e1a57723b23d021ea32f

Observation 99891245-cc99-4ab7-9a17-510f3a59453f · inbound

Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs cites this paper.

Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs Scalable Vision Language Model Training via High Quality Data Curation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:25:16.423936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T01:23:01.892612Z digest=sha256:a6e7c248814f806df649d412ee9634854ae0fe8cd8533a35751ab42fa0757446

Observation b2028d01-687a-41f2-a5b2-7f76926d2d2b · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Scalable Vision Language Model Training via High Quality Data Curation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.726967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:a83873a174fbe02ca02975d6a14f0fd449f8fcb7347583e2f1335cb4fee60b35

Observation 4b18d18c-0f14-454d-8fe6-7f82253a753c · inbound

Affordance Benchmark for MLLMs cites this paper.

Affordance Benchmark for MLLMs Scalable Vision Language Model Training via High Quality Data Curation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:52.599631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:59:52.599631Z digest=sha256:c930c562e712b7b02fa5d4a64c085f117016baf02ba8bfb715b80b384d943510

Observation 53938199-ee10-439f-95ff-bb1581701a47 · inbound

VGR: Visual Grounded Reasoning cites this paper.

VGR: Visual Grounded Reasoning Scalable Vision Language Model Training via High Quality Data Curation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:12:14.410055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T09:11:00.295700Z digest=sha256:4d3a74c3f5a6a7eedc27f23d7f7efacb2c8a92168057eda87d4c86d844d2e0e4

Observation e768521e-f731-4390-ba9b-14a09ac76714 · inbound

MMSearch-R1: Incentivizing LMMs to Search cites this paper.

MMSearch-R1: Incentivizing LMMs to Search Scalable Vision Language Model Training via High Quality Data Curation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:27:04.398232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T15:27:04.228144Z digest=sha256:ad44e6b25d9df71ec17a3effc7872bd9542c473e0eac03f4c2f515b94f42fe3d

Observation cda295d1-e7e2-4768-962b-fe867de61575 · inbound

Ovis-U1 Technical Report cites this paper.

Ovis-U1 Technical Report Scalable Vision Language Model Training via High Quality Data Curation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:19.113047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:19.113047Z digest=sha256:7549a57fda16c6688797a8e111d4eb84b670dd7a78734c55ce354cd672ca2728

Observation bd0b9c54-068a-4ccd-8543-fb75e41684bd · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scalable Vision Language Model Training via High Quality Data Curation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.977919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.977919Z digest=sha256:68570728cb4aacd062eb891ece4a8eccd1008e10990ddf2bd9af2abc424a1d7c

Observation e726fda4-961a-49f9-9f94-11b812202f0d · inbound

Describe Anything Model for Visual Question Answering on Text-rich Images cites this paper.

Describe Anything Model for Visual Question Answering on Text-rich Images Scalable Vision Language Model Training via High Quality Data Curation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:23.196657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:23.196657Z digest=sha256:d2f5882268e71785b054ffcb5264fe82456ce7d1e29fd22d89800f561db20408

Observation b047048d-af23-46fa-a388-0d7e3eb87d20 · inbound

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs cites this paper.

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs Scalable Vision Language Model Training via High Quality Data Curation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:54:20.305212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T19:51:04.983299Z digest=sha256:072a001cf8e4e3305c4b8dfca38919d158537b3f078081ac181149c227487bad

Observation ffa74552-aa60-46b4-9ef2-edd80efcf464 · inbound

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs cites this paper.

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs Scalable Vision Language Model Training via High Quality Data Curation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T21:43:59.233138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:43:59.233138Z digest=sha256:4805b2243fcc1a2a3aff7b193b1cd07e16a42eaad3afab2548abd2c74633ac83

Observation d7a1f2b6-c17e-4c71-8291-0ba6d6dbdf98 · inbound

S-GRPO: Unified Post-Training for Large Vision-Language Models cites this paper.

S-GRPO: Unified Post-Training for Large Vision-Language Models Scalable Vision Language Model Training via High Quality Data Curation

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:22:37.500851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:18:59.432486Z digest=sha256:6a260285cf27d15cbcd72846bc54ba6a71fa6732b24b694e214d49854e1c9580

Observation b7b0dfae-f74c-4dfc-90af-709908f83d66 · inbound

S-GRPO: Unified Post-Training for Large Vision-Language Models cites this paper.

S-GRPO: Unified Post-Training for Large Vision-Language Models Scalable Vision Language Model Training via High Quality Data Curation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T16:10:54.128765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:10:54.128765Z digest=sha256:e60e07b643c8018973bfd9d6be7e416a3eae02f0afb88a79ce5e435347c1e064

Observation 41f30715-eb9e-4cd9-a640-6df4c649d18f · inbound

Erase Persona, Forget Lore: Benchmarking Multimodal Copyright Unlearning in Large Vision Language Models cites this paper.

Erase Persona, Forget Lore: Benchmarking Multimodal Copyright Unlearning in Large Vision Language Models Scalable Vision Language Model Training via High Quality Data Curation

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:11:18.015960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-07T17:58:17.879345Z digest=sha256:a4a7ac34d61de5b02cf7aa2f511c4a343828944aa812a9afe1b25fa0d7bb03f9

Observation 44e3a5ec-46ba-42eb-af05-5908e7a706a5 · inbound

Shaping Schema via Language Representation as the Next Frontier for LLM Intelligence Expanding cites this paper.

Shaping Schema via Language Representation as the Next Frontier for LLM Intelligence Expanding Scalable Vision Language Model Training via High Quality Data Curation

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:27.331648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T05:01:38.118237Z digest=sha256:59ab2575d07d82f848848835701540fce110c9ca06be267560388a1f8014d52d

Observation daf12f4d-9e6c-4abf-8a53-bfab3084d48e · inbound

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval cites this paper.

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval Scalable Vision Language Model Training via High Quality Data Curation

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T05:49:36.502408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T15:34:55.062016Z digest=sha256:d8f393d27a055f1854dd766cd4129a36385a60ac8e4d2018754976a7e904382f

Observation 6cf7d13e-e151-4f27-9c49-c120b048f620 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Scalable Vision Language Model Training via High Quality Data Curation

Reference 231

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.278258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:eeef7fa8b8b008726e97619f8bb023764db85c75a956a70d156454d9f8f2af04