Pith. sign in

Paper Citation Record · LEDGER

Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2410.18558.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.18558 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.508137Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:59:44.997303Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5fae675e-72b8-49f3-8c13-71fc57915129 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.280731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:e8e1ab5462bf09cae4b6d046828564fd0ef790706f5e79216251e3eb810debe0

Observation 882b0012-713c-4514-a4f7-c7e1319e0ba2 · inbound

V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding cites this paper.

V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T16:58:03.063676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:58:03.063676Z digest=sha256:d17cf1c6082a5424383c30e86bac674ab70ce417c45aea261005a7aa1dfe3c36

Observation 5d007b92-3e16-4868-a8ea-e67c3e9e1f6e · inbound

Jasper and Stella: distillation of SOTA embedding models cites this paper.

Jasper and Stella: distillation of SOTA embedding models Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T01:03:27.326386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T01:03:27.326386Z digest=sha256:8a3f1f51fa4016758119904277a257b0f8e3441c1cf2bf16e51f4c22845a95f6

Observation 7417e222-25cc-4ed3-a0a5-430725dcbf4c · inbound

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining cites this paper.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.470574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.470574Z digest=sha256:b50435762e63273f3a3a873ff73bbe3f8c3a67f40b4bd425a647047dd71196e2

Observation 8ef304d7-de66-4712-b7b1-e647386d48be · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:39:22.490274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:7e2ca5f619304df990ce1b5eea9dd63c900aad12b6b6732f1f568e5aaa56c164

Observation c03a9b22-2881-4bda-b7ae-883bbcd2a9d4 · inbound

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design cites this paper.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:18.919339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:18.919339Z digest=sha256:fece5ba2c72701a8f3993f8a7d6ff0743b65e58dbb803dc7ed8f0abc0b501476

Observation 51840bca-4d2a-4d65-ae23-6428312a90f4 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.665218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.665218Z digest=sha256:960b4b496e4763b67374442eb1332e00d48d8a01b17bc45eaf32771fa743eaa5

Observation 62830ee8-01c9-4fcd-a88b-38ec23fbcc33 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.311987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:540da7e67311f475e26153e7fc0682445be7ff7fa67c992260a399d9b3fcefae

Observation 913fb9ae-d4bf-462a-a344-78099a022d86 · inbound

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark cites this paper.

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:40.508137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:40.508137Z digest=sha256:02b3f9a933540862051e7e135e90b697756da9102ff3c04c99f457f41c2e6e75

Observation a9a8c5b9-0124-4481-8765-c5224935fc06 · inbound

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs cites this paper.

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:19.937612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:19.937612Z digest=sha256:59343620336b3075f1b2b9040c491799a9c70d3e7d479a34b02720bbb3a2c191

Observation 12ed2e58-d17a-4825-87f8-63a45ae8839a · inbound

OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation cites this paper.

OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:32.510127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:32.510127Z digest=sha256:4010c54ed55d03929f68887983f721d4471abae97692ad1c77c2ca749f231ace

Observation 80bdcfef-4607-4095-b1a4-cd1470bc8cd7 · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:22.710649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:22.710649Z digest=sha256:2d49317be45337d48191353635782d914ca924a83fb8f945e0410ed4a34cc135

Observation 4cb5fe32-b9e9-42ac-9d5d-df4e41a384bb · inbound

Unified Multimodal Understanding via Byte-Pair Visual Encoding cites this paper.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:57.327402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:57.327402Z digest=sha256:5d3f6be7e9f07e2bf912ad2c1a4fec3219ad1391635e90101e9e8e2febb57932

Observation a1e2ecc9-2e50-4db3-95b5-b9ad56713b35 · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.556702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.556702Z digest=sha256:5b40e1c0f6f13a32c78bee86001b496a949357f40642bf66a4919faa13d0d916

Observation aa497178-df82-495f-98b7-c896c5ffffa8 · inbound

Kwai Keye-VL Technical Report cites this paper.

Kwai Keye-VL Technical Report Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:07.165354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:07.165354Z digest=sha256:944b9ad284a7e8e14855fbdbf727b4442c40b850387eee94e8948ecb3c487ce0

Observation 04379f03-058a-4e4f-a9fe-3456385065b4 · inbound

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning cites this paper.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:18.071945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:18.071945Z digest=sha256:ac253196a3927ed010de28dc1d268f13fa922f99ba847b9cb2fade1919e5ea51

Observation 10c461ad-573b-46a6-96e0-aeeac1fce22d · inbound

Egocentric Instruction-oriented Affordance Prediction via Large Multimodal Model cites this paper.

Egocentric Instruction-oriented Affordance Prediction via Large Multimodal Model Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T17:05:14.904841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:05:14.904841Z digest=sha256:d6071aee9858a7dcbae7caea604e8fd5506ba68e1537dc2412a4cc1aee27f973

Observation 24a1d6e9-b071-4b2f-8f28-822816425e6a · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.104407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:cf5bb7d33fb5a0bc2028861c52bbec8aef88a075d282d77cf8e771436df86867

Observation 933d34f6-4ffb-4fd8-ae93-ffb88ebe1745 · inbound

Emu3.5: Native Multimodal Models are World Learners cites this paper.

Emu3.5: Native Multimodal Models are World Learners Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.653172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:8315899245d4c90e07ad2683f9128ddf8cdb4c3ae6cab23b5e0968179fca27d7

Observation 26c0b9b7-d9d2-4ae6-ac04-737f0737d4fe · inbound

AutoNeural: Co-Designing Vision-Language Models for NPU Inference cites this paper.

AutoNeural: Co-Designing Vision-Language Models for NPU Inference Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T18:57:01.402053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:57:01.402053Z digest=sha256:767625a78e7b338ea21825eaf3501c86133e75e97383028c85845ca7e1387289

Observation ae888698-4138-4df4-bfa0-ca4c4b4bc9d8 · inbound

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs cites this paper.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:52.540683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:52.540683Z digest=sha256:74d7b2d9a2b01c686cc3e585e3818237124abdc1e6e926be7c319104bfc52471

Observation ede44cd1-57c2-4562-a8c1-97466c5d50cc · inbound

Kwai Keye-VL-2.0 Technical Report cites this paper.

Kwai Keye-VL-2.0 Technical Report Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:37.038777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T13:53:10.352603Z digest=sha256:7a719aff4df131e9c1c43d770a6b1050baca314cbb7b4c380891808d29dd0d20

Observation cc87095f-161e-4e89-99d4-b1481cb90e30 · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 261

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:59:44.998736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T09:19:50.623741Z digest=sha256:0c110492e7345031b52cd1e9575189607393d03c9c84c1fe6d30f30fd0d47091

Observation 75d95d23-8d3b-47aa-9446-f0d82ee2c2ee · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 260

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:55:59.576681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-29T01:18:19.195007Z digest=sha256:d141e1a7789e4f5053b293fab5fe504efd5d4735ce500b937856106a33f3a355

Observation 0f101feb-80fc-4c41-b4f0-0be78a664ba1 · inbound

Bridging Video Understanding and Generation in a Unified Framework cites this paper.

Bridging Video Understanding and Generation in a Unified Framework Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:05:40.337986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-01T05:57:54.653504Z digest=sha256:bd807f49aaaced4b73945572a720adec3badb165fa935d37ad93918db5decd40

Observation 57187ae8-b034-49f7-b94c-ce96b817a597 · inbound

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models cites this paper.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.485161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.485161Z digest=sha256:cf21d92efdc765b4245f935b39ac27075c16cdf70a8df7a9139c6be2b40e4b95

Observation 8f29ceba-d232-424b-8231-d13a2bacb445 · inbound

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs cites this paper.

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T02:46:49.649902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:46:49.649902Z digest=sha256:81825d59851b309384d67daff2da4458dca4e43410fe7d0d605287cb93b7d6ff

Observation 2f7fa112-b604-4d32-9496-53e11777f5d4 · inbound

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs cites this paper.

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T04:16:08.298412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:16:08.298412Z digest=sha256:2fa84d1a1b710bba3bb2dde3e0311c474fb4b2e2767d6a4acd7c01080e5aa90c