Pith. sign in

Paper Citation Record · LEDGER

InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2503.21307.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.21307 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:18:06.939468Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T14:09:53.831006Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e08e65ee-056b-403a-9f2a-8b6b957b00c7 · inbound

R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO cites this paper.

R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:33.396784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:01:33.396784Z digest=sha256:358d4277b0a7060d10efa3ded09bbd49dc4339d1a7d5ce31a6d7663dc7716203

Observation a11f7e75-7d53-4f32-8fc8-0fb0710c46c6 · inbound

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts cites this paper.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:16.681525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:16.681525Z digest=sha256:1c97f627e2704c1ab45a5ba515317d6f72053abe81df9aaf6efcf724c9b4575a

Observation 48ae9154-6455-4740-9750-a32c2ad19818 · inbound

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models cites this paper.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.064975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.064975Z digest=sha256:6ea48f19855252a7d5f35af93676e3383f467d4bc66b023299c686f8af4c983c

Observation 65c1534e-1b89-400c-b96a-1bbc0500d2d5 · inbound

R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model? cites this paper.

R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model? InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T05:06:45.171417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:06:45.171417Z digest=sha256:f57decac75e6bb0b07d005f91eb8f414870385c6ab2943fbe2c81201ff7dc6fb

Observation b8713d11-63f4-4866-8f3a-31267e02be27 · inbound

An Efficient Token Compression Framework for Visual Object Tracking cites this paper.

An Efficient Token Compression Framework for Visual Object Tracking InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:35.749787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T01:27:18.535624Z digest=sha256:9d6c524c492d27b2a67f1b70e8737eaa9151f83f9b14e03434cb09534b30e578

Observation 9caa1331-a672-4e77-995d-2ea219c237d4 · inbound

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs? cites this paper.

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs? InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:11:16.340053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T02:06:45.858231Z digest=sha256:d2cd848958c03f705fd5eed2483e7b0328a60c345dda5a90d49ac7c0abce49fd

Observation 0e34dc88-bf54-4c21-bf7f-74d5a6de833e · inbound

CIVIC: End-to-End Sequence Compactness for Efficient Vision-Language Models cites this paper.

CIVIC: End-to-End Sequence Compactness for Efficient Vision-Language Models InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:13:26.442979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T12:12:51.867760Z digest=sha256:88887bc53f8dc8723296a03ca425372a314e81baa4869ea76510b506969936f9

Observation 8bd093a2-dbd9-4379-a643-7dbba2c8a816 · inbound

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference cites this paper.

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T14:09:53.832588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-26T04:24:38.917137Z digest=sha256:42fe7e5bba72b79f40489e03bea466f338df0d5aeb8ecfa4259e4e8417aad5ab

Observation 9ba06075-77b6-45c8-97a1-15ca8e42cf37 · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.576340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.576340Z digest=sha256:602e3d3a326fb5105e2dfb9eb9af92c2150e1a81d94886e28a0ea2e3642ae5a4

Observation 28c532bd-5bee-4187-a525-2142fc9a8462 · inbound

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding cites this paper.

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T04:28:23.606938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:28:23.606938Z digest=sha256:244895ca58146136c3437666667f0134733ff4d4aa41d34a561f3bb1101ea2b2

Observation 6921e4da-165f-4615-a4bd-f248cc995554 · inbound

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding cites this paper.

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:40.925614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:30:40.925614Z digest=sha256:16fbf877a67e9dcd13041d5f4da17344f43ed9915a8b34a37e416aa83ad4f426

Observation a8788006-6be9-4c2b-a03c-b18ed1b5c7ef · inbound

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding cites this paper.

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:06.939468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:18:06.939468Z digest=sha256:2291bdd173fafa82a606d143836de32a0c35c5866fbcc9e94685070a274a8f07

Observation 3e278458-c443-4484-aac3-85b23c675074 · inbound

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin cites this paper.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.475211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.475211Z digest=sha256:a1bfd362b8fa495916474c793c5f3e3b881743a628648d34b40bfb95d8cb61ea