Pith. sign in

Paper Citation Record · LEDGER

Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2410.16261.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.16261 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:27:32.357748Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:39:44.698268Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7fedb350-47bb-4f9a-9359-3913c8e6be31 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.094306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:4842a85a62b0f2b8c78993b8003638d5d50e651eb9b101e319d37c430d2c35fc

Observation 74a09442-949e-43fa-b0e9-8dcd2b66185e · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.305987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:bac9713d41bd99e80fa6af92d2161ccb0dba86bd90e14cc81a30d2808d9b7625

Observation 25d78dc3-1e8c-4e0e-b101-df0023bd965b · inbound

MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs cites this paper.

MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:32.357748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:32.357748Z digest=sha256:54e1492346dea58d1fb497e5da6bd5d0cac899cb2d93beb1a5808d0c36992fbb

Observation fd55f037-e5a8-499c-80cb-7ce337155cc7 · inbound

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study cites this paper.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:41.004736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:41.004736Z digest=sha256:022270235b65582dbcf7ddb190a0eed8b06e3e18b19e1f743b7c3bee49026325

Observation 05f04817-a770-4ae9-a563-be7b0d121e73 · inbound

Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025 cites this paper.

Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025 Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:13.702701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:13.702701Z digest=sha256:7a14b4bbe328bf9fcc64a1f0339c8b8e04fbfc055ba885e257c107c3d8ba5f98

Observation 7883666e-4ec6-4b91-a616-a090dfa88957 · inbound

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets cites this paper.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.913989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.913989Z digest=sha256:e1d1343fccebd4bdb358ed79094c1c8db71d46df1ab874e675636f6ff227fbcf

Observation b00b1574-7563-444b-bf4e-6c77be7efac6 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:00.817293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:00.817293Z digest=sha256:f73d3139431572ae983cf3c6f006afb3f0700583f6ad8fd6ffbf3d464f6fed6e

Observation 4cd8fa66-5bfa-4f4f-a128-d1cd0d32135b · inbound

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems cites this paper.

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T15:43:16.211195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:43:16.211195Z digest=sha256:12aab9481026260441463aa08a6fdaa8758a0989c2fea319af57eda0b1590a8d

Observation 2cb98090-3eb1-4a65-8b1f-1d91e71b65e4 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.097039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:8b8fe5cf5f27b3b6b94ff96462f1bbf5b87fe76f8ceed468fddc4b3bd3b89157

Observation 00a4bd24-b0d4-4288-8287-cf1da5b5f214 · inbound

The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving Models cites this paper.

The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving Models Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:15:50.128171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T20:04:46.144856Z digest=sha256:e60ee8ac09af441baa56d5db6fe891188ed9798b2ef7c7509be46c1891965680

Observation 2764177a-b161-4f0c-a723-9ac3cb1d0fd5 · inbound

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability cites this paper.

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:38:09.440345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:35:32.225708Z digest=sha256:f1e82a6ddbf9d262d3b60f8eb12cd33ef675938aa1acc2426a72e9fc8095a6f1

Observation f8b9de0f-89ce-459f-a406-dd62ab82d8f8 · inbound

EventDrive: Event Cameras for Vision-Language Driving Intelligence cites this paper.

EventDrive: Event Cameras for Vision-Language Driving Intelligence Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:57.083781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T01:26:43.752335Z digest=sha256:28d55cb44ba369c3888a540372c4b961f5a54e1ab0b61d0f0c38960dad98a677

Observation 6ef7b290-08a6-4782-9174-8a312e6a2b37 · inbound

LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement cites this paper.

LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:39:44.700307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T08:43:17.902876Z digest=sha256:65a7cc5c4ebf41d37ada9a8cf3113b6c1bc331f4aa1a42a95a86b1c10c6b16cf

Observation 2d5289ad-e642-4ab7-8f38-abb0500efc6f · inbound

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video cites this paper.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:068c6c40d1ef3d196c40f9e16252cd2840c0476ccdb4173ffb37ee62f9db4fee