Pith. sign in

Paper Citation Record · LEDGER

Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2410.16261.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.16261 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:27:32.357748Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:39:44.698268Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7fedb350-47bb-4f9a-9359-3913c8e6be31 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.094306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:cd7095b8c522e457ee72c4e3d65951366606a5719a3da847752f0be5ae53d3a6

Observation 74a09442-949e-43fa-b0e9-8dcd2b66185e · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.305987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:f5e85cabdffc96b4a821c26125f947a8da883ab80eacfb93afcd5050821e418c

Observation 25d78dc3-1e8c-4e0e-b101-df0023bd965b · inbound

MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs cites this paper.

MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:32.357748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:32.357748Z digest=sha256:3c5291d849def6c080afd94004c616cec01df583e738cbf7278cd80192d5f05e

Observation fd55f037-e5a8-499c-80cb-7ce337155cc7 · inbound

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study cites this paper.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:41.004736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:41.004736Z digest=sha256:d0c8e2aa5ab1b9f5b9ffb8b11b414d004703eebbfd13c6002afd8beee64a7393

Observation 05f04817-a770-4ae9-a563-be7b0d121e73 · inbound

Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025 cites this paper.

Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025 Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:13.702701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:13.702701Z digest=sha256:7a14b4bbe328bf9fcc64a1f0339c8b8e04fbfc055ba885e257c107c3d8ba5f98

Observation 7883666e-4ec6-4b91-a616-a090dfa88957 · inbound

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets cites this paper.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.913989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.913989Z digest=sha256:869977f56479fe606b04be2ab433bee1080da6f021fe022e2abb95ea1ca629f6

Observation b00b1574-7563-444b-bf4e-6c77be7efac6 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:00.817293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:00.817293Z digest=sha256:cf4d1d06515852a406ca04ee5b18ca8ae5193c02999cba5be53949c78950e199

Observation 4cd8fa66-5bfa-4f4f-a128-d1cd0d32135b · inbound

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems cites this paper.

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T15:43:16.211195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:43:16.211195Z digest=sha256:0a59e4adee1159797fd0de7188768c7775641e27878eaf528f685ce2ea742e08

Observation 2cb98090-3eb1-4a65-8b1f-1d91e71b65e4 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.097039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:ca8169142a483342141a3eded3d6e827a44e68b863f73a849dac996aceb11eed

Observation 00a4bd24-b0d4-4288-8287-cf1da5b5f214 · inbound

The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving Models cites this paper.

The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving Models Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:15:50.128171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T20:04:46.144856Z digest=sha256:a6d48a5fc29ef2fa1637bd652b7ac871b35e2976ed60060e325b5febd4665a42

Observation 2764177a-b161-4f0c-a723-9ac3cb1d0fd5 · inbound

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability cites this paper.

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:38:09.440345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T07:35:32.225708Z digest=sha256:4143ae077ce2d5a60940272ed22fe7ab5227e7f95e3f330b2c964b618ef423a0

Observation f8b9de0f-89ce-459f-a406-dd62ab82d8f8 · inbound

EventDrive: Event Cameras for Vision-Language Driving Intelligence cites this paper.

EventDrive: Event Cameras for Vision-Language Driving Intelligence Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:57.083781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T01:26:43.752335Z digest=sha256:c869d5d7c7352745da2ad46e7ef1420f2f82233490ed915ba40025a2579139e6

Observation 6ef7b290-08a6-4782-9174-8a312e6a2b37 · inbound

LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement cites this paper.

LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:39:44.700307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T08:43:17.902876Z digest=sha256:8725c47951ea6290424b37296977a2aceb35ca4fd769c72e7732d2e082620e59

Observation 2d5289ad-e642-4ab7-8f38-abb0500efc6f · inbound

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video cites this paper.

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T07:12:52.433947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:12:52.433947Z digest=sha256:068c6c40d1ef3d196c40f9e16252cd2840c0476ccdb4173ffb37ee62f9db4fee