Pith. sign in

Paper Citation Record · LEDGER

Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2410.06169.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.06169 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:11:47.631053Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T00:02:17.737999Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9fe4f246-6e6f-4a33-a99c-7c132b2b21eb · inbound

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts cites this paper.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.844982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.844982Z digest=sha256:d9c5138b9768a733add4b14399dc39fc9f1f80732727b69ae95c0a525ecdac05

Observation 0e68b2d4-546a-4dd0-8c91-07175928aee8 · inbound

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning cites this paper.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:35.570387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:35.570387Z digest=sha256:6bd7d18c9c3597dfd8394492236f625eec4b0259e576aab2f835726d3f09dc3c

Observation 15d8b9b0-f2f0-40a0-a0ee-27ca5b52ac8a · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:02:17.741875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:c45bc0289707a53a53565e92503d1a44ea33852b139308d97f292aed69f5f32f

Observation 587c20a4-2cdb-4edc-a1be-098a234b51c7 · inbound

LOP: Learning Optimal Pruning for Efficient On-Demand MLLMs Scaling cites this paper.

LOP: Learning Optimal Pruning for Efficient On-Demand MLLMs Scaling Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:47.631053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:47.631053Z digest=sha256:1da8bbe045cd7006639411ae268e5df7a35fd3dd4e23e77f6c01c3a924f94480

Observation 342ec6a2-1905-450e-b7f2-2146ea7721b6 · inbound

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models cites this paper.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:06.117105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:06.117105Z digest=sha256:ecd1e730ba60806d24b4ea49150b4d45c4b99f9e6f6c7796d13fbcf905ad39b3

Observation b42a0d17-7f87-4a78-8598-5098a9b2a63e · inbound

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation cites this paper.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:38.229729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:38.229729Z digest=sha256:4c5416a0678c5cc9d85b50b10796d56c35d7941162c0d84cef0297f53cef8a30

Observation c88f045b-b0bf-4adc-8807-714f44e7571d · inbound

Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning cites this paper.

Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:48:11.181219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T19:48:03.278822Z digest=sha256:d2d1c1a8e14f22b3c316997255a6f4c11f08847f500c2d3550307a8d2c1a7b6f

Observation 0455d4c3-b8ee-4ffa-aa19-99096315cec7 · inbound

Counting to Four is still a Chore for VLMs cites this paper.

Counting to Four is still a Chore for VLMs Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:03.307146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:39:06.327888Z digest=sha256:c39232fd482b04c848f84f333b776327da505500f8312c4d773d18290fe8c0b7

Observation d331525c-62db-428e-af46-17f1a622b9c0 · inbound

EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling cites this paper.

EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:01:49.126917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T07:00:36.870817Z digest=sha256:f4e1f1b4a686d92a8f1bf5396aca91c9639154a5af416a6fb395199595b2ca06