Pith. sign in

Paper Citation Record · LEDGER

Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2410.06169.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.06169 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:11:47.631053Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T00:02:17.737999Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9fe4f246-6e6f-4a33-a99c-7c132b2b21eb · inbound

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts cites this paper.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.844982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.844982Z digest=sha256:ec168543ff648a9e69b894f6693e1ad772c70fb7901b14b8d33808286b567921

Observation 0e68b2d4-546a-4dd0-8c91-07175928aee8 · inbound

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning cites this paper.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:35.570387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:35.570387Z digest=sha256:995517c0d9218e9b6028b1dbbd28ed6bfb4ce36fbc920ce44f43bc6cab70f85b

Observation 15d8b9b0-f2f0-40a0-a0ee-27ca5b52ac8a · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:02:17.741875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:91e4d02e331a867267594f93c771e8911ecf62eb2650a04dda21d278da5c5a23

Observation 587c20a4-2cdb-4edc-a1be-098a234b51c7 · inbound

LOP: Learning Optimal Pruning for Efficient On-Demand MLLMs Scaling cites this paper.

LOP: Learning Optimal Pruning for Efficient On-Demand MLLMs Scaling Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:47.631053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:47.631053Z digest=sha256:10c6d23daabb1255ec7c06da845f738341b27fe814e993ecb2b577202eba3d05

Observation 342ec6a2-1905-450e-b7f2-2146ea7721b6 · inbound

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models cites this paper.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:06.117105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:06.117105Z digest=sha256:b26922f3e6043a8356d32797108ece1350925d4691ae13a6f85a26460b1a71f3

Observation b42a0d17-7f87-4a78-8598-5098a9b2a63e · inbound

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation cites this paper.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:38.229729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:38.229729Z digest=sha256:ecef6f73023445c1e567e66a7510379ec2367a07398d400e8773b5d449ecb6bd

Observation c88f045b-b0bf-4adc-8807-714f44e7571d · inbound

Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning cites this paper.

Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:48:11.181219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T19:48:03.278822Z digest=sha256:3ee638d8f9c4fd99ad80a37ca7bb67265d231a23736fbe2d4ed7232b4664d141

Observation 0455d4c3-b8ee-4ffa-aa19-99096315cec7 · inbound

Counting to Four is still a Chore for VLMs cites this paper.

Counting to Four is still a Chore for VLMs Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:03.307146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:39:06.327888Z digest=sha256:247ce25fcc98c4a296fdb75ecd946656fe722193424cfba6846de88a0c66b500

Observation d331525c-62db-428e-af46-17f1a622b9c0 · inbound

EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling cites this paper.

EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:01:49.126917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T07:00:36.870817Z digest=sha256:518c1585805ab97c6c95eb6fbe32c2682938f5bb2a01f62d16b298b865bab74c