Pith. sign in

Paper Citation Record · LEDGER

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering

As of 14 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2607.04163.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04163 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T21:15:25.010442Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T22:30:28.401425Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4b7259a3-4a16-4092-b3d4-2f59411556bd · outbound

This paper cites Qwen Technical Report.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:9f85ae1d132744448f85db1a00d2f46e10db5771c2580b3100a1cdded620e351

Observation 10e41f4b-2d86-4135-a636-d290c04d5718 · outbound

This paper cites Token Merging: Your ViT But Faster.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering Token Merging: Your ViT But Faster

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:24b7fd3bfd5d32d7d4bacc3fb3d6325d2bb5fc8c0c90a89cd04d11e9675c64a6

Observation 4eaef454-e64d-4848-b478-eb1ab57e35ef · outbound

This paper cites Hallucinatory Image Tokens: A Training-free EAZY Approach on Detecting and Mitigating Object Hallucinations in LVLMs.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering Hallucinatory Image Tokens: A Training-free EAZY Approach on Detecting and Mitigating Object Hallucinations in LVLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:75a3b5498bcc0213d7f6939b5412695ebf37ed3ebf2e13adc50a72be54a616f2

Observation d0b7326d-2383-42c9-9fb2-519686d2e957 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision- language models.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision- language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:5a53631c5d6e1be0163f8f37c2a7f34929331bf7e787aa4a6d72d44fdb3e5571

Observation d907c1e1-c2ba-4734-8e1c-cbe88d796015 · outbound

This paper cites DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:93a0f4bd709a50a39ed6bc9dfc383607c1791dbe1d3310fd3adf9c38682dedd2

Observation c0f44307-c052-4974-9a52-f26a5dcc9a25 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for lan- guage understanding.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering Bert: Pre-training of deep bidirectional transformers for lan- guage understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:ac9433806d729bd359d74ba3c4f4fa5850e20dbb4e2b9451b7d878a113dbf99a

Observation 02d8d75a-719c-4320-9cea-51f7ea77b814 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:6f942c80939870a78d0134fc87c7947f174393365b03ad6d0c7776fea6275500

Observation 19c7353c-3bf3-4ee6-a47c-bd1674acfd41 · outbound

This paper cites STAR: Stage-Wise Attention-Guided Token Reduction for Efficient Large Vision-Language Models Inference.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering STAR: Stage-Wise Attention-Guided Token Reduction for Efficient Large Vision-Language Models Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:867fbed2304369324ed081963fdd6f8fdc6f3bb6816beb8d373446789c0aa120

Observation a008a3cb-2cf6-46fd-9ab4-74ec8ed2dc2b · outbound

This paper cites FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:0c8a07bacf5d8fad56a758eb01256152b43bb4d5202174e333fe54ae99af77ea

Observation 44e06353-3e91-496d-b4e4-68ba6ab95e51 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:c66bae540c7c4cd8a81ffc0022695e6c199d44e8531d4abb0c010374a8d198c0

Observation 560e52da-ad60-4c77-811e-890006c97e45 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering Evaluating Object Hallucination in Large Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:1c9cacf7a0c4e596c069281c85a040fb6276459ff7eae101d5e77b8927b5f21b

Observation 0cad0e59-4ada-494f-a6f0-dbbbdfd61b08 · outbound

This paper cites an unresolved cited work.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:8ca050bd2589fc1776afd412517ce028f49d5d8ce4334d10d4933a41f4bed7a5

Observation 2fccedcd-8608-4de3-b7d7-cca237206623 · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:130367a4fe7cb8e0e90e14b5be77bbfc9bf96b80df9ee6f09de5aa89f415319c

Observation 0378b2c9-4a19-40b0-82d5-41e3ae4afd07 · outbound

This paper cites A Survey on Vision-Language-Action Models for Embodied AI.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering A Survey on Vision-Language-Action Models for Embodied AI

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:c5dcc556453fd0dde1275dbf51d8320fb760d22f5537ddddec0377c1490ad2e9

Observation 958f2fa3-56c7-466c-a839-649fd0ad9f5e · outbound

This paper cites A., and Kundu, S.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering A., and Kundu, S

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:d2328d6127a3578532ce3304096b8a62d7b9262d61225628cbd758c43a036224

Observation 7bcb6fd0-3701-4cc2-a328-71e433d80097 · outbound

This paper cites J., and Yan, Y.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering J., and Yan, Y

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:7373cf9b3017d5e6a049378022c7331884bb69bce54a7ba741bf4f285a603243

Observation 281d4bee-1cdd-495e-acc0-9192d1101549 · outbound

This paper cites Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:c33ddf8d587a7901fae766cb0177ec9b91eb29a6f9d17de2aa8cd12b09304178

Observation 61c767f6-271d-412d-a0af-c84cc599ffe5 · outbound

This paper cites AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:dc3dbdfa3f3729aa6e2512c65d355c36547b494b86bc21abc5092db90043f294

Observation 1a135e4f-65fb-4554-bc10-866a59161538 · outbound

This paper cites Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:f4743cf83916bd337fe1477cc93db7f3982df7617e805af9e4f46c7070f1082a

Observation 3dbee225-cf8d-46f4-b824-5c33a77ee4f7 · outbound

This paper cites Hallucination is Inevitable: An Innate Limitation of Large Language Models.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering Hallucination is Inevitable: An Innate Limitation of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:ab0e88096d88defe7e38931a4c669ad17d16e57ec82250433c93fab1d5d095f8

Observation d222690d-2e51-4425-b6df-b17441f817d5 · outbound

This paper cites LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:797a24d1e02aefabac7f8df5bf2d795d47713be5c50ef8c0bc14db2081eab938

Observation 1b35f242-db33-4035-8c48-f71bcdfba914 · outbound

This paper cites Not all errors are created equal: Ascot addresses late-stage fragility in efficient llm reasoning.arXiv Prepr.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering Not all errors are created equal: Ascot addresses late-stage fragility in efficient llm reasoning.arXiv Prepr

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:a3e5c399a3e60510221edf7a6f872988ee68678488d6299049cb037605729918

Observation c5c597f6-df0e-4cdf-af84-05820f8bfb73 · outbound

This paper cites Not all queries need deep thought: Coficot for adaptive coarse-to-fine stateful refinement.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering Not all queries need deep thought: Coficot for adaptive coarse-to-fine stateful refinement

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:127cd6661eebfbf22d299243f421f44c2992ce79420ff041d6a414b043bb9737

Observation 18157f53-17f8-435c-8620-248a523d780a · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:217482423f9f638b077baa9dded8227c2bbccf0673846501a5f756bf00c3c8b5

Observation f236b25a-a134-43b3-a47d-85bbc32c7c0a · outbound

This paper cites InfMLLM: A Unified Framework for Visual-Language Tasks.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering InfMLLM: A Unified Framework for Visual-Language Tasks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:61f858df89938939c7c12be79739bbea1ba1a887a1d3dff126106787129d68d0

Observation 2066eb10-ace4-4c3c-abf2-95c0f4d1aaf8 · outbound

This paper cites Analyzing and Mitigating Object Hallucination in Large Vision-Language Models.

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering Analyzing and Mitigating Object Hallucination in Large Vision-Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T21:15:25.010442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:15:25.010442Z digest=sha256:ad652d6b6abb39477a4f07169773b8e4fab9112707fcc19d2f92e8d9f5beb94d

Pith citing papers

Observation 4ac47558-83b9-4470-a5cc-20aa60767907 · inbound

SPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large Language Models cites this paper.

SPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large Language Models SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T12:50:33.254161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:50:33.254161Z digest=sha256:6148110e6be9d95e04933a57000476768a7a2e341bac5771f4f563d2ff0f404f

Observation 38483ab5-3502-41d5-8e6b-e19e2bfebca4 · inbound

Better Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed Reasoning cites this paper.

Better Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed Reasoning SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T22:30:28.401425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:30:28.401425Z digest=sha256:b0d8829b876990dedc9f609cca0189bb5b0a9bcfd18cdc55ff6a36c1e10bcbc9