Pith. sign in

Paper Citation Record · LEDGER

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning

As of 16 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 4 inbound Pith citation observations for arXiv:2509.06461.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06461 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:19:43.691285Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:23:29.666893Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T23:55:52.903831Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved20
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d2a5b67c-64f8-4def-9291-00163c139acd · outbound

This paper cites Star Attention: Efficient LLM Inference over Long Sequences.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning Star Attention: Efficient LLM Inference over Long Sequences

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.546665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.546665Z digest=sha256:b28e081d6e1e01cf84053cbf414f085f27693b1f1deb6e372f45265463ce343b

Observation 8cac38d7-b8da-48eb-a52e-f7f3d7bafba0 · outbound

This paper cites Dp-iqa: Utilizing diffusion prior for blind image quality assessment in the wild.arXiv preprint arXiv:2405.19996,.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning Dp-iqa: Utilizing diffusion prior for blind image quality assessment in the wild.arXiv preprint arXiv:2405.19996,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.576633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.576633Z digest=sha256:da7ccac87b950f969942e436c6951e40a723c41a9940d8809ec3f0037c1d1ef1

Observation 14dc7873-6181-4b8f-bdf6-adac22422ce5 · outbound

This paper cites Vistawise: Build- ing cost-effective agent with cross-modal knowledge graph for minecraft.arXiv preprint arXiv:2508.18722,.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning Vistawise: Build- ing cost-effective agent with cross-modal knowledge graph for minecraft.arXiv preprint arXiv:2508.18722,

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-15T16:19:44.188584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:19:43.583822Z digest=sha256:d856e6db5b37707e731dbb03d7ffee755d7a863e7703890dd6d86135e42d2d20

Observation 646f8f1b-9e21-45f4-b514-4813b9ae1d7d · outbound

This paper cites Mrfd: Multi-region fusion decoding with self-consistency for mitigating hallucinations in lvlms, 2025a.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning Mrfd: Multi-region fusion decoding with self-consistency for mitigating hallucinations in lvlms, 2025a

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.589372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.589372Z digest=sha256:dbf1bf9f945d4bc1b3f26a8f80e6326f157ff5dcbb4de511c7d1fe95bfdabc3e

Observation 187f8fbc-ac5b-4954-9e73-4c1c971c5b0c · outbound

This paper cites DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.604994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.604994Z digest=sha256:86e33fb4a36eb3828dae64328b96241f0a5adf7d88a1051d86105924b74da7f3

Observation 130a646e-6e67-4bc7-a161-6c610a1a8fea · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:19:44.445640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:19:43.611233Z digest=sha256:ec3e40e238d5e686aa71645f950f0810e2edbe839071598577331b997c951d6c

Observation 4a47b18b-0307-4338-85c3-34e184585a06 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning Evaluating Object Hallucination in Large Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.616653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.616653Z digest=sha256:09df444c2790ec7245e164abf7bad54031676ea2934bb79540f85f381623bafa

Observation aacfdf11-dec1-4b6b-9b27-c0d10752de47 · outbound

This paper cites Structured Attention Matters to Multimodal LLMs in Document Understanding.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning Structured Attention Matters to Multimodal LLMs in Document Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.621986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.621986Z digest=sha256:b8f54668a901546ea2392127f85d7a24928c1873aaf723b7cd16f80952732b78

Observation d5f598cc-4646-4262-ba5c-cb7eec314d56 · outbound

This paper cites Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Tasks.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.627310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.627310Z digest=sha256:a061170a485aa98d6ad069bd94c06f0908c2941269fbb80a1fc88b91d10e248d

Observation 7f98f819-6a45-47af-a677-3df58c7dc331 · outbound

This paper cites SLANG: New Concept Comprehension of Large Language Models.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning SLANG: New Concept Comprehension of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.632137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.632137Z digest=sha256:1e42817e0475493b497344961e1481d0add4351e78321cb5562ed1c3140b2152

Observation 124d6a33-14e5-470f-bffd-895ff40ab780 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.637935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.637935Z digest=sha256:d4f2abf7700b9d56b1085b23dc0085e7dbb35dc8791908353fbe7cb4bc6277a3

Observation 98b367f3-9f44-48ee-b5ea-cbcc583a7662 · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.649105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.649105Z digest=sha256:8d42a517ce2425298dc9b432d681ad9c8421bce5220921b39e65d52d982aa545

Observation 61082bcc-5e41-4261-b7f1-0ced4349f1d5 · outbound

This paper cites V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.654520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.654520Z digest=sha256:6b845146e79715a093081eb6e573c3dc97af1dd6511c54707e9c8fe5c863d094

Observation 29a3692e-d7aa-40c3-95ec-cbad89f9a31a · outbound

This paper cites Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.660111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.660111Z digest=sha256:6710bdf1ec84c5a66e1727ae80755ea7011dd8947b534d1870412f3fb2d04a71

Observation 15710224-9697-4d87-abdd-dde34cb5df7c · outbound

This paper cites Who is in the Spotlight: The Hidden Bias Undermining Multimodal Retrieval-Augmented Generation.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning Who is in the Spotlight: The Hidden Bias Undermining Multimodal Retrieval-Augmented Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.665830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.665830Z digest=sha256:5298e0f1d5048b19ffc4df70eb1a459b22fb343235b4a5b84e5c77b334a1956a

Observation 2b0f0b57-98ec-409d-908f-de89a1a19015 · outbound

This paper cites MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.672057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.672057Z digest=sha256:45a1ad4694f5042f57ea498a7522c561a1df90f73164c7c764efc16b57ebe117

Observation 35db9617-d28d-4013-84a5-d329bcca5bd0 · outbound

This paper cites TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.679677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.679677Z digest=sha256:71103b093b878e1ecd497f9e299b33954cf4113ea6a2d3a1843e5d674186ae9b

Observation 1131f406-f9cc-4d71-8de2-db2359a9c6ee · outbound

This paper cites Factual Dialogue Summarization via Learning from Large Language Models.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning Factual Dialogue Summarization via Learning from Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.685611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.685611Z digest=sha256:80141fd2f625f1dbebf2f1cc9e4b07244543b8d29b5b3ac94cb0f4682a058001

Observation 5407da0c-93f8-4459-8bde-453a7b9205a0 · outbound

This paper cites Write a general description of the image.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning Write a general description of the image

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:19:44.406142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:19:43.691285Z digest=sha256:7d28dda1f359360918a282168d92a73b9f8f67788de5b1b92e4f4689570f7b05

Observation a71acdbd-f22b-4bde-a763-d0a731939cb2 · outbound

This paper cites PIS: Linking Importance Sampling and Attention Mechanisms for Efficient Prompt Compression.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning PIS: Linking Importance Sampling and Attention Mechanisms for Efficient Prompt Compression

Reference 1986

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.565364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.565364Z digest=sha256:80afed6f61a0a3ec2de972b691be07a471317b1569f05539dec6dcf89556b363

Observation fd73f359-b963-4e7b-9e00-7cb609001b6d · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning A-okvqa: A benchmark for visual question answering using world knowledge

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:19:44.428222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:19:43.643196Z digest=sha256:e6e3ef24a4aa719a5c03f94d479f09a94186f1f1b6ecb8928e71804354e78508

Observation 81252269-68f7-455b-acf3-fd21917f19c9 · outbound

This paper cites Context-DPO: Aligning Language Models for Context-Faithfulness.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning Context-DPO: Aligning Language Models for Context-Faithfulness

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.553312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.553312Z digest=sha256:f9664865c9d80791767aae833e981e7bad56ca3dcfb9c7597c2b3cdfd798e976

Observation 8ef3bc0a-87fe-44a9-9505-dca87e350ed1 · outbound

This paper cites Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.595045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.595045Z digest=sha256:b62f681932a29f64bb066c2bf3ce801ec7c51d53a5961c6ece2aba3ebef70788

Observation a88f0c62-f252-43fa-b6c5-9ca316a74632 · outbound

This paper cites RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.559195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.559195Z digest=sha256:546c984ec19dbb339926e538f0d0580804f6d4b7ed75b88d5595aaf84a5bc87e

Observation 8b865cef-0eee-4883-9fe5-13a6aeb0a3a0 · outbound

This paper cites Related Knowledge Perturbation Matters: Rethinking Multiple Pieces of Knowledge Editing in Same-Subject.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning Related Knowledge Perturbation Matters: Rethinking Multiple Pieces of Knowledge Editing in Same-Subject

Reference 2025

Resolution
malformed identifier
no resolver link, observed 2026-08-15T16:19:43.570711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.570711Z digest=sha256:1a18136776379583fe0aa88bc5559bf379fdfb79489a3a75603a60dcc7e80f59

Pith citing papers

Observation 5fda0e9e-2c53-4a1f-9083-8218ce0a7761 · inbound

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG cites this paper.

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:50.082045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T19:15:02.124035Z digest=sha256:f7a95db8b47eb9eca7d677cdea3ff6c6bce08edc2e70c204f9da2a51436f7cea

Observation 883c2d02-7a73-435e-96b7-daa1a9bd304a · inbound

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models cites this paper.

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:55:52.906909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T18:46:26.869644Z digest=sha256:0375332033dd2fac519e61e392c14583054313677184d2125e64f0525bb5dd63

Observation 4049c009-9922-4234-bd67-1e1ff8eebe64 · inbound

SAGE: Selective Attention-Guided Extraction for Token-Efficient Document Indexing cites this paper.

SAGE: Selective Attention-Guided Extraction for Token-Efficient Document Indexing Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:32:52.107088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T08:32:02.222528Z digest=sha256:f041b25334433f4b2457c44031611ba1d76ba6b07d15995d5f97411d9b884cbd

Observation 4b4b0dab-6a3f-4a77-a3dd-70cbd4b3efb0 · inbound

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure cites this paper.

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T08:23:29.666893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:23:29.666893Z digest=sha256:c855f00bdd17007f9683078950896da602f48aa7edd540e0c8f0f103db58c1ea