Pith. sign in

Paper Citation Record · LEDGER

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models

As of 10 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 9 inbound Pith citation observations for arXiv:2502.01419.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01419 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T15:26:19.007163Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T21:09:22.538353Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:40:07.123416Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dc454b96-53da-4674-9971-29ee93a38a47 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.616973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.616973Z digest=sha256:26031c2c3f874777b047043fa8f2a6aad51d28d927f5e516a87696b6351735ab

Observation 929745e4-1eb3-46e7-926d-d7eee55e8359 · outbound

This paper cites Qwen Technical Report.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.622787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.622787Z digest=sha256:9a55774dc442d226f8bbc5166d6d9351a7767c3dfb89c832c566439633486577

Observation c6b79058-3286-4746-ba23-fa5799509d55 · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Hallucination of Multimodal Large Language Models: A Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.749075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.749075Z digest=sha256:3c22758cc9f1e617672495b66c0ce79652050f4f819ef1c8cf5c26fa804c624d

Observation 5655d819-507b-4aff-b2ca-1bde8bc3465c · outbound

This paper cites Understanding Information Storage and Transfer in Multi-modal Large Language Models.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Understanding Information Storage and Transfer in Multi-modal Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.754601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.754601Z digest=sha256:e1980a9323412be4fb0ea99c8ec3372c043bff559c0770f1352f6d8254c4aee1

Observation dac0f0e0-37fd-4cc7-85ff-8b6013bf3e8d · outbound

This paper cites Y., and Furlotte, N.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Y., and Furlotte, N

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:26:20.279117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:26:18.759715Z digest=sha256:e3212744330c8713a04562f26124ed20df395c2bac14885bcf8527d57f92991b

Observation 768ff524-1948-4ea3-8a64-8d127762fdbc · outbound

This paper cites CLAIR: Evaluating Image Captions with Large Language Models.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models CLAIR: Evaluating Image Captions with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.764360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.764360Z digest=sha256:73f06dae3d8bb882917f72492421e9fbc8f9a1c729c7a2877c140a9fbe0fae3c

Observation 36792edd-90c0-4ebc-a640-d0fb42ae9d61 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.769342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.769342Z digest=sha256:4a667eb620eaa4c0b3e847a19aa17a94db272fcc8ff55736be38ee33995412af

Observation d5ba4fad-8213-4420-bc26-8d22e767bea7 · outbound

This paper cites Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.773637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.773637Z digest=sha256:0bda7d8be8ae761887396d961608c549c8ca8640a544fec3de167144dc8f4d8d

Observation fc52342b-b5e1-4817-99b3-ffa0fde8affe · outbound

This paper cites Vision transformers need registers.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Vision transformers need registers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.777967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.777967Z digest=sha256:bfcaf8f698c7767fa21e401f1d41e82edb291b4a80204b6f16572c55c2c6ff61

Observation 94b37f1b-95ef-49bf-811f-ac77e0ac460a · outbound

This paper cites The Llama 3 Herd of Models.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.782254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.782254Z digest=sha256:9a2a222032e359d31f2173118d740e582f34f904517484b3cf1c9bacd080688c

Observation 6d611715-887e-40cf-9c2c-94bfbcb03611 · outbound

This paper cites Multi-modal hallucination control by visual information grounding.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Multi-modal hallucination control by visual information grounding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.787091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.787091Z digest=sha256:de236d2b2730a01027d998042554d413be02edb83b1b97218c431b6030b8f2ba

Observation a828e6d3-4d90-4496-be4d-97b563001f75 · outbound

This paper cites ImageInWords: Unlocking Hyper-Detailed Image Descriptions.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models ImageInWords: Unlocking Hyper-Detailed Image Descriptions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.791790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.791790Z digest=sha256:460d0792d767ba835c593c986e61d77a495c76a63987134cb478a116d066a113

Observation 09708e40-51b5-41fc-ace0-ea80c40a84f1 · outbound

This paper cites Damro: Dive into the attention mechanism of lvlm to reduce object hallucination.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Damro: Dive into the attention mechanism of lvlm to reduce object hallucination

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.796443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.796443Z digest=sha256:a764156efdb793eb751c8367c38bef3163d105b816c33fe562870038135c3dfb

Observation 28e1978c-33b5-4dbd-9fda-f3d2f367da4f · outbound

This paper cites Detecting and preventing hallucinations in large vision language models.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Detecting and preventing hallucinations in large vision language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:26:20.222254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:26:18.800701Z digest=sha256:01fc16b83d0b1a59b0b59689f72696b34949fd2e6dd7e6a6352d9d2a574444a7

Observation c2292dcf-936d-46f4-8e90-083cb94d6c31 · outbound

This paper cites ChartLlama: A Multimodal LLM for Chart Understanding and Generation.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models ChartLlama: A Multimodal LLM for Chart Understanding and Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.805012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.805012Z digest=sha256:2be45a83bf7bde3e1fd386f4265ced021480de7e774f1c34ce3598dc6b9b93c3

Observation fa9046c4-4bb0-40b7-b71e-7e79621c16c6 · outbound

This paper cites A multi-modal foundation model to assist people with blindness and low vision in environmental interaction.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models A multi-modal foundation model to assist people with blindness and low vision in environmental interaction

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:26:20.206745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:26:18.809785Z digest=sha256:0e9bb36e966295d9c7fd1b5cc02aa8710811f77814f8ea4589973c65cf24a3d3

Observation b4a7800c-c682-4ff8-806f-a28293abd7c3 · outbound

This paper cites Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.814189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.814189Z digest=sha256:5d3cecc68b0c43f5c704f065b3323f78526baa19751a71c656125676a9986b8f

Observation e24ae048-8936-45a1-9ce8-d59201f9c993 · outbound

This paper cites Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.818734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.818734Z digest=sha256:2731f585134c3a20227323c46c0f4edfbde127b9614f2d13f6d58728765477d5

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.823683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.823683Z digest=sha256:f87290a07d4b80722c981bc5c6cb4ee9491870ce12525f77fa6643cde62b2821

Observation 0d31e5fd-2734-4a17-b51e-1d22a8f52548 · outbound

This paper cites Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.828246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.828246Z digest=sha256:af24172ed118acb9533b601e606b8e6fc18a377626fa48966c8b310c1a1ba880

Observation 163df10a-8836-4206-ae1b-7df13092ee82 · outbound

This paper cites See What You Are Told: Visual Attention Sink in Large Multimodal Models.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.833482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.833482Z digest=sha256:340bdeb5e5a13f14e6de24213fc0264297ea475df41410550defc459013f2a3d

Observation 1da64ded-5870-471f-85d3-d2c441a3b1fe · outbound

This paper cites Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.838137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.838137Z digest=sha256:d60142a83de2da9bcd1e4c1450e9f688cf4586632f8f6f96787ae3cf2ede40e0

Observation 4984a3af-918b-41e9-a7cf-d6276e863dae · outbound

This paper cites Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.843059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.843059Z digest=sha256:6b749dde23b557e6e4cbf25f11c3cb09487390f31d8e6df700a87da36dc6cc42

Observation 703cbab1-369b-4076-beb2-c3100d8d8972 · outbound

This paper cites Mitigating object hallucinations in large vision-language models through visual contrastive decoding.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Mitigating object hallucinations in large vision-language models through visual contrastive decoding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.847626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.847626Z digest=sha256:75af863f12f0318ff90e2d000f20d94756a37e44125c08fd3abbb808edf652aa

Observation 6db63981-1b8b-49ee-8a9a-716399238624 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.852478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.852478Z digest=sha256:a45e40f80f4700f10f4dbadd61c435f2133103cc20b97a625e9d98165c99b9ce

Observation 830131d7-cf91-473e-a02c-34758493ff7b · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.858826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.858826Z digest=sha256:99370a43ea3a45518413ab3fd5d92740a43d68b406145cce2c3a5f5bb3235249

Observation 7ac09cf9-c142-4a9a-a90f-443ce8d0678d · outbound

This paper cites Cross-Modal Attention Calibration for LVLM Hallucination Mitigation.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Cross-Modal Attention Calibration for LVLM Hallucination Mitigation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.863554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.863554Z digest=sha256:aa29ffd32d0af131b02a871500bbb61f62017c79474462170100f33a344ae0c3

Observation e7444099-b997-455c-bbf0-93ca8acba103 · outbound

This paper cites Inference-time intervention: Eliciting truthful answers from a language model.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Inference-time intervention: Eliciting truthful answers from a language model

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:26:20.146464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:26:18.868930Z digest=sha256:2f0adbf39ca110ca6037601e0ce926621cb33d6af62cd5f8fa13c3a2fb8ca156

Observation 31722197-d145-45df-8b98-1cbe9753e336 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Evaluating Object Hallucination in Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.873607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.873607Z digest=sha256:0e251a4363bd00c0d2eeb50fee792a886c4b97cbbd3d28fec099b36fa1974228

Observation 39c2733b-45a4-4007-a569-74f0c70ab65d · outbound

This paper cites Vila: On pre-training for visual language models.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Vila: On pre-training for visual language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.878445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.878445Z digest=sha256:ed9922ad4cd85e12076572b3c2abf4e9cabd6590e91be47a205b4f8c6a0d7dbb

Observation b918d1fc-8e51-486a-9e60-da4b57d37268 · outbound

This paper cites an unresolved cited work.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.882657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.882657Z digest=sha256:658e5f6c3f8162178b0a74f2feb2eb268fc07327e670ffb4ecf62fd97d33515b

Observation f6d22130-cf54-47ba-8e94-2817bdad97c7 · outbound

This paper cites an unresolved cited work.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.886450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.886450Z digest=sha256:134c5f18f8a43f2cd4e012172098d1aad3452d76d2af6b83dc5c2c7ba0bf88b5

Observation 759121a0-e4b1-4609-8e3f-c84a130b11d5 · outbound

This paper cites an unresolved cited work.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.890273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.890273Z digest=sha256:b82b74eeed7cca82c8d7d116a8a1fb8644901e2988d6c3b15d9a224005ee5b5b

Observation a76229f8-5a87-49f7-8822-70c208e58f1b · outbound

This paper cites an unresolved cited work.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.894434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.894434Z digest=sha256:b95c24377be0384169ca2219268edc87aaa7bff815aa35b37ea2e5307ddcb577

Observation 8212c0e5-0e29-41b6-bd50-b7fec29fcac1 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models A Survey on Hallucination in Large Vision-Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.898320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.898320Z digest=sha256:e9893724787fb75fee9e41d3874d3008fb3b9c48a67c81b834ca19f5de784258

Observation 8b20465f-2f18-4cee-9756-0fc0f4a19943 · outbound

This paper cites Paying more attention to image: A training-free method for alleviating hallucination in lvlms.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Paying more attention to image: A training-free method for alleviating hallucination in lvlms

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:26:20.059606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:26:18.902420Z digest=sha256:4eea1c3dbc20a3738adf41d2454196c9af184634200d5c807afe3e0fc4cf4e2c

Observation 3a066575-8af0-4f62-a543-226beb93ca90 · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models NVILA: Efficient Frontier Visual Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.906231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.906231Z digest=sha256:29ca2b27e152bf8c19f50df54a175aa72dba8707a5bd9bca8bb55b56c44c1384

Observation 00fee5ee-15ef-43a5-a228-e0ff08d8719b · outbound

This paper cites Compositional chain of thought prompting for large multimodal models.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Compositional chain of thought prompting for large multimodal models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:26:20.039403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:26:18.910203Z digest=sha256:f805fdfdf61c734e3ad2426c863fc5a0aee8f426f05f55d282f8a35b31910bdd

Observation e474e85a-abd5-412c-8a32-68f7b29ab44c · outbound

This paper cites Docci: Descriptions of connected and contrasting images.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Docci: Descriptions of connected and contrasting images

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:26:20.023284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:26:18.914484Z digest=sha256:d92b9119b4262d6e5192a417f72092fff11debbaf9aad0318a7c6870263fdbd4

Observation b9cedbf7-9d93-4f6f-931a-54537ea03e37 · outbound

This paper cites A., Shalaby, M.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models A., Shalaby, M

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:26:19.998650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:26:18.918841Z digest=sha256:773fadf8c483312a5c252bb5b3b0d82a9d804f2145c009ce40e6d550d7961e23

Observation f15fb0b3-9fd9-49a0-b3d1-33912980296a · outbound

This paper cites Instruction Tuning with GPT-4.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Instruction Tuning with GPT-4

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.923315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.923315Z digest=sha256:00778499b7a9112a080999716c6e6d45a06e67bbf8055ccaf74047560beb4cb4

Observation d12fc426-50c4-4c63-bda3-5e8f53725a1c · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models SAM 2: Segment Anything in Images and Videos

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.928214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.928214Z digest=sha256:4910b72eba5832afaf46d05082944a73016d2f8fbfb60cc7e0cde5e6ee51bba5

Observation a68c285a-3519-4933-8a84-461668832f04 · outbound

This paper cites Grounded sam: Assembling open-world models for diverse visual tasks, 2024.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Grounded sam: Assembling open-world models for diverse visual tasks, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.932776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.932776Z digest=sha256:0c4d5e5c840e9709ce8a1046baf820d9abc2863175a6cdf6d401f3b1b919ce7c

Observation 562cea6c-bcdf-4aca-8b3d-66c29f6dce0f · outbound

This paper cites Object Hallucination in Image Captioning.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Object Hallucination in Image Captioning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.937139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.937139Z digest=sha256:8e7c7598671202ac74a98ec6a5b385a46739c4482b1446fb6375f98934aa00f9

Observation 46c47705-ddd7-4769-9df1-31e11b0dfdc4 · outbound

This paper cites Wasserstein distance guided representation learning for domain adaptation.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Wasserstein distance guided representation learning for domain adaptation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.941824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.941824Z digest=sha256:1878304f0333dc23636b5b2173c7c867c1f75d5407840ee777c10e012af0c8c5

Observation c426fc89-9ac4-4a64-920b-6b22da6e79ee · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.946105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.946105Z digest=sha256:1a0a61dd8f36fa3cbf7f656039b9d4f68663fe90d41fc4b4dd74b616348bec6c

Observation e6bd7613-9db6-43af-a8b4-4ef9388fed1b · outbound

This paper cites Calculation of the wasserstein distance between probability distributions on the line.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Calculation of the wasserstein distance between probability distributions on the line

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:26:19.952507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:26:18.950544Z digest=sha256:6abedc05f2968390f8d48106ca20cfa91ff8a4664c6b6f1238808d1ca52e36df

Observation 4572e567-32f1-4456-972c-1433bbfa4185 · outbound

This paper cites Attention is all you need.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Attention is all you need

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.955068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.955068Z digest=sha256:5fe17be0a13cf3a0baeecafda666fd775661b67b6efbd52f400ad2f719e3e3fc

Observation b7d85cb7-4d73-4e51-bdfd-5863dc4c2aa4 · outbound

This paper cites Efficient Large Language Models: A Survey.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Efficient Large Language Models: A Survey

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.959612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.959612Z digest=sha256:17cd5ce2bd9434bc676853e7babf2dfdb184881edf5f75b4c8e504d79b87cb0c

Observation f1133279-1091-4a56-b01d-5ffc383e9e32 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.964914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.964914Z digest=sha256:c1593f1c6bf05d3a6fe8b9576a7fbb49c6186c0e3f47822121016b2f2b6b45b0

Observation 03d01345-741e-4ae5-9a86-9a68d8519573 · outbound

This paper cites Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.969743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.969743Z digest=sha256:da0db674711912c8907bf3879f7461102a2e6abdf66f8f83aa28119e91d4b53a

Observation dfd7ffc4-6e19-4861-9f2e-dfd7aee0f512 · outbound

This paper cites Mitigating Object Hallucination via Concentric Causal Attention.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Mitigating Object Hallucination via Concentric Causal Attention

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.974611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.974611Z digest=sha256:7a9bb41e64b5ca5b20a324cac5bad4a286b999bdee4fd2dc3bc06aba8e845622

Observation a03e7f78-8efd-483f-b46f-3dc9e7866ee4 · outbound

This paper cites Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.979301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.979301Z digest=sha256:55073c01d379c9c12d0a17a9ed0c5f88e489b8bd3e353dbb254ea936dd7f7a13

Observation 253cf50f-d1b4-4129-a141-da335b064f5e · outbound

This paper cites MLLM s know where to look: Training-free perception of small visual details with multimodal LLM s.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models MLLM s know where to look: Training-free perception of small visual details with multimodal LLM s

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:26:19.926146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:26:18.984229Z digest=sha256:8a775de2cdb86413651e853f3db696e634ae477372612e3167a25a5aadce416f

Observation ec6a6d8e-3533-4188-af1d-8de9b73f68c0 · outbound

This paper cites Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.988905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.988905Z digest=sha256:0849cb4fec6801eba299842a786cb2199426cc558233159653c2e773b8f3b1e4

Observation 3bbe3254-050b-4a0d-8abe-bd0229721ed8 · outbound

This paper cites Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.993526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.993526Z digest=sha256:cf6d2e78c647724d1489ca7aec686d95263cb102c9ef743f36679759d3d37856

Observation c38059a1-b5b1-42dd-a0a5-60afd10374e5 · outbound

This paper cites Analyzing and Mitigating Object Hallucination in Large Vision-Language Models.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Analyzing and Mitigating Object Hallucination in Large Vision-Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.998158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.998158Z digest=sha256:f38eb353a9c475524eb474b91ca731f5f1582be69e3c4cee8b7114b1c0a0e1b6

Observation 42e5d5f6-d04e-408f-8854-9b108b5ce77f · outbound

This paper cites IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:19.003177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:19.003177Z digest=sha256:88cf4f9d6a8f13be739eff30de4e6eb28321073a4f0519ff43727bf525405baf

Observation 39c70211-9980-4c13-adc6-20dfedbd2f98 · outbound

This paper cites write newline.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models write newline

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:19.007163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:19.007163Z digest=sha256:95ee641e70a5af5ff9258028bd4cf1f555ced2b189e62535a42f61ee402353cd

Pith citing papers

Observation 341e1f61-fa3e-416f-970b-95f992795572 · inbound

Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework cites this paper.

Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:32:36.355166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T12:31:25.257879Z digest=sha256:28354f9bd8c29e3c87892a51fcc621cae5a95291bd1c855b096da36b2919ef0f

Observation 25c8c43b-135f-4629-87e9-c85888966d5d · inbound

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards cites this paper.

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:22.538353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:22.538353Z digest=sha256:6d27335b6e91de99614095f0142c9610453748b1209a69edb6a4a8a7fda2fea4

Observation cfb9597e-fe8a-42ce-abaf-2b0832b018fb · inbound

Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models cites this paper.

Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:50.600673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:06:45.417233Z digest=sha256:52c3249c2379ab086506956536719effceda76322d9980f101632f779162d849

Observation dd292eed-8c43-4279-927b-c44726795d69 · inbound

TraversalBench: Challenging Paths to Follow for Vision Language Models cites this paper.

TraversalBench: Challenging Paths to Follow for Vision Language Models Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:36:02.729236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:25:38.215013Z digest=sha256:994421b16d41115f17d21fa68ecdb3613696581357ec0575c37391c4d861ef71

Observation 1e295fe0-14a8-46f4-81ee-d81be756772d · inbound

Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval cites this paper.

Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:13.633902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T16:56:52.714346Z digest=sha256:12286c0ac49c316d29d998a518d43169aa6d3158e2fb75ff0da55e8a552adc02

Observation 406b86ce-0591-4441-967b-0ab2d5884f90 · inbound

Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learning cites this paper.

Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learning Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.851487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:28:35.155099Z digest=sha256:337fbb5a20a180e020394388d3fadb883c58321276d67679bb253ec30079dcd0

Observation a2a9f9a6-797d-441c-a065-0f67309b9ebf · inbound

Steer Where It Matters: Token-Level Visual-Sensitivity Steering for LVLMs Hallucination Mitigation cites this paper.

Steer Where It Matters: Token-Level Visual-Sensitivity Steering for LVLMs Hallucination Mitigation Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:06:29.650511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:22:40.055153Z digest=sha256:672dbfcf5c08e4a712d6112493abd7d23823a313bfdb38aa59addd5bad51e545

Observation a95a305f-3762-4453-8a21-7134a1e655c8 · inbound

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation cites this paper.

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:07.125342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-25T21:05:36.836361Z digest=sha256:d8955500e7297e7b54900486a8863b2173e04f81908e25dcec573264916c0860

Observation 6272d759-cbc3-4095-bf0c-48fdcb8124ff · inbound

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation cites this paper.

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:39.546947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:30:27.178950Z digest=sha256:1d6196621e6ab5cc23e80eb20c4d7dd326dd1b1b5e098d45680df31f4e4b245f