Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-08T19:49:20.020874Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2607.05978.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-08T19:49:20.020874Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 49ca3f7d-9d52-4e68-9e85-9533f3f7ada9 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 65a341be-a14d-4d58-a01b-06044049e8f1 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention End-to- end object detection with transformers
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3768151d-15c9-49b2-a0f1-695000a92cfe · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation aefe5c54-efa5-487d-b89a-7d89ae8a04d2 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention BEATs: Audio pre-training with acoustic tok- enizers
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8223a1a4-b59f-46e4-b110-a41da26b81d0 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ef046333-4b51-45c5-a0d8-e50586390d8a · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Lookback lens: De- tecting and mitigating contextual hallucinations in large lan- guage models using only attention maps
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e3d90b48-35f9-4d10-9cd0-de7588d00934 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Fu, Stefano Ermon, Atri Rudra, and Christopher R´e
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation df80284b-8a5d-4368-a0b1-c228898df0d1 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention coco-gemini: Zero-shot COCO detec- tion with Gemini.https://github.com/simedw/ coco-gemini, 2025
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 67814fe2-40d1-45b6-8552-1652a7985482 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Multi-modal hallucination control by visual information grounding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation eb0b4114-781a-463c-a197-5e1d513d4e63 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention TALL: Temporal activity localization via language query
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 52dcb6a6-ebf9-4184-aaed-7871279815dc · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Gemma 3 Technical Report
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 92ea6d40-80d9-4b7e-9eb1-b033106f1a52 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Gemmeke, Daniel P
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3f087581-5dfa-440e-a7ca-ef0f46246b8d · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c0cc97b4-24f1-4ba0-ad02-2566d353699b · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention DAMRO: Dive into the attention mechanism of LVLM to re- duce object hallucination
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation be4929bb-018b-4874-93be-351e5d8346ac · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Making the V in VQA matter: El- evating the role of image understanding in visual question answering
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a4eff65d-50b6-47a3-b525-57bcfdd5b2ad · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6b426b24-5376-44d6-bb0d-fdd9f36d3f9a · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention OPERA: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 54ddb34c-2b88-45bb-9f4d-3ee8ef27d513 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Interpreting and editing vision-language rep- resentations to mitigate hallucinations
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 324b8e7c-c3b1-4a52-ad62-526928ef1beb · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Devils in middle layers of large vision- language models: Interpreting, detecting and mitigating ob- ject hallucinations via attention lens
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 229bcd6c-8178-4afd-8a8c-ce2525cfcc1c · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Shamma, Michael S
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 05d4d1e9-b467-4b78-96f3-624df10f30b7 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Berg, and Mohit Bansal
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3c7a3251-c031-4d6a-9865-6bf63d160993 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 62a7be5d-a2b1-4828-9052-a9cf2542b54a · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Grounded language-image pre-training
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8870ead8-958a-473e-af59-edec0dd466d6 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Evaluating object hallucination in large vision-language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bac235d7-7159-4f89-868b-7e955c3dcbca · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Lawrence Zitnick
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8dd7a3de-7b9d-4f7f-aba3-ae3fc92ce1ae · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 535ec38c-64a7-4aeb-b415-ca87ea3989d2 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Paying more atten- tion to image: A training-free method for alleviating halluci- nation in LVLMs
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c5c98b0b-11c7-46c1-a9ae-c10bdc58afcf · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention MMBench: Is your multi-modal model an all-around player? InECCV, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 142b4287-43b9-4fcf-9dd3-77db617e1d8e · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Simple open-vocabulary object detection with vi- sion transformers
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1b3ccc82-5771-4cb9-adb4-dc273502c279 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Query-Dependent Video Represen- tation for Moment Retrieval and Highlight Detection
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cf384d7b-f0d5-46c9-930c-b400c45887cc · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Verjans, Phi Le Nguyen, and Vu Minh Hieu Phan
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 281d80a6-1c69-4e76-8a24-9ef93bbb5bf4 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention GLSim: Detecting ob- ject hallucinations in LVLMs via global-local similarity
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 156f1f7c-eabb-4f3c-9bdb-8b84acedf8f2 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Kosmos-2: Grounding multimodal large language models to the world
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ba38589b-a41b-4b48-ba66-9c86bb97a9ba · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Beyond logit lens: Contextual embeddings for robust hallucination detection & grounding in VLMs
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6e499d5b-46c1-46ad-aad6-93a29f9e9197 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Effective pre- training of audio transformers for sound event detection
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4c1a8bf1-36d2-4fdc-beac-e8c39e32b0f5 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c01575d3-9dd5-44a0-8ce1-cc07aa9206cf · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Berg, and Tamara L
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a484ad0f-67ce-4a29-a1c8-e12d11fc78e6 · outbound
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention bbox_2d": [x1,y1,x2,y2],
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
No inbound Pith citation observations are available.