Pith. sign in

Paper Citation Record · LEDGER

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention

As of 22 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2607.05978.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.05978 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T19:49:20.020874Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact6
  • verified fuzzy30
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 49ca3f7d-9d52-4e68-9e85-9533f3f7ada9 · outbound

This paper cites Qwen2.5-VL Technical Report.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:55:33.910545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:35dc3cb9cc7f5ad433ccf5392e016a1fa70766c5eed6d04c642fff3df7169dc3

Observation 65a341be-a14d-4d58-a01b-06044049e8f1 · outbound

This paper cites End-to- end object detection with transformers.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention End-to- end object detection with transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.264918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:96148af88a415091ac46f94d4a2d26c7ec335e6f8cd135da4ed5ba0d8919e284

Observation 3768151d-15c9-49b2-a0f1-695000a92cfe · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:55:33.905378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:0ccf0764539583c48646845ac7be713a3f982f50fbe8bf7ff6f96d441a6047ed

Observation aefe5c54-efa5-487d-b89a-7d89ae8a04d2 · outbound

This paper cites BEATs: Audio pre-training with acoustic tok- enizers.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention BEATs: Audio pre-training with acoustic tok- enizers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.273355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:9d64f0bdf0bca1014f05d03caf9d7b2fc3b9716cb9f01bd25efa56b502a3a930

Observation 8223a1a4-b59f-46e4-b110-a41da26b81d0 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:55:33.907904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:1c3f692dc2f08073a3d51202cfc773697bceceff82fcfd4d1af0b897336bd417

Observation ef046333-4b51-45c5-a0d8-e50586390d8a · outbound

This paper cites Lookback lens: De- tecting and mitigating contextual hallucinations in large lan- guage models using only attention maps.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Lookback lens: De- tecting and mitigating contextual hallucinations in large lan- guage models using only attention maps

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.266727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:b7c23f1f3fbd6782252e58571a49203b6c017689b824137b1e2028767efbe0ab

Observation e3d90b48-35f9-4d10-9cd0-de7588d00934 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R´e.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Fu, Stefano Ermon, Atri Rudra, and Christopher R´e

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.292144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:a30e2d9a6de4d91f62de689dcdee9fa3ea4a3dc9c2bfe3415d2791c8155cd154

Observation df80284b-8a5d-4368-a0b1-c228898df0d1 · outbound

This paper cites coco-gemini: Zero-shot COCO detec- tion with Gemini.https://github.com/simedw/ coco-gemini, 2025.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention coco-gemini: Zero-shot COCO detec- tion with Gemini.https://github.com/simedw/ coco-gemini, 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.276645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:879318efb6e1c4e602b411ea39f79d5da4f0779250c5566ba38e2d2486485f90

Observation 67814fe2-40d1-45b6-8552-1652a7985482 · outbound

This paper cites Multi-modal hallucination control by visual information grounding.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Multi-modal hallucination control by visual information grounding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.310543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:4c58a75057830bb577d5d89813d46dd87f1a64dbe99ee2f351b42409b66fa2ed

Observation eb0b4114-781a-463c-a197-5e1d513d4e63 · outbound

This paper cites TALL: Temporal activity localization via language query.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention TALL: Temporal activity localization via language query

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.280076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:87340656cc5448c205f9c5d3c95436fdd3242b154df4fba3c9fb521db4e58c71

Observation 52dcb6a6-ebf9-4184-aaed-7871279815dc · outbound

This paper cites Gemma 3 Technical Report.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Gemma 3 Technical Report

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:55:33.902913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:8b78962044e9e8cafdfdf4ace6d7987255e5e9dbe758f490a19187d43aba4df9

Observation 92ea6d40-80d9-4b7e-9eb1-b033106f1a52 · outbound

This paper cites Gemmeke, Daniel P.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Gemmeke, Daniel P

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.308839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:314d29bf5e673226290150f47f21e5f94dd2c8cfbb3b4f5d6d50cddb4f0b46b7

Observation 3f087581-5dfa-440e-a7ca-ef0f46246b8d · outbound

This paper cites Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:55:33.901819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:24d4979c4792939c05b7c918643e27330aba6b2913c7166c26ad2b83d952420b

Observation c0cc97b4-24f1-4ba0-ad02-2566d353699b · outbound

This paper cites DAMRO: Dive into the attention mechanism of LVLM to re- duce object hallucination.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention DAMRO: Dive into the attention mechanism of LVLM to re- duce object hallucination

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.286983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:4488b67849642cb25b9a07029bd1fdaa5932efd643f7c731c055588af1e02d8e

Observation be4929bb-018b-4874-93be-351e5d8346ac · outbound

This paper cites Making the V in VQA matter: El- evating the role of image understanding in visual question answering.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Making the V in VQA matter: El- evating the role of image understanding in visual question answering

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.317666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:24c228ed5445f6b3a41c2de3f77cc03b4a312f6ba2593ed34ad4c188318aa3da

Observation a4eff65d-50b6-47a3-b525-57bcfdd5b2ad · outbound

This paper cites an unresolved cited work.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-07-08T20:45:37.305482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:181fa875a53b5e3e04774b72af51c9a1090ed70158eee43932c2bf3263bfb893

Observation 6b426b24-5376-44d6-bb0d-fdd9f36d3f9a · outbound

This paper cites OPERA: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention OPERA: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.283399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:a9e56d7a2286df84a88e8a62e2164440cf8f0103ea6448fb9e54d959b4489d0f

Observation 54ddb34c-2b88-45bb-9f4d-3ee8ef27d513 · outbound

This paper cites Interpreting and editing vision-language rep- resentations to mitigate hallucinations.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Interpreting and editing vision-language rep- resentations to mitigate hallucinations

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.290443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:8c4d286aa36d34cbf92450528ec0843256febb31303952b3f2877bc9ca0a56ca

Observation 324b8e7c-c3b1-4a52-ad62-526928ef1beb · outbound

This paper cites Devils in middle layers of large vision- language models: Interpreting, detecting and mitigating ob- ject hallucinations via attention lens.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Devils in middle layers of large vision- language models: Interpreting, detecting and mitigating ob- ject hallucinations via attention lens

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.285189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:87224b3243a34d677081ef3c392b08ccc4a14bcd36489c61038a836de186f0a7

Observation 229bcd6c-8178-4afd-8a8c-ce2525cfcc1c · outbound

This paper cites Shamma, Michael S.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Shamma, Michael S

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.307082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:a102460bc8bb93c126efa8016e929591a342881e01ee2ad5cc94816331a6d71d

Observation 05d4d1e9-b467-4b78-96f3-624df10f30b7 · outbound

This paper cites Berg, and Mohit Bansal.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Berg, and Mohit Bansal

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.293761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:03da532ea6043b96fa11b9e7442630a1cb48b3be531f26e41e7110ddd99eba45

Observation 3c7a3251-c031-4d6a-9865-6bf63d160993 · outbound

This paper cites Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.268615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:e9eace4074eb6db02cc9bdd904f401cfdedb1f2337d85f6dff8d2195859307fd

Observation 62a7be5d-a2b1-4828-9052-a9cf2542b54a · outbound

This paper cites Grounded language-image pre-training.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Grounded language-image pre-training

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.300407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:0331229b4413d8fd1a4b3d4ff495ded98b5acc4602d9b5b311526b5b77dd4b10

Observation 8870ead8-958a-473e-af59-edec0dd466d6 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Evaluating object hallucination in large vision-language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.281676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:00e4f740e7079eb3610b61662e8c7c0a0b8e61bafd015c1ecf022971a5840c01

Observation bac235d7-7159-4f89-868b-7e955c3dcbca · outbound

This paper cites Lawrence Zitnick.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Lawrence Zitnick

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.263281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:214330763fc399ec1e4816b04b19bb1c6831c07c8deac0b10f47edeb85e68618

Observation 8dd7a3de-7b9d-4f7f-aba3-ae3fc92ce1ae · outbound

This paper cites Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.312378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:3f55f3bd268161ee54529547d0e2b51284913f03dea716086e6079250d1028e7

Observation 535ec38c-64a7-4aeb-b415-ca87ea3989d2 · outbound

This paper cites Paying more atten- tion to image: A training-free method for alleviating halluci- nation in LVLMs.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Paying more atten- tion to image: A training-free method for alleviating halluci- nation in LVLMs

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.261517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:03a88e2b31fdebde2cb1408f90529436dfbadd60b0316f42adec609300fe07ec

Observation c5c98b0b-11c7-46c1-a9ae-c10bdc58afcf · outbound

This paper cites MMBench: Is your multi-modal model an all-around player? InECCV, 2024.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention MMBench: Is your multi-modal model an all-around player? InECCV, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.314104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:1bacf10eee5da403f89199628984a87fbdc39a456d9aedfeb4d28a0a2d2abf3d

Observation 142b4287-43b9-4fcf-9dd3-77db617e1d8e · outbound

This paper cites Simple open-vocabulary object detection with vi- sion transformers.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Simple open-vocabulary object detection with vi- sion transformers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.288680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:92331dec82c47a2c71fe3b0395f6d8731bdebb94b110b57120b82c2d450fe4f8

Observation 1b3ccc82-5771-4cb9-adb4-dc273502c279 · outbound

This paper cites Query-Dependent Video Represen- tation for Moment Retrieval and Highlight Detection.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Query-Dependent Video Represen- tation for Moment Retrieval and Highlight Detection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.297122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:1c2ac5293d4a49beedb6d02cf38245a8a8a8f0a6f1fb6e787b87e55acd67de87

Observation cf384d7b-f0d5-46c9-930c-b400c45887cc · outbound

This paper cites Verjans, Phi Le Nguyen, and Vu Minh Hieu Phan.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Verjans, Phi Le Nguyen, and Vu Minh Hieu Phan

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.298811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:2d2b04ef9c63594d603d4a0bc436d8b56e3f8440ace98a738a13f43ea2798764

Observation 281d80a6-1c69-4e76-8a24-9ef93bbb5bf4 · outbound

This paper cites GLSim: Detecting ob- ject hallucinations in LVLMs via global-local similarity.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention GLSim: Detecting ob- ject hallucinations in LVLMs via global-local similarity

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.278345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:44af8767e15b8d74f0f610d4502cdd29c0d4842923ff0e435b88a1ae6d02d1a9

Observation 156f1f7c-eabb-4f3c-9bdb-8b84acedf8f2 · outbound

This paper cites Kosmos-2: Grounding multimodal large language models to the world.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Kosmos-2: Grounding multimodal large language models to the world

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.295387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:73cb20e9a18b26216b4ea821d5f4191888387098d5824d9b95652b0026c525c3

Observation ba38589b-a41b-4b48-ba66-9c86bb97a9ba · outbound

This paper cites Beyond logit lens: Contextual embeddings for robust hallucination detection & grounding in VLMs.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Beyond logit lens: Contextual embeddings for robust hallucination detection & grounding in VLMs

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.303962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:97f1d7679343f909555c8f21d2febb3175269a97355cbd553170bd77152d8d90

Observation 6e499d5b-46c1-46ad-aad6-93a29f9e9197 · outbound

This paper cites Effective pre- training of audio transformers for sound event detection.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Effective pre- training of audio transformers for sound event detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.302250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:2dbb91903def3b9b1ae0b30385fbe7073cf6385e0d761fc8eb57d7fedaffac86

Observation 4c1a8bf1-36d2-4fdc-beac-e8c39e32b0f5 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:55:33.897014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:2fa519c51ef9593715f07729627b9040a49e3dc5518e4cbe901a834168d0e7a4

Observation c01575d3-9dd5-44a0-8ce1-cc07aa9206cf · outbound

This paper cites Berg, and Tamara L.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Berg, and Tamara L

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.315775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:02b8b6db62dd98e156dfc00ad304f53d4200e20019a24173c55b723470c2672d

Observation a484ad0f-67ce-4a29-a1c8-e12d11fc78e6 · outbound

This paper cites bbox_2d": [x1,y1,x2,y2],.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention bbox_2d": [x1,y1,x2,y2],

Reference 38

Resolution
malformed identifier
raw_fallback, observed 2026-07-08T20:45:37.274934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:d1bb31c621518ca4fff42136756e90fb4ee14ed6ff76c34aa36bc7d6c0272cbb

Pith citing papers

No inbound Pith citation observations are available.