Pith. sign in

Paper Citation Record · LEDGER

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs

As of 20 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2506.17901.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17901 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:27:19.422402Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T04:55:07.037081Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 273b3927-c791-4d73-ae2b-d1bc481b4586 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.137072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.137072Z digest=sha256:99c38bf1a0ed7b78e182f8a623ddf39f85cb785d7ace51e8e9d5b34bf8b7c948

Observation aa400c8d-d4ca-4635-95fa-5316f5e2134a · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.142549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.142549Z digest=sha256:c6c9d2a7630898c13297f1be28cf786b6892949bd346bf373e93a69e84d34d12

Observation 32175eb1-23fc-440a-8e50-8f07e4a0e8bf · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.147296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.147296Z digest=sha256:5c30fbd70388edc3b5c14f73fbfb43d76830ae917d78e0ebf179a57ad51ede05

Observation d4b16d96-276a-42a5-9803-672a34bcd566 · outbound

This paper cites GPT-4 Technical Report.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.153898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.153898Z digest=sha256:859725bf2128351d9d58534de54f957e030588d3c14b1681ace295f802e2eabe

Observation d19aa91a-eb45-4ff2-ad7c-43d7846ad7cd · outbound

This paper cites Improved baselines with visual instruction tuning.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Improved baselines with visual instruction tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.159315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.159315Z digest=sha256:786aad4c714988075c8d254bdfffbffc1b49df2b1262c95ca3a99305b3415f31

Observation ca2a5fca-638c-4e0f-9852-7cb5c097dad5 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.164043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.164043Z digest=sha256:20b8faeb4079c09059138f4c79eb99cc18a4ea7e7e45951140ab832fbd960511

Observation f82fde58-886e-4f91-9968-e168b67e6bee · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.168676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.168676Z digest=sha256:c1c167d2b91a5aa062aa34ac1cbeac5c8cc74ed1cda5bc0b4515d7c9dff7439c

Observation 93ceceb7-6b52-404b-913f-734d14c999be · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs A Survey on Hallucination in Large Vision-Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.173386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.173386Z digest=sha256:86adb2cb57a37ce1506f7a4cbbbf7ca48ce135175845643f927a8d7edde80e3e

Observation 02ba2f93-d8a2-4594-a6b1-07d36ce67313 · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Hallucination of Multimodal Large Language Models: A Survey

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.178110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.178110Z digest=sha256:44d5faaca6fc722acffdf4d0cb3a6efeb285d1c8217ea42e56d8bb0d88681d8d

Observation 40e6b8fc-6543-44f4-93fa-f8dc726fa650 · outbound

This paper cites Analyzing and Mitigating Object Hallucination in Large Vision-Language Models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Analyzing and Mitigating Object Hallucination in Large Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.182352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.182352Z digest=sha256:4514e29bccee58f9a1b8985ad65ddbc1861063479bb412760df24bcaad866a82

Observation bbe7b840-b758-452b-ad9d-111b7f59d2b0 · outbound

This paper cites The deluge of spurious correlations in big data.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs The deluge of spurious correlations in big data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.186869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.186869Z digest=sha256:f83e955799a83c362dfaf5b4e627654f5a66ebf28baf4846e29ec9fe067a22db

Observation dd9ca390-6541-4b3a-9a8a-5272bd001bbe · outbound

This paper cites Spurious correlations in machine learning: A survey.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Spurious correlations in machine learning: A survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.191298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.191298Z digest=sha256:81923b75731763a85ecf2c3e077a81206f17acbaa1b3b3fd143dd161fd919bff

Observation 48703dd0-3329-485b-9045-e787be222265 · outbound

This paper cites Mitigating object hallucinations in large vision-language models through visual contrastive decoding.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mitigating object hallucinations in large vision-language models through visual contrastive decoding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.248345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.196105Z digest=sha256:4f2b1206419e580490d58124c2808dcf7229796f587b7fd16c6e811bc24ad56f

Observation 52ba5597-4e00-4d45-bd0d-02cfdc97634e · outbound

This paper cites Described object detection: Liberating object detection with flexible expressions.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Described object detection: Liberating object detection with flexible expressions

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.233649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.200189Z digest=sha256:cfdf1666d3dc908f2b177a7989e2bba20b04610791da4267de1b44215b251898

Observation b42d05f7-f51c-4955-a2e2-1cbf62a10090 · outbound

This paper cites Mattnet: Modular attention network for referring expression comprehension.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mattnet: Modular attention network for referring expression comprehension

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.220220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.205034Z digest=sha256:1d6f3aadaec181c464cfe315e5b8c6db5734ed5b335f6449c08872f978889afc

Observation 9f7685f1-b46c-4210-9ca0-edab7be1a164 · outbound

This paper cites A fast and accurate one-stage approach to visual grounding.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs A fast and accurate one-stage approach to visual grounding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.207560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.209588Z digest=sha256:96abfa5087f956197b29c62c45ab03aa6922bc67f22833bdd5bf722ea578ef8d

Observation a588206f-1b12-4e05-bdeb-c69053877f28 · outbound

This paper cites Mdetr: Modulated detection for end-to-end multi-modal under- standing.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mdetr: Modulated detection for end-to-end multi-modal under- standing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.194646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.213891Z digest=sha256:5118a290aaeaa864d664f75166375287c238ade5f4db88fc199f55a1c15bc20a

Observation 45b28f3a-55ae-4190-ad07-926b025c1485 · outbound

This paper cites Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.181167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.217883Z digest=sha256:46b1fc99d4bdc5dbbddef96f233231751a8da2e4e53bb83c81ac01e6a5cf3545

Observation d3c8b111-2b48-48c6-ba44-67752411513b · outbound

This paper cites Grounded language-image pre-training.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Grounded language-image pre-training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.222225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.222225Z digest=sha256:018d517729f9cd98dc6dcb80439d76feba859426f66aaf85b4d41f9b8e2cc0d6

Observation 18aadb9d-8643-4a34-9cc2-8673b052e883 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.226753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.226753Z digest=sha256:e8d8d951099023dc957db39b0d9ee9e5594b81304e741a69731cf8e31db2dcaa

Observation 0827b9be-23cc-4579-8685-d2ad90942133 · outbound

This paper cites Phrasecut: Language- based image segmentation in the wild.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Phrasecut: Language- based image segmentation in the wild

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.151548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.230979Z digest=sha256:6b14dc01396a43bc9df5d7525952b6e90d0762e0138cb2c77b83040cc3f495d3

Observation a711b74c-fd5d-4930-b514-ea8bfed4ece3 · outbound

This paper cites Advancing referring expression segmentation beyond single image.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Advancing referring expression segmentation beyond single image

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.138869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.235705Z digest=sha256:ed2cc6347740efe09522cd8f47cfb356d5704447eaa2126a4976be46b6d8a438

Observation a943a3ac-de6c-4415-8207-6933e1295116 · outbound

This paper cites Language as queries for referring video object segmentation.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Language as queries for referring video object segmentation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.125731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.239704Z digest=sha256:5f32f7c9383bf61a017f9728a168837d71f37c1336079c8bede9c6fd6818d351

Observation 67ce0251-1079-440a-83eb-03fc44b7c557 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.244352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.244352Z digest=sha256:82236be08c7899dc7a6251da41bd88cd61b9920db7ed028875fee9d8a8086a73

Observation 89fc4323-c67c-43ff-a2a0-cf1dfbfb5186 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.249169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.249169Z digest=sha256:cbb6f247289b1b18dc5af70beb2af72b3fbbbe2b974677648c5671b1926fd17f

Observation e9eb3e5b-e6f3-4f26-84dc-86866fc8b19a · outbound

This paper cites Llava-grounding: Grounded visual chat with large multimodal models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Llava-grounding: Grounded visual chat with large multimodal models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.111553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.254479Z digest=sha256:698185f88826299c98a66d22b06ebdd5294418183a1a87787035194bd8513215

Observation 49669868-d3c4-4eed-8c4e-67a765c3624d · outbound

This paper cites BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.259475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.259475Z digest=sha256:f05a9eed00801e52360f982785f016d89d5d9fa7201dcdd0b58067bf905f2859

Observation c608b4a2-7148-4b59-8f83-29cd6fb8de7a · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Lisa: Reasoning segmentation via large language model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.264438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.264438Z digest=sha256:97d26b0d0ff56b3c211eea5373ef7950d1e03ffb62d52d7123a5afd34caf1977

Observation 2df227a7-fe8c-42da-9e7c-5d795981635f · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Gsva: Generalized segmentation via multimodal large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.085128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.268906Z digest=sha256:102ef7f11772c7e93621e690bcc93149b24ef88d5e8214a70a8aee5b3e9568ed

Observation 3d5202c9-d0da-4c77-b193-e2d61653cb87 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Glamm: Pixel grounding large multimodal model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.274135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.274135Z digest=sha256:34dc5cf56d514729fdd3a469db0fb76edb581b339eddd05eddfae492d614ae21

Observation eac4a933-149b-4c56-8259-00a0d42a7e3a · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Pixellm: Pixel reasoning with large multimodal model

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.063536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.278494Z digest=sha256:f31e240d1e217cbf6bfd4dee37bbf8935821f35aaf78628a08ba2fae8bb2d904

Observation c6ec7099-6cf0-421d-883e-b59151bb6fe6 · outbound

This paper cites Ground- hog: Grounding large language models to holistic segmentation.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Ground- hog: Grounding large language models to holistic segmentation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.049650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.282891Z digest=sha256:b60734b578110d22d4ecfbca197a052c611aed68c3426eb26a3b9fd39cfbd614

Observation 0038e964-8c30-4998-bc21-3b7d57f2bec4 · outbound

This paper cites Woodpecker: Hallucination correction for multimodal large language models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Woodpecker: Hallucination correction for multimodal large language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.035923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.287227Z digest=sha256:e47b1190ce03f8fba703a363fe304b2337b6787577674cdb86ada18ea2184a72

Observation ffdafa19-6a83-46f6-8203-d7442500db88 · outbound

This paper cites Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.291980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.291980Z digest=sha256:2462b848ed78692b37b989699393169a4b60e13b9559a3b1c8cf6a900d31984f

Observation 2d8a7ba5-f0f8-42b4-9f20-5a7d4346ee61 · outbound

This paper cites Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.020444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.297724Z digest=sha256:cd0df1a56acf8791f93986845edd0bc71cc144f653dfeef81d3ae1fef38e5431

Observation 8e69ed7f-7f4b-442c-afca-f7e6eafabd88 · outbound

This paper cites HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.302520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.302520Z digest=sha256:c4733a37114ee7de86d12ced99c571ba4c30bd2d9ca3610bbee0f5a3c5316183

Observation a788739c-d432-430e-ab17-4578f08fca09 · outbound

This paper cites IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.308057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.308057Z digest=sha256:5f939586056c65c39ed182dcbaa9e536682b6040cce675060850a944424edb33

Observation 26c07d63-eb7f-4cfe-b766-28ba79cd7a93 · outbound

This paper cites Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.313207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.313207Z digest=sha256:cff547cea37f81cc09e935ab52789d3ad3fa418a7e776820bb89f6d274f94e30

Observation 941699a5-a68d-4549-b759-f20cd207bae8 · outbound

This paper cites Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.317381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.317381Z digest=sha256:b4b87e198aa0d0bbceefcbbe0eeb91cac87b4dccc25e753c2e3a6ced78d55412

Observation a0289201-04d8-4b11-abd7-52c0fd50dc5c · outbound

This paper cites MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.322266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.322266Z digest=sha256:6add729ab0e1cd77636a8e2276ab8f8d1659a773cafabb2aed581d19653d0945

Observation 83809ade-bc12-4f4c-a9f3-9d932fe0c3fa · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.327237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.327237Z digest=sha256:cb7ecb92923cd5a5e01b0480e108cf76abe60024e254883c0d33a59ff0fd7d78

Observation 9266776b-4b02-4827-9325-d3ce05895f3b · outbound

This paper cites Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.331587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.331587Z digest=sha256:06afc3b3cee9571db3ac92e928cd8bd42575f0e9b8efc80b552fb1735b970a7a

Observation 44f50dc7-c9ee-46c5-a2bb-9acd531102e8 · outbound

This paper cites Mitigating fine-grained hallucination by fine-tuning large vision-language models with caption rewrites.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mitigating fine-grained hallucination by fine-tuning large vision-language models with caption rewrites

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:19.998872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.335920Z digest=sha256:cc8252328f5dfea42ad031efcdad5255c0c22525911955893175e895258bd206

Observation 5a4157b9-c626-4e58-ae53-43ce1a26264a · outbound

This paper cites Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.340811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.340811Z digest=sha256:f7699059886bb8b223e6707022bd78bda91d2aeaa18b89e650760432baa78ceb

Observation a79bad53-b962-41ef-921b-072ee2ab4e2b · outbound

This paper cites Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.345738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.345738Z digest=sha256:908a0aba9f7bd8a43829d63d3815ee617ac602451c2d1fcf4969eda0cba79b69

Observation c32024ec-ed48-4892-8f1a-7e24be6b056e · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Silkie: Preference Distillation for Large Visual Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.350818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.350818Z digest=sha256:5557d2796af336d4c0e2cc061aa23213c28fc22c3c02fbb443dabb4bc5710420

Observation 2e1ece5e-0ad4-4f77-827f-bbcf6daad7bb · outbound

This paper cites Detecting and preventing hallucinations in large vision language models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Detecting and preventing hallucinations in large vision language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:19.985299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.355853Z digest=sha256:eda7d55f1a0f22963c6e00c0a3fa8773b8bddd4d273c20e304c131f1582818ac

Observation 73579177-224d-48bb-b450-acc716d508dc · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.360031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.360031Z digest=sha256:e1964f0777db5575bda7f991a0892009633ec3f890cd6372bc13c6d8e5f9338d

Observation df92342c-82be-4b17-ab21-c721e7ac7c18 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Multimodal Chain-of-Thought Reasoning in Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.364377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.364377Z digest=sha256:c3e8fb8c78b0d876f9bd1cf6bb8e85924fc0bc3d3e978fbd0f4051f6ccb27a5b

Observation 270c8a6e-257e-4f8a-88e7-8247fbfbe69f · outbound

This paper cites Lora: Low-rank adaptation of large language models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Lora: Low-rank adaptation of large language models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.368667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.368667Z digest=sha256:5eadd9c7f0a55d56a5f92d2e4bd43e1a77e2e698b96420eb467fe7b31078b616

Observation a16614ee-0a09-4571-a35c-29ce190a0800 · outbound

This paper cites Segment anything.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Segment anything

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.373049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.373049Z digest=sha256:a76bc3c0bb13d76d2325f2b5c9f83c3c4d364c46a89c3aba4c6de5a723f4b912

Observation 6087d896-7570-4182-b910-32cb740d880a · outbound

This paper cites Decoupled Weight Decay Regularization.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Decoupled Weight Decay Regularization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.377682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.377682Z digest=sha256:13ca5c92b4665c696a591c32ea1aff260ad56b8aa5a439c1cad938ba5e5748b7

Observation 721e8b60-a827-4b75-bf89-3337f723d851 · outbound

This paper cites Haloquest: A visual hallucination dataset for advancing multimodal reasoning.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Haloquest: A visual hallucination dataset for advancing multimodal reasoning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:19.954610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.381749Z digest=sha256:927a6414776bbdde981bbc7f203262a7f67b29f9bbee2c75b9134be708b20f07

Observation 8cc0add7-bea0-48fe-baad-d63090dd7ee1 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Evaluating Object Hallucination in Large Vision-Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.385833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.385833Z digest=sha256:c39434ebc8d2d6d24f92914ad3ed0a6e284eba8c3617e9ed50c9a0525aab09df

Observation c82111ec-6b2e-4b60-9ff0-465307f7a97b · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.390047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.390047Z digest=sha256:fc6066a3902eb3d6b1872d7af338ddde35d18d5ab90f54df0a0786736403a6d6

Observation 9220180b-39ae-4622-aa9e-460ed9cef1f2 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.395003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.395003Z digest=sha256:d2e007b00608a454be3fd67e400a14e0a816772cf950ce73712d6030f0b3bc77

Observation 4ca48f0e-6a07-4bc5-ad2c-1bd431309bfb · outbound

This paper cites A survey of multimodel large language models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs A survey of multimodel large language models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:19.923518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.399460Z digest=sha256:1aca1e0090ba330de17c9f49265892e17873d20babf7d5c8b347630674994a81

Observation be649093-e7c6-4ac7-bd39-c3a49928230a · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Referitgame: Referring to objects in photographs of natural scenes

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.403801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.403801Z digest=sha256:6da8f0c116cedf94e996a399daac398704209c86da546f91863f8ab81b297107

Observation 517d6065-62a0-40aa-91b7-816888cdbd26 · outbound

This paper cites Empowering Segmentation Ability to Multi-modal Large Language Models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Empowering Segmentation Ability to Multi-modal Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.407889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.407889Z digest=sha256:29e8e77e4b18f79267083caff4be9f1948f48aeef8807dea054e748e04061a49

Observation 5d59454f-1ad6-49e5-9138-13c7cecfad0e · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.413192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.413192Z digest=sha256:3c255aec26601a1c5f10eff2cf9b4afbdfa69c30452bcd51013818018953d9e0

Observation e0cc1135-b213-4530-bd64-d317a063904d · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.418235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.418235Z digest=sha256:5bf369d810cd7bdf15879df6da6fe695206a003dd9c16cdc5e9aff204b4c6932

Observation b6a2fb75-2e9a-4ac6-b8cc-131aba5b33ac · outbound

This paper cites Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:19.899780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:27:19.422402Z digest=sha256:400950723f06073a8d6ff7c742d2659e16de5d2f38bb1cdebb00c6adf653b350

Pith citing papers

Observation aabe33ca-8363-46f4-84c2-e6edcf9dd92d · inbound

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding cites this paper.

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T04:55:07.037081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T04:55:07.037081Z digest=sha256:fae793c0c3a9fb5befb5d4ef98ed600f83343a3b1f0b27fd6bb9ae6e1b710982