Pith. sign in

Paper Citation Record · LEDGER

Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2408.04594.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.04594 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:01:12.370094Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T11:55:33.372242Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d87a4c97-c318-4cce-b567-56c9cff7cff2 · inbound

EventGPT: Event Stream Understanding with Multimodal Large Language Models cites this paper.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.370094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.370094Z digest=sha256:870d45994c6dde7d4daabc6078e5687c09a8b23c942c9920b04d554b74ccdbef

Observation f2529405-57f4-4820-b8f5-ee92f68e3190 · inbound

Progress-Aware Video Frame Captioning cites this paper.

Progress-Aware Video Frame Captioning Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:58.381034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:58.381034Z digest=sha256:ddde1b6a3b8ddb2e0d484832b3680fdd00e549c4f7773e7193087247c10837a4

Observation 6e343e37-5f03-41bf-b747-fca572f1783c · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.925517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:247e136dc170d0dcc2c672c9d3e2dc8abde371c80a343e3268cf1efbffb7d645

Observation 2ef8ce1a-999d-44fb-bdec-d0ef11d773ad · inbound

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models cites this paper.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:10.953813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:10.953813Z digest=sha256:c9b81163d6ccda421ae552349cecd3177a90e6e10667bd848fe3503e29f4db6c

Observation 6ccdf34e-d9cf-4a89-841f-9bb7359e47e8 · inbound

Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence cites this paper.

Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:18.289420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:18.289420Z digest=sha256:92bb7c1d34b4054e65f0841bec7dc60aea4cd1dca87f79d2c03986255d058f52

Observation d45ef472-ecbd-4580-88e2-d8349dac5d33 · inbound

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models cites this paper.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.553252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.553252Z digest=sha256:1a11ed89506012648205ed169200aa7c15b6f57c9637f726eee23403f803c389

Observation 6ecfd414-2ba0-4de2-bf97-3d0dddc260f4 · inbound

Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models cites this paper.

Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T16:44:00.638413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:44:00.638413Z digest=sha256:1ec6d8ef26b429e29c63e969442e8e8261ba1b30a48a0cf248392e9b81281ec5

Observation 0a710233-2290-482c-8601-ea476edc3b1f · inbound

AD-Copilot: A Vision-Language Assistant for Industrial Anomaly Detection via Visual In-context Comparison cites this paper.

AD-Copilot: A Vision-Language Assistant for Industrial Anomaly Detection via Visual In-context Comparison Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:55:33.374124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T11:54:18.587529Z digest=sha256:4b5b6ff774954e35a2127db8e7f98c1798859651a394651d3f8f3bb0140d4ec8