Pith. sign in

Paper Citation Record · LEDGER

ImageInWords: Unlocking Hyper-Detailed Image Descriptions

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2405.02793.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.02793 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T15:26:18.791790Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T14:41:30.125046Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a828e6d3-4d90-4496-be4d-97b563001f75 · inbound

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models cites this paper.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models ImageInWords: Unlocking Hyper-Detailed Image Descriptions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.791790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.791790Z digest=sha256:460d0792d767ba835c593c986e61d77a495c76a63987134cb478a116d066a113

Observation d6fbdd1b-17b8-4566-bbc1-4cb0689ba3bf · inbound

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation cites this paper.

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation ImageInWords: Unlocking Hyper-Detailed Image Descriptions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T11:41:44.364946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:41:44.364946Z digest=sha256:a18066c994fb368c1b5f91fe721dec8b4972a70b1453645ff8c82cdaaf6d881b

Observation 865be4a2-489c-47a7-8a66-36dddf56bc2b · inbound

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World cites this paper.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World ImageInWords: Unlocking Hyper-Detailed Image Descriptions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.220822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.220822Z digest=sha256:467a7903514669c8b04ef1758e857b4a8a0c7b97d3dc5c294222363e3ff9f77e

Observation dea49971-4fc1-4e00-81e8-0d2aecc88b25 · inbound

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text cites this paper.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text ImageInWords: Unlocking Hyper-Detailed Image Descriptions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.976190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.976190Z digest=sha256:82ba20e75a1662d783e9005d576b06e75d8b201587c32e2239a48c4b35468935

Observation e6dc0a98-1f1e-4dd8-af7e-2e8d392853db · inbound

ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation cites this paper.

ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation ImageInWords: Unlocking Hyper-Detailed Image Descriptions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T05:59:15.495125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:59:15.495125Z digest=sha256:dc3abb352e669d27ed38a780e861c945762984d461acb93564f54d3935e9fe87

Observation 6700b83e-be5f-4701-bb27-bef5698197a2 · inbound

SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation cites this paper.

SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation ImageInWords: Unlocking Hyper-Detailed Image Descriptions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:55.764871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:38:55.764871Z digest=sha256:b2a88c6798231ad4ccf3ed886a0de481ab77fb610bb999056427379ceb988569

Observation c27dfd70-2e5f-4be8-bb32-0c6aa7b97957 · inbound

Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs cites this paper.

Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs ImageInWords: Unlocking Hyper-Detailed Image Descriptions

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:41:30.128181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T14:41:20.403259Z digest=sha256:33baea894f508e6963e105b2be2af8d24f91814356b731c5f722de338908bfbb