Pith. sign in

Paper Citation Record · LEDGER

Benchmarking and Improving Detail Image Caption

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2405.19092.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.19092 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:26:11.913594Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T22:49:01.206350Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 813805c8-f1ca-4920-bd4d-4238b150c20f · inbound

Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation cites this paper.

Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation Benchmarking and Improving Detail Image Caption

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:11.913594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:11.913594Z digest=sha256:d789636ab72627ef5f8eb48ee3390961369e91507c62392ad14134841951d8b9

Observation 60695a16-2667-4633-927a-546143c50a99 · inbound

Exploring The Visual Feature Space for Multimodal Neural Decoding cites this paper.

Exploring The Visual Feature Space for Multimodal Neural Decoding Benchmarking and Improving Detail Image Caption

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:16:40.200691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:16:40.200691Z digest=sha256:fee20d04bafa40a4c41d5d8888bab1853f838e009d4449b262bf673beb0645c3

Observation a33f1fad-b67d-4fe2-9011-183d95200ed5 · inbound

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM cites this paper.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Benchmarking and Improving Detail Image Caption

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.933843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.933843Z digest=sha256:60afac6b7c47d177af465ba7dd1c471239c8e1c76eebe68b87fd4b5f7e09ef98

Observation 23690255-5bef-4f42-9c2d-01ee12a1f55b · inbound

GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning cites this paper.

GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning Benchmarking and Improving Detail Image Caption

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:05.762074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:28:05.762074Z digest=sha256:3b42264582893c5bdabfa1ac2b086b2312c73c5b25a073ea98255f6efc4bd375

Observation 285fe182-3d78-4df4-860c-79512ddbfdd4 · inbound

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction cites this paper.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Benchmarking and Improving Detail Image Caption

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.153714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.153714Z digest=sha256:96de7c9a8420b1ed5de7aedf0805778050efe50a2f5d9112dd7f863e72062c18

Observation 93ed0f7c-a450-4603-a1d2-da3006019d51 · inbound

Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline cites this paper.

Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline Benchmarking and Improving Detail Image Caption

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:07.461098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:07.461098Z digest=sha256:4de06769a55587929655d1a799625cb274c7cdcf28862bbd6aeac1e7c1f2d8bf

Observation 1652daa6-5db0-47f8-9260-53ec472bc50c · inbound

VGR: Visual Grounded Reasoning cites this paper.

VGR: Visual Grounded Reasoning Benchmarking and Improving Detail Image Caption

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:12:14.524347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T09:11:00.295700Z digest=sha256:8159031f073cae73995813f30b940ddd65e5e4bd51c2d7663287b4922911d3cb

Observation b9a8a1c8-dff9-438b-a88b-5f5ccc91fee6 · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Benchmarking and Improving Detail Image Caption

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.136534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.136534Z digest=sha256:b85973f8b5fb037c3b26417b0fbd46c0bb869c5df64c5f6c5208f4bbd4793b36

Observation f4df095d-74b4-4c46-ba70-32982931c397 · inbound

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning cites this paper.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Benchmarking and Improving Detail Image Caption

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:34.074648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:34.074648Z digest=sha256:9554a337a70a601356c34a9dae5816ead67a90695119dea07c6e0b66746bc09f

Observation 7efa60df-1ec1-42e0-922b-b3e191ecf6f9 · inbound

SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation cites this paper.

SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation Benchmarking and Improving Detail Image Caption

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:55.536773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:38:55.536773Z digest=sha256:ad9db368e58cebd6e7a36f7ca9654015a4784febf02d821efebe159ea6010512

Observation 57d04ba4-f38c-4f2f-aa8e-40cba7f54fb5 · inbound

CaptionQA: Is Your Caption as Useful as the Image Itself? cites this paper.

CaptionQA: Is Your Caption as Useful as the Image Itself? Benchmarking and Improving Detail Image Caption

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:44:07.252387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T05:42:57.288054Z digest=sha256:efc8eef5eb46b02b476d6fe9b14fe4489ac47b17e59e746b1d89cf3b50f73cea

Observation bd380b50-7238-4f25-b968-bf86e303ea03 · inbound

DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions cites this paper.

DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions Benchmarking and Improving Detail Image Caption

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:10:52.908202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:39:29.592129Z digest=sha256:75d7762a934c5c9675b16070c2489e3872eb5dd142b3aa58f5dc3ec7ee3f3316

Observation 2b4fd3c9-f59f-41f6-8b87-4bb3d60f7392 · inbound

Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction cites this paper.

Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction Benchmarking and Improving Detail Image Caption

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:55.072265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:16:31.701345Z digest=sha256:84b136a9bb0f84dbfff0d074444e611a2e74390993a34040f635bb82d0c090ec

Observation 95564b76-b7e8-4276-997c-26f3b7849598 · inbound

Evaluating Remote Sensing Image Captions Beyond Metric Biases cites this paper.

Evaluating Remote Sensing Image Captions Beyond Metric Biases Benchmarking and Improving Detail Image Caption

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T01:10:09.388979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T01:06:35.604862Z digest=sha256:b2f3fd34c40c4ecdcfc33ccba2ccd02aadf63f0f9d5545b06ff4af5d652a0a01

Observation b52990be-c408-4fb4-abbe-760ec644682e · inbound

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation cites this paper.

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation Benchmarking and Improving Detail Image Caption

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:07.565923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T14:10:28.941565Z digest=sha256:c68f73dbd192eee2f00916e68612421a5307df5c7afc2d8850b314e19c9c917a

Observation 15ce065b-da16-4908-a860-9868c414d8f5 · inbound

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models cites this paper.

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models Benchmarking and Improving Detail Image Caption

Reference 124

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:19:28.558057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T20:13:18.813131Z digest=sha256:9ba173be1383dd38ea7e49553b24f31ce5318c0ecc75face2a4d22100d355677

Observation 53f6c6cc-1921-43c5-81d7-76619efc7981 · inbound

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models cites this paper.

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models Benchmarking and Improving Detail Image Caption

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:18:31.277022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:14:59.487486Z digest=sha256:5e1b76daafe711345fe716ebaded3a07e355ca2e0446760234510e676bbeb0ce

Observation a73d521d-0fd6-4a4d-aa45-aca8f4a986fb · inbound

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models cites this paper.

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models Benchmarking and Improving Detail Image Caption

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:49:05.326127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T21:47:32.112057Z digest=sha256:f967314facbdc42026500f5ed358c3cd88e9a05c1a5c2dffdfb550436f30f765

Observation 7309b0d8-e611-4360-b512-130b7e3a228f · inbound

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison cites this paper.

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison Benchmarking and Improving Detail Image Caption

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.742970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T08:36:27.676888Z digest=sha256:dcbc4e8fc3d5c08e528135677509533cb1edddec1c51739fa288143ac61aa980

Observation 01a168a1-f316-41bb-815b-d6fce0b9e254 · inbound

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison cites this paper.

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison Benchmarking and Improving Detail Image Caption

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.330839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T17:57:47.409741Z digest=sha256:b4afd9c401846b5b3c43451ac972719d0099d8bafba1963919e50b63a942d8ed

Observation 1c8fa18b-14d9-41df-b6cd-c5e96f593863 · inbound

ChartFI: Benchmarking Faithfulness and Insightfulness of Chart Descriptions from Multimodal Large Language Models cites this paper.

ChartFI: Benchmarking Faithfulness and Insightfulness of Chart Descriptions from Multimodal Large Language Models Benchmarking and Improving Detail Image Caption

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:15:20.104167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T04:13:11.756347Z digest=sha256:029e99d71b943b2ad820e85c658809996cadd8140d4bd49f585ba0f32d5b7f83

Observation 2056ba08-dead-4a32-9d61-f7822c5dc32f · inbound

ChartFI: Benchmarking Faithfulness and Insightfulness of Chart Descriptions from Multimodal Large Language Models cites this paper.

ChartFI: Benchmarking Faithfulness and Insightfulness of Chart Descriptions from Multimodal Large Language Models Benchmarking and Improving Detail Image Caption

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:14:53.358294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T16:10:01.774725Z digest=sha256:2cf0a533d740db24f57c045eaa2173b4a8d1c9a02f5f7385a68ec819b993a15c

Observation bf29fd71-c785-4c46-85ff-7997601171a5 · inbound

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models cites this paper.

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models Benchmarking and Improving Detail Image Caption

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:06:59.615769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T01:30:36.491137Z digest=sha256:6dfbd0b3921c77be003428b91d2350d0f60d7a911eb0e8adfe9b86989a69bfa3

Observation 3f29d0d0-7245-4052-b647-458683eae725 · inbound

Reliability-Prioritized Fine-Grained Generation in Multimodal Large cites this paper.

Reliability-Prioritized Fine-Grained Generation in Multimodal Large Benchmarking and Improving Detail Image Caption

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:14:20.718552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T07:13:43.265960Z digest=sha256:e21b330b35ab2bca8c16cc0bd2557f7759c2aa0a549f6f3b5fd3eb526a86a3a5

Observation db0fdf84-57ac-476d-9aee-8c1109c2c26f · inbound

Reliability-Prioritized Fine-Grained Generation in Multimodal Large cites this paper.

Reliability-Prioritized Fine-Grained Generation in Multimodal Large Benchmarking and Improving Detail Image Caption

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T22:49:01.212984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T22:39:40.186814Z digest=sha256:fc188bbf0b9b6f74780a71b85fbf286750cdaa2ab2f28fe2f2d213a87fe87938

Observation df35cb17-457a-4714-8dcf-25e9a5cf77bc · inbound

Reliability-Prioritized Fine-Grained Generation in Multimodal Large cites this paper.

Reliability-Prioritized Fine-Grained Generation in Multimodal Large Benchmarking and Improving Detail Image Caption

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T09:41:13.256572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:41:13.256572Z digest=sha256:f8239384044d1c567559b570fced52b83ad6fabcd06d187497badb7c919728d8

Observation f4e26d98-e229-4de5-8b26-f342ab24b51f · inbound

Holo-Captioning: Toward the Text Equivalent of 3D Scenes cites this paper.

Holo-Captioning: Toward the Text Equivalent of 3D Scenes Benchmarking and Improving Detail Image Caption

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T06:12:48.722467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:12:48.722467Z digest=sha256:a894d72daf2f4bbe5264011f4d876609fc6b17c87026154a1f6b2e655423608a

Observation 4e184ca3-34f5-4e1a-b777-799addbfa806 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World Benchmarking and Improving Detail Image Caption

Reference 250

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:76c048c6f714d7242082c33dc7e85c26815e60aec748252a55030ba0de18541c

Observation 3405f8cb-2315-4bf4-a71e-fe7e5c686028 · inbound

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation cites this paper.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Benchmarking and Improving Detail Image Caption

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:49.411968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:49.411968Z digest=sha256:abf30019320d7acec62c8632614b4f58d213edf3591e9ca8ae4981d22c9dbf31