Pith. sign in

Paper Citation Record · LEDGER

Benchmarking and Improving Detail Image Caption

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2405.19092.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.19092 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:26:11.913594Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T22:49:01.206350Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 813805c8-f1ca-4920-bd4d-4238b150c20f · inbound

Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation cites this paper.

Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation Benchmarking and Improving Detail Image Caption

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:11.913594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:11.913594Z digest=sha256:d789636ab72627ef5f8eb48ee3390961369e91507c62392ad14134841951d8b9

Observation 60695a16-2667-4633-927a-546143c50a99 · inbound

Exploring The Visual Feature Space for Multimodal Neural Decoding cites this paper.

Exploring The Visual Feature Space for Multimodal Neural Decoding Benchmarking and Improving Detail Image Caption

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:16:40.200691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:16:40.200691Z digest=sha256:fee20d04bafa40a4c41d5d8888bab1853f838e009d4449b262bf673beb0645c3

Observation a33f1fad-b67d-4fe2-9011-183d95200ed5 · inbound

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM cites this paper.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM Benchmarking and Improving Detail Image Caption

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:10.933843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:10.933843Z digest=sha256:1f1a936ef669a122c451f977ef8786d66819005bbd634672cbd08dd0ff56ffc1

Observation 23690255-5bef-4f42-9c2d-01ee12a1f55b · inbound

GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning cites this paper.

GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning Benchmarking and Improving Detail Image Caption

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:05.762074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:28:05.762074Z digest=sha256:b915e3406c710e42978031092d6eb1e3d20bfd09e61578e7b5125ee568095db3

Observation 285fe182-3d78-4df4-860c-79512ddbfdd4 · inbound

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction cites this paper.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Benchmarking and Improving Detail Image Caption

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.153714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.153714Z digest=sha256:96de7c9a8420b1ed5de7aedf0805778050efe50a2f5d9112dd7f863e72062c18

Observation 93ed0f7c-a450-4603-a1d2-da3006019d51 · inbound

Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline cites this paper.

Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline Benchmarking and Improving Detail Image Caption

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:07.461098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:07.461098Z digest=sha256:4de06769a55587929655d1a799625cb274c7cdcf28862bbd6aeac1e7c1f2d8bf

Observation 1652daa6-5db0-47f8-9260-53ec472bc50c · inbound

VGR: Visual Grounded Reasoning cites this paper.

VGR: Visual Grounded Reasoning Benchmarking and Improving Detail Image Caption

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:12:14.524347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T09:11:00.295700Z digest=sha256:e103d28c86a4d6cc322c321fef2fe266aa0a8f3a00a69b0603a5124000937dee

Observation b9a8a1c8-dff9-438b-a88b-5f5ccc91fee6 · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Benchmarking and Improving Detail Image Caption

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.136534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.136534Z digest=sha256:cb449acf47c43f6ac629277f0ec71b2ebcc3d18573ee3ccc16ae633ea6e489bd

Observation f4df095d-74b4-4c46-ba70-32982931c397 · inbound

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning cites this paper.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Benchmarking and Improving Detail Image Caption

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:34.074648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:34.074648Z digest=sha256:9554a337a70a601356c34a9dae5816ead67a90695119dea07c6e0b66746bc09f

Observation 7efa60df-1ec1-42e0-922b-b3e191ecf6f9 · inbound

SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation cites this paper.

SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation Benchmarking and Improving Detail Image Caption

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:55.536773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:38:55.536773Z digest=sha256:c639db0afc3ab37846a77901b57a520f9bdc15e859d425ea68b04a7e9782585b

Observation 57d04ba4-f38c-4f2f-aa8e-40cba7f54fb5 · inbound

CaptionQA: Is Your Caption as Useful as the Image Itself? cites this paper.

CaptionQA: Is Your Caption as Useful as the Image Itself? Benchmarking and Improving Detail Image Caption

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:44:07.252387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T05:42:57.288054Z digest=sha256:5c00be2712e00efb3ca4388df87af3335d1f13c4cf2e663d0d97e83b5b991f16

Observation bd380b50-7238-4f25-b968-bf86e303ea03 · inbound

DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions cites this paper.

DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions Benchmarking and Improving Detail Image Caption

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:10:52.908202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:39:29.592129Z digest=sha256:9ab7587c353cf3b15afec6d08daa7658c4eab23aaf79d594715f5d323e2efd5c

Observation 2b4fd3c9-f59f-41f6-8b87-4bb3d60f7392 · inbound

Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction cites this paper.

Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction Benchmarking and Improving Detail Image Caption

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:55.072265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:16:31.701345Z digest=sha256:59ca46a76168fd64c38363eedc26bacb3f1f04b94e91ee87ee2c8155696fc807

Observation 95564b76-b7e8-4276-997c-26f3b7849598 · inbound

Evaluating Remote Sensing Image Captions Beyond Metric Biases cites this paper.

Evaluating Remote Sensing Image Captions Beyond Metric Biases Benchmarking and Improving Detail Image Caption

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T01:10:09.388979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T01:06:35.604862Z digest=sha256:4ebc4c74a7143d1139a1d73b71630073016b072adb03555e6e5c7576734fb09e

Observation b52990be-c408-4fb4-abbe-760ec644682e · inbound

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation cites this paper.

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation Benchmarking and Improving Detail Image Caption

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:07.565923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T14:10:28.941565Z digest=sha256:796bdf1396884e00da72c670087ab5806c2967146dbcb17be8cfd192249d76b0

Observation 15ce065b-da16-4908-a860-9868c414d8f5 · inbound

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models cites this paper.

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models Benchmarking and Improving Detail Image Caption

Reference 124

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:19:28.558057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T20:13:18.813131Z digest=sha256:3ffb904b1c431ca6c3aa496b1665c74ce5e64c2dee4541079d354e66d7d6d4e4

Observation 53f6c6cc-1921-43c5-81d7-76619efc7981 · inbound

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models cites this paper.

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models Benchmarking and Improving Detail Image Caption

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:18:31.277022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:14:59.487486Z digest=sha256:e023a3c2ef0e2a496c4f3c1e058fd9786b1325e7a3ec84e9c26011cee96518a4

Observation a73d521d-0fd6-4a4d-aa45-aca8f4a986fb · inbound

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models cites this paper.

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models Benchmarking and Improving Detail Image Caption

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:49:05.326127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T21:47:32.112057Z digest=sha256:505701c2d18538689ebb4b5a9bf4c0120f6e03dff54bb32dc4bd19b4a9f66115

Observation 7309b0d8-e611-4360-b512-130b7e3a228f · inbound

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison cites this paper.

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison Benchmarking and Improving Detail Image Caption

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.742970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T08:36:27.676888Z digest=sha256:fc9706f11ed1514b1c4e11fbc709abeec09013a5b94e1df5e4a5e367b106c9c1

Observation 01a168a1-f316-41bb-815b-d6fce0b9e254 · inbound

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison cites this paper.

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison Benchmarking and Improving Detail Image Caption

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.330839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:57:47.409741Z digest=sha256:c55900f9040483d020baccc72cf48927a645b299d260ddad103d72e05145771d

Observation 1c8fa18b-14d9-41df-b6cd-c5e96f593863 · inbound

ChartFI: Benchmarking Faithfulness and Insightfulness of Chart Descriptions from Multimodal Large Language Models cites this paper.

ChartFI: Benchmarking Faithfulness and Insightfulness of Chart Descriptions from Multimodal Large Language Models Benchmarking and Improving Detail Image Caption

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:15:20.104167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T04:13:11.756347Z digest=sha256:4955639afc52c8324f09bdcffc3b24c4a5b2d42116cfb1670c1f292646a1ded2

Observation 2056ba08-dead-4a32-9d61-f7822c5dc32f · inbound

ChartFI: Benchmarking Faithfulness and Insightfulness of Chart Descriptions from Multimodal Large Language Models cites this paper.

ChartFI: Benchmarking Faithfulness and Insightfulness of Chart Descriptions from Multimodal Large Language Models Benchmarking and Improving Detail Image Caption

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:14:53.358294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T16:10:01.774725Z digest=sha256:73c778ad9ea3efda12fca4fc61d329b0246c6a520d810052d29e3ceeb94f0bd0

Observation bf29fd71-c785-4c46-85ff-7997601171a5 · inbound

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models cites this paper.

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models Benchmarking and Improving Detail Image Caption

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:06:59.615769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T01:30:36.491137Z digest=sha256:3c34c9f578715e305e689189780928df3273656e7d06b4d6d4c7181d48d82111

Observation 3f29d0d0-7245-4052-b647-458683eae725 · inbound

Reliability-Prioritized Fine-Grained Generation in Multimodal Large cites this paper.

Reliability-Prioritized Fine-Grained Generation in Multimodal Large Benchmarking and Improving Detail Image Caption

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:14:20.718552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T07:13:43.265960Z digest=sha256:b77923070f68b35172523710119f59d0e0a7bcc0767d9195229f4d6d4bba4ece

Observation db0fdf84-57ac-476d-9aee-8c1109c2c26f · inbound

Reliability-Prioritized Fine-Grained Generation in Multimodal Large cites this paper.

Reliability-Prioritized Fine-Grained Generation in Multimodal Large Benchmarking and Improving Detail Image Caption

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T22:49:01.212984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T22:39:40.186814Z digest=sha256:a27739f4c4bca172b1d34cd1995b3eea106d75721c2fd56fbbdba8da3cd33d7d

Observation df35cb17-457a-4714-8dcf-25e9a5cf77bc · inbound

Reliability-Prioritized Fine-Grained Generation in Multimodal Large cites this paper.

Reliability-Prioritized Fine-Grained Generation in Multimodal Large Benchmarking and Improving Detail Image Caption

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T09:41:13.256572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:41:13.256572Z digest=sha256:f8239384044d1c567559b570fced52b83ad6fabcd06d187497badb7c919728d8

Observation f4e26d98-e229-4de5-8b26-f342ab24b51f · inbound

Holo-Captioning: Toward the Text Equivalent of 3D Scenes cites this paper.

Holo-Captioning: Toward the Text Equivalent of 3D Scenes Benchmarking and Improving Detail Image Caption

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T06:12:48.722467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:12:48.722467Z digest=sha256:a894d72daf2f4bbe5264011f4d876609fc6b17c87026154a1f6b2e655423608a

Observation 4e184ca3-34f5-4e1a-b777-799addbfa806 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World Benchmarking and Improving Detail Image Caption

Reference 250

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:76c048c6f714d7242082c33dc7e85c26815e60aec748252a55030ba0de18541c

Observation 3405f8cb-2315-4bf4-a71e-fe7e5c686028 · inbound

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation cites this paper.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Benchmarking and Improving Detail Image Caption

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:49.411968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:49.411968Z digest=sha256:abf30019320d7acec62c8632614b4f58d213edf3591e9ca8ae4981d22c9dbf31