Pith. sign in

Paper Citation Record · LEDGER

CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2503.12329.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.12329 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:35.033889Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:59:53.203155Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0c5da5d0-b350-4d61-90a5-265e4b549417 · inbound

Towards Scalable SOAP Note Generation: A Weakly Supervised Multimodal Framework cites this paper.

Towards Scalable SOAP Note Generation: A Weakly Supervised Multimodal Framework CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:35:41.711143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:35:41.711143Z digest=sha256:e0f8d2363f411949f904b51b980dfd45dd1abee3d2b4e6065b8e5bc81abb92fc

Observation d7974b22-fb85-43b9-b84e-fead96685d62 · inbound

Skin-SOAP: A Weakly Supervised Framework for Generating Structured SOAP Notes cites this paper.

Skin-SOAP: A Weakly Supervised Framework for Generating Structured SOAP Notes CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:40:43.812735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:40:43.812735Z digest=sha256:ed16179f5faf73cd05560c5897124d30c86e30cbd3696f368f0cf46a88bcb418

Observation dbf31a79-d4f5-4590-a3bf-91095cc033ad · inbound

CaptionQA: Is Your Caption as Useful as the Image Itself? cites this paper.

CaptionQA: Is Your Caption as Useful as the Image Itself? CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:44:07.225915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T05:42:57.288054Z digest=sha256:463cbac37180ad16f532e9da597fb4f3f8728625a7b2bfdb04e94a820eb62bec

Observation 574ce0d5-7389-4f12-ae23-a6af18a398a0 · inbound

Vision-aligned Latent Reasoning for Multi-modal Large Language Model cites this paper.

Vision-aligned Latent Reasoning for Multi-modal Large Language Model CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:50:44.238024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T07:48:23.272002Z digest=sha256:3f0d787236441a90d11569041de6a3a438133a8396911a5871fef823bd422ab0

Observation c230eead-ff46-483a-99ae-30ad6364a4d3 · inbound

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation cites this paper.

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:07.552941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T14:10:28.941565Z digest=sha256:d4fc0a79815bc7b70d0ac3436ce16b0550ce2c7789f86e7e8c7eaffad7d4bed6

Observation 9f292d24-b66f-4b10-95f7-9c31d2ab9caa · inbound

BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning cites this paper.

BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:15:58.030863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T01:52:27.764121Z digest=sha256:38b9648ead2c41501e0efd5c5ee69465542740a148b72e29f7419d798bd4e0a7

Observation 21347e65-6998-4b47-b65a-2b7dec7cc69e · inbound

Context-Aware Spear Phishing: Generative AI-Enabled Attacks Against Individuals via Public Social Media Data cites this paper.

Context-Aware Spear Phishing: Generative AI-Enabled Attacks Against Individuals via Public Social Media Data CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:52:05.356803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:48:46.380615Z digest=sha256:41740d1b4b5b2ac8761a9134734f87281ab866686eeba4493c0830d64ccab42d

Observation d10a42fa-dc51-4dee-af00-fd5b4db9b0f0 · inbound

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models cites this paper.

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:06:59.633495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T01:30:36.491137Z digest=sha256:403a068ba754c63bb78d475b3ca7869829e7a3e41ce52d886b90e60cb1e26c02

Observation fee9f304-a5db-4605-aa0a-b1b01eb3acef · inbound

Robust Onion: Peeling Open Vocab Object Detectors Under Noise cites this paper.

Robust Onion: Peeling Open Vocab Object Detectors Under Noise CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:59:53.204581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T05:33:25.373918Z digest=sha256:57144662aa3fbe2af5c69a7ce5acf3576a0f1969b58054d7d16548fdb8c6266c

Observation 05c19696-2c8d-430e-8cec-d4eb7a2b22f7 · inbound

Robust Onion: Peeling Open Vocab Object Detectors Under Noise cites this paper.

Robust Onion: Peeling Open Vocab Object Detectors Under Noise CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:15:50.059824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T00:42:36.605869Z digest=sha256:9b735fc83115918c60142b592e5508b1594e50ea9e233ef0960c0ff612dc6297

Observation 4053493b-0866-4e49-b640-edcdf83ca423 · inbound

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions cites this paper.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.834315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.834315Z digest=sha256:2f8cae8bb5851a02983edae13353e2ac02058f2cc8227d97844e08a4d392b3fa

Observation e6977576-efe3-4033-b901-1ae2ef49a408 · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.033889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.033889Z digest=sha256:5f2718b8d590f452815776b82a38410f3a0fe07b8edff6186f92f1c9d92a1ea7