Pith. sign in

Paper Citation Record · LEDGER

CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2503.12329.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.12329 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:35.033889Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:59:53.203155Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0c5da5d0-b350-4d61-90a5-265e4b549417 · inbound

Towards Scalable SOAP Note Generation: A Weakly Supervised Multimodal Framework cites this paper.

Towards Scalable SOAP Note Generation: A Weakly Supervised Multimodal Framework CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:35:41.711143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:35:41.711143Z digest=sha256:fa49b5d37276c38b4c05753f89d65b2c62dc4615645634c2aa46d3e5a9cffb41

Observation d7974b22-fb85-43b9-b84e-fead96685d62 · inbound

Skin-SOAP: A Weakly Supervised Framework for Generating Structured SOAP Notes cites this paper.

Skin-SOAP: A Weakly Supervised Framework for Generating Structured SOAP Notes CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:40:43.812735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:40:43.812735Z digest=sha256:ed16179f5faf73cd05560c5897124d30c86e30cbd3696f368f0cf46a88bcb418

Observation dbf31a79-d4f5-4590-a3bf-91095cc033ad · inbound

CaptionQA: Is Your Caption as Useful as the Image Itself? cites this paper.

CaptionQA: Is Your Caption as Useful as the Image Itself? CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:44:07.225915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T05:42:57.288054Z digest=sha256:7a8373b883bbc3ab916ab64e7a56da96ca101d349ee239e9ecb2a88057ec7c5b

Observation 574ce0d5-7389-4f12-ae23-a6af18a398a0 · inbound

Vision-aligned Latent Reasoning for Multi-modal Large Language Model cites this paper.

Vision-aligned Latent Reasoning for Multi-modal Large Language Model CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:50:44.238024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T07:48:23.272002Z digest=sha256:d8e290f002c76b471235af0fa28d84418b389d2d3cc5875369a771410d150784

Observation c230eead-ff46-483a-99ae-30ad6364a4d3 · inbound

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation cites this paper.

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:07.552941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T14:10:28.941565Z digest=sha256:0bbf00bc5ed703d3f306d677396d651d2b429016af110393005f0578d855a5c2

Observation 9f292d24-b66f-4b10-95f7-9c31d2ab9caa · inbound

BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning cites this paper.

BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:15:58.030863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-11T01:52:27.764121Z digest=sha256:9c0aaad8dbb22d739705357e370fb8442bdf16153ec12060f49c9602b89c663b

Observation 21347e65-6998-4b47-b65a-2b7dec7cc69e · inbound

Context-Aware Spear Phishing: Generative AI-Enabled Attacks Against Individuals via Public Social Media Data cites this paper.

Context-Aware Spear Phishing: Generative AI-Enabled Attacks Against Individuals via Public Social Media Data CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:52:05.356803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T01:48:46.380615Z digest=sha256:0fe9b615b5159b7f767e9543b41fc4adb335a065d70b703f31f6d6c513aa8fc4

Observation d10a42fa-dc51-4dee-af00-fd5b4db9b0f0 · inbound

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models cites this paper.

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:06:59.633495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T01:30:36.491137Z digest=sha256:cdfd6e7d79392bfa5558730cd351c92055185f64b98b9520e359c033771d6d5a

Observation fee9f304-a5db-4605-aa0a-b1b01eb3acef · inbound

Robust Onion: Peeling Open Vocab Object Detectors Under Noise cites this paper.

Robust Onion: Peeling Open Vocab Object Detectors Under Noise CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:59:53.204581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T05:33:25.373918Z digest=sha256:e93d3b61dc6f4da88fd566944378c36a5facd273ed9d0ebe675353bee74f452d

Observation 05c19696-2c8d-430e-8cec-d4eb7a2b22f7 · inbound

Robust Onion: Peeling Open Vocab Object Detectors Under Noise cites this paper.

Robust Onion: Peeling Open Vocab Object Detectors Under Noise CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:15:50.059824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T00:42:36.605869Z digest=sha256:95217a7ff9a46f6ce45e147ae7253d3f3f8def42d8ec9f39558b0a6ba01863d0

Observation 4053493b-0866-4e49-b640-edcdf83ca423 · inbound

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions cites this paper.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.834315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.834315Z digest=sha256:7eac93d752f3891ed021a8ee97cbcc9ceedcad4d56d21c93a6daec232d300902

Observation e6977576-efe3-4033-b901-1ae2ef49a408 · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.033889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.033889Z digest=sha256:19699d2aaa236b63df9eb373c51e1a639e8e4470603344b13be4e3bfa9dcc926