Pith. sign in

Paper Citation Record · LEDGER

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

As of 10 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2608.02589.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02589 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:09:52.617504Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1d799694-479a-44de-ac2e-97c61fc1fa56 · outbound

This paper cites 19 Table 4: Normalized understanding and generation benchmark scores.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation 19 Table 4: Normalized understanding and generation benchmark scores

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:52.617504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:52.617504Z digest=sha256:bae1a7bc37f37511f95150578f6c1c69e1e985ed3540000195cc052bfe99cd45

Observation 3935fad3-6107-46ce-8ffa-4315b12c3992 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:49.214324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:49.214324Z digest=sha256:2398530f787f9cde4a542c93dc57c16a7ad0c67f9ca1bf42521e204f7dbb853d

Observation 3906d2c0-9003-4264-af0f-a124724ca9e9 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:49.326199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:49.326199Z digest=sha256:87892de532d121822be555623213e121bce449d7d5c83241db65a70e62cce48c

Observation 3405f8cb-2315-4bf4-a71e-fe7e5c686028 · outbound

This paper cites Benchmarking and Improving Detail Image Caption.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Benchmarking and Improving Detail Image Caption

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:49.411968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:49.411968Z digest=sha256:abf30019320d7acec62c8632614b4f58d213edf3591e9ca8ae4981d22c9dbf31

Observation a87f0171-fe34-42dc-86dc-ead1f635dfe0 · outbound

This paper cites FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:49.613004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:49.613004Z digest=sha256:a294573925c78feb886a564b960584055de23c9f6e68f4e05efa8db8671ed2a8

Observation 3fb70424-9216-48a8-b94d-89b097a08d86 · outbound

This paper cites DataComp-VLM: Improved Open Datasets for Vision-Language Models.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation DataComp-VLM: Improved Open Datasets for Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:49.764855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:49.764855Z digest=sha256:92cf1d0719fc34e51ee179c66708c7ea0dd92ef1683a1d519b95309733755997

Observation 6e22bf7f-8d30-4241-994b-3a824c549664 · outbound

This paper cites GAVEL: Grounded Caption Error Verification and Localization.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation GAVEL: Grounded Caption Error Verification and Localization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:49.932057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:49.932057Z digest=sha256:c24350e092e86831912d21d38db6b3eb94f49172a5ebcff4fa549cb71c72be72

Observation 12a90096-8617-4d26-b368-00fcd32d5fa1 · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:50.080348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:50.080348Z digest=sha256:f472f7f0bf2859c6c733f4945b56c2061b9b880e62cca92f4359b898fdc84db6

Observation 0db60075-53c5-4cfe-a953-90b48ac25c80 · outbound

This paper cites Clipscore: A reference-free evaluation metric for image captioning.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Clipscore: A reference-free evaluation metric for image captioning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:50.202425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:50.202425Z digest=sha256:61db0db663bcd02af4abfce6e8ab7800aba53027c833bc32727033bb04d15b4d

Observation 45891aaa-9f93-4b23-9477-eb715d4018a4 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:50.455520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:50.455520Z digest=sha256:014efe523e20b571fdff20cd33a707184c470609a98fec6eb4a8163430572e27

Observation c4b31fb8-47c9-4716-b54c-af61adc38e27 · outbound

This paper cites Qwen2.5-Coder Technical Report.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Qwen2.5-Coder Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:50.567390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:50.567390Z digest=sha256:c21f81faa9f673142eecdb35ad4bf0cdc313b65f28e614934ea3eb8b334135fe

Observation 893d933d-6f00-4da1-84f3-d36db0bd92c8 · outbound

This paper cites Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:50.687652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:50.687652Z digest=sha256:2cfb4d63dc5b18f8debe119ea287508e8fbaf9ec147fe75590a272fbf78cc0fe

Observation 7e788c91-99e1-4caf-aa69-bd9a5349ba1a · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:50.853167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:50.853167Z digest=sha256:7599b6e2e45d3c54fe8ad1f9b1a3ce646e92f7b006c286b220f23cb7a0aa51e7

Observation 1ed8f6ab-86f3-4701-abe3-ae7d6c2326eb · outbound

This paper cites Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and Generation.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:51.216578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:51.216578Z digest=sha256:68da865cfd2f000eaa94e7e540b4fa2e618e10dfb9bed4f7b917950957ecde4b

Observation de52a59e-038c-49af-a0ef-08c5fbf428d4 · outbound

This paper cites Accessed: 2026-07-26.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Accessed: 2026-07-26

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:51.341587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:51.341587Z digest=sha256:6e35fc08bc0ac2aa025d59b27bbaedf3f3414e8c1f6b218517058c69aaa096da

Observation 1cd060d6-39e7-47a5-a709-cf54e8de246e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:51.781179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:51.781179Z digest=sha256:bf15e4d253c07ef0861fd87d6c5918f7fb1d633a9d8eb56880178800c3fc28ea

Observation 935f03fd-bc32-4a4f-b872-69c45a516335 · outbound

This paper cites Grasp any region: Towards precise, contextual pixel understanding for multimodal llms.arXiv preprint arXiv:2510.18876, 2025a.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Grasp any region: Towards precise, contextual pixel understanding for multimodal llms.arXiv preprint arXiv:2510.18876, 2025a

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:51.857417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:51.857417Z digest=sha256:7a9e4e58b6a728016831747d3345e0244d6c535a77487b0d0cff856c57677cbe

Observation 2bef3bad-7c14-4c1b-b6d0-41b722d51922 · outbound

This paper cites VGR: Visual Grounded Reasoning.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation VGR: Visual Grounded Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:51.972270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:51.972270Z digest=sha256:4dce907b4604f5707075e88b40df1ccdab89d07faa503a15eb79f6fedbdc5e59

Observation 676ce45b-8c4e-4278-b58b-8e8fc6186f8d · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:52.095666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:52.095666Z digest=sha256:1692dceaf204f7940a22fddd2ca38cd086d6f1e450577d0fa8555b83a4e8c4a9

Observation ab956a87-61b2-4c77-b0e1-79f09a48c5f6 · outbound

This paper cites Qwen-Image Technical Report.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Qwen-Image Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:52.218566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:52.218566Z digest=sha256:99db7f6d2a320f5bcf1d9aaf532725385313c025be02310d2b20813b0461c567

Observation 225c35c2-f0cb-40f4-9ff8-f32086799fb2 · outbound

This paper cites Qwen3 Technical Report.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Qwen3 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:52.342656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:52.342656Z digest=sha256:04e3c689c845b8e6b1589fe27603aa448394b707ebaa801a8c333c28aae8b54e

Observation 89035bb5-ecf1-4149-8a0f-3d30bd5422da · outbound

This paper cites CaptionQA: Is Your Caption as Useful as the Image Itself?.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation CaptionQA: Is Your Caption as Useful as the Image Itself?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:52.429597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:52.429597Z digest=sha256:f21347b2eec8b8bea10c20ee250691f3e3a592383ccbf4a0134dd99c113cc194

Observation d98d1b9c-efa6-44e1-beb6-f1e6ceb109eb · outbound

This paper cites Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:52.487467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:52.487467Z digest=sha256:39a05c0b7209131c659c29ef3ca90860844508c171867980aa35efb542bfd610

Observation 3a734a36-931b-4707-b5bf-109b8efadd8c · outbound

This paper cites Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:52.550837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:52.550837Z digest=sha256:1b52faee71ccdc898933ef7c28c46839de7157b2d4fcde1d64826e577a65703d

Observation 6acf51c1-8315-4717-9c22-00ed4c529ebe · outbound

This paper cites Aloha: A new measure for hallucination in captioning models.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Aloha: A new measure for hallucination in captioning models

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:51.455472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:51.455472Z digest=sha256:4d9ae478eba007b51aa9968788caa182da5119d41d9572248b2d4f440cecfc6b

Observation cdc5b4cf-ae0f-4aa0-a13a-72f89c33a572 · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:50.353673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:50.353673Z digest=sha256:fa34f10bc0f47cf869f73ddeb090e24af60d2ff3595bef4bca39428d652e48e9

Observation 53469e09-0e28-4021-aacc-7be29cebff5f · outbound

This paper cites Caparena: Benchmarking and analyzing detailed image captioning in the llm era.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Caparena: Benchmarking and analyzing detailed image captioning in the llm era

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:49.263896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:49.263896Z digest=sha256:e0195c21c4f0a2494a2359296bc633c25c17ea7b8b9553575d281b4f025320be

Observation 446143d6-f499-4551-b330-86af913451f2 · outbound

This paper cites LongCat-Image Technical Report.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation LongCat-Image Technical Report

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:51.542639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:51.542639Z digest=sha256:eb6302bb70f2990b38b1a070c85ac524127e8e4eab68e9bcf075ca825aa892b3

Observation 36e6ca72-5551-497d-9d24-a924091063ad · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:49.012960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:49.012960Z digest=sha256:3f50f5f2dec8debdaa7b4585ce67bf4ecc110a8dfc1fb8452646598c2bb03e3e

Observation 4c72255b-dd99-4530-908c-e847c85a552b · outbound

This paper cites Qwen3-VL Technical Report.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Qwen3-VL Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:49.178595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:49.178595Z digest=sha256:494b00a8540035cef0b8def3d44cd2cf69de960e919eb8c1581a5ddfe1fda653

Observation 6736a50e-288a-4389-bdf0-5818b74812f8 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:48.916885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:48.916885Z digest=sha256:1eafc3512c71ab2cb101a1d5ac1dbce176c285eedd31563f2f5ac1e328485f54

Observation 47486a42-f38a-45fc-9a4b-b1aec395b30b · outbound

This paper cites Qwen Technical Report.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Qwen Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:49.068148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:49.068148Z digest=sha256:d55555397bd808893e2f57e1da296aad6befb617adf11d5fbc4bd42aec706861

Observation 4e90874d-8934-4c6e-ab7f-c67569e93003 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Evaluating object hallucination in large vision-language models

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:51.136197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:51.136197Z digest=sha256:8291e61d5d3c48060df861754ab8152793b16ae69a5c0bd2177cfe5bcf289f9b

Pith citing papers

No inbound Pith citation observations are available.