Pith. sign in

Paper Citation Record · LEDGER

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions

As of 21 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2607.23235.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.23235 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T00:03:55.551725Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de9bb82a-8b21-466b-896b-75f2732a3774 · outbound

This paper cites Image captioning: Transform- ing objects into words.Advances in neural information processing systems, 32, 2019.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Image captioning: Transform- ing objects into words.Advances in neural information processing systems, 32, 2019

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:53.826066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:53.826066Z digest=sha256:932d5bc62e5bab03b5a38c2d234fa3da068461a20bb9a8585f1efb04be44e733

Observation ec69123c-8d40-43d7-8397-04aa21157e27 · outbound

This paper cites A comprehen- sive survey of deep learning for image captioning.ACM Computing Surveys (CsUR), 51(6): 1–36, 2019.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions A comprehen- sive survey of deep learning for image captioning.ACM Computing Surveys (CsUR), 51(6): 1–36, 2019

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:53.911123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:53.911123Z digest=sha256:36668b42160476e4463162405f8e22b04ece688989d2e5f7405e727eafe73134

Observation 1b3f5a94-b55a-4720-98f0-129e4b7a8758 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.018826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.018826Z digest=sha256:b2320ee92b708abcb19813635572682a634110b30a1246eba2884dcf2cc922d9

Observation c9cb565e-3ac3-4d85-b09e-16220d3f368f · outbound

This paper cites Surveying the landscape of image captioning evaluation: A comprehensive taxonomy and novel ensemble method.arXiv e-prints, pages arXiv–2408, 2024.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Surveying the landscape of image captioning evaluation: A comprehensive taxonomy and novel ensemble method.arXiv e-prints, pages arXiv–2408, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.132944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.132944Z digest=sha256:ccb63afe87dec804478ce6a089f4e439d55b9c66a5784b34a48eccc5971f79ef

Observation 1592e1d7-4a1c-4987-96b0-93cab0612cbe · outbound

This paper cites Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.285204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.285204Z digest=sha256:8fd6c941a445c0f7ed0904e5d1b1c5e2eba34473b90512782479f56c31b816df

Observation f67b4f26-d368-4036-8a39-a626c20d35d1 · outbound

This paper cites Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.390830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.390830Z digest=sha256:48fb42198e139200c11f17e51c709ab49c1a7573a4bd76d2b365b02b12d02619

Observation 4bacc46c-d093-4197-94d7-7e97b0c39f0e · outbound

This paper cites Show and tell: A neural image caption generator.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Show and tell: A neural image caption generator

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.555733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.555733Z digest=sha256:a7ca737ba714998a9afbf5b7e02c96dbd9a18d2ef061d045f624f9dcbdd2f0e7

Observation edc2adb1-0f1c-4218-a72e-a33ea1749f93 · outbound

This paper cites End-to-end transformer based model for image captioning.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions End-to-end transformer based model for image captioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.723648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.723648Z digest=sha256:a4869948b92fe803f1ee3361a5489093c6cd67f973e24c0ae01587f799fcc368

Observation 06e431af-6471-48f8-b589-2aa773f0d7a0 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.799819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.799819Z digest=sha256:2b33f251e0b4ad95faab171e471304727cc5d93150a9d3857f3c9f131630e444

Observation ca182a09-b1f8-4bc1-ba63-74f61bfd27d5 · outbound

This paper cites Cccaption: Dual-reward reinforcement learning for complete and correct image captioning.arXiv preprint arXiv:2602.21655, 2026.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Cccaption: Dual-reward reinforcement learning for complete and correct image captioning.arXiv preprint arXiv:2602.21655, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.803504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.803504Z digest=sha256:a80d2c399e77e89648b2a585901958749881c11afaa97816b7e48d3f77d60464

Observation e1c1aa8d-138b-46ab-9229-c9bfa5ca1e76 · outbound

This paper cites Learning transferable visual models from natural language supervision.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Learning transferable visual models from natural language supervision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.807372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.807372Z digest=sha256:3e8f4648eed29443d404c1bc2c9454e88562d4e71a35355b1371699b4369664e

Observation 1924cb55-3517-4cc0-93de-8c1ca9103f66 · outbound

This paper cites Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.811305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.811305Z digest=sha256:e6a00300029388fa3a2fe2905398586c891c272d2e7aa8b043e7c2b98c0e1e47

Observation 7d61b1ac-53bd-400b-8b04-be6af738f114 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.814700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.814700Z digest=sha256:548553b08eff192795780a723ce9bd594de17925076c5858c370d1b1a2d42919

Observation bbc2ebdd-6316-4446-bbdc-66241cca246c · outbound

This paper cites Qwen2.5-VL Technical Report.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Qwen2.5-VL Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.818105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.818105Z digest=sha256:e8d412645711c68ead6781e12b666ad27443f5bdb8b76ef5c09612e17c9747d5

Observation ab0b9b6a-df27-4efd-a05d-bba5c842678c · outbound

This paper cites Qwen3 Technical Report.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Qwen3 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.822192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.822192Z digest=sha256:40cfd82d807f44778219d80407922e5f7bcbdced337ce18fef479d53713819fe

Observation 6222624d-ce3d-41f3-853c-e876c5b631ce · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.825892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.825892Z digest=sha256:02bbe194e47ed5e820bd03eaf1d53030d7b6f84b7026cd5cb1b0b0cf0e137fb2

Observation 6e255aaf-3065-4dcc-95b7-10398ba50bf6 · outbound

This paper cites FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.830417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.830417Z digest=sha256:ce627d0af8040f7ea3b6b4740c46dbb5e44d97db21fa5a4682e85636cf8f8085

Observation 4053493b-0866-4e49-b640-edcdf83ca423 · outbound

This paper cites CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.834315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.834315Z digest=sha256:b43e2496975edc839a5e78cc84890af6cbdbe1e5b41301cee33e2311ee320b9c

Observation fcee4a54-64e3-4c9c-8660-633eec882e8c · outbound

This paper cites Prometheus-vision: Vision-language model as a judge for fine-grained evaluation.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Prometheus-vision: Vision-language model as a judge for fine-grained evaluation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.838788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.838788Z digest=sha256:546980b394c60ba360d56a925790d8c7fe9256e7e4a07d0e9fc8e2a3634c5931

Observation 824807e0-7ee9-4a89-995f-a8dc7096958a · outbound

This paper cites Caprl: Stimulating dense image caption capabilities via reinforcement learning.arXiv preprint arXiv:2509.22647, 2025.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Caprl: Stimulating dense image caption capabilities via reinforcement learning.arXiv preprint arXiv:2509.22647, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.842474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.842474Z digest=sha256:f048485053ece6182f5d0ee4e8e29cea22060fa6a429d330137cbc364ab524cd

Observation 5601d520-b86d-4d48-8d07-d7162a776818 · outbound

This paper cites Learning transferable visual models from natural language supervision.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Learning transferable visual models from natural language supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.846076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.846076Z digest=sha256:56a58d5826d622fbdc76a4a793328de8dc5c9abdcaf1a4f3157b063c85b0d823

Observation 0f825394-57c9-4c17-8cac-a67cbc569acf · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.849522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.849522Z digest=sha256:1589f58aef9c67dd40022a0c7bb59eb817afb2a15ce019b7dbd2eae51fb3536e

Observation ed5e168c-6f53-4ece-bef1-d084656422e4 · outbound

This paper cites Positive- augmented contrastive learning for image and video captioning evaluation.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Positive- augmented contrastive learning for image and video captioning evaluation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.853171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.853171Z digest=sha256:7b9546f6b3384c7b9a2060f537dc82950d41636c3a1bfc62f9e98c5fa8dfe83a

Observation d0c275ef-fe7a-421c-957a-f7f46703118c · outbound

This paper cites Evaluating Image Caption via Cycle-consistent Text-to-Image Generation.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Evaluating Image Caption via Cycle-consistent Text-to-Image Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.856531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.856531Z digest=sha256:0337706dea8195d2cd4bd4b62ebd3d7f6073fa11f9700f14d274278f6cfc4680

Observation 86224940-f088-419c-b747-c63cdb19c656 · outbound

This paper cites Im- age2text2image: A novel framework for label-free evaluation of image-to-text generation with text-to-image diffusion models.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Im- age2text2image: A novel framework for label-free evaluation of image-to-text generation with text-to-image diffusion models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.860280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.860280Z digest=sha256:514f99be3ac4ab1b65d7bf52b598e23d100938c89e2df70c4628c84174bc6f3e

Observation d760d816-c7ee-4a39-b9eb-39787985c465 · outbound

This paper cites Visual fact checker: Enabling high-fidelity detailed caption generation.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Visual fact checker: Enabling high-fidelity detailed caption generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.863659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.863659Z digest=sha256:c7239b45e856fa7e67ece878dc1065d1ae25f5eafbf2106832c4b7b4147a1f75

Observation b557d234-f48b-4829-8f35-10d73131098d · outbound

This paper cites TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.867139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.867139Z digest=sha256:5359a5060968bb32afb37476af1ca60a30aebc6d7b378dc30b06bd9bb4e4d837

Observation a4454d93-4993-4a2a-ad70-aca87f72da10 · outbound

This paper cites Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-to-image generation.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-to-image generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.874458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.874458Z digest=sha256:6867991ba1c06c4231b2f6e5d35c76cce7c3e8da4c20ac677846c5c195b4ff75

Observation a06b21d3-1251-4760-9ed6-0a91ad2de85a · outbound

This paper cites Microsoft coco: Common objects in context.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Microsoft coco: Common objects in context

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.936994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.936994Z digest=sha256:dd25e9f8ddcf8bd2ee34e9d9b352810e1be1be114deab0db05ba1591be20d283

Observation a2459e5d-a224-4bd9-8fbf-24c56d8981de · outbound

This paper cites Vqa: Visual question answering.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Vqa: Visual question answering

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.046200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.046200Z digest=sha256:49d4466d3e1aa180a0cd2a005fd218882aaf9eebbbf1ed9f3de3b06a3582b7de

Observation ea811959-84c1-4db9-9c08-3fa46ba4dd24 · outbound

This paper cites Beyond quantity: Distribution- aware labeling for visual grounding.arXiv preprint arXiv:2505.24372, 2025.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Beyond quantity: Distribution- aware labeling for visual grounding.arXiv preprint arXiv:2505.24372, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.197279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.197279Z digest=sha256:81231e85c957ce3db3c888baab28c4cf3c480b95ffd4660fd90490538ab203e9

Observation cdce35e3-d0d8-497e-811f-e4419971087c · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean Conference on Computer Vision, pages 216–233.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Mmbench: Is your multi-modal model an all-around player? InEuropean Conference on Computer Vision, pages 216–233

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.355646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.355646Z digest=sha256:345a5dd7a4d0deeaefddc9bc7e0f2807a116beead01d0349c7a6af604821cc44

Observation 6ad42cf7-0956-48a8-a8c7-940d7755b153 · outbound

This paper cites Are we on the right way for evaluating large vision- language models?Advances in Neural Information Processing Systems, 37:27056–27087, 2024.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Are we on the right way for evaluating large vision- language models?Advances in Neural Information Processing Systems, 37:27056–27087, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.461172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.461172Z digest=sha256:8ddb344b9586a7f1cf552a5787624d0108eb2f8b1d8e87da1fd28dfca57eea7f

Observation 4dab1268-75c8-4190-adef-9f4ae9020ae4 · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.464973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.464973Z digest=sha256:2238563bacb0d2a10d5c436b7a47120466d536361eee162775dcf8cde80592cf

Observation 9ab00de5-cb71-4dbc-8162-53162fca8cd4 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.469092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.469092Z digest=sha256:86ac129f8561dcf548d0ce898e78cd301f923f596d817e4086e100f52284c0e8

Observation c6675f83-2f1f-48b5-89c5-2e52fc679dcb · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.472736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.472736Z digest=sha256:b6389d238f39a3148cc7c9937a77e43290b6d547f3a57dc5cc962dc98f81dcc0

Observation 97a2ed12-8357-4d50-acde-0bb217817342 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InEuropean Conference on Computer Vision, pages 169–186.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InEuropean Conference on Computer Vision, pages 169–186

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.476457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.476457Z digest=sha256:93f4f374a4a82da0d9679e8aca6f727f0fb9a62fd93abd0ae394bf92350511f1

Observation 6e5f531b-3f8f-4612-9638-7a2ff6a6df15 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Measuring multimodal mathematical reasoning with math-vision dataset

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.480121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.480121Z digest=sha256:ce83daea0b97e30459e71f2a745f7221d53d09600393e85c4e03b6743d8f7767

Observation da2fef2f-9f97-468f-84a1-07e11fc165c8 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102, 2024.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.483851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.483851Z digest=sha256:8642110d044e6152ecbe2394c3b99852d8a5b812a50696701f0c40e12bf57204

Observation 636f817c-2c60-4728-be89-ab62bbac3110 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Ocr-vqa: Visual question answering by reading text in images

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.487616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.487616Z digest=sha256:526f740c8b32945a3b4b8aa4d94a2b54bd7f3fe9d0021964dfe3ad527eb3a8f1

Observation cdd9693b-fd8f-4670-be20-27f3e3eac1a1 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Referitgame: Referring to objects in photographs of natural scenes

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.491028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.491028Z digest=sha256:5500a26877b14cd556181932edff9b635540df42991800c9a433c2d4cc6ae018

Observation b6d4ddd4-45b7-4beb-a220-523c16d3a8bd · outbound

This paper cites Improved baselines with visual instruction tuning.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Improved baselines with visual instruction tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.494520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.494520Z digest=sha256:b20dedeea30a21f1e829c39c7b86da6880b340537ce9e637cbbd964d020728ac

Observation 61ac7eb2-943a-4d9b-acc2-967d7f0b56b4 · outbound

This paper cites Llama-3.2-11b-vision – multimodal large language model (text + image → text).

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Llama-3.2-11b-vision – multimodal large language model (text + image → text)

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.498441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.498441Z digest=sha256:9ae32d8e982f5b027deb129de0aaef8d97f89ee80bf660134a365878ef56808d

Observation a4e25302-67e0-48c7-ae84-8151916f4a1e · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.502062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.502062Z digest=sha256:ed469451682f332d3e686f87ecb62e4ca2f1b902e52333da5b5fe45a70433f79

Observation 6399c4b2-7c03-46d9-8ed9-d68d57647d12 · outbound

This paper cites Finite mixture models.Annual review of statistics and its application, 6(1):355–378, 2019.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Finite mixture models.Annual review of statistics and its application, 6(1):355–378, 2019

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.505878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.505878Z digest=sha256:e12579604213dcdf4ee247c1a5588b1b934d9926cbb671cb7dce570be8085371

Observation 1ebe9983-979c-4dad-8819-c3f9b26d089d · outbound

This paper cites A density-based algorithm for discovering clusters in large spatial databases with noise.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions A density-based algorithm for discovering clusters in large spatial databases with noise

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.509506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.509506Z digest=sha256:98520229589228c88ceaf059b840b323e9d620f11fbebaa5c2efc5fb96656de5

Observation 09792c1d-2976-4d88-a576-7955c28e6100 · outbound

This paper cites Qwen-Image Technical Report.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Qwen-Image Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.513125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.513125Z digest=sha256:7bb2947bb10a717ad40f5fd98fa2769b6db953ef5f089081e9fe2c28bde92c8c

Observation f010210b-df38-4485-9d34-cd89dcb7f9fc · outbound

This paper cites Unitbox: An advanced object detection network.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Unitbox: An advanced object detection network

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.517181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.517181Z digest=sha256:b9b467a152287d519007235574d74bfa075703c30a4aae1a5b6abe7c51e54a8d

Observation 62414ae0-c443-4f4c-869b-226b1ea7010e · outbound

This paper cites Opensearch-ai / ops-mm-embedding-v1-7b.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Opensearch-ai / ops-mm-embedding-v1-7b

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.520716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.520716Z digest=sha256:4ffca8f4a3b381cfc26970eb799f39ddd98007bdf51651d5e4145a660052f36b

Observation 2606616d-bfa6-47bc-ad81-aa284a0ae4df · outbound

This paper cites Gpt-5: A new era in language models.OpenAI Technical Report, 2025.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Gpt-5: A new era in language models.OpenAI Technical Report, 2025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.524213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.524213Z digest=sha256:06207e318eb6dc3ce5284ca4eed717696cab3bc5096aa71938493325dc232782

Observation 86dd6d9b-d379-489b-b898-047473b38ff3 · outbound

This paper cites Scaling Laws for Neural Language Models.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Scaling Laws for Neural Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.528338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.528338Z digest=sha256:cdc24ca86e086e52d46ecd09343c2bfcfd20c44007f25d3ef98ac474268c1160

Observation 5181199c-25e9-4ae8-9847-8d597757ba8d · outbound

This paper cites Prism: A framework for decoupling and assessing the capabilities of vlms.Advances in Neural Information Processing Systems, 37:111863–111898, 2024.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Prism: A framework for decoupling and assessing the capabilities of vlms.Advances in Neural Information Processing Systems, 37:111863–111898, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.531910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.531910Z digest=sha256:b0f54c290f6500d4c0574966bda31ec19c7fa76d1c1c8c03bfd079897f29ca24

Observation bc17697f-a386-4aa6-84c2-6440d632210d · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions High- resolution image synthesis with latent diffusion models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.535566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.535566Z digest=sha256:2d524faae7cc6605607d63f4affe4cb735ce9c5d80257b959a90ecef09ce5af6

Observation e7066cee-b6ef-45f8-8ae7-29fb0fed6d7a · outbound

This paper cites question.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions question

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.539199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.539199Z digest=sha256:c78354130e1a4584ae91ecd35048a401fc18838579df236b1c2a2e5212b30eb3

Observation 26925ec0-cdb4-40ee-af5f-f3595014d1df · outbound

This paper cites an unresolved cited work.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.543775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.543775Z digest=sha256:499fc73a93528fe37170e7c1cc82713f15555e036ab53f93b7ff972e7c2949d1

Observation 472851b9-bdb6-470e-9596-53bf21ca11e9 · outbound

This paper cites gold standard.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions gold standard

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.547632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.547632Z digest=sha256:04b1f319cdd30f9495e24ba45b9e5fa78f2a350daf0cb079f210560de128f54f

Observation 9fc606de-a5fc-4260-bbec-4aca9748345b · outbound

This paper cites Compares the performance of two judgers, Qwen2.5 (Qwen2.5-VL-3B) and Qwen3 (Qwen3-VL-8B).

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Compares the performance of two judgers, Qwen2.5 (Qwen2.5-VL-3B) and Qwen3 (Qwen3-VL-8B)

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:55.551725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:55.551725Z digest=sha256:a376b845ed074a5c11bfc4b57956c016776fc24b474461d82b484bb723b4e904

Pith citing papers

No inbound Pith citation observations are available.