Pith. sign in

Paper Citation Record · LEDGER

OmniCaptioner: One Captioner to Rule Them All

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2504.07089.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.07089 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:22:42.176796Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T03:17:51.543739Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cac583b4-7728-4211-be92-4a421430d8f9 · inbound

Chimera: Improving Generalist Model with Domain-Specific Experts cites this paper.

Chimera: Improving Generalist Model with Domain-Specific Experts OmniCaptioner: One Captioner to Rule Them All

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:47.621108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:47.621108Z digest=sha256:3bc33b6f1b998cd51c111633e95a6490e081193ab990bd023af8231e4ca4fb8f

Observation e6661b07-756e-4a3a-9b42-92a283aab301 · inbound

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs cites this paper.

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs OmniCaptioner: One Captioner to Rule Them All

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:58.716151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:58.716151Z digest=sha256:951d2be37682a8d8cb757e136364e0677670485e06b99f842891b83ddcd49b3e

Observation dfb51c60-7593-4a3c-a539-9c0f8e795978 · inbound

Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models cites this paper.

Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models OmniCaptioner: One Captioner to Rule Them All

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:21.482694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:21.482694Z digest=sha256:eeb708de321183db1347d588e7ce5c37dee1f0eb7878e6c0b74cebad1ecbbc06

Observation 0f97bf16-8931-4f66-9981-b06ac3b36d9e · inbound

Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling cites this paper.

Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling OmniCaptioner: One Captioner to Rule Them All

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T18:22:42.176796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:22:42.176796Z digest=sha256:0427b0f24f8c3da452249d1898a2d510e6375cb78c891e32fd75632d112a1085

Observation e91735c6-df67-4819-b970-c7a0857b77f5 · inbound

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning cites this paper.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning OmniCaptioner: One Captioner to Rule Them All

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:36.852405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:36.852405Z digest=sha256:433427b7f95a13a6b16851632d232e249a62a1e43c41c381a55b4bf5072b8b2e

Observation a7ed7b5c-c64b-4b29-89e0-41dc41e350d9 · inbound

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards cites this paper.

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards OmniCaptioner: One Captioner to Rule Them All

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:23.473947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:23.473947Z digest=sha256:c574de6ae15c5c2f578e62469b0f3602ff8d48e0d29205024ff2bb33bb022277

Observation c51d8ef9-8ad0-4d17-9ee5-a3d61fc27f41 · inbound

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer cites this paper.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer OmniCaptioner: One Captioner to Rule Them All

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.265735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T14:08:36.801359Z digest=sha256:97a53282baa9bbd4cea93b0706eee52fd0fd691ded34f7739290ccf65cd597f2

Observation c21bfeb1-2a69-48f2-8e4d-de3b6d031cfa · inbound

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer cites this paper.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer OmniCaptioner: One Captioner to Rule Them All

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.786122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.786122Z digest=sha256:9c74e8356593f3e6e7f4f23336409f95e84be4caf554f6049e956eacd7d84480

Observation 661e7ce4-ad9a-45d9-bd79-4b1ae5053281 · inbound

CodePercept: Code-Grounded Visual STEM Perception for MLLMs cites this paper.

CodePercept: Code-Grounded Visual STEM Perception for MLLMs OmniCaptioner: One Captioner to Rule Them All

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T23:22:13.847876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:22:13.847876Z digest=sha256:3ed59e0c42626b63a635df6465a2e4bde1a33d9ac6a61d5d888d4a80fb89d3fc

Observation 30f33aa0-b206-4d6c-93c1-2d1309a7affe · inbound

DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions cites this paper.

DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions OmniCaptioner: One Captioner to Rule Them All

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:10:52.847901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:39:29.592129Z digest=sha256:84c9c2589c6317f55f70f92a1995a0ef76a98f82450ca0398fe8ab96ae14bff7

Observation 7fbaec1f-7bdb-4f93-a552-67b33733d204 · inbound

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning cites this paper.

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning OmniCaptioner: One Captioner to Rule Them All

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:55:44.238614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T17:56:47.621050Z digest=sha256:7cb2beb1a912af22615c4fee80e0abdf7e81add50cc70bb642c7b67bd6a9b665

Observation f4e8a51d-2404-4008-af5e-42bb65278d08 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs OmniCaptioner: One Captioner to Rule Them All

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.021888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:93b96053c274945f15e4b67276e10c829d534ac35c5ebe67387d04310c01dd97

Observation 2f302196-1973-4509-88fc-adeaf37cf13d · inbound

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models cites this paper.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models OmniCaptioner: One Captioner to Rule Them All

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:17:51.572452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-11T03:11:49.713677Z digest=sha256:0095f63866b8a57c2bbd36a39e9896b3c3d759b5a870d820a248cab4e824b3e7

Observation 28cde4be-7dcf-4d3d-893c-76029b9f4b60 · inbound

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models cites this paper.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models OmniCaptioner: One Captioner to Rule Them All

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:1c25127ac8e6c50e203739bcff1a0859d5fdea5cf6a545f4c8b8aedc19cbce7c

Observation ead60c4d-46c1-4950-977f-efa478de054f · inbound

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward cites this paper.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward OmniCaptioner: One Captioner to Rule Them All

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.157674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.157674Z digest=sha256:252f7810083edcee6f84cd2e0ef9cde3d03222640974409023b9529b11661ee4