Pith. sign in

Paper Citation Record · LEDGER

From Pixels to Prose: A Large Dataset of Dense Image Captions

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2406.10328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.10328 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:41:44.532731Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:23:28.438005Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d15b7e7e-f44d-4dd6-9422-b149b8c24e0a · inbound

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation cites this paper.

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:09:16.246786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:09:16.001309Z digest=sha256:de6e2d38fbb5a21e4e54f48896b6211ef44d9df17fd5ab24509990471d4eeb64

Observation fdf7a16c-0e59-4fc1-91d5-130f6ccbcf13 · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:09:25.282416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:5d14d126c12ec2f6339ea297f482fbdfa8e870d5a679be46b0061d0d7a7e54ff

Observation 0d768d3b-8d18-42c7-bbc8-72919175e8b8 · inbound

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation cites this paper.

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T11:41:44.532731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:41:44.532731Z digest=sha256:ceabd4d2de2c2385651f376678a9929594ce09dc1b304a19073ec15c3f9606db

Observation 3e95b839-3773-4b95-99c6-bf0a01e15116 · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:52:01.832523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:1641eb261b12c226804fdc47c8bee6033efb82921cd96b42893ca72284157c32

Observation 8f4a2872-484e-4da9-85af-e0d56c56e405 · inbound

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions cites this paper.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:56.862819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:56.862819Z digest=sha256:9cac112f2b2affdfc1af13b6cf1623c83300330016e85b488148e92df443cec6

Observation 2cc224ce-97fb-4208-a204-b9c5ec103c97 · inbound

Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal Transport cites this paper.

Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal Transport From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.258130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.258130Z digest=sha256:dfc7b220364005ef63445ea343f436e54c52f027b700747f777504a977446ff4

Observation 4aaaa547-1faf-4365-b5ba-e27ee147e874 · inbound

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings cites this paper.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.893267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.893267Z digest=sha256:ae2fe6b2ce316ceef36d0c1097db6ee5ea6eccaf34933819b2b0764bc69f96fc

Observation becd85ce-93b1-4790-9592-8da2c45b33fb · inbound

MS-DPPs: Multi-Source Determinantal Point Processes for Contextual Diversity Refinement of Composite Attributes in Text to Image Retrieval cites this paper.

MS-DPPs: Multi-Source Determinantal Point Processes for Contextual Diversity Refinement of Composite Attributes in Text to Image Retrieval From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:06:03.980611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:06:03.980611Z digest=sha256:fd6d481060d8995731f77f5b251b6b52e8cc5e38b4afec823eca11360b1502bd

Observation ce87387f-6774-4174-adf5-76bfa3e64914 · inbound

ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation cites this paper.

ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T05:59:15.603444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:59:15.603444Z digest=sha256:7dcf34c823412295fefc3aaadeb4c24f56156d724f19ced004c0da0bb7c0e9c1

Observation 93220cf9-883b-4f79-8d90-b1cfa87f3d93 · inbound

Transition Models: Rethinking the Generative Learning Objective cites this paper.

Transition Models: Rethinking the Generative Learning Objective From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T10:19:54.462627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:19:54.462627Z digest=sha256:5d551206f105375ca061b4efa9344559962ecddad7769056ddbe69451b186888

Observation a40919a6-6724-4fa7-b093-3f5f9d14a557 · inbound

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark cites this paper.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.803853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.803853Z digest=sha256:4aa5b0635eca543f5b09970f65eee32928e64dcf6d4f054b278fe09efbfcd830

Observation 1e3a1767-f545-4f45-b653-e5e7622b1d96 · inbound

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens cites this paper.

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:08:15.863682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T12:04:19.761430Z digest=sha256:cc89a44b4f682fd81d86ff11cfda364a8845c5e3f4be8ee2e14612f267a5a68b

Observation 6fb820e6-d82a-455e-ac0b-fcf50011016b · inbound

BEiTScore: Reference-free Image Captioning Evaluation with an Efficient Cross-Encoder Model cites this paper.

BEiTScore: Reference-free Image Captioning Evaluation with an Efficient Cross-Encoder Model From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:04:45.988241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T09:01:24.453821Z digest=sha256:9e71c51ad7e7b9d6e6e575563d7efe93f7adb155fbb9a9ce1836e37d30d17c7c

Observation 9416758b-e8e1-4ecd-9f1b-7af77abf22f4 · inbound

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning cites this paper.

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.439444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T13:13:57.599970Z digest=sha256:7122b48863e38c039560f7c1f57428e6dc668eca3d8fdb8035a09b58eabf422d

Observation f2f6443d-f7f4-4dff-bcb1-acd6bcf3d080 · inbound

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget cites this paper.

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-02T06:14:02.725335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:14:02.725335Z digest=sha256:567646a5bc2d229ab94e5f6f8c641ca3f461b623c82c8519fde979f99bc6b9ad

Observation 99b7a9c9-e0ce-4fd2-9a40-4f126a3beb0d · inbound

MIDAL: A Dataset of Math Image Descriptions for Accessible Learning cites this paper.

MIDAL: A Dataset of Math Image Descriptions for Accessible Learning From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T00:49:46.472687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:49:46.472687Z digest=sha256:f828fa67367ba331a864432aa7addea4d32639b8a9abf135a6ad2eb691f6c98a