Pith. sign in

Paper Citation Record · LEDGER

AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2410.03051.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.03051 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:06:35.088592Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:54.928668Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 831294d4-84b3-4832-bc3a-b50709e32390 · inbound

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks cites this paper.

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:05:29.239902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T07:05:08.716223Z digest=sha256:e37edc17d9fb845dd59c773042b708b35b1e39da2d17b54d4507937288d859b3

Observation 68652cac-f1a8-4700-8de1-5f779442999b · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.152534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:22625d6e4c7f32fd62884354da5b6eeaaf296f16cb6f89bc084ace5ebb533dd1

Observation b5837546-afe1-4b76-bd4d-509e4deaf5f0 · inbound

VideoRoPE: What Makes for Good Video Rotary Position Embedding? cites this paper.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.088592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.088592Z digest=sha256:7ab4d5fe938098166d521302a0febb5850f946d1f558d7884cf67c25165547d2

Observation 2537b212-0d3b-463a-bec1-b35f1d15dcc5 · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:02:17.878492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:3bbf23a3fc84d33c538ce7f4b13c2ee4bee1e86c06a5eee10b8e2658f3f5ea0f

Observation caca77dd-0c7f-436f-a4a7-f8dbfe66e7db · inbound

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval cites this paper.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.686789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.686789Z digest=sha256:d290236572883983d91c7254d4b176e93b90c96a1f0c64befc9757d92f875e9e

Observation 286408a4-2bf5-439d-b8c2-c1f170a644df · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:59.661583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:02:59.661583Z digest=sha256:f021fcbbd7668ed721e0fa5b5cafa8034cb526dbd2a164704a19b40512150a96

Observation 14d775da-00a6-4a49-9a83-208061a1d3d8 · inbound

Vid-SME: Membership Inference Attacks against Large Video Understanding Models cites this paper.

Vid-SME: Membership Inference Attacks against Large Video Understanding Models AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:00.159271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:00.159271Z digest=sha256:52bcdba901086138c93b2715b23ef54821ed893618d376495d80af1e7bbc3261

Observation ea527a28-a167-4066-bada-0ccd3e18f566 · inbound

ARGUS: Hallucination and Omission Evaluation in Video-LLMs cites this paper.

ARGUS: Hallucination and Omission Evaluation in Video-LLMs AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:39.591343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:39.591343Z digest=sha256:65258dcc12d5341cb4e96a5a33525fb014ae41681b303e515120b4fdca3d281b

Observation 3bf9bb16-7765-40f5-b839-9587f6f7bfcb · inbound

ToSA: Token Merging with Spatial Awareness cites this paper.

ToSA: Token Merging with Spatial Awareness AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.947783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.947783Z digest=sha256:e96635bc2e9caaad844179d5eb468f566e98df0e7903117bb660935ef185be16

Observation 2e94a8c9-1420-4ad4-97ce-400db8aba0f9 · inbound

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization cites this paper.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:19.764060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:19.764060Z digest=sha256:9c89726d52bacfd437872d251476d6cc6397396305b9977b38f09f9c339fe1d5

Observation 7b21e852-8f3b-4f2d-9372-9a4dce0f8077 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.227081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.227081Z digest=sha256:43ddc95e9b626518ccafebe07e494fb2e86d047214477f093ed5994e3099d69c

Observation 50b612f2-6f4e-4f4d-b273-6ace6548816b · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:26.518534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:26.518534Z digest=sha256:bb9674e052ff642a79116a27afadb76a6b847f29928fb797a529f93c5c7f8b45

Observation 307b969c-1c54-4663-83d5-2209560690f9 · inbound

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models cites this paper.

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:26:27.013397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T18:25:21.621268Z digest=sha256:60341d4f219d0aa6837772df571c2e0b6f227c399bed8b010b75292a0872e74b

Observation cdf41a37-fada-4535-901d-9ea0ad180cda · inbound

ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling cites this paper.

ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:25.635849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T00:56:47.841355Z digest=sha256:989364862a5fd73b3750cf7b54813e227450c3129cab4a9fde1852a8d65662c1

Observation 9fdf966d-545e-41f9-a686-acde1a8d4700 · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.576239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:72efa971f327dc4fe45c6d9e593c685923333d2305a84be2634cc6fcd95adce1

Observation 1f19ee89-936f-4124-b487-3852d4baad26 · inbound

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation cites this paper.

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:07.544729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T14:10:28.941565Z digest=sha256:f270ae77bc2d72856e10c09f86fc5d75de6a967dd9c7e566e63797bb57ed0f5a

Observation 5715f666-a5d1-451f-928e-f5e25c060ee4 · inbound

Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation cites this paper.

Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:16.313840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T15:47:33.435653Z digest=sha256:0ae606b9298f454451f912739f3f911b45e75ed82323dce3eefb993886bc0138

Observation 9c1e0d0d-2b0d-4e9c-a200-53bc6e838f59 · inbound

Balancing Image Compression and Generation with Bootstrapped Tokenization cites this paper.

Balancing Image Compression and Generation with Bootstrapped Tokenization AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:55.406296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T03:07:33.054518Z digest=sha256:4d6fbb4901f6eefc5e1f7d23e2b4d70933db359a910df2c743a6736a7eb87d8a

Observation 6098b20b-c083-409d-ae59-491209b36e95 · inbound

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning cites this paper.

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:56:47.382456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T06:33:32.090913Z digest=sha256:5f4ac9e456e9e80bec741bdfe3f37ca82319ac12073e4e03d508a53b7069b66d

Observation 00fc8319-181c-4837-ad2f-97b1f9e2570d · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.575211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:c28e9e70e15006290db220b471f1326dea9e6e9274d7712f7b942eb5159860ab

Observation 7aff1fca-7e12-45ca-a8f0-9ba6eb33cb73 · inbound

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales cites this paper.

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:09:41.272203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T12:09:01.026544Z digest=sha256:012d9ea42f61f8de4f1014b0394e7540330b449667fc521a1e71760ae4a09802

Observation 48128038-df63-4ab8-b3bf-897687c62909 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.930565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:abae8b2b04dec54319581a05412b7c5125bd147669cf4da4f132aa177f3b45ad

Observation 4a2ab20c-32cd-4207-afa5-4298ace80d12 · inbound

PercepCap: Video Captioner with Structured Spatio-Temporal Perception cites this paper.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.141536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.141536Z digest=sha256:550ced918ac146b8b665ce6c29b2b828a48641880858d358f499c2274c729025

Observation 2f5f4e4f-4e40-44e3-8325-75d037f4915d · inbound

Visual Token Compression Enhances Robustness of MLLMs cites this paper.

Visual Token Compression Enhances Robustness of MLLMs AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T13:10:37.202730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:10:37.202730Z digest=sha256:c74799b6d2e00aacc8096790d44d6b840bf2054955dd1a274761939a75e67bb0