Pith. sign in

Paper Citation Record · LEDGER

AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2410.03051.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.03051 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:37.686789Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:54.928668Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 831294d4-84b3-4832-bc3a-b50709e32390 · inbound

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks cites this paper.

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:05:29.239902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T07:05:08.716223Z digest=sha256:79c8de9fe637585db8598ae823e846959900ef0672f83860f481f244bbd21a53

Observation 68652cac-f1a8-4700-8de1-5f779442999b · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.152534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:04d6d58cc1a881dcd899e560955a4255450deb60b808afa318b54cdc2498ff1c

Observation 2537b212-0d3b-463a-bec1-b35f1d15dcc5 · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:02:17.878492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:e78532723e522435ae003dd339d747b7c1051e32bd2a5f71b66b660922d76c79

Observation caca77dd-0c7f-436f-a4a7-f8dbfe66e7db · inbound

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval cites this paper.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:37.686789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:37.686789Z digest=sha256:c65103cec573fe019cc7327d535083a845d2f87fdfb5fd80a15ee96eaec719bd

Observation 286408a4-2bf5-439d-b8c2-c1f170a644df · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:59.661583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:02:59.661583Z digest=sha256:9fb69bb01ba0fdd5ab3c3136f4b27e33d20f1c2e1f21a501b76b925f58c308fe

Observation 14d775da-00a6-4a49-9a83-208061a1d3d8 · inbound

Vid-SME: Membership Inference Attacks against Large Video Understanding Models cites this paper.

Vid-SME: Membership Inference Attacks against Large Video Understanding Models AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:00.159271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:00.159271Z digest=sha256:52bcdba901086138c93b2715b23ef54821ed893618d376495d80af1e7bbc3261

Observation ea527a28-a167-4066-bada-0ccd3e18f566 · inbound

ARGUS: Hallucination and Omission Evaluation in Video-LLMs cites this paper.

ARGUS: Hallucination and Omission Evaluation in Video-LLMs AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:39.591343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:39.591343Z digest=sha256:9d78abcacd3ff5ce30e63389297cd41fad42185940d799c73868eeaf49efef7a

Observation 3bf9bb16-7765-40f5-b839-9587f6f7bfcb · inbound

ToSA: Token Merging with Spatial Awareness cites this paper.

ToSA: Token Merging with Spatial Awareness AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.947783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.947783Z digest=sha256:e96635bc2e9caaad844179d5eb468f566e98df0e7903117bb660935ef185be16

Observation 2e94a8c9-1420-4ad4-97ce-400db8aba0f9 · inbound

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization cites this paper.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:19.764060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:19.764060Z digest=sha256:3baa42ef7b8ed4b1732e4551de042f0796f49723cdde90c9c2c22dff739f8d86

Observation 7b21e852-8f3b-4f2d-9372-9a4dce0f8077 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.227081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.227081Z digest=sha256:43ddc95e9b626518ccafebe07e494fb2e86d047214477f093ed5994e3099d69c

Observation 50b612f2-6f4e-4f4d-b273-6ace6548816b · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:26.518534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:26.518534Z digest=sha256:bb9674e052ff642a79116a27afadb76a6b847f29928fb797a529f93c5c7f8b45

Observation 307b969c-1c54-4663-83d5-2209560690f9 · inbound

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models cites this paper.

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:26:27.013397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:25:21.621268Z digest=sha256:400e505405e70c89501c2fa43300e0574c2d40cc81a0e5b157d1b110a557eec0

Observation cdf41a37-fada-4535-901d-9ea0ad180cda · inbound

ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling cites this paper.

ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:25.635849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T00:56:47.841355Z digest=sha256:c494f57d807e359ac713a424b7cf97e6cad70a913827a5739ca87e0dff359f03

Observation 9fdf966d-545e-41f9-a686-acde1a8d4700 · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.576239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:0bc79a04c573246e6c3e481ddca8059cb4723b12b54da1aafcf8fc85c85bbac8

Observation 1f19ee89-936f-4124-b487-3852d4baad26 · inbound

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation cites this paper.

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:07.544729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T14:10:28.941565Z digest=sha256:d46a17384265c3b5d5b2c4744a6e418c066bd2fe9250da25e07ba43b8aca1e76

Observation 5715f666-a5d1-451f-928e-f5e25c060ee4 · inbound

Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation cites this paper.

Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:16.313840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T15:47:33.435653Z digest=sha256:17ca34643f0199f54c13832edacc249554bdd3b610cf1d8bc70c0ef803c34199

Observation 9c1e0d0d-2b0d-4e9c-a200-53bc6e838f59 · inbound

Balancing Image Compression and Generation with Bootstrapped Tokenization cites this paper.

Balancing Image Compression and Generation with Bootstrapped Tokenization AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:55.406296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T03:07:33.054518Z digest=sha256:962fc3f7af9de7814bed9a2aacf9df70c2a26a39d21a66451b9ba21612c19289

Observation 6098b20b-c083-409d-ae59-491209b36e95 · inbound

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning cites this paper.

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:56:47.382456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T06:33:32.090913Z digest=sha256:ccf6c7aec4fa45db02bd2ad2431a7c2662cb4025cff21d234149748494816e3c

Observation 00fc8319-181c-4837-ad2f-97b1f9e2570d · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.575211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:ac2964771fc2b647ebf67d1176fbf2aee0aa9c86359b680384a69a6b18898ed5

Observation 7aff1fca-7e12-45ca-a8f0-9ba6eb33cb73 · inbound

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales cites this paper.

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:09:41.272203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T12:09:01.026544Z digest=sha256:6ed2d9ec4062658b1b636e15beb87ac7d38a18fa2b325212297f65ef30f2ca7d

Observation 48128038-df63-4ab8-b3bf-897687c62909 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.930565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:db048504b9511d2c42bd5bb71a893e1877c091b63387660d7daab7effeb18f47

Observation 4a2ab20c-32cd-4207-afa5-4298ace80d12 · inbound

PercepCap: Video Captioner with Structured Spatio-Temporal Perception cites this paper.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.141536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.141536Z digest=sha256:550ced918ac146b8b665ce6c29b2b828a48641880858d358f499c2274c729025

Observation 2f5f4e4f-4e40-44e3-8325-75d037f4915d · inbound

Visual Token Compression Enhances Robustness of MLLMs cites this paper.

Visual Token Compression Enhances Robustness of MLLMs AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T13:10:37.202730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:10:37.202730Z digest=sha256:c74799b6d2e00aacc8096790d44d6b840bf2054955dd1a274761939a75e67bb0