Pith. sign in

Paper Citation Record · LEDGER

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization

As of 15 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 1 inbound Pith citation observation for arXiv:2507.01492.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01492 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:53:21.073886Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T01:05:26.926438Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-08T01:05:27.182138Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9fe64d67-b6fd-4161-9c67-9f5f0b0eb188 · outbound

This paper cites Qwen2.5-VL Technical Report.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:19.709202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:19.709202Z digest=sha256:637e885e51c03e536e7ff1559b26367c800cb4fffaec90a05d1d558a2b416271

Observation 2e94a8c9-1420-4ad4-97ce-400db8aba0f9 · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:19.764060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:19.764060Z digest=sha256:625490a05ab660aa77df1a1416b7ed1c089c41bc883583032a647668b0926118

Observation c4068880-fffb-454c-9764-ff3ea91bdf1b · outbound

This paper cites Personal- 4 ized video summarization by multimodal video understand- ing.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Personal- 4 ized video summarization by multimodal video understand- ing

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:22.379581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:53:19.833575Z digest=sha256:c22bd38b9dbda9f2a8a03c7b76f053432f5d098df7bbb62cdfa851e988715fa0

Observation 740fe1f0-4bfb-4316-960a-00379e54b628 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions, 2024.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Sharegpt4video: Improving video understanding and generation with better captions, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:22.208046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:53:19.908823Z digest=sha256:97db2f1a623164071298a5d5df3bf6933c35dc349d8cc90a28f308c31fe8eb09

Observation 8f8b76c6-569e-4441-a756-8921c0ef814e · outbound

This paper cites Versavid-r1: A versatile video understanding and reasoning model from question answering to captioning tasks.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Versavid-r1: A versatile video understanding and reasoning model from question answering to captioning tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:19.971746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:19.971746Z digest=sha256:c1408b7185c6b18d5cebb1579573b2af8fab986460522119a55a093b8d99e333

Observation 42df691b-53f9-4751-a864-fee302343cbc · outbound

This paper cites Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.012906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.012906Z digest=sha256:1e41303dee95f6d788b7f39cc5842abf941e8846540035e61fa1d939013fd8ac

Observation 81842245-3584-42af-815f-d44d8e3626c9 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.062315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.062315Z digest=sha256:f3b478b92768da5385370c3088ae087f24a279243dbdc6e4dd8f385eb31646ae

Observation 632099fc-9a4c-47e4-b513-325dddaaf3e7 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.138597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.138597Z digest=sha256:b173997e3158cb3e44508e146575fc5e31e54704c94f3eb615a1620dd4cc8cae

Observation 5e285e60-b5a0-4e43-b69b-06dd47471ff1 · outbound

This paper cites Describe Anything: Detailed Localized Image and Video Captioning.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Describe Anything: Detailed Localized Image and Video Captioning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.191292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.191292Z digest=sha256:2a7c78438bcdbbce644d1bcee5ce81e3008a762226cbff123935b4d9b86052c2

Observation aa56aee0-a092-4d00-bf56-8ac3e8d230f4 · outbound

This paper cites Inference-time scaling for generalist reward modeling.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Inference-time scaling for generalist reward modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.256177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.256177Z digest=sha256:e6d700f5887bfe22972ddddc7e6f8ed2b1fbe8fc81fca2f89b1e7ad517767343

Observation 5a27968d-73f0-4706-a707-88aff6f7073a · outbound

This paper cites Videocap-r1: Enhancing mllms for video captioning via structured thinking, 2025.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Videocap-r1: Enhancing mllms for video captioning via structured thinking, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:22.081458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:53:20.295026Z digest=sha256:0399aeb9057b1052aace33c877a1fdcb959668b41709c06dcf65104c04d042f4

Observation 25fe8ebe-aa16-489f-aea8-7827e1cd60b3 · outbound

This paper cites ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.338147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.338147Z digest=sha256:46ec5c63aba0c0f9e817afcbf17ff2ed865df995aaaaf5df4b5162505d19d083

Observation 8a90b50f-be42-4b50-b8ce-caff318d76b8 · outbound

This paper cites Cockatiel: Ensembling synthetic and hu- man preferenced training for detailed video caption, 2025.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Cockatiel: Ensembling synthetic and hu- man preferenced training for detailed video caption, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:21.865289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:53:20.381184Z digest=sha256:db32b3b779e21960fd5943e41bf05636404fb6b7dd104080a24bb8707d14ddc3

Observation 9db24de2-15f7-4aee-a32a-5d542f5de9c8 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Direct preference optimization: Your language model is secretly a reward model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:21.766181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:53:20.441583Z digest=sha256:44b655a7b557be192b4681e8745e7e07d78ac698b303daa560b9884572662c9f

Observation e7e616bd-45a4-4437-83a7-612507711469 · outbound

This paper cites Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.489681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.489681Z digest=sha256:2ac047ecc9c2871d4c8a3b6414ac44a339f83167105b2ba1ca06982cef9901a7

Observation 0daa0044-fe45-4da8-873d-8fa6a4394704 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.539317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.539317Z digest=sha256:2c815918c742f67241a8d071d3cfd7468e718db1f6ed6a41dd3b3408d8447c70

Observation 467c9247-8a1a-4fd7-953e-12fe389bb211 · outbound

This paper cites Progress-aware video frame captioning.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Progress-aware video frame captioning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:21.674173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:53:20.617223Z digest=sha256:9c00e9c5ecdb54fb992f289e4aa5301e91e4f1480b32b96a12dc617d33c8180b

Observation b44abbe4-452e-4abe-804d-ad4bb0f2d803 · outbound

This paper cites VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.756229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.756229Z digest=sha256:56773841b4f9a6322d0d58a950e49be5e2393f5694cb2eb17e071dee03eb3dc4

Observation 9d1e3cba-0d00-4e50-a1e7-d1699e39c8d3 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:21.033536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:21.033536Z digest=sha256:706f5a47fe45ac1ec791fc2d0092ac7b8f09322ba3860038f95d89ea5e7eb535

Observation f140f3be-3eb9-4c52-a2d4-ff48d67292d1 · outbound

This paper cites Llamafac- tory: Unified efficient fine-tuning of 100+ language mod- els.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Llamafac- tory: Unified efficient fine-tuning of 100+ language mod- els

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:21.559531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:53:21.073886Z digest=sha256:6cd1f25eebc578fb754194152c99519f0174d0c9c18e136a5d85392eaf443268

Pith citing papers

Observation 4d026c89-1521-4c95-a135-b6cd914c956b · inbound

CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generation cites this paper.

CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generation AVC-DPO: Aligned Video Captioning via Direct Preference Optimization

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-08-08T01:05:27.188251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T01:05:26.926438Z digest=sha256:fe7de60b120eed0d88fccd6a1ea05a502c80917f4aaf8cbbf5d4385a6aaf98c0