Pith. sign in

Paper Citation Record · LEDGER

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization

As of 7 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2507.01492.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01492 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:53:21.073886Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9fe64d67-b6fd-4161-9c67-9f5f0b0eb188 · outbound

This paper cites Qwen2.5-VL Technical Report.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:19.709202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:19.709202Z digest=sha256:8749013619ad2100518200d939394d6355ce2588bb03004bcbe2414a6286a718

Observation 2e94a8c9-1420-4ad4-97ce-400db8aba0f9 · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:19.764060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:19.764060Z digest=sha256:3baa42ef7b8ed4b1732e4551de042f0796f49723cdde90c9c2c22dff739f8d86

Observation c4068880-fffb-454c-9764-ff3ea91bdf1b · outbound

This paper cites Personal- 4 ized video summarization by multimodal video understand- ing.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Personal- 4 ized video summarization by multimodal video understand- ing

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:22.379581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:19.833575Z digest=sha256:9badc3e2fc4ad072362b3a783d38934aa4a712af268ecf2b61381fd514dda789

Observation 740fe1f0-4bfb-4316-960a-00379e54b628 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions, 2024.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Sharegpt4video: Improving video understanding and generation with better captions, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:22.208046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:19.908823Z digest=sha256:3189b539415b8b3b4375666313e0241e427d09202bf26b94c52b46a4c38a603a

Observation 8f8b76c6-569e-4441-a756-8921c0ef814e · outbound

This paper cites Versavid-r1: A versatile video understanding and reasoning model from question answering to captioning tasks.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Versavid-r1: A versatile video understanding and reasoning model from question answering to captioning tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:19.971746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:19.971746Z digest=sha256:35307b020030afc764f6c15015a1e72527c8ece56ee12eea472860eda3d6d8a1

Observation 42df691b-53f9-4751-a864-fee302343cbc · outbound

This paper cites Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.012906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.012906Z digest=sha256:b7c84d4610b9aa32945732111e9e3570b413c90b832e3e3db53d30003faac2e1

Observation 81842245-3584-42af-815f-d44d8e3626c9 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.062315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.062315Z digest=sha256:fff221c0b85fd439e860c48f4bbc0d805c7bf6cc74a3ec32f6298f2feb47fedd

Observation 632099fc-9a4c-47e4-b513-325dddaaf3e7 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.138597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.138597Z digest=sha256:22f8e5297446f34b9684b450ce8035a68eed6a5fbc9c6ccefdf33fb6188a575b

Observation 5e285e60-b5a0-4e43-b69b-06dd47471ff1 · outbound

This paper cites Describe Anything: Detailed Localized Image and Video Captioning.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Describe Anything: Detailed Localized Image and Video Captioning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.191292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.191292Z digest=sha256:3f7c9be8741b2384ec1e59134c3afa5e7300f924e1b74051950379e34c3d4023

Observation aa56aee0-a092-4d00-bf56-8ac3e8d230f4 · outbound

This paper cites Inference-time scaling for generalist reward modeling.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Inference-time scaling for generalist reward modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.256177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.256177Z digest=sha256:911296f44a7049282981bbab786da4523890301b0159d01f5c9fb897894e3cf4

Observation 5a27968d-73f0-4706-a707-88aff6f7073a · outbound

This paper cites Videocap-r1: Enhancing mllms for video captioning via structured thinking, 2025.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Videocap-r1: Enhancing mllms for video captioning via structured thinking, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:22.081458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:20.295026Z digest=sha256:3026642c16d857f7932f70fa8baaa81492859a8b36663768fda16c79c1af3643

Observation 25fe8ebe-aa16-489f-aea8-7827e1cd60b3 · outbound

This paper cites ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.338147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.338147Z digest=sha256:db4fc0cc4c56e0f5089217d19b27a3405643d74b312b2a79d0e6f5e371eace91

Observation 8a90b50f-be42-4b50-b8ce-caff318d76b8 · outbound

This paper cites Cockatiel: Ensembling synthetic and hu- man preferenced training for detailed video caption, 2025.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Cockatiel: Ensembling synthetic and hu- man preferenced training for detailed video caption, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:21.865289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:20.381184Z digest=sha256:d10714e3937d65b0204a92f4fd8ba55ed2e6d9901002a0a95ed1610531359bf5

Observation 9db24de2-15f7-4aee-a32a-5d542f5de9c8 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Direct preference optimization: Your language model is secretly a reward model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:21.766181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:20.441583Z digest=sha256:563ab020fe49a476ba381f731cc591b555fb68e7c1b692a3a69de20f1dab5b93

Observation e7e616bd-45a4-4437-83a7-612507711469 · outbound

This paper cites Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.489681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.489681Z digest=sha256:0cf80dd1dc51d2b4e46ef0e771fc78340a8e1c8e9cf77b12380569aa88c72d76

Observation 0daa0044-fe45-4da8-873d-8fa6a4394704 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.539317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.539317Z digest=sha256:5a45dd982ef5259467ac2abd3a72d518668918b29a2d93dd4a7839d4474e79a8

Observation 467c9247-8a1a-4fd7-953e-12fe389bb211 · outbound

This paper cites Progress-aware video frame captioning.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Progress-aware video frame captioning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:21.674173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:20.617223Z digest=sha256:bb1089f7c0e41aedf0b76818c04bbb472e221e0bb3d11d68c63513cb9381fd2e

Observation b44abbe4-452e-4abe-804d-ad4bb0f2d803 · outbound

This paper cites VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.756229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.756229Z digest=sha256:eb3b0486d8e4d09d8c213de45680ad1351e1568f268fdd61847c14ba5a247d8f

Observation 9d1e3cba-0d00-4e50-a1e7-d1699e39c8d3 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:21.033536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:21.033536Z digest=sha256:0c00d3ef727ec1fce6b495c4132887a7f99d0837866a5355403ffc6e94b61605

Observation f140f3be-3eb9-4c52-a2d4-ff48d67292d1 · outbound

This paper cites Llamafac- tory: Unified efficient fine-tuning of 100+ language mod- els.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Llamafac- tory: Unified efficient fine-tuning of 100+ language mod- els

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:21.559531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:21.073886Z digest=sha256:746f0e310a17bbddd7e1168e33558c47285cbf0c71c601b067ad1c9c9600ff76

Pith citing papers

No inbound Pith citation observations are available.