Pith. sign in

Paper Citation Record · LEDGER

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models

As of 18 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 10 inbound Pith citation observations for arXiv:2504.13122.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.13122 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:43.580227Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:19:03.525559Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:27:56.144016Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved42
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa1715a4-5412-48fe-878e-a379bcbbc447 · outbound

This paper cites Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.327925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.327925Z digest=sha256:a79f0a86e94cfb50fbbe343f94c71f1a9950f1779780b53681ccb6edccd7b3ec

Observation a14fddc4-8cd3-4e66-8a28-aa9c270d0795 · outbound

This paper cites IS IN VIDEO?.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models IS IN VIDEO?

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:44.581214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:19:43.559546Z digest=sha256:0674c2ef78cf4fafff35b844245cf58d6f7751cbacb093db4b9c35b26411c164

Observation dc1855ff-4341-4f1e-b0c9-6337d2494712 · outbound

This paper cites The Llama 3 Herd of Models.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.356656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.356656Z digest=sha256:ce30024bdf9ff4cfeaf0e4fb8ee0c1ef73bc06e437ee71369c7722f919876e1c

Observation a414dfc4-0f7d-48a7-9136-632206189ff2 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models KTO: Model Alignment as Prospect Theoretic Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.362082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.362082Z digest=sha256:66acc0023a1d09cdb8abbf0ce91c4a3ddeddd9660fdfbec2e2ab2916b47be8a0

Observation 4da6417e-9ebc-4de9-a98d-121cf9a32f72 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.367110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.367110Z digest=sha256:0f72ed9a5f67973a383843a75ac17dc1886feddfce2504a0ca5465174d9bfbb1

Observation 0d678e26-6420-4724-95e6-d52c75fb19c9 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.373003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.373003Z digest=sha256:76c2916e8b5b205112a7de87f9d4a51807c990ea8d5e6a3ea7e768c0563e1bb2

Observation 6ea9a593-c438-455f-a9e1-6afe101ede42 · outbound

This paper cites Recent Advances in Multimodal Affective Computing: An NLP Perspective.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Recent Advances in Multimodal Affective Computing: An NLP Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.378645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.378645Z digest=sha256:3b28a48cc2bfe5bf544328616026ff47ea605c5eb202013c4de017a9c8d11b12

Observation 014862bd-4d94-40a8-8061-f7fba16d91bc · outbound

This paper cites Evidential Deep Partial Multi-View Classification With Discount Fusion.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Evidential Deep Partial Multi-View Classification With Discount Fusion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.384521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.384521Z digest=sha256:39d023ec97df092b3a6d3db6fa70b8e65089c7276e84d435848b4523aa547baf

Observation 4ff65b27-49f2-4ed6-a12f-31ac24b22024 · outbound

This paper cites is_in_video.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models is_in_video

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:44.507219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:19:43.580227Z digest=sha256:776f9449341c254e08b63b061a473e78623a4012418dd4f5d30f152ad6b5c75d

Observation e42c3810-321a-495b-b2b7-8e8fab72f62c · outbound

This paper cites A Survey of Hallucination in Large Visual Language Models.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models A Survey of Hallucination in Large Visual Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.394520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.394520Z digest=sha256:25bb79997164505b4ec0dd2a72226e8fece933136e5e65b5d7108bd6c3809e5a

Observation d99dc812-2306-41f2-a83b-ec9ed90bc5ef · outbound

This paper cites The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.400125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.400125Z digest=sha256:f375976a9bb7a7fb4586b36a160a4141d569e18c8d8ae4c2a5487f1c4ebeed1d

Observation b7e47254-d162-4fa4-9d5f-f15ef6ceea48 · outbound

This paper cites VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.405426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.405426Z digest=sha256:6da1ee2f399f623e80a657c682919e59183887ab3b62efc56806223554f4862e

Observation 7711dc94-2044-41e8-9bd6-28b039097fb5 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.410517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.410517Z digest=sha256:3587ddac75ae8771c2ba5140afe31d042b8462f0563da18c866258fc3c6fdfbc

Observation 1c3cb0b1-4dc1-4976-8cd0-348a781071c6 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models A Survey on Hallucination in Large Vision-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.416665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.416665Z digest=sha256:c68605422a7be57dc5442b40a6ac6705ecab48a57ec97f90decc827f87263779

Observation 86a5c270-8311-488c-b405-56b6c3d4b9f6 · outbound

This paper cites Eliminating Biased Length Reliance of Direct Preference Optimization via Down-Sampled KL Divergence.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Eliminating Biased Length Reliance of Direct Preference Optimization via Down-Sampled KL Divergence

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.422381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.422381Z digest=sha256:05fbf4327d231258fa47e77bcd1165f430453ec2c8ac69d429cd924cf1de152c

Observation e8ddf8e9-9d1c-47a8-902b-dbaccfe1bccf · outbound

This paper cites NUS-Emo at SemEval-2024 Task 3: Instruction-Tuning LLM for Multimodal Emotion-Cause Analysis in Conversations.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models NUS-Emo at SemEval-2024 Task 3: Instruction-Tuning LLM for Multimodal Emotion-Cause Analysis in Conversations

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-16T12:19:44.196154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:19:43.427485Z digest=sha256:1d76f789123d26d1423c27cbf675cac1fc1bc9feb67b7544662ddc026df799c5

Observation 2bd36449-a311-4a77-abfa-d0f6b8488a9f · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.433823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.433823Z digest=sha256:c6b0184fadff684d25aeafa057ad9aea7d0303e2068ac42cea1e7c16988f84be

Observation 110a5a09-9512-4b78-b90d-8bb86694db2a · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.438655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.438655Z digest=sha256:51f3478f7ade9f11d4730ceb3e87b98f1487de07c937c15e378a36e863ebb3b2

Observation ef4b2a90-6878-43f8-85c3-d64409e6a09d · outbound

This paper cites STAR: A Schema-Guided Dialog Dataset for Transfer Learning.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models STAR: A Schema-Guided Dialog Dataset for Transfer Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.443399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.443399Z digest=sha256:1ca01a79dbf3392cb507a0b6d63c23a0ea87fff3aed5980de179831ff0804a0e

Observation 14f20c35-5b03-4bee-a0e2-520c1cd02e99 · outbound

This paper cites Instruction Tuning with GPT-4.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Instruction Tuning with GPT-4

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.454744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.454744Z digest=sha256:db2c264f19e02d1f80f453117d528163f02087f29c73e00e524a00feb1eeaa19

Observation deb317cd-67e6-4ec0-accf-f3f8d3cc0e5c · outbound

This paper cites Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.459735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.459735Z digest=sha256:1e27f8c4fb019ff14de0cf48faa3d4687802adefef76f2b1b1b8f5a383b72c0a

Observation 2002e216-17be-4da8-b4f8-d82623194de1 · outbound

This paper cites Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.464663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.464663Z digest=sha256:a68d8402d7f9cf9ccd36139b45b56efd735651273a167b4c25b9061b02eca951

Observation bbba619d-6583-490e-9106-74a68b0958d7 · outbound

This paper cites TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.469684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.469684Z digest=sha256:27f5514c9504dfa3ca8e063f61d855ac06fd7b92b52071c669aa12f8434b03a1

Observation 8ba5f19e-68cb-4b59-9a54-ef6b9b60e739 · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Reason-rft: Reinforcement fine-tuning for visual reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.474298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.474298Z digest=sha256:94baceb5b202b71b885ad7f2d8944bf4a3b61af7b849685aef26efe9c294d244

Observation 1fdcb6b7-89d2-4790-bb73-d83abf680ca4 · outbound

This paper cites EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.479239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.479239Z digest=sha256:3f603d262075f3922a385c17526d59b1e5226675b13dab37ccde608fd363020f

Observation 4d5c43f6-d0db-42c0-9abb-f0708063362b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models LLaMA: Open and Efficient Foundation Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.484326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.484326Z digest=sha256:3273ebb7ade760eebe30ffec0ecaf9cc225f4549c2351c9c43807251ec65a088

Observation 6e1df478-a7ee-4c6b-b415-6aa98a6eb868 · outbound

This paper cites LightGen: Efficient Image Generation through Knowledge Distillation and Direct Preference Optimization.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models LightGen: Efficient Image Generation through Knowledge Distillation and Direct Preference Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.493875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.493875Z digest=sha256:0b1acf63a4f3497fad61bec9ddbe0c4666fec34afe27762cda40e491131fac06

Observation 2e1ba7be-3f4c-405c-9874-7c507e6fa610 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.498984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.498984Z digest=sha256:8f7caac896f1551aff4fc509b1d5ab53343558b0290088568ab65bddcdd07c1c

Observation 0549031e-ccc2-4243-8199-e326ae49d9b3 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.514613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.514613Z digest=sha256:8bc3d0db9feda379d87a88849d6b59d38d7b930e189efa939f26b1a683b45c61

Observation 8b9bff73-a945-4c15-8c34-54044d7aba17 · outbound

This paper cites Urbanclip: Learn- ing text-enhanced urban region profiling with contrastive language-image pretraining from the web.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Urbanclip: Learn- ing text-enhanced urban region profiling with contrastive language-image pretraining from the web

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:44.599255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:19:43.519672Z digest=sha256:280f1861607159e5826bee2b6f3a3b2cad889ecdd866821bf62e1ac9935644ee

Observation d99d4add-c2ac-4770-9530-6da7ede684fb · outbound

This paper cites Selective preference optimization via token-level reward function estimation.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Selective preference optimization via token-level reward function estimation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.524714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.524714Z digest=sha256:068aca6d1a40d900c43815411cf492b1c86e5297b586a39b6213137b2f1ee1b5

Observation 1d3f4ab6-0bb8-4fc6-90b1-4a39f5bcdeb7 · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.529144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.529144Z digest=sha256:80e56b387063a3182382c37b8cae558253b97cd0b29d6eacdb330fe297cd2899

Observation f4727efe-4204-49e4-a05d-e8b74aae3abe · outbound

This paper cites VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.534105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.534105Z digest=sha256:df8f35e0143aa5cf00344f219061089b0f5c55317d7efc65f8829e30e133a829

Observation 10e4deab-3929-4688-bbd8-84862fb74085 · outbound

This paper cites Token-level Direct Preference Optimization.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Token-level Direct Preference Optimization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.539104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.539104Z digest=sha256:1482c223806139be354d9aeb1b81d01843b8641fe870a17ac144fcab359f7299

Observation 67e4f955-643e-40d8-a0a2-fd62db5e5569 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.544103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.544103Z digest=sha256:1502b0fb5527bc283d57ac9654664c922f380ff38a2dc4bba29cf84739be5423

Observation ed24308b-da7a-43e9-a597-b2c6e10df56b · outbound

This paper cites Two critical distinctions are summarized as follows: • Spatial-Temporal Video Preference Optimization: Previous DPO methods predominantly focused on language-level alignment.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Two critical distinctions are summarized as follows: • Spatial-Temporal Video Preference Optimization: Previous DPO methods predominantly focused on language-level alignment

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:44.563284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:19:43.564296Z digest=sha256:076239d92f6ce8de3f1b576a1cb1aecce56c57a8621ed245d2afe70cdef800d6

Observation db0fe6a9-35ca-4622-8952-afd3770a5c7c · outbound

This paper cites While some works incorporated image-level visual alignment, these approaches remained limited to static images.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models While some works incorporated image-level visual alignment, these approaches remained limited to static images

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:19:44.544280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:19:43.569520Z digest=sha256:3e37539dfb21a5aebc348a6b85e706f33da60e8014ae0290f3a689c891bedb67

Observation e240e4f2-7486-419e-9b1e-8627bf3537f3 · outbound

This paper cites is_in_video.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models is_in_video

Reference 48

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T12:19:44.526485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:19:43.574735Z digest=sha256:0cf10575a8a70fc9da8d6d42eea70e874c87c939d1cdbd3667f0470c046a9718

Observation 7cecc59b-4976-4416-8162-2a511b63c0a4 · outbound

This paper cites GaussianVTON: 3D Human Virtual Try-ON via Multi-Stage Gaussian Splatting Editing with Image Prompting.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models GaussianVTON: 3D Human Virtual Try-ON via Multi-Stage Gaussian Splatting Editing with Image Prompting

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.339819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.339819Z digest=sha256:defc93637d0a26b331b8870801fa23f2aef691a20470af9caaafe416b857a394

Observation 5c7795b5-c494-4620-bcb2-5ed9ee909201 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.509698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.509698Z digest=sha256:5f1158f3a69a27cd4a4216b25fce16e7fdd9121627e18bfaf0504f4c2494acd4

Observation 27e4f9ef-9ed5-4a54-a994-7026b829a785 · outbound

This paper cites Supervised Fine-tuning in turn Improves Visual Foundation Models.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Supervised Fine-tuning in turn Improves Visual Foundation Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.389601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.389601Z digest=sha256:8bea4489f901104d338749214275c5286cb3a7ff4b36f18fc108083bf1531339

Observation 16b16208-ad9f-48bb-afb3-1aeb5829791f · outbound

This paper cites HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.553849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.553849Z digest=sha256:10f4e93511823444c4e49585ace8e41ada3b55eba825785d879a3ce97d394249

Observation ac853aea-44ef-4262-8e15-0707ff2148e7 · outbound

This paper cites VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.488922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.488922Z digest=sha256:7092f14ac224b6528ac64eb457e3c721e4b873922491eca26dde48cacb86cf4b

Observation c8e471bc-555a-41ba-81f2-98b146f7e896 · outbound

This paper cites Disentangling Length from Quality in Direct Preference Optimization.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Disentangling Length from Quality in Direct Preference Optimization

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.448976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.448976Z digest=sha256:97d7143e4f839e79502a0091911cd0b04ddd85a535121f78dd8fbf41c3775bfa

Observation c5b44fcb-65cd-4d81-b90d-48fddb055aac · outbound

This paper cites V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.504039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.504039Z digest=sha256:0e45e8bcb40531fc9545c81cc0aab5e46306a68f38f34c65531e5c5d4774871e

Observation fafe68ba-0152-4ec0-8b70-8d69d4b66504 · outbound

This paper cites Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.548831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.548831Z digest=sha256:d1b905b64ce6eaef740f849b1a35cea6b2ece7d3dad0781e2ea7e84d2d785670

Observation 3d24b363-648c-4be6-9afb-da25fd9ca4cd · outbound

This paper cites DependEval: Benchmarking LLMs for Repository Dependency Understanding.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models DependEval: Benchmarking LLMs for Repository Dependency Understanding

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.350498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.350498Z digest=sha256:f1f38c44aa511f91a199d109c15f073f8d369ac97c40ae9474b2b0034c5fef36

Observation d3bade13-e637-4519-b9a0-c3a4c5785031 · outbound

This paper cites Qwen Technical Report.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Qwen Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.334265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.334265Z digest=sha256:3b1f3ed1c4fe3ba0b94a5a0f79ff47edaeb3e71f79f93554ae0ed6dc850af7ac

Observation 520d7d3c-ee8d-4c49-a34a-f4604b25a23e · outbound

This paper cites Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment.

VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:43.345242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:43.345242Z digest=sha256:702815490e9438f8abfdda47e369851a0bf3e7244f914b0a8144e77c3984ccc5

Pith citing papers

Observation 862f5644-7c48-4c7b-a1d9-9ceb880ba262 · inbound

FinePhys: Fine-grained Human Action Generation by Explicitly Incorporating Physical Laws for Effective Skeletal Guidance cites this paper.

FinePhys: Fine-grained Human Action Generation by Explicitly Incorporating Physical Laws for Effective Skeletal Guidance VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:19:03.525559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:19:03.525559Z digest=sha256:080b430fa98b0783a73e1fcb183fc4e826f787b3a096d9ebe2d3f1635cdb9855

Observation 3051702f-4258-486d-9388-5c856b21aee1 · inbound

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs cites this paper.

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:05.050891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:14:05.050891Z digest=sha256:de76c4665bf084183a88964183f7309a713ae53b03912ea90e6482a87a4e2e55

Observation da0bf6ba-2227-4950-a54e-0e4d03af8456 · inbound

Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation cites this paper.

Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:21.374295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:20:21.374295Z digest=sha256:73e645fed7e5fedca3d5cf2e8e4371821ceea86d56aa34e231ef42b350c0dc4e

Observation 0fde41f4-4c7a-4303-82c6-f3da424f2f48 · inbound

From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms cites this paper.

From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T20:17:46.217418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:17:46.217418Z digest=sha256:4677985c2684ff4396ded75acac38be5fbcaa4b754990ba1cade9ef875ddd4c2

Observation 41f780d3-8f56-4fd1-b112-2ccb19bdf0e7 · inbound

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models cites this paper.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:25.474447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:25.474447Z digest=sha256:b6abc71a2321ff3d21a82b2a7b67ba22cac4f929a12eadc406b369a20870222f

Observation bc1bc1d1-0492-4066-98a4-af3552b1a420 · inbound

SignDPO: Multi-level Direct Preference Optimisation for Skeleton-based Gloss-free Sign Language Translation cites this paper.

SignDPO: Multi-level Direct Preference Optimisation for Skeleton-based Gloss-free Sign Language Translation VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:01:06.484603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T04:17:32.036750Z digest=sha256:f998dabeb3124f9ed65c72842f242f74ad689f78dfc49e0040200e8e4ab24369

Observation 96dc0c62-b9e9-4410-9d98-cc57cc1ff24d · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models

Reference 185

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.670866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:5d90a43856758cb3862fcebae2ba0460e5a3046cfe085ed881c168f9bae2b2eb

Observation c24e1fca-75e9-4209-a322-ffa3ec4e6d3e · inbound

MultiToP: Learning to Patch Visual Tokens to Mitigate Hallucinations in Video Large Multimodal Models cites this paper.

MultiToP: Learning to Patch Visual Tokens to Mitigate Hallucinations in Video Large Multimodal Models VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.145426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T10:02:58.341050Z digest=sha256:37358fd3b93c22d5b9158b6fa2ca49cb922ce1e9c1f70b6c5a6b7eddb7a88584

Observation f372314a-ea01-4b55-8836-f58116685c68 · inbound

No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs cites this paper.

No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:41.369041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-01T05:35:08.200219Z digest=sha256:f983291ba28ca9a1b89d6a1608dd02bcc9f150ad3662e63a7bf4bc1093e05264

Observation 05809b2c-d457-4012-bab5-d524384f3d09 · inbound

VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models cites this paper.

VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:48.781355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:35:48.781355Z digest=sha256:413d69cc6d7de79a95b9097df068420fe4c14d1607a8e195ae8da6773cc5be36