Pith. sign in

Paper Citation Record · LEDGER

Temporal Preference Optimization for Long-Form Video Understanding

As of 15 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 8 inbound Pith citation observations for arXiv:2501.13919.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13919 v3

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:35:30.340013Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:35:47.754884Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T13:45:46.078854Z

Reference resolution

73 of 73 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6388fa38-6206-498b-8586-de772a5ed90c · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Temporal Preference Optimization for Long-Form Video Understanding Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.052266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.052266Z digest=sha256:254d29005ad92183eb60df30a62d0beebe6f918edb80d79e8e242bfb608998ab

Observation af86cbbf-dd04-49a9-8b09-c85557555cc3 · outbound

This paper cites GPT-4 Technical Report.

Temporal Preference Optimization for Long-Form Video Understanding GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.058831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.058831Z digest=sha256:33de9cc9ecc47dbb2e861fc23602e39273afe80fabc3bb409bea9f54eb8632ef

Observation ccd4f21c-0009-444f-95cc-21a91a44285f · outbound

This paper cites ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO.

Temporal Preference Optimization for Long-Form Video Understanding ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.065401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.065401Z digest=sha256:36e789e9c19793d7c957c550ab9d7bd03b608315c4a4fb886ee22b19b7f49cf6

Observation 244a4b93-45c9-4d73-bedf-b8b6ef16fd29 · outbound

This paper cites Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback.

Temporal Preference Optimization for Long-Form Video Understanding Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.073317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.073317Z digest=sha256:d5c182e253b1bd229ea64d66ae069c1f71ad7ad4526679f0866505a9571d6b6f

Observation 50d9d4a1-7f72-4fc6-aebf-a59309f2e270 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Temporal Preference Optimization for Long-Form Video Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.078522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.078522Z digest=sha256:dc24d6e4c006db92a21037707edd17c5a8b6774a5097bd05be99486024b12283

Observation 431f0539-0b9d-447c-9712-5c0ea0be68f2 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Temporal Preference Optimization for Long-Form Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.082895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.082895Z digest=sha256:656f05bf98fbeb40302883ce9ca28348237ca9f438fbd3d2a0fb304c35270180

Observation 178594a7-eaa1-444d-a2c1-a127bef70c95 · outbound

This paper cites TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability.

Temporal Preference Optimization for Long-Form Video Understanding TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.088602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.088602Z digest=sha256:0c776d53566a98d69c3b9bee586ec975a3a20ac1b21486aa758ceddf1d316245

Observation 0278fd76-9f7e-46f5-897f-4362c793b0fc · outbound

This paper cites Large-margin contrastive learning with distance polarization regularizer.

Temporal Preference Optimization for Long-Form Video Understanding Large-margin contrastive learning with distance polarization regularizer

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:35:31.058537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T15:35:30.092665Z digest=sha256:ebd7c4681428bd3b0b697bff9c46fe2f50073238db6fe44e69fe38205971c622

Observation 89af25b5-0305-4b2b-9175-803d0c9c4187 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Temporal Preference Optimization for Long-Form Video Understanding InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.097928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.097928Z digest=sha256:95e338cb6a726c386924448d1cf9840a5ad4e85515d1496c0893871dd6b50363

Observation b6701107-79a5-4fcc-8e50-569efb463f4e · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Temporal Preference Optimization for Long-Form Video Understanding Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.102854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.102854Z digest=sha256:1c6770fe19e1f852bf8b035ba091c856577e2d04a1d05a191765cdd568baaad5

Observation 45e4da70-ee5b-4608-a319-f59ec92a4418 · outbound

This paper cites Understanding Transferable Representation Learning and Zero-shot Transfer in CLIP.

Temporal Preference Optimization for Long-Form Video Understanding Understanding Transferable Representation Learning and Zero-shot Transfer in CLIP

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.107038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.107038Z digest=sha256:829d8cdca25ba8d3beca45836c5784c12f1386267833d7b9dce157076936f78d

Observation 02aee7bf-126b-48ca-9cd2-908c8021e7b0 · outbound

This paper cites Enhancing large vision language models with self-training on image comprehension.

Temporal Preference Optimization for Long-Form Video Understanding Enhancing large vision language models with self-training on image comprehension

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:35:31.040312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T15:35:30.111836Z digest=sha256:d0a502dee0e880961cdc707e46332c642c2f8278f552408353d7724104732131

Observation c8260679-dd87-454a-84e7-b9c06c9110d9 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Temporal Preference Optimization for Long-Form Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.116631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.116631Z digest=sha256:ac386903af74485f5a5b9c39d9128f25a3ca9de83ecee45b0c7801978e77ebd9

Observation 90129829-1761-4e0d-8c4b-7998284f5cde · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Temporal Preference Optimization for Long-Form Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.121345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.121345Z digest=sha256:872692b6bf8d0f988a14573766dfef7cb7b593b2ce5fd7b3aeb78638cf04e12c

Observation be63f7ac-5fce-4683-9ddd-1052420ac8cc · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Temporal Preference Optimization for Long-Form Video Understanding VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.125217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.125217Z digest=sha256:8e08f2098bb1fadc1fe718af57f28af842d3dc41718ee0ed8dceedd65599d28e

Observation 3005a8aa-3314-4506-8c76-f3a7083a177c · outbound

This paper cites Tall: Temporal activity localization via language query.

Temporal Preference Optimization for Long-Form Video Understanding Tall: Temporal activity localization via language query

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:35:31.028835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T15:35:30.129019Z digest=sha256:456083a0b6b1aa0d0a42f9a31529f4ca9ac2be59b36b63d9bcfc444a45ab1d17

Observation 7fd242ed-fd9e-41b5-870a-6cb38cf9afcb · outbound

This paper cites Detecting and preventing hallucinations in large vision language models.

Temporal Preference Optimization for Long-Form Video Understanding Detecting and preventing hallucinations in large vision language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.132541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.132541Z digest=sha256:1b03bedb8d7adad1550e6922fe65363d3c858461b74d6195cf74e52f778175d1

Observation 0ae9fa5b-c640-4e58-930a-cc0726152325 · outbound

This paper cites Large Language Models Are Reasoning Teachers.

Temporal Preference Optimization for Long-Form Video Understanding Large Language Models Are Reasoning Teachers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.136285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.136285Z digest=sha256:4a78453b8b812b9d4a47ded4b4e2a06ada63fb711353ff0372c264acf26c8b8e

Observation 5f20bb0b-ca90-4602-96a2-3f33908ac49d · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

Temporal Preference Optimization for Long-Form Video Understanding CogVLM2: Visual Language Models for Image and Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.140139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.140139Z digest=sha256:5aa6af80a00ae55e36c87560855a2a9f14adb6cd30d607507d8a91049325ee7b

Observation 9edc5c85-44d2-4ab4-baae-56f8187c86ac · outbound

This paper cites Lita: Language instructed temporal-localization assistant.

Temporal Preference Optimization for Long-Form Video Understanding Lita: Language instructed temporal-localization assistant

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:35:31.010533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T15:35:30.143440Z digest=sha256:6a615a28c641a0bbcfb9b9de67f9a5417a13a0ca7ef6c9525b40a191bf1437bc

Observation b4a7e62e-a4b8-46d8-8c5d-4fda27ffe16c · outbound

This paper cites Large Language Models Can Self-Improve.

Temporal Preference Optimization for Long-Form Video Understanding Large Language Models Can Self-Improve

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.146705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.146705Z digest=sha256:dfcc1e80b7ac40739b4ae98e138b4b84c67b550b35411987819bd437c321084f

Observation 9f19911a-5b01-4695-93f5-d749ae93e01c · outbound

This paper cites Bimba: Selective-scan compression for long-range video question answering.

Temporal Preference Optimization for Long-Form Video Understanding Bimba: Selective-scan compression for long-range video question answering

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:35:30.998517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T15:35:30.150125Z digest=sha256:49a01f3dd95d2f8740c0299c199e1327e0e48c03565ff87b87ea713634d1a65f

Observation 78f59d88-b517-4a00-b2be-32e334da12ca · outbound

This paper cites What matters when building vision-language models?, 2024.

Temporal Preference Optimization for Long-Form Video Understanding What matters when building vision-language models?, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.153194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.153194Z digest=sha256:9b1fa642dca76f06c80ea4c85cc8ebc1c672f05b0bdbb1aedf8a9e971d7c26f6

Observation 50c563c7-d614-4ab2-861d-1e102fba0451 · outbound

This paper cites Detecting moments and highlights in videos via natural language queries.

Temporal Preference Optimization for Long-Form Video Understanding Detecting moments and highlights in videos via natural language queries

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:35:30.980478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T15:35:30.156129Z digest=sha256:95fe12a2b856b3eeb9a5a51c76114f09012808fe033c001b96691275bea2a17a

Observation 6d969358-0c24-4406-bef4-d42986ae8cd2 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Temporal Preference Optimization for Long-Form Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.158917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.158917Z digest=sha256:972201f0cdf7c8616bd8e1d30572c9a57df5bacbca5248621075e66194f501df

Observation 4205c792-fcd8-44b6-aa00-cdcc428a21c9 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

Temporal Preference Optimization for Long-Form Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.162277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.162277Z digest=sha256:ad4b2cff7e85171b1ccc9c368b3c3154926b204e42c87f345f8f419122f00e0d

Observation 3e83448f-87fd-4b04-a40e-c81e2a3900a3 · outbound

This paper cites Llava-next: Tackling multi-image, video, and 3d in large multimodal models, June 2024.

Temporal Preference Optimization for Long-Form Video Understanding Llava-next: Tackling multi-image, video, and 3d in large multimodal models, June 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:35:30.968842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T15:35:30.165888Z digest=sha256:900f43992cd1547e52c230b8a467337103b96c46409fdf23335cdb53538427df

Observation a4755349-b04c-4fda-8bd6-7c526c9e1545 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Temporal Preference Optimization for Long-Form Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.172948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.172948Z digest=sha256:7ee466e452bfbbfb38f74297f78d00a56caadfa3c926dcca26363790adc86c13

Observation 8e534273-46d3-4802-b811-757a84634a33 · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

Temporal Preference Optimization for Long-Form Video Understanding Silkie: Preference Distillation for Large Visual Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.176415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.176415Z digest=sha256:350249d822b449f4062181f774cd388e8cfa4c2cdcd6b3a95e95897b5546de3d

Observation 43647c39-46b5-4224-9489-403e3bbd2b17 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Temporal Preference Optimization for Long-Form Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.180773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.180773Z digest=sha256:13bebf292e07bd5eba0ab8ed78bb7d48d307687482d97cf6aee2ec6317ef4c69

Observation 75b6d04b-cfdc-41b8-bc66-749917494c59 · outbound

This paper cites Vila: On pre-training for visual language models.

Temporal Preference Optimization for Long-Form Video Understanding Vila: On pre-training for visual language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.184563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.184563Z digest=sha256:c53960dc6b1de5afc4f5b0fc08fe44795f8d9986cbf7aaa92678d35691c5efa2

Observation 2fa9da1c-f4f2-4c5b-aa50-9e24f619ce23 · outbound

This paper cites Visual instruction tuning.

Temporal Preference Optimization for Long-Form Video Understanding Visual instruction tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.187975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.187975Z digest=sha256:afa269776349143525134cfadb6b99e5132f510fb09c04be144470b11a364aad

Observation f1414b7e-854b-45fe-b45c-5f502fa584fb · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

Temporal Preference Optimization for Long-Form Video Understanding Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:35:30.944875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T15:35:30.191942Z digest=sha256:66dd99c7b7cc691dc77bc09fc87804398d58eeec156a53166a7ac9126c0655ef

Observation c71c0487-de8e-48a9-ac5f-a66dde5d4b27 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Temporal Preference Optimization for Long-Form Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.195645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.195645Z digest=sha256:08920204536f69ce26ed6ccb764aa528d2caba98281666b08fb602d194f89319

Observation d33a54a7-9b41-4224-8a20-4010fe75f9eb · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

Temporal Preference Optimization for Long-Form Video Understanding NVILA: Efficient Frontier Visual Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.199392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.199392Z digest=sha256:e20416ae298b13b7ecc5568ee4c2bfd26f1a6397de596d44c0147076fa5b9b3a

Observation 83e485c8-b2cd-4efe-bf0c-123b9c521a66 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Temporal Preference Optimization for Long-Form Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.203011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.203011Z digest=sha256:fc4bf05527edf02030f8c3a7a28babbc07f942d8015922b2c2bf878aa02d29f2

Observation 6f5b5fd2-90fd-42e7-88d8-8f0f30c49746 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Temporal Preference Optimization for Long-Form Video Understanding SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.207074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.207074Z digest=sha256:e84583013f0dbe3bbae91de6b22aebaafbeeff20c49068524184ace014526942

Observation d4def448-4799-4251-9c34-6e43bffc9157 · outbound

This paper cites Deepseek-vl: Towards real-world vision-language understanding, 2024.

Temporal Preference Optimization for Long-Form Video Understanding Deepseek-vl: Towards real-world vision-language understanding, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.210839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.210839Z digest=sha256:705ef6940859be2fe10b78dda95bd31fb5238544fbeef0a6c9a2e6343049735e

Observation 8e65e415-d391-4ba6-8474-01141e1bd17b · outbound

This paper cites Query-dependent video representation for moment retrieval and highlight detection.

Temporal Preference Optimization for Long-Form Video Understanding Query-dependent video representation for moment retrieval and highlight detection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:35:30.927008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T15:35:30.214485Z digest=sha256:2f1c556578cc92c31c771e01fb974e43bd6b88096982b3d9297525d0f759588f

Observation 7150d3db-2156-4f82-b1c1-53c00ae1c930 · outbound

This paper cites Training language models to follow instructions with human feedback.

Temporal Preference Optimization for Long-Form Video Understanding Training language models to follow instructions with human feedback

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.217991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.217991Z digest=sha256:6735aff28b6f96a7a011adfab4cd786b8c342e24dcefc03217152c98042132db

Observation 608135bd-2053-4086-82d0-8c3c25d82d2f · outbound

This paper cites Generative agents: Interactive simulacra of human behavior.

Temporal Preference Optimization for Long-Form Video Understanding Generative agents: Interactive simulacra of human behavior

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.221878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.221878Z digest=sha256:be6a96ba36d65fb9cc4298b47434403720459577ffda6635563b5c4162b5cd48

Observation 9bff6494-6963-4327-9d17-c0c336f63cc3 · outbound

This paper cites Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization.

Temporal Preference Optimization for Long-Form Video Understanding Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.225400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.225400Z digest=sha256:3ebd77f525de90ea8d6f5fb285962c1c610866d58a564fd956a97bcd270f5d83

Observation e2dc9ba7-19b0-4187-92a4-e105f8c5bc76 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36, 2024.

Temporal Preference Optimization for Long-Form Video Understanding Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36, 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.229228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.229228Z digest=sha256:ab659720e497338d2163de00752b38f8c23b2194f9ab43e981303f5832f75227

Observation d43587ee-2bb1-4c20-80ae-0a42361f05f9 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Temporal Preference Optimization for Long-Form Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.232886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.232886Z digest=sha256:36cb4afdc1f5cd60a8b6548294278bb48651286524d2e5996af6a1e6d36ce91c

Observation 252355a5-fb54-4144-841f-c37f1284835a · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

Temporal Preference Optimization for Long-Form Video Understanding Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.236757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.236757Z digest=sha256:0bddfd54585657392329ad960d217e93e8424fa19eab5b3661dc5d4e31eba997

Observation 9e670a45-50be-4d6a-beb9-9f45701052e8 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Temporal Preference Optimization for Long-Form Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.240530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.240530Z digest=sha256:903e8888668c6ab65390975dfc037d78e8420d1ef77ef1b9339d9951b565e319

Observation 7b71ba94-2146-45e8-8a0f-07d06ad2a63e · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Temporal Preference Optimization for Long-Form Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.244486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.244486Z digest=sha256:93e23797ce4f7a844dc2aac92da9d67274aba2034143ea09a54beb599faa2df4

Observation d6fc0a06-e86d-44eb-a151-a975c8419f00 · outbound

This paper cites Learning to summarize with human feedback.Advances in Neural Information Processing Systems, 33:3008–3021, 2020.

Temporal Preference Optimization for Long-Form Video Understanding Learning to summarize with human feedback.Advances in Neural Information Processing Systems, 33:3008–3021, 2020

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.250453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.250453Z digest=sha256:5c1ea5813a199f638b1b18f04d9494984291eaaa1e141ec42d05d4eebb359b47

Observation 4bca69db-3c72-4332-a9d5-349878114619 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Temporal Preference Optimization for Long-Form Video Understanding Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.254242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.254242Z digest=sha256:584efbd0d7dd1a7f59bced733747133c0012ed18dfdae1dc7dabbdfadcc0e51f

Observation ce014ffa-038b-4798-a758-e213929b9342 · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

Temporal Preference Optimization for Long-Form Video Understanding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.257411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.257411Z digest=sha256:f8c74d6db436c2610d418bee5cc23ace8d6184b7e6ea323fe6c79b952b207f82

Observation 7c611c7e-9fc9-4b3a-b831-d5b91c8a66be · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Temporal Preference Optimization for Long-Form Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.261192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.261192Z digest=sha256:551b222d0be27da0f20d19f6d32e08229d41030f88f161e52f8e6e488dfbec26

Observation e66732ff-36fd-445a-95a2-600bd9949f16 · outbound

This paper cites End-to-end dense video captioning with parallel decoding.

Temporal Preference Optimization for Long-Form Video Understanding End-to-end dense video captioning with parallel decoding

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:35:30.884727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T15:35:30.264527Z digest=sha256:6bb78ddd8b4557d0d94828611392f74856dc59478b7294af63545b6d08dfe5b5

Observation 40077cef-3a59-4d9b-94d4-828df56a57a2 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

Temporal Preference Optimization for Long-Form Video Understanding Videoagent: Long-form video understanding with large language model as agent

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:35:30.873713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T15:35:30.267901Z digest=sha256:07f1412a44159489fe53460245baf43e89db2c6499082eb74aba1c78b190f19e

Observation 41e79b09-b4b4-4c5d-a509-1a2bda6a9979 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

Temporal Preference Optimization for Long-Form Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.270844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.270844Z digest=sha256:0f57329bee7256900750e7a75e3a4fef08c6ca2df924803e810874b3b1ca2359

Observation 78e2b051-dc7c-4cb7-8613-23b21e81a885 · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

Temporal Preference Optimization for Long-Form Video Understanding Can i trust your answer? visually grounded video question answering

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.274095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.274095Z digest=sha256:b0b7020345e2e61c0604212251b764ca58a23da5ecfc1a490335bd64d3501816

Observation 71444874-7339-4630-913f-031016faf40a · outbound

This paper cites Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024.

Temporal Preference Optimization for Long-Form Video Understanding Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.277584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.277584Z digest=sha256:bc4acc849addca315909e4c71befdc9368ec89c02d92601e3d3d67a99bc2688e

Observation ec269b35-c878-4a5c-ac71-0163e9a852a7 · outbound

This paper cites Vid2seq: Large-scale pretraining of a visual language model for dense video captioning.

Temporal Preference Optimization for Long-Form Video Understanding Vid2seq: Large-scale pretraining of a visual language model for dense video captioning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.281216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.281216Z digest=sha256:689ef3418cc93a496a3f123cfff21e3e108075913f0bbd145e846b612e8010ca

Observation 7f1bde46-37a7-46b4-92ce-81e5ab375542 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Temporal Preference Optimization for Long-Form Video Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.284542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.284542Z digest=sha256:5c6ac7c1ba8d27c4d0ab5c9169c4ee6cde3e341914ad4303f9a863ee0799d521

Observation ac5b31d3-66bf-4961-9f14-d033c55cc6dd · outbound

This paper cites Every moment counts: Dense detailed labeling of actions in complex videos.International Journal of Computer Vision, 126: 375–389, 2018.

Temporal Preference Optimization for Long-Form Video Understanding Every moment counts: Dense detailed labeling of actions in complex videos.International Journal of Computer Vision, 126: 375–389, 2018

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:35:30.843308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T15:35:30.288209Z digest=sha256:fecd44c32a544128beabf7611e6b05bab78c55642f85e1fe992fd5bc4cfa502f

Observation 1f5289d5-c8d3-493b-813a-3d85fde8b2fe · outbound

This paper cites aws-prototyping/long-llava-qwen2-7b, 2024.

Temporal Preference Optimization for Long-Form Video Understanding aws-prototyping/long-llava-qwen2-7b, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:35:30.833065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T15:35:30.291607Z digest=sha256:8c58c41305db6b05a4eacba267fd989b228f8e81a1e5cdd0bb81b35f3cb2d1aa

Observation 8bfbdaa9-e89c-44bf-ba55-5dd07e84a98f · outbound

This paper cites Semantic conditioned dynamic modulation for temporal sentence grounding in videos.Advances in Neural Information Processing Systems, 32, 2019.

Temporal Preference Optimization for Long-Form Video Understanding Semantic conditioned dynamic modulation for temporal sentence grounding in videos.Advances in Neural Information Processing Systems, 32, 2019

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.295157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.295157Z digest=sha256:79635ec6bb947b862ecbb5eeb10ff658b469247df9d473d80a4a9c618f09a440

Observation 6c296809-5d39-478b-a984-d512c1931a7b · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.

Temporal Preference Optimization for Long-Form Video Understanding Star: Bootstrapping reasoning with reasoning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.298352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.298352Z digest=sha256:678cadaf70479de4cb74d1cee9c6f154ada543a25e3b6014b465b6a15b52d051

Observation 5b9854d1-d8a6-4e5e-b9a3-3af0b15ddcbb · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Temporal Preference Optimization for Long-Form Video Understanding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.301738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.301738Z digest=sha256:65f827e6074c97f6cdc108060a70a75ec2d4857d35a7c911b21999fca3caa85a

Observation cfb261d3-c08a-4034-8835-35aad62cee14 · outbound

This paper cites Long Context Transfer from Language to Vision.

Temporal Preference Optimization for Long-Form Video Understanding Long Context Transfer from Language to Vision

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.305277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.305277Z digest=sha256:e53eb2f0d6fa0f2655c0fddb67c7958dceae16365c0b98947c9b66d0b92d9edf

Observation 908c1f38-1639-40e4-aa1a-2ecadfba7fd6 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

Temporal Preference Optimization for Long-Form Video Understanding Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.309108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.309108Z digest=sha256:5edf9c722190769da87390562e59d5e1d20ec6301191880da9bcc7ed9d6f9451

Observation efe6e47c-5173-4d10-9d5c-a440411f6d01 · outbound

This paper cites Llava-next: A strong zero-shot video understanding model, April 2024.

Temporal Preference Optimization for Long-Form Video Understanding Llava-next: A strong zero-shot video understanding model, April 2024

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:35:30.807930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T15:35:30.312880Z digest=sha256:3d3d5de0447dc84ea6fbf605f16c408fb033f5ee6666291021daf6b4b8f15dd1

Observation 10ba334a-bb4c-4f00-b439-b1f38cf09b18 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Temporal Preference Optimization for Long-Form Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.317413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.317413Z digest=sha256:941ef80efef84591eb98a2f8973b2d4eea76a1fa7d86294b3bb3f19d36d009a3

Observation 4fb06a6a-0640-451b-9394-8c264991c1fd · outbound

This paper cites Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization.

Temporal Preference Optimization for Long-Form Video Understanding Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.320936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.320936Z digest=sha256:8f956fb33c4cfa4470373df4216dbe9e2c6a317bd70b2701c8ec1b1bb51a756c

Observation 3218d2ec-b6a7-41f5-9998-e47532f8c3ff · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Temporal Preference Optimization for Long-Form Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.324482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.324482Z digest=sha256:e9e15c20cb36f0e78895a9babdf2696c60039d66d3b084eb40c99ab5f193aca6

Observation 0b26ef39-f39e-4129-bef5-86dbb0056aa0 · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

Temporal Preference Optimization for Long-Form Video Understanding Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.328632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.328632Z digest=sha256:4bb0c07f16903fe5d34a33d603a641680f22079739c8ac0e9c76db126bed8e36

Observation c659258e-d00f-41d3-8a28-946ddb101ae7 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Temporal Preference Optimization for Long-Form Video Understanding Fine-Tuning Language Models from Human Preferences

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.332097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.332097Z digest=sha256:371a45f9895b5003eb7b03ae636ba49cb2d7e81c5ee2f338f4c33886912afc10

Observation cbd74da5-b010-4ae6-9fc1-70794fdbb2eb · outbound

This paper cites Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision.

Temporal Preference Optimization for Long-Form Video Understanding Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.335661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.335661Z digest=sha256:c91b9153e8c7747ee1be23fe77c35babbd1d0cfb1fee3dc051e7a3b8a75437b7

Observation bae030c5-b38e-4907-8317-21555e5d6524 · outbound

This paper cites Apollo: An Exploration of Video Understanding in Large Multimodal Models.

Temporal Preference Optimization for Long-Form Video Understanding Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.340013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.340013Z digest=sha256:8ba83e53285c56ae23134a513ddd8bec06d7a3b3b70dc6de18f2b90bafa5a036

Pith citing papers

Observation 07820ea5-d273-419d-b0f9-1712d3f3882c · inbound

Unified Reward Model for Multimodal Understanding and Generation cites this paper.

Unified Reward Model for Multimodal Understanding and Generation Temporal Preference Optimization for Long-Form Video Understanding

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:44:30.698805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T00:44:30.558048Z digest=sha256:c91149a113bd66bb919adf85494ac93ef8ad7cb92862cea5576961e41b094b8f

Observation 94a971de-4074-476a-9cf7-63b5e04448ba · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs Temporal Preference Optimization for Long-Form Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:46.708160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:46.708160Z digest=sha256:72399b6d39472825b1bb1412a185502b9fff595514122aa060ef520ac83f58cf

Observation ad71e2fe-5895-4c17-814c-8f8992cfa555 · inbound

How Important are Videos for Training Video LLMs? cites this paper.

How Important are Videos for Training Video LLMs? Temporal Preference Optimization for Long-Form Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.841588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.841588Z digest=sha256:b56f63de2ab06b5950b56ad83fdcf583ac4f475bb3b33426330b5c123cbd9dd6

Observation c4e058ca-8cf5-4bea-9fd7-7106af43756d · inbound

Towards Temporal Compositional Reasoning in Long-Form Sports Videos cites this paper.

Towards Temporal Compositional Reasoning in Long-Form Sports Videos Temporal Preference Optimization for Long-Form Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T19:27:29.843866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:27:29.843866Z digest=sha256:4e1050f22e9636f378f84c6750d853f8be84cbb4260387a9fc7be53995701723

Observation ae89cc00-7ed8-4799-a24e-f763409a485b · inbound

CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding cites this paper.

CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding Temporal Preference Optimization for Long-Form Video Understanding

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T13:45:46.080486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T22:47:19.742035Z digest=sha256:15b38e39799d9bb9330e1bf7f9404b95d1b8e3f71243720e004c7eb722213faa

Observation ea400a07-fc86-4d12-83fd-776d02899ee7 · inbound

CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding cites this paper.

CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding Temporal Preference Optimization for Long-Form Video Understanding

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T00:18:46.708657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:18:46.708657Z digest=sha256:9a0c429f0301d245d6c5b77e92e8123dabacfd073c2d50bdbc96fad52fa8ecf4

Observation 80fa6c11-be87-49d7-ac79-ab920f783b35 · inbound

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction cites this paper.

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction Temporal Preference Optimization for Long-Form Video Understanding

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:54:22.386718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T07:48:01.719339Z digest=sha256:db4c8ceb5f22980a0cce7786b8c7c513866d95ac86d9a73866e25080e8c54b30

Observation fa3d5b0d-9bba-4c10-99f3-c8a6c601656f · inbound

VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models cites this paper.

VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models Temporal Preference Optimization for Long-Form Video Understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:47.754884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:35:47.754884Z digest=sha256:fd97aaebafb07ad05193b2ba331a554cdd2507d45fa574fb45e96f22d7dafd0d