Pith. sign in

Paper Citation Record · LEDGER

LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2412.04814.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04814 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:06.660518Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:09:29.415956Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 26a2e067-1bc8-4ef8-b328-94d6cc290264 · inbound

Improving Video Generation with Human Feedback cites this paper.

Improving Video Generation with Human Feedback LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:30:02.820925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T15:30:02.578430Z digest=sha256:fe179e135c1b0c7ac5053a7d8a7799c0147ec59bf821949250f70424f849a834

Observation fb551883-4bf8-404d-84e4-118056d81d74 · inbound

Unified Reward Model for Multimodal Understanding and Generation cites this paper.

Unified Reward Model for Multimodal Understanding and Generation LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:44:30.907784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T00:44:30.558048Z digest=sha256:580fcc8eb2b114cde56cc4234402cd45ac00fc34d30e4866b6d8bc2e1d739151

Observation 51bb15f3-bdcb-4d6f-8df0-9038d79f7a30 · inbound

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation cites this paper.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:06.660518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:06.660518Z digest=sha256:a003bee398d66133d9d4c481131366f26bd988c14f9df171705df3644b714eac

Observation 3a0ac913-9b07-423f-aaf0-9bc2807e32a9 · inbound

GigaVideo-1: Advancing Video Generation via Automatic Feedback with 4 GPU-Hours Fine-Tuning cites this paper.

GigaVideo-1: Advancing Video Generation via Automatic Feedback with 4 GPU-Hours Fine-Tuning LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:29:35.464597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:29:35.464597Z digest=sha256:fdde2278404fb9ab802e9c2483994bdf8a73e3948e92df93c3aa879a0903cb77

Observation 05c183fa-f768-4f88-b5e6-19ca219cb2c7 · inbound

VQ-Insight: Teaching VLMs for AI-Generated Video Quality Understanding via Progressive Visual Reinforcement Learning cites this paper.

VQ-Insight: Teaching VLMs for AI-Generated Video Quality Understanding via Progressive Visual Reinforcement Learning LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:49.304243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:49.304243Z digest=sha256:f8a89b65307948f493d8a6074c8fe84c0be7b0e539d42b2ac42f68d9ce9a94b8

Observation c6b52557-2d14-4cfb-b891-2bb68efe2717 · inbound

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation cites this paper.

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:49.370416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:49.370416Z digest=sha256:ba7eefc30d7deebc11e69d025bc29f287cbe819ec66b40921a6a69306192ca23

Observation 7223992c-bdde-4e4c-92d6-60682d147d0b · inbound

RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling cites this paper.

RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:15:54.437321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:13:42.934115Z digest=sha256:6558dc768957d4027b8634842812180037771e777ab7155675e12b720bbe6db2

Observation 3f4938c0-35dd-41a5-9dac-d47cbd1d8b98 · inbound

PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models cites this paper.

PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:00:27.260501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T17:57:57.263574Z digest=sha256:e955c5a7406167bf6694f1b2f8243c653c17ddd3744b2e9ef4fdf2d28186bae0

Observation dc25a1a6-9e7f-4971-a8d6-4e92b8c3e659 · inbound

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models cites this paper.

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:28:05.557128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T16:25:03.743594Z digest=sha256:7d8dd093f2a73b378d17f31fee19c12a274b9902a6da546155196e78196a5b39

Observation 30d9a368-faee-43d0-a259-42251d648c52 · inbound

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models cites this paper.

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-21T16:04:14.695189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T16:01:52.150950Z digest=sha256:4fb1eb482dffcb9830c24966764b8a743f196876ba87640c3c09bdc84fe36ff8

Observation b0eead6d-ecf0-4ee5-91ca-cdc3b2f7d1b6 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:44.499036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:44.499036Z digest=sha256:56301512b82d1c9886251c93c8d5284a59400b5ba7557c80077f7d60ca0dedeb

Observation 40da0efd-d045-4131-9956-2a02ba3bc2e7 · inbound

Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling cites this paper.

Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:42:30.755627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T07:37:52.346280Z digest=sha256:794e71eb185acab1f89ae6674a015511a2daa895f32931919454eb6371e4aabe

Observation c7b756ed-008d-4e65-9680-f1c02f06a5ca · inbound

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating cites this paper.

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:22.083007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:00:31.582714Z digest=sha256:568a78e6930779654422ec273787d42ba009bef137d1973254b921b6608989a8

Observation 4a59fe8f-2958-4793-829f-2ccdbf7ac32a · inbound

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating cites this paper.

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:55:45.202918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T22:38:43.102769Z digest=sha256:4aa6931644aca735d47dc998b51f58ed60c6df0c002069c7276347b5a89d606d

Observation e72140b6-6d2d-4569-9803-5c41a7bb1ea0 · inbound

DRM: Diffusion-based Reward Model With Step-wise Guidance cites this paper.

DRM: Diffusion-based Reward Model With Step-wise Guidance LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.429079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:12:39.225557Z digest=sha256:38575ccf36bbfd8f998e79f714837463e2e2e369990292a6dec87f3d7d1fd371

Observation 1b0fc3d7-38f8-46a3-97a6-ea35b1e814cd · inbound

Through the PRISM: Preference Representation in Intermediate States of Video Diffusion Models cites this paper.

Through the PRISM: Preference Representation in Intermediate States of Video Diffusion Models LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:09:29.417783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T18:25:49.041639Z digest=sha256:a01844e4b6c5e68fef2648e5d579c279a5aa8c2337c61a4fc2a0bb9e9ec34f44

Observation e06c7cb8-e6e0-4222-94cc-022811bc940e · inbound

Reward Lightning: Fast Video Generation via Homologous Preference Distillation cites this paper.

Reward Lightning: Fast Video Generation via Homologous Preference Distillation LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-11T22:44:24.796514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:44:24.796514Z digest=sha256:8bcd1da81c77e3c8e0a56782b70703ca2ca9c8b5958e94f7d8933f25a4cd1590

Observation 547c4f4d-b0eb-40c6-8583-dadb4b70a32b · inbound

Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion cites this paper.

Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T19:09:45.675824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T19:09:45.675824Z digest=sha256:aceeb14a60922392a59c3229654bc54987a52fadb4fdbf138f94cd915f500677