Pith. sign in

Paper Citation Record · LEDGER

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation

As of 21 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 3 inbound Pith citation observations for arXiv:2506.11144.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11144 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:55:28.944960Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:57:29.607694Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T22:20:22.357042Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fd8d86da-6153-4bef-a97b-4378a606b53a · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation A general theoretical paradigm to understand learning from human preferences

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.694299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.694299Z digest=sha256:21bdda038f1ab86337788f7690da54be1bef819c1710a0626d411fbc856313f2

Observation 1fa5b9a3-079e-49ef-af72-95a0684e3a38 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.699706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.699706Z digest=sha256:4beb6992db308e105de13218f4da5d677779ed63424ca6f13cf2c54f142695b2

Observation 59b2f9db-ad79-4432-bd5b-8233d8d2d7e9 · outbound

This paper cites SkyReels-V2: Infinite-length Film Generative Model.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation SkyReels-V2: Infinite-length Film Generative Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.705104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.705104Z digest=sha256:56d8e969b5c7440f1c7281c7dbf7b52a81b55626427ceb2fdd7629ea9feb51dd

Observation 4b8ab91f-1943-4a55-96c5-781a00fd4a7b · outbound

This paper cites Echomimic: Lifelike audio- driven portrait animations through editable landmark conditions.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Echomimic: Lifelike audio- driven portrait animations through editable landmark conditions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:55:30.991079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:55:28.710520Z digest=sha256:972c2b9552a5ea62722a2cb17dfad95be0530d450e684ad9b7cf08f983771225

Observation 31c0ca70-889b-4879-8337-4a1f37d8e2dd · outbound

This paper cites Out of time: automated lip sync in the wild.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Out of time: automated lip sync in the wild

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.715503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.715503Z digest=sha256:f606cee079a5108815cb5a304e179bf3a45eadf619797f0148094f6b23fc2df5

Observation 8fa74461-f8fc-4e81-820f-7b24069f5283 · outbound

This paper cites The Llama 3 Herd of Models.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.720484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.720484Z digest=sha256:e52d0617846dcf599be82230a8894e82f9561db8d1e977f6de04c9631c78f57d

Observation ccbba998-7073-4628-b767-729d3c09c679 · outbound

This paper cites VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.726179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.726179Z digest=sha256:8cb9106d60fcd58af180163fe6e165e2015dc71208819450a858bb071eb41da3

Observation 332b2d9a-3fe5-4db6-b829-f34eee398ab6 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.731300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.731300Z digest=sha256:99c2cb87d24add2b7bdfd511dec860b3efc5f9219a7eb2c262f654a893899512

Observation 0a64a6aa-ae2a-4e2d-ac61-1cd4ad56309a · outbound

This paper cites Diffted: One-shot audio- driven ted talk video generation with diffusion-based co-speech gestures.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Diffted: One-shot audio- driven ted talk video generation with diffusion-based co-speech gestures

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:55:30.764884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:55:28.736426Z digest=sha256:f27f555db573bc26e0a322bbb8ca8e4fc9c2ba6ae2ad2c4eebc41c15dbac6061

Observation f3cb1880-620a-49d7-abef-6e4645e6b613 · outbound

This paper cites Loopy: Taming audio-driven portrait avatar with long-term motion dependency.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Loopy: Taming audio-driven portrait avatar with long-term motion dependency

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:55:30.549854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:55:28.741288Z digest=sha256:5f88c1d922ae7f40fd39bfcee3e3f1c1fb449ca0556cbd5b0c181928fee8fab7

Observation 4de104e3-24dc-47c3-9f22-8e2c4e85cdc7 · outbound

This paper cites Elucidating the design space of diffusion-based generative models.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Elucidating the design space of diffusion-based generative models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.746899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.746899Z digest=sha256:e7c66b4eb5d75a6b396f70899018ad4655568f553dd03d59f382ba15d723ed0d

Observation 14601b1c-aea7-419d-99d0-34d6dd353218 · outbound

This paper cites Auto-encoding variational bayes, 2013.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Auto-encoding variational bayes, 2013

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.751488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.751488Z digest=sha256:b08c6f916ba241b302f9c0828ed62b18ff648ba377edb2697e4de92ba66fb384

Observation 40ef71f1-9dc5-4609-b324-a1291948ba06 · outbound

This paper cites OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.756180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.756180Z digest=sha256:74324ad9944fcf016575d69321f1b285a2d2f29ed671868446710ad568b9ba1b

Observation fd030be6-b309-445a-b29d-afe56d43af43 · outbound

This paper cites T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.761207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.761207Z digest=sha256:f795cad5c27d3c845a3ab1b0033c625efe18e059cfee5f6ce22e2061c58345ba

Observation 55c403f1-9e1d-4dc4-8708-d78ecdbdda0d · outbound

This paper cites Cyberhost: A one-stage diffusion framework for audio-driven talking body generation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Cyberhost: A one-stage diffusion framework for audio-driven talking body generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:55:30.352728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:55:28.766234Z digest=sha256:f597ea9527192c0fbc1c10784ec875ef59858000aa7c4bb937b29e9f6f59608f

Observation 6e47c66c-d5ac-4cfb-81e4-cd383f6c223d · outbound

This paper cites OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.771375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.771375Z digest=sha256:5111a071cb943aa931f8d33e1fe56d907caa580a80742850b0d1cadf830e1ce4

Observation 3ec72267-a68f-48f7-bd25-d5955335aa6a · outbound

This paper cites Flow Matching for Generative Modeling.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Flow Matching for Generative Modeling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.776489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.776489Z digest=sha256:6bc5d40c6128040e3bb088b0705343ed6bbd3782277a4f8ab5cad9f8524a7fa8

Observation a37f8ded-f857-4f57-9170-1f9ee3402428 · outbound

This paper cites Improving Video Generation with Human Feedback.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Improving Video Generation with Human Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.781462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.781462Z digest=sha256:97d3ecd7e0c05c670eebb99fbd49286609eacea92b850dd655f352455633ff66

Observation b3b91f74-c925-4196-a87f-0d3bf56f71de · outbound

This paper cites VideoDPO: Omni-Preference Alignment for Video Diffusion Generation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation VideoDPO: Omni-Preference Alignment for Video Diffusion Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.786930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.786930Z digest=sha256:e8b857849f845d229d9e7a5f2bcf9ea2d362a4a14b4c2060913821ad97256212

Observation eaee10a4-b814-4bed-be3d-ab01b6ab6454 · outbound

This paper cites Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.791865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.791865Z digest=sha256:13c90bee8683b5f12eadf8203c52498436081a369be2c1c4c762b3b5a533a5d1

Observation 67473cb3-1792-452b-b3c9-f9add8a36d7b · outbound

This paper cites OpenELM: An Efficient Language Model Family with Open Training and Inference Framework.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation OpenELM: An Efficient Language Model Family with Open Training and Inference Framework

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.796700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.796700Z digest=sha256:c7ceb888c8eb501bab2712251ff80cec475d711dac9a6c456e7a551b86b77785

Observation 5df5be53-4302-4b5e-aa28-5d2486ae2e29 · outbound

This paper cites Echomimicv2: Towards striking, simplified, and semi-body human animation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Echomimicv2: Towards striking, simplified, and semi-body human animation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.801724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.801724Z digest=sha256:04bcb610f1246c82d9931e91f0df32a4ca1d286d6341c4d6e9be3426a1387f62

Observation 55458c22-2b77-4943-ab82-be5c5862081f · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Simpo: Simple preference optimization with a reference-free reward

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.806487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.806487Z digest=sha256:d0564b3cddd7020b3b8b20c9b5cb839d64095ba808ffb6e11c42ebaba95c2536

Observation 560225d4-646b-4d12-8ba8-9bf455d38ca6 · outbound

This paper cites Clip-dpo: Vision-language models as a source of preference for fixing hallucinations in lvlms.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Clip-dpo: Vision-language models as a source of preference for fixing hallucinations in lvlms

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:55:30.211552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:55:28.811290Z digest=sha256:4857a083f4d756b05c508273445148adb391115057ba4501e92c9ac13403ed9b

Observation d0e261d7-f504-4929-bf90-a71fd56def66 · outbound

This paper cites Scalable diffusion models with transformers.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Scalable diffusion models with transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.816015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.816015Z digest=sha256:dc3b6d9694c235b7a447c30034d72376dfcea66fce7f8136435a20cabb2e0c90

Observation f412f38c-0a59-4fe2-a4e8-94f7499ef526 · outbound

This paper cites SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.820779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.820779Z digest=sha256:91a99a69a08ec3e62bfd6efe3636ee450967e3643fd790d5c2383ac0661880b8

Observation 2f1fd874-d2a5-4f25-aaa8-6e7570cbc331 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Direct preference optimization: Your language model is secretly a reward model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.826201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.826201Z digest=sha256:a75d7805068575add6946f1da1407ccecaa9ba9273a44906a5a8db8ccc3cd160

Observation 53c8f23a-562d-4125-af33-9fd5bf8e15ed · outbound

This paper cites Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.831101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.831101Z digest=sha256:6c15927b0074e7e13c8c78fe398378eea31273428c4bd5fd077aa253ca314ed6

Observation 84fdc69f-6180-43d1-82f3-e60177f829ef · outbound

This paper cites First order motion model for image animation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation First order motion model for image animation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.836232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.836232Z digest=sha256:74e781990e13ad1c88ab313496ec935f818fb3a5699869744ae8b25d5a23631b

Observation eb78e81a-9ec2-45fa-aa05-8f363e9ed10a · outbound

This paper cites Motion repre- sentations for articulated animation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Motion repre- sentations for articulated animation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:55:30.020083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:55:28.841312Z digest=sha256:20b69c0a7e660f7e9b02bfb44c771c5c2549cbcd121f6432512d8826d2548926

Observation c0e90711-c221-422b-9c4c-c10a9716e5ee · outbound

This paper cites Learning to summarize with human feedback.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Learning to summarize with human feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.846144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.846144Z digest=sha256:980c2c033b0178326a668b1003af44bfc810bb5a23754a197baa465a13fb4166

Observation 71ee290a-41c2-40a6-a35c-52be9b948044 · outbound

This paper cites EMO2: End-Effector Guided Audio-Driven Avatar Video Generation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.850843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.850843Z digest=sha256:a60da420d8a55e2b808789446d66b4bdff41e31b14562e95dd64ad9974a4c8db

Observation 3317897e-905e-4580-b747-3b0fcb39c634 · outbound

This paper cites Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.855934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.855934Z digest=sha256:919cd3dcb2130175ee3a6401f6f7c1470e51e0bafe7e4d1a4db20ef15ffa7bdd

Observation 3b28877d-56fe-4a91-9b86-b40291324f9c · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.860592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.860592Z digest=sha256:194ddb1fee855ea3705c17709567ee5b5e282e394570132226178540e34ef8a7

Observation 812fa1c1-c70e-4371-a564-bfb6b70afd5c · outbound

This paper cites Fvd: A new metric for video generation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Fvd: A new metric for video generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.865734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.865734Z digest=sha256:499f2f7353d34ad9703f202976c2ac3e5d6f9ef3f2b464a43bd83aa1a3793614

Observation 9de4a076-298c-4f18-845f-79c9c37e7b38 · outbound

This paper cites Diffusion model alignment using direct preference optimization.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Diffusion model alignment using direct preference optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.870532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.870532Z digest=sha256:ca7e447838d2dcd770aa7ffaa6f4049bb8666cba8ddc25550109dfc0c5292b64

Observation ab4b4deb-7c2c-47d2-8177-452149814295 · outbound

This paper cites FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.875127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.875127Z digest=sha256:340cd5a95dc34ac776a4d7bfd61112270ec8f62d8165632abf5c3ef1822d7675

Observation 9c039a11-5378-406b-8f97-09dfc2d9857c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.880281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.880281Z digest=sha256:dd98713ae36195ae16eb317422d14c90a9a01cd1a08020cb1fe4bf9ccb3fce8d

Observation 66d15514-4ce7-4003-b194-d045990134d5 · outbound

This paper cites Video-to-Video Synthesis.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Video-to-Video Synthesis

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.884977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.884977Z digest=sha256:4659cbd04b7e10f54ec84f93881e6db88a4986c00598897d8034811d77f567e2

Observation 31529f22-2b38-45ac-bb66-f124081eb677 · outbound

This paper cites MoCha: Towards Movie-Grade Talking Character Synthesis.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.890139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.890139Z digest=sha256:231ac15eaf8fa4295e76992cc39b180b2d0f85bf658e33f829960e67bb6d490a

Observation 3623e821-9b74-4eb0-8bca-2c23fbf8d40c · outbound

This paper cites Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.894940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.894940Z digest=sha256:3c5eb4f0c5c1e070867ee6f8576a70a8c5e2216529e14771aa109d313a6509e5

Observation b96501ee-2d9b-4a1d-b447-6e00bd365e02 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.900227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.900227Z digest=sha256:f01effe8e8398e0cae22133e32b5a394f727d5e1310e8d8765782f3432f39e89

Observation 02dfe58c-2446-4726-9098-ea9aa4e97b5b · outbound

This paper cites VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.905535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.905535Z digest=sha256:05cd93ece005bcf269e4bd963f9290a41781b92d70f693b2fa53e05716079c62

Observation 973dce22-a8ec-4689-adcc-bebf3488a4f9 · outbound

This paper cites Imagereward: Learning and evaluating human preferences for text-to-image generation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Imagereward: Learning and evaluating human preferences for text-to-image generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.910579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.910579Z digest=sha256:8d55cd9f4c3bfd4306b8930850df843ba703eb978fe36c062dcd3a0cc8079540

Observation f8b87c68-d613-4395-a5e6-d84a82c53b51 · outbound

This paper cites Qwen2.5-1M Technical Report.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Qwen2.5-1M Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.915211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.915211Z digest=sha256:361fb6beb73cecfa5f523d401dd81f85dac3dd2309c4a7b6ac2b09210529151f

Observation 38fb32d1-51cb-41d7-a3e7-27574f913fd6 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.919921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.919921Z digest=sha256:9fa69b18d9de4c38f6287054138bd9eb2e90840ce8e4700071458ad0a6ec718e

Observation 91a34f10-0f46-42b2-b700-0a2d7e954b6b · outbound

This paper cites MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.924931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.924931Z digest=sha256:7ae963bd55d23b922fa2b108903da462b60c6b992a5daf7029b91ef65ea1ba71

Observation 43bb5afe-d99e-421f-b48d-a980d648cfd6 · outbound

This paper cites MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.929697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.929697Z digest=sha256:ed6f7ae742a0d5898045c4c32a3ea635ec634c37d87ef0aa7e8163d0f9be1d22

Observation cf331162-f1b4-4f4e-a36a-2d3f7c10ba00 · outbound

This paper cites Learning to compare for better training and evaluation of open domain natural language generation models.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Learning to compare for better training and evaluation of open domain natural language generation models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:55:29.765277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:55:28.935045Z digest=sha256:24646656b555ac4b84d98017f20724ff6126a9df597256bc7875885e67a0cc77

Observation 81a56989-01fb-48eb-b3bd-a3f7b69ee758 · outbound

This paper cites Taming diffusion models for audio-driven co-speech gesture generation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Taming diffusion models for audio-driven co-speech gesture generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.940019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.940019Z digest=sha256:493174a804b330a1b7358ab528f1e8783e5f16a7de29cd5fbd1c568e2c1e89b9

Observation 4b5ee12e-bb9b-47aa-84d2-20ad52b4d177 · outbound

This paper cites Vlogger: Make your dream a vlog.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Vlogger: Make your dream a vlog

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:55:29.620891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:55:28.944960Z digest=sha256:33f0040d7b452204318806b842c020d98ed32d0916536e8f662fb67ff4286a23

Pith citing papers

Observation 80b09f0d-f1bb-4a0f-88ba-5793b7779ac2 · inbound

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation cites this paper.

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:47.301483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:47.301483Z digest=sha256:215093977545c0a90b1dddbe976d900c3bda72d7a9c874bc64ab11cbbab70d0d

Observation 80078491-1f0c-4640-bab4-1ab7cf469d17 · inbound

OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation cites this paper.

OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T16:57:29.607694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:57:29.607694Z digest=sha256:9b16ce95abc97e16b91a0428adf5cabeb876f80bb30f9d176aa30069286a6878

Observation 2220945c-701f-4271-b0de-e9a73bb81630 · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:20:22.360892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:1d229b4ce25b9ce8a10b91758f44c4b0cdeb199386405a30e4cb009737d35fec