Pith. sign in

Paper Citation Record · LEDGER

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation

As of 10 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2506.11144.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11144 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:55:28.944960Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:06:47.301483Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T22:20:22.357042Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fd8d86da-6153-4bef-a97b-4378a606b53a · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation A general theoretical paradigm to understand learning from human preferences

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.694299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.694299Z digest=sha256:21bdda038f1ab86337788f7690da54be1bef819c1710a0626d411fbc856313f2

Observation 1fa5b9a3-079e-49ef-af72-95a0684e3a38 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.699706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.699706Z digest=sha256:4beb6992db308e105de13218f4da5d677779ed63424ca6f13cf2c54f142695b2

Observation 59b2f9db-ad79-4432-bd5b-8233d8d2d7e9 · outbound

This paper cites SkyReels-V2: Infinite-length Film Generative Model.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation SkyReels-V2: Infinite-length Film Generative Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.705104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.705104Z digest=sha256:164d8bb90dcf0804d0ac10705009bfd7cf53b21cda31a19d78f3175f50338080

Observation 4b8ab91f-1943-4a55-96c5-781a00fd4a7b · outbound

This paper cites Echomimic: Lifelike audio- driven portrait animations through editable landmark conditions.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Echomimic: Lifelike audio- driven portrait animations through editable landmark conditions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:55:30.991079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:55:28.710520Z digest=sha256:aa82829de2c34a4b131f193a125479af3505d24d2984a2f034c7d88a013ee737

Observation 31c0ca70-889b-4879-8337-4a1f37d8e2dd · outbound

This paper cites Out of time: automated lip sync in the wild.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Out of time: automated lip sync in the wild

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.715503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.715503Z digest=sha256:f606cee079a5108815cb5a304e179bf3a45eadf619797f0148094f6b23fc2df5

Observation 8fa74461-f8fc-4e81-820f-7b24069f5283 · outbound

This paper cites The Llama 3 Herd of Models.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.720484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.720484Z digest=sha256:e1aee47e00cbe7f1260b63f73ca745d1e7aba31a65675d5f8036d7a7b9c85ba9

Observation ccbba998-7073-4628-b767-729d3c09c679 · outbound

This paper cites VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.726179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.726179Z digest=sha256:f2626539c726a457581d0332b9e8675ca365dee7a0372c40955ce39617c3bb6e

Observation 332b2d9a-3fe5-4db6-b829-f34eee398ab6 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.731300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.731300Z digest=sha256:99c2cb87d24add2b7bdfd511dec860b3efc5f9219a7eb2c262f654a893899512

Observation 0a64a6aa-ae2a-4e2d-ac61-1cd4ad56309a · outbound

This paper cites Diffted: One-shot audio- driven ted talk video generation with diffusion-based co-speech gestures.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Diffted: One-shot audio- driven ted talk video generation with diffusion-based co-speech gestures

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:55:30.764884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:55:28.736426Z digest=sha256:05a113c859756bee26c892d772960122ad683ffea1da61d2e4a067f291f02f50

Observation f3cb1880-620a-49d7-abef-6e4645e6b613 · outbound

This paper cites Loopy: Taming audio-driven portrait avatar with long-term motion dependency.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Loopy: Taming audio-driven portrait avatar with long-term motion dependency

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:55:30.549854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:55:28.741288Z digest=sha256:e4ac7966a60d36fb70b4c776a2f3ce2ecbaf287236278b1c3e473f52982fe845

Observation 4de104e3-24dc-47c3-9f22-8e2c4e85cdc7 · outbound

This paper cites Elucidating the design space of diffusion-based generative models.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Elucidating the design space of diffusion-based generative models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.746899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.746899Z digest=sha256:e7c66b4eb5d75a6b396f70899018ad4655568f553dd03d59f382ba15d723ed0d

Observation 14601b1c-aea7-419d-99d0-34d6dd353218 · outbound

This paper cites Auto-encoding variational bayes, 2013.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Auto-encoding variational bayes, 2013

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.751488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.751488Z digest=sha256:b08c6f916ba241b302f9c0828ed62b18ff648ba377edb2697e4de92ba66fb384

Observation 40ef71f1-9dc5-4609-b324-a1291948ba06 · outbound

This paper cites OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.756180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.756180Z digest=sha256:62e5bfff01b24e96b0eb0917d1d33b11b262f689d6ed3903713d4c4f533d1081

Observation fd030be6-b309-445a-b29d-afe56d43af43 · outbound

This paper cites T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.761207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.761207Z digest=sha256:db736bca84e513951ea1f6277f2cad564fe9b0794c00b66c26f8c7ad7fec4ebf

Observation 55c403f1-9e1d-4dc4-8708-d78ecdbdda0d · outbound

This paper cites Cyberhost: A one-stage diffusion framework for audio-driven talking body generation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Cyberhost: A one-stage diffusion framework for audio-driven talking body generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:55:30.352728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:55:28.766234Z digest=sha256:0dddeebda8fddb7f7177d4bc7fa8553ac363aa5ca37444a5a33181afeabca36c

Observation 6e47c66c-d5ac-4cfb-81e4-cd383f6c223d · outbound

This paper cites OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.771375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.771375Z digest=sha256:2a2e10c33c26529ac59e639be88f1fe7d5a55e61424f45189543f897805b26da

Observation 3ec72267-a68f-48f7-bd25-d5955335aa6a · outbound

This paper cites Flow Matching for Generative Modeling.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Flow Matching for Generative Modeling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.776489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.776489Z digest=sha256:1097fcd8bb01987e29bd05ca7284649c1648bb43c72327509064c3590a8925c2

Observation a37f8ded-f857-4f57-9170-1f9ee3402428 · outbound

This paper cites Improving Video Generation with Human Feedback.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Improving Video Generation with Human Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.781462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.781462Z digest=sha256:97d3ecd7e0c05c670eebb99fbd49286609eacea92b850dd655f352455633ff66

Observation b3b91f74-c925-4196-a87f-0d3bf56f71de · outbound

This paper cites VideoDPO: Omni-Preference Alignment for Video Diffusion Generation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation VideoDPO: Omni-Preference Alignment for Video Diffusion Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.786930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.786930Z digest=sha256:1fbdefdc562896f78f4eefa388f9a31836dff6e2cbf738f4486b5dac1c0540df

Observation eaee10a4-b814-4bed-be3d-ab01b6ab6454 · outbound

This paper cites Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.791865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.791865Z digest=sha256:13c90bee8683b5f12eadf8203c52498436081a369be2c1c4c762b3b5a533a5d1

Observation 67473cb3-1792-452b-b3c9-f9add8a36d7b · outbound

This paper cites OpenELM: An Efficient Language Model Family with Open Training and Inference Framework.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation OpenELM: An Efficient Language Model Family with Open Training and Inference Framework

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.796700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.796700Z digest=sha256:fdcd3828e61fb8d156b8e461c3446955d644b4cc8cf8a9ea0394abc966c50217

Observation 5df5be53-4302-4b5e-aa28-5d2486ae2e29 · outbound

This paper cites Echomimicv2: Towards striking, simplified, and semi-body human animation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Echomimicv2: Towards striking, simplified, and semi-body human animation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.801724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.801724Z digest=sha256:04bcb610f1246c82d9931e91f0df32a4ca1d286d6341c4d6e9be3426a1387f62

Observation 55458c22-2b77-4943-ab82-be5c5862081f · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Simpo: Simple preference optimization with a reference-free reward

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.806487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.806487Z digest=sha256:d0564b3cddd7020b3b8b20c9b5cb839d64095ba808ffb6e11c42ebaba95c2536

Observation 560225d4-646b-4d12-8ba8-9bf455d38ca6 · outbound

This paper cites Clip-dpo: Vision-language models as a source of preference for fixing hallucinations in lvlms.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Clip-dpo: Vision-language models as a source of preference for fixing hallucinations in lvlms

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:55:30.211552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:55:28.811290Z digest=sha256:5db22ee4587ffb46bc06f4d2151ffa636aee98eaba5483f4d737a2fe2003273b

Observation d0e261d7-f504-4929-bf90-a71fd56def66 · outbound

This paper cites Scalable diffusion models with transformers.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Scalable diffusion models with transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.816015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.816015Z digest=sha256:dc3b6d9694c235b7a447c30034d72376dfcea66fce7f8136435a20cabb2e0c90

Observation f412f38c-0a59-4fe2-a4e8-94f7499ef526 · outbound

This paper cites SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.820779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.820779Z digest=sha256:33a32929f9b5cff3505011c656521fd0839423fb9bc2e9166b7ff139dbedce92

Observation 2f1fd874-d2a5-4f25-aaa8-6e7570cbc331 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Direct preference optimization: Your language model is secretly a reward model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.826201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.826201Z digest=sha256:a75d7805068575add6946f1da1407ccecaa9ba9273a44906a5a8db8ccc3cd160

Observation 53c8f23a-562d-4125-af33-9fd5bf8e15ed · outbound

This paper cites Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.831101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.831101Z digest=sha256:ce15d977bde18aa89603a375ab92d8f4299584749b5ed7ff580c1dbbb8357f53

Observation 84fdc69f-6180-43d1-82f3-e60177f829ef · outbound

This paper cites First order motion model for image animation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation First order motion model for image animation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.836232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.836232Z digest=sha256:74e781990e13ad1c88ab313496ec935f818fb3a5699869744ae8b25d5a23631b

Observation eb78e81a-9ec2-45fa-aa05-8f363e9ed10a · outbound

This paper cites Motion repre- sentations for articulated animation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Motion repre- sentations for articulated animation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:55:30.020083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:55:28.841312Z digest=sha256:8b6a1933eaee4f5682e0ad5d51ca2c02b21cd80a4e57895deeabf37795d28b3a

Observation c0e90711-c221-422b-9c4c-c10a9716e5ee · outbound

This paper cites Learning to summarize with human feedback.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Learning to summarize with human feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.846144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.846144Z digest=sha256:980c2c033b0178326a668b1003af44bfc810bb5a23754a197baa465a13fb4166

Observation 71ee290a-41c2-40a6-a35c-52be9b948044 · outbound

This paper cites EMO2: End-Effector Guided Audio-Driven Avatar Video Generation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.850843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.850843Z digest=sha256:94e764e1c3a6fb1acfb5aca3c0f9283efb886806c1864ba22e306999e83b90fb

Observation 3317897e-905e-4580-b747-3b0fcb39c634 · outbound

This paper cites Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.855934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.855934Z digest=sha256:919cd3dcb2130175ee3a6401f6f7c1470e51e0bafe7e4d1a4db20ef15ffa7bdd

Observation 3b28877d-56fe-4a91-9b86-b40291324f9c · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.860592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.860592Z digest=sha256:194ddb1fee855ea3705c17709567ee5b5e282e394570132226178540e34ef8a7

Observation 812fa1c1-c70e-4371-a564-bfb6b70afd5c · outbound

This paper cites Fvd: A new metric for video generation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Fvd: A new metric for video generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.865734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.865734Z digest=sha256:499f2f7353d34ad9703f202976c2ac3e5d6f9ef3f2b464a43bd83aa1a3793614

Observation 9de4a076-298c-4f18-845f-79c9c37e7b38 · outbound

This paper cites Diffusion model alignment using direct preference optimization.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Diffusion model alignment using direct preference optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.870532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.870532Z digest=sha256:ca7e447838d2dcd770aa7ffaa6f4049bb8666cba8ddc25550109dfc0c5292b64

Observation ab4b4deb-7c2c-47d2-8177-452149814295 · outbound

This paper cites FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.875127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.875127Z digest=sha256:ba8496855bc764e054ddc8a0b851b106467300e70df11b89a32d385132483eb8

Observation 9c039a11-5378-406b-8f97-09dfc2d9857c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.880281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.880281Z digest=sha256:dd98713ae36195ae16eb317422d14c90a9a01cd1a08020cb1fe4bf9ccb3fce8d

Observation 66d15514-4ce7-4003-b194-d045990134d5 · outbound

This paper cites Video-to-Video Synthesis.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Video-to-Video Synthesis

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.884977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.884977Z digest=sha256:bde6fdd395543ddf1c4086d4717239814ef68f8ba55c9cf12de1009a973e1cea

Observation 31529f22-2b38-45ac-bb66-f124081eb677 · outbound

This paper cites MoCha: Towards Movie-Grade Talking Character Synthesis.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.890139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.890139Z digest=sha256:dcbcb81d12ed5f45012c59865bc418677cca47d92f8495b38f76ae1100392806

Observation 3623e821-9b74-4eb0-8bca-2c23fbf8d40c · outbound

This paper cites Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.894940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.894940Z digest=sha256:3c5eb4f0c5c1e070867ee6f8576a70a8c5e2216529e14771aa109d313a6509e5

Observation b96501ee-2d9b-4a1d-b447-6e00bd365e02 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.900227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.900227Z digest=sha256:6ee24c31b5affaae941e2fb4d28b87685b10d2492f3670835647470eb5da363d

Observation 02dfe58c-2446-4726-9098-ea9aa4e97b5b · outbound

This paper cites VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.905535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.905535Z digest=sha256:27699a9c022fd84f9b5c6fab1f612639a0b95fc6fb6a314d180a7056ea5d33bf

Observation 973dce22-a8ec-4689-adcc-bebf3488a4f9 · outbound

This paper cites Imagereward: Learning and evaluating human preferences for text-to-image generation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Imagereward: Learning and evaluating human preferences for text-to-image generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.910579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.910579Z digest=sha256:8d55cd9f4c3bfd4306b8930850df843ba703eb978fe36c062dcd3a0cc8079540

Observation f8b87c68-d613-4395-a5e6-d84a82c53b51 · outbound

This paper cites Qwen2.5-1M Technical Report.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Qwen2.5-1M Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.915211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.915211Z digest=sha256:fcf1bd8e9aadda0748d4f4db339207e2db500c4b1357fac6e5453099fe90031a

Observation 38fb32d1-51cb-41d7-a3e7-27574f913fd6 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.919921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.919921Z digest=sha256:9fa69b18d9de4c38f6287054138bd9eb2e90840ce8e4700071458ad0a6ec718e

Observation 91a34f10-0f46-42b2-b700-0a2d7e954b6b · outbound

This paper cites MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.924931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.924931Z digest=sha256:e288e988b1ab3a4ffaf60d2bba558bc947881dfe5889a35bbd8074285dec1ab4

Observation 43bb5afe-d99e-421f-b48d-a980d648cfd6 · outbound

This paper cites MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.929697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.929697Z digest=sha256:bb0862457a34cea5062a7cc30dec7d63dc45f2b739f1ddba5c5c2a6cc7aafce8

Observation cf331162-f1b4-4f4e-a36a-2d3f7c10ba00 · outbound

This paper cites Learning to compare for better training and evaluation of open domain natural language generation models.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Learning to compare for better training and evaluation of open domain natural language generation models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:55:29.765277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:55:28.935045Z digest=sha256:f00e3d80686459dfaf39d71e34c571455d1f646c9639392e76bc42af59311a28

Observation 81a56989-01fb-48eb-b3bd-a3f7b69ee758 · outbound

This paper cites Taming diffusion models for audio-driven co-speech gesture generation.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Taming diffusion models for audio-driven co-speech gesture generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:28.940019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:28.940019Z digest=sha256:493174a804b330a1b7358ab528f1e8783e5f16a7de29cd5fbd1c568e2c1e89b9

Observation 4b5ee12e-bb9b-47aa-84d2-20ad52b4d177 · outbound

This paper cites Vlogger: Make your dream a vlog.

AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation Vlogger: Make your dream a vlog

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:55:29.620891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:55:28.944960Z digest=sha256:615aaad5e362d51820c4fe20f8460344e4b2d133c566341aa7b39251afc7c28d

Pith citing papers

Observation 80b09f0d-f1bb-4a0f-88ba-5793b7779ac2 · inbound

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation cites this paper.

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:47.301483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:47.301483Z digest=sha256:def29c05727a767258559369ce647579bd26c31c4757ffd4d03b60a6e96d53bc

Observation 2220945c-701f-4271-b0de-e9a73bb81630 · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:20:22.360892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:d7676d83c35f6ebfff1e3984a21e608ac65d2626f4abc232bef1062936b39718