Pith. sign in

Paper Citation Record · LEDGER

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

As of 21 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2608.11752.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11752 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:34:44.921475Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e0a2bd94-66fd-41ab-82c9-92f5384b7982 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Emerg- ing properties in self-supervised vision transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.617921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.758406Z digest=sha256:5eac4f6eb0057a3e55291c4884c4e1bcdba58d30e9755f44e1ef40cec0590fde

Observation 131cc951-9f92-4247-b836-801728c96427 · outbound

This paper cites Diffusion forcing: Next-token prediction meets full-sequence diffu- sion.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Diffusion forcing: Next-token prediction meets full-sequence diffu- sion

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.608749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.762274Z digest=sha256:5c59ecac9214e75731a3291523adba07698fc9918c5e3c2f2bcbe006776e8865

Observation 80487082-2489-480d-aafb-69e303f873f6 · outbound

This paper cites Simswap: An efficient framework for high fidelity face swapping.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Simswap: An efficient framework for high fidelity face swapping

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.599587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.765808Z digest=sha256:a0913c8d4e6d0b382529a7b1a670f57206e9d2c842f14eff25eecb88d4c7fd46

Observation 4b5412f7-ae97-47f5-a7b4-c92da1f5e8f4 · outbound

This paper cites Wan-animate: Unified character animation and replacement with holistic replica- tion.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Wan-animate: Unified character animation and replacement with holistic replica- tion

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.769542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.769542Z digest=sha256:d64af122e88e6045646f662f033e6db432779524ca14e722eed083fca4b56022

Observation 41245199-c981-429f-b2e9-d543c56e4958 · outbound

This paper cites Out of time: Auto- mated lip sync in the wild.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Out of time: Auto- mated lip sync in the wild

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.590449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.772765Z digest=sha256:0f1355a7407bac97951bc51f711ea54553c0f5469bbaf7ae2451c8ffec09c566

Observation 4928abc3-0ffe-454e-b5c4-bfa568fe2e5a · outbound

This paper cites High fidelity neural audio compression.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos High fidelity neural audio compression

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.581075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.776021Z digest=sha256:03b6fb31def32185011641ccc21234a4cf618dbd8b722865dc902af5f2f7bee3

Observation 1e77cd2b-6f56-4497-96f8-6ca223385895 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.779593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.779593Z digest=sha256:c24fa7b3c0d4f5713c360a7c2b1526e37078e56bb8c3b7682849798bca2dcb65

Observation ce80ddc6-bd10-4157-8407-f312b4e86e03 · outbound

This paper cites Freeman, and Michael Rubinstein.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Freeman, and Michael Rubinstein

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.571614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.783139Z digest=sha256:4223dc831947d26d50f2f08b400dc6da45c666f54707f885484755f0d794c010

Observation ea42b9c8-20c1-4395-a209-de4a99b83469 · outbound

This paper cites Video Diffusion Transformers are In-Context Learners.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Video Diffusion Transformers are In-Context Learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.785950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.785950Z digest=sha256:a62c4de2f00acb711d830fb8ce164547f79563a6a04cf0b9b6c1f13f9248d573

Observation 13ae0261-a23d-4d0b-b3b5-7e85f30771a4 · outbound

This paper cites Infoswap: Information bottleneck disentanglement for identity swapping.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Infoswap: Information bottleneck disentanglement for identity swapping

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.561843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.789439Z digest=sha256:64db0a0829201b58819ce2dc137b744b4dc73694a48057ca6b2f5a88f0a23b3f

Observation d2593158-ce82-4d20-a5fd-2699eebc81c1 · outbound

This paper cites LTX-2: Efficient Joint Audio-Visual Foundation Model.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos LTX-2: Efficient Joint Audio-Visual Foundation Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.792702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.792702Z digest=sha256:924b65a96829ad0959a9ad5e82a11f4e28f630482a7d01b48fded58300d13195

Observation 84d79091-442b-4a65-9d08-8155f7354844 · outbound

This paper cites FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.795982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.795982Z digest=sha256:92dc2045ba54bc7a0aceff97e607d93bb220d5ef813fc21c073d75ff839aeee4

Observation ef0750c1-638e-483d-8947-82f86d380629 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.552326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.799943Z digest=sha256:c8a96c9a72937f1e96e20eaa3e1033500a02b9a568b72fac6d0465d8fba52eaf

Observation 1aad171a-0f5d-4b06-a868-d2a11f8e1f47 · outbound

This paper cites HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.802842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.802842Z digest=sha256:9feaa526676ebed100d97fddd4f181215b84f4c6246ff1953aac3dd96c51bd9e

Observation 45c9d40e-5cc9-42ac-aa28-f3ff1e0ebb9f · outbound

This paper cites In-Context LoRA for Diffusion Transformers.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos In-Context LoRA for Diffusion Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.806956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.806956Z digest=sha256:f2c9abb02222137a791af42a0f9b30d50ce4fc641bf41054eaa90cf4f5fdd100

Observation ca726dfa-3ba5-4a90-8266-6570709a4c23 · outbound

This paper cites Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.810403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.810403Z digest=sha256:9b11071db726073fa54a075a8045240385485baa72e9622d74a77f9d7e22b908

Observation 44c1d740-c104-4cd3-b809-a566076147bc · outbound

This paper cites REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.813743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.813743Z digest=sha256:711e6e9bc4dcaa1312ba3a4bc170100a669199098d26e0e94ba1264daaacdfab

Observation c48a0d8a-00f4-4601-bf67-0ed00ea735c8 · outbound

This paper cites VACE: All-in-One Video Creation and Editing.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos VACE: All-in-One Video Creation and Editing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.817988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.817988Z digest=sha256:e360b1517edf460b75f16a4bb75d265bc4cb3163ae6b2fcca4216b34a72d6573

Observation 6483e68a-ab90-48ca-88f1-d0ec1eea21fa · outbound

This paper cites Faceshifter: Towards high fidelity and occlusion aware face swapping.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Faceshifter: Towards high fidelity and occlusion aware face swapping

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.543317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.821578Z digest=sha256:2b7cb9743cb10fb4a910402287f0117cd23bbb5a4b41c9e16815a0f0f8ddbc4f

Observation 6515f221-9296-4b83-9bc6-6fba7e6d4e29 · outbound

This paper cites Rolling forcing: Autoregressive long video diffusion in real time.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Rolling forcing: Autoregressive long video diffusion in real time

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.533176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.824578Z digest=sha256:e65708c30998ed2a57a9d860d8c01de62251a0139698e8a6f5c80932de6a3502

Observation 1de3ba52-1241-4d3e-888c-d7004100fa8e · outbound

This paper cites JavisDiT: Joint audio-video diffusion trans- former with hierarchical spatio-temporal prior synchroniza- tion.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos JavisDiT: Joint audio-video diffusion trans- former with hierarchical spatio-temporal prior synchroniza- tion

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.521525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.827507Z digest=sha256:7bf051ab7205e1363bc25523eb0d9815c14a53063f69c1de95f2c3b505b2bb7e

Observation 02750d3b-f5d3-46e8-8522-9a31d5734a70 · outbound

This paper cites Zero-shot Voice Conversion with Diffusion Transformers.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Zero-shot Voice Conversion with Diffusion Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.830514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.830514Z digest=sha256:e427d6a95caccde011ba519d06d8370eb60a412257c105f739522a3f79e46b84

Observation f19a9fc5-caa5-49ba-9085-a4bbcfb44244 · outbound

This paper cites Scalable diffusion models with transformers.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Scalable diffusion models with transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.511740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.833696Z digest=sha256:b4170f71c80dc4a797d9efbcc59e3eb1e3b7862596f42c03b2444bb6c1ae4c75

Observation c0e569ce-d332-4391-a1f8-166049140720 · outbound

This paper cites OpenVoice: Versatile Instant Voice Cloning.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos OpenVoice: Versatile Instant Voice Cloning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.836725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.836725Z digest=sha256:97d45d3cd595b037d70c3d6d8a3ea322afc3346e41cf9653986e86c29d51cbd4

Observation 7fea2ab3-3036-4c28-b478-673e796c16bc · outbound

This paper cites Sam 2: Segment anything in images and videos.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Sam 2: Segment anything in images and videos

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.501215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.840178Z digest=sha256:2c122567dd4dccebfdafe0a1bb09e9205629e54dc7ff619f638fd33076ba3194

Observation 4b19e922-6906-4475-b323-9d24855ffcfd · outbound

This paper cites an unresolved cited work.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:34:45.491605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.843641Z digest=sha256:a13fcd8a0b1efbc9bff9fce04a90c5e4fdd9df64d74b3e8884b884263901bf19

Observation 53991683-2ff0-40f4-9991-609ffbaceea2 · outbound

This paper cites MM-Diffusion: Learning multi-modal diffusion models for joint audio and video generation.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos MM-Diffusion: Learning multi-modal diffusion models for joint audio and video generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.481419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.847306Z digest=sha256:2afc170ec1cfbbe75342bd5dafaa565f94fde84d8ac5acf540afcbf33a62be31

Observation 257e803d-5781-4ce5-9c0b-659b4e0d74de · outbound

This paper cites Consistency models.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Consistency models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.850720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.850720Z digest=sha256:d84c56259e4d6c4312c3500fbd5e8ec6a418ec31a9562ce43d389fec7cf11981

Observation 9fd3a51e-658e-4533-bf93-f464a06f43ef · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Roformer: Enhanced transformer with rotary position embedding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.466367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.854000Z digest=sha256:b3cb7ce3d1da67863f8fce6a2f12bb3c6005beafec48bdb95bb7331717dfcbce

Observation 99ab2679-559d-454a-a029-d90548a7a6b3 · outbound

This paper cites Omniforcing: Unleashing real-time joint audio- visual generation.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Omniforcing: Unleashing real-time joint audio- visual generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.857252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.857252Z digest=sha256:2a96f851079290fcf7167a6da0bda9ba64725873836feddb3154fab432278504

Observation 89528c1c-f894-4ead-a3dc-9ca6444e4839 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Wan: Open and Advanced Large-Scale Video Generative Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.860514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.860514Z digest=sha256:b752b30691ef2b8349943708e9b32a036d5a6198bab6971b33196791959249a0

Observation fe34031a-b2e8-43af-b62a-73bcf18a7490 · outbound

This paper cites Generalized end-to-end loss for speaker verification.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Generalized end-to-end loss for speaker verification

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.456924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.864312Z digest=sha256:b12b673abda7556280199ce1bd1c7449d46f47deb8e007bd780da8e9ce59631f

Observation 58307b3f-a855-4825-a056-cf32436d8b5c · outbound

This paper cites Bovik, Hamid R.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Bovik, Hamid R

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.447706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.867781Z digest=sha256:613f0df5c967f8e47f3a6c30351a87f778441e5313e769b10a6c575480eccbf8

Observation cf83fc06-4075-4d70-aa5f-d8ce9cba181e · outbound

This paper cites Williams and David Zipser.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Williams and David Zipser

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.438474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.871050Z digest=sha256:c5dd680fab67be0565ca0e46eb0d712bd2703e52c9e85a958add4ebde508ec31

Observation bb0d50fc-89bc-41b0-a29b-4eedf4bd61e7 · outbound

This paper cites Q-Align: Teaching LMMs for vi- sual scoring via discrete text-defined levels.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Q-Align: Teaching LMMs for vi- sual scoring via discrete text-defined levels

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.429404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.875183Z digest=sha256:6979331b4470f64eb6ef0a436903b6b24a57e1bc414637e99d4be69fbe410b2f

Observation 9a94d4ea-e773-4e0f-b6b3-5873782d1241 · outbound

This paper cites Efficient streaming language models with attention sinks.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Efficient streaming language models with attention sinks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.419609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.878806Z digest=sha256:25fcadc2e1bf97444f80fc3209d4a74cc046fae29fa30aad01cf79dec8804da5

Observation afb43cf0-ccee-4f73-add5-d82decaa9806 · outbound

This paper cites Vit- pose: Simple vision transformer baselines for human pose estimation.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Vit- pose: Simple vision transformer baselines for human pose estimation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.409786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.882068Z digest=sha256:9d902b1185ed8ad01d1b9d270ee27c762d90031af8e0bd7b6b66e3dc8547c081

Observation fd0df51c-5499-4fe8-b510-0302c302c4da · outbound

This paper cites Mocha: End-to-end video character re- placement without structural guidance.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Mocha: End-to-end video character re- placement without structural guidance

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.885408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.885408Z digest=sha256:5f251a13aaf64fda47bf233028bc2866bf418caafee870bfd51e077d349a3373

Observation 34d99a8a-e375-42f7-846f-178ac4403970 · outbound

This paper cites SCAIL-2: Unifying Controlled Character Animation with End-to-End In-Context Conditioning.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos SCAIL-2: Unifying Controlled Character Animation with End-to-End In-Context Conditioning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.888795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.888795Z digest=sha256:96b8ddd8c32fd557bf9b866ee9fe512127e8c2fd5cbdc557bcfa681842cb3718

Observation 075238d9-9d63-4afa-9e59-aabaf7e7e446 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:44.892364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:44.892364Z digest=sha256:7c0e8380ba4d87244af05c812d5ecac57f9edad2b112851e0616f3cc56118b3c

Observation 92a404c2-381d-4a3c-ba4d-5b6e4d836660 · outbound

This paper cites an unresolved cited work.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:34:45.399087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.895952Z digest=sha256:c837346644090a76240c16b89b043045589aec914b434061657889fa8646ed29

Observation ff9bd18d-c2e3-4c90-b0a7-b8a03b906b33 · outbound

This paper cites Freeman, and Taesung Park.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Freeman, and Taesung Park

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.388288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.899017Z digest=sha256:896fa5e9a46e401299aa7360a587e495d9d5626e47c32f2d263e615453f189f7

Observation 4492e9ed-aa5a-494e-b2ff-17a36c0cc83b · outbound

This paper cites Free- man, Fr´edo Durand, Eli Shechtman, and Xun Huang.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Free- man, Fr´edo Durand, Eli Shechtman, and Xun Huang

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.377929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.902081Z digest=sha256:bcbfb8054f440afcaaa6b8239ee78b968f571a3a5a36797d62e851baeca3f1fe

Observation 0c368f77-fe29-47d4-9c41-cd7a5dc58c3e · outbound

This paper cites The ablated variants exhibit increasing identity drift and vi- sual artifacts in later segments, whereas the full model re- mains more consistent.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos The ablated variants exhibit increasing identity drift and vi- sual artifacts in later segments, whereas the full model re- mains more consistent

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.367030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.905261Z digest=sha256:c350368fc5ba846782d14a7224d96dd7e9ab2495b2cb65fa2111038b3db46173

Observation de58b078-ca90-434e-af6a-44373ebb30cc · outbound

This paper cites The reference cache persists throughout generation, source keys and values are tem- porary, and completed target blocks are committed to the clean-history cache.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos The reference cache persists throughout generation, source keys and values are tem- porary, and completed target blocks are committed to the clean-history cache

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.355170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.908491Z digest=sha256:cbb0ef8a1c63ad11b6f50891289aa669eb4171f8b98349ea81198f4f4a617854

Observation 896a7ddd-61ad-43c7-bf44-608589ad03ca · outbound

This paper cites The study compared UniSwap with four video-replacement baselines, each paired with Seed-VC following the cascade protocol in Table 1.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos The study compared UniSwap with four video-replacement baselines, each paired with Seed-VC following the cascade protocol in Table 1

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.345225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.911900Z digest=sha256:ffe33b6961a76b58d4a48315ce21cab5e554b9fb2a68e7789d1f27029267293e

Observation bc9a1f1d-3ab5-4c1e-a6e3-b7bcdeca3c31 · outbound

This paper cites In all figures, each example con- tains a reference image and reference voice clip, a source video and its audio, and the joint audio-video output pro- duced by UniSwap.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos In all figures, each example con- tains a reference image and reference voice clip, a source video and its audio, and the joint audio-video output pro- duced by UniSwap

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.335051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.915155Z digest=sha256:4c80d72d5751c486c7bf62cbed37d34540d0576b76bdf4b05388501be2b3d73e

Observation a15be8c3-6fd1-496e-9ff5-4214f51def9c · outbound

This paper cites an unresolved cited work.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:34:45.323486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.918493Z digest=sha256:d72809ea49665a994d32a709db90394d5132fff4355c7c6455eb93a64f60719b

Observation fd69542b-0ccd-45b2-8e7a-6dea3a7af39e · outbound

This paper cites Deployment should require consent and provenance mechanisms, visible disclosure where ap- propriate, access controls, and compatibility with forensic detection tools.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Deployment should require consent and provenance mechanisms, visible disclosure where ap- propriate, access controls, and compatibility with forensic detection tools

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:34:45.312445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:34:44.921475Z digest=sha256:e552cea72df76ceba69d2008d889639ae1260164c72f3fc33abe7c10b9f9074c

Pith citing papers

No inbound Pith citation observations are available.