Pith. sign in

Paper Citation Record · LEDGER

Vorch-Omni: Multi-Task Orchestration of Sight and Sound

As of 14 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.05803.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05803 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:16:52.552901Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact5
  • verified fuzzy2
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 085f7559-aa81-4101-b8b4-4d893902df8a · outbound

This paper cites SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.478464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.478464Z digest=sha256:1812f9e6c8f07449f65f1315846f4992232fdfe882d7578cc33c6a974e7a60fa

Observation b7705bf0-d6d7-4bb3-a837-1f6f546c3682 · outbound

This paper cites Dreamid-omni: Unified framework for controllable human-centric audio-video generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Dreamid-omni: Unified framework for controllable human-centric audio-video generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.482667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.482667Z digest=sha256:37da4dae6cd7bb55f7b13a42d34ee630424aad664f13026a11f323052f46ac8d

Observation efce4dc6-b3b2-4491-864c-6a11ebde837e · outbound

This paper cites Mmdisco: Multi-modal discriminator- guided cooperative diffusion for joint audio and video generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Mmdisco: Multi-modal discriminator- guided cooperative diffusion for joint audio and video generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:16:53.584280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T23:16:52.491684Z digest=sha256:87f28a58e2edc581ad9c14a52c046336f0deecf2437e422c5abece8a8e688274

Observation c560d014-c578-457c-b133-377d5f55c84e · outbound

This paper cites HunyuanVideo 1.5 Technical Report.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound HunyuanVideo 1.5 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.495601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.495601Z digest=sha256:2710ef90475dcc041d1ebb47f368c4d1ebc589639bf009ea76e13642e164ca33

Observation 3da2db8c-0201-4506-a4fd-aed8acb3510d · outbound

This paper cites Native Audio-Visual Alignment for Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Native Audio-Visual Alignment for Generation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:52.959820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T23:16:52.499515Z digest=sha256:45d0780c998a0d082d905edbd926e24a51de391c2e75364d1986a319f07249e1

Observation 5b81f74d-1820-4e9d-8c3f-7ca0ab47edb7 · outbound

This paper cites Kling-Omni Technical Report.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Kling-Omni Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.502930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.502930Z digest=sha256:6f30a88e63133956fa282e2871ab539dba33b3cf411624d92f903a1176f07415

Observation 9ccd18d3-a82e-4af8-9f5d-da02fe9e350a · outbound

This paper cites MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:52.929357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T23:16:52.506326Z digest=sha256:3c5042a3fc9c96b8c8b2e4e4708b3a5e4286a05bdaf77809612a2d8715b086cf

Observation 787248e6-93ad-49ca-9c0f-a24ddfdf72b8 · outbound

This paper cites Flow Matching for Generative Modeling.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Flow Matching for Generative Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.509929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.509929Z digest=sha256:fa10a1109730cdb200528b39da4481f6092ab0256884aeabf5ad9829a96a316e

Observation 9468b1d8-23a1-490a-82c1-a3b1ef9b346b · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.513807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.513807Z digest=sha256:a649e2d1f010242358ee9f9d7b776ddbbb5649108b431854f220711797471e59

Observation 89ef2161-5831-4a26-a182-418b358c793b · outbound

This paper cites Team OpenMOSS.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Team OpenMOSS

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.521689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.521689Z digest=sha256:de05ea3eaa9bdb17ab2f29d81878df9a12694d5324e0603f3371bfdff7be8357

Observation ae50f3a9-5f51-43d7-b4ea-c0c1fc095286 · outbound

This paper cites Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.529134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.529134Z digest=sha256:4854682338d9219bcb6d95a3546e94c4a9479356805d7b9413d880d337e77f81

Observation 7d54ed40-72b2-4164-9ea7-829bc0695111 · outbound

This paper cites Baton: Explicit Semantic Blueprints for Joint Video-Audio Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Baton: Explicit Semantic Blueprints for Joint Video-Audio Generation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:52.635253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T23:16:52.532992Z digest=sha256:9a1b62d9b5f7688ef80de6341d3bfc92fdb5415092d5803d8e4ee4686aa0f10f

Observation b30009e5-84ca-4f87-a714-af268271ee5b · outbound

This paper cites UniVerse-1: Unified Audio-Video Generation via Stitching of Experts.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound UniVerse-1: Unified Audio-Video Generation via Stitching of Experts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.540959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.540959Z digest=sha256:b6ae72e6e136e99e3dceac132c0dc9b862045ad929f6f12492bd6190f07c8cc7

Observation d2b55d41-71fa-4823-8db6-cac76c6ffe71 · outbound

This paper cites Uniavgen: Unified audio and video generation with asymmetric cross-modal interactions.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Uniavgen: Unified audio and video generation with asymmetric cross-modal interactions

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:16:53.571568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T23:16:52.544882Z digest=sha256:d181e413e5de9256dffae6b19a67c3b95386ad381714e60e31e8afcf1d07e0ae

Observation dd650d56-eda4-4d06-80ef-7bbfaa557f56 · outbound

This paper cites UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.548806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.548806Z digest=sha256:d7a93099eb580c1db95275e00d452708e3ff5ef2345848d9414de53819ac00c5

Observation d846b8d5-983e-4fe4-b01d-7996cf6fe358 · outbound

This paper cites InstructAV2AV: Instruction-Guided Audio-Video Joint Editing.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound InstructAV2AV: Instruction-Guided Audio-Video Joint Editing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.552901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.552901Z digest=sha256:ebd05693f08cec3bef6a01ae31261cf94188caff7904471b5d806a10c4f9a5d6

Observation 088567b2-ec9d-4d18-bd2c-453312f10e2e · outbound

This paper cites Guibin Chen, Dixuan Lin, Jiangping Yang, Youqiang Zhang, Zhengcong Fei, Debang Li, Sheng Chen, Chaofeng Ao, Nuo Pang, Yiming Wang, et al.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Guibin Chen, Dixuan Lin, Jiangping Yang, Youqiang Zhang, Zhengcong Fei, Debang Li, Sheng Chen, Chaofeng Ao, Nuo Pang, Yiming Wang, et al

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.464969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.464969Z digest=sha256:e10a76f2de22de987d2146b3a78a42a696355493544f4b54bb725ac70dde7296

Observation 87f681e6-8702-48ec-9471-5b10076424dd · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Wan: Open and Advanced Large-Scale Video Generative Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.536889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.536889Z digest=sha256:bc0d44ddd698fe7590e3b03024eaea90e4554933e8e00cc88a1d4727b662c67b

Observation 8454eb56-0efb-475a-93c2-af9c3b5f0718 · outbound

This paper cites Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.517841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.517841Z digest=sha256:d1981886bd290cc3d6910fd4f8ce0aae0445a7c396e00b1fb34f812991bfc9ad

Observation d3063727-59cd-4966-925a-49b47f222a7c · outbound

This paper cites AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

Reference 2023

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:52.664679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T23:16:52.525251Z digest=sha256:dde55da102fb74b8c21b5388bb3935c1728f0237a33b0f4f4dbf664376d7e49a

Observation 47bbb723-024c-4a6a-8dc9-451b43a86f7e · outbound

This paper cites CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:53.347134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T23:16:52.469576Z digest=sha256:7eca9b7a3963301ea4e9a9be9282435babe456c7304ac10a8191e1df47e2e219

Observation 0fadbedf-3701-49b7-866b-674715952e60 · outbound

This paper cites Speed by simplicity: A single-stream architecture for fast audio-video generative foundation model.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Speed by simplicity: A single-stream architecture for fast audio-video generative foundation model

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.474289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.474289Z digest=sha256:28985848a63e96568e700398ca4cd0000c225a0ee17f7dfe6ee8014c36ebbf6f

Observation ec011e65-83f0-4560-aa54-334be2c026ab · outbound

This paper cites LTX-2: Efficient Joint Audio-Visual Foundation Model.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound LTX-2: Efficient Joint Audio-Visual Foundation Model

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.486493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.486493Z digest=sha256:1e4a80ec587d718e340d3845bc4a9794ae2b57379bd2d90faae8dc0d5ab931ae

Pith citing papers

No inbound Pith citation observations are available.