Pith. sign in

Paper Citation Record · LEDGER

Vorch-Omni: Multi-Task Orchestration of Sight and Sound

As of 9 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.05803.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05803 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:16:52.552901Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact5
  • verified fuzzy2
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 085f7559-aa81-4101-b8b4-4d893902df8a · outbound

This paper cites SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.478464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.478464Z digest=sha256:e47ec8122a4a152b520f7e660fc91d527e64b6d6ecd835f0ac46a22b5a9f9c6f

Observation b7705bf0-d6d7-4bb3-a837-1f6f546c3682 · outbound

This paper cites Dreamid-omni: Unified framework for controllable human-centric audio-video generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Dreamid-omni: Unified framework for controllable human-centric audio-video generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.482667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.482667Z digest=sha256:9810a017138f44e28007e7c116343d261e2661393baec5782a4034dc9af38f50

Observation efce4dc6-b3b2-4491-864c-6a11ebde837e · outbound

This paper cites Mmdisco: Multi-modal discriminator- guided cooperative diffusion for joint audio and video generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Mmdisco: Multi-modal discriminator- guided cooperative diffusion for joint audio and video generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:16:53.584280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.491684Z digest=sha256:3fbc2c6a91dd36188844e4d80fc7d16c63c776574535cf980f24196bda0d8ce1

Observation c560d014-c578-457c-b133-377d5f55c84e · outbound

This paper cites HunyuanVideo 1.5 Technical Report.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound HunyuanVideo 1.5 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.495601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.495601Z digest=sha256:ed0f70a621a055d33bee0915e03fafdf45969ff3dd07f1428da8ce4fe7d7be5b

Observation 3da2db8c-0201-4506-a4fd-aed8acb3510d · outbound

This paper cites Native Audio-Visual Alignment for Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Native Audio-Visual Alignment for Generation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:52.959820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.499515Z digest=sha256:a9ee139027d480a32bd30298d0d771767d4325a942443ca0bd4d263eae895f09

Observation 5b81f74d-1820-4e9d-8c3f-7ca0ab47edb7 · outbound

This paper cites Kling-Omni Technical Report.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Kling-Omni Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.502930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.502930Z digest=sha256:95304d3937dfad259d6eceb8bd5f22d57be90f3bfd93a5dd6c8bc06594bc6137

Observation 9ccd18d3-a82e-4af8-9f5d-da02fe9e350a · outbound

This paper cites MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:52.929357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.506326Z digest=sha256:4dfae1c1d977335909be1f32c4e2dade3e13ec83a6906a94e504f91ca73eb1c4

Observation 787248e6-93ad-49ca-9c0f-a24ddfdf72b8 · outbound

This paper cites Flow Matching for Generative Modeling.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Flow Matching for Generative Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.509929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.509929Z digest=sha256:bfe4f31227a4aa597ba49f3169f064c419ef8c7a176fd814909b18ec4817fb86

Observation 9468b1d8-23a1-490a-82c1-a3b1ef9b346b · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.513807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.513807Z digest=sha256:bd4767a3c4496de84fe808261618c97117430816b31c0b563016c4ca639e133b

Observation 89ef2161-5831-4a26-a182-418b358c793b · outbound

This paper cites Team OpenMOSS.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Team OpenMOSS

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.521689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.521689Z digest=sha256:d138b41cda3e10be4a72fbf398e5710115a0601f967b63801c19896fe20e6464

Observation ae50f3a9-5f51-43d7-b4ea-c0c1fc095286 · outbound

This paper cites Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.529134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.529134Z digest=sha256:bb964e9ff3d90e55950e8a1482644efdf2cceeb01544018ad977592439f7bb79

Observation 7d54ed40-72b2-4164-9ea7-829bc0695111 · outbound

This paper cites Baton: Explicit Semantic Blueprints for Joint Video-Audio Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Baton: Explicit Semantic Blueprints for Joint Video-Audio Generation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:52.635253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.532992Z digest=sha256:2e2d7644b30592d6ac3c6f6f1129497df2a3feef9c51c85f1db9d76c2ca1b236

Observation b30009e5-84ca-4f87-a714-af268271ee5b · outbound

This paper cites UniVerse-1: Unified Audio-Video Generation via Stitching of Experts.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound UniVerse-1: Unified Audio-Video Generation via Stitching of Experts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.540959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.540959Z digest=sha256:6306843eae8a216d49517d4791166ee751c357e40f360581c9ba9301fbda73e9

Observation d2b55d41-71fa-4823-8db6-cac76c6ffe71 · outbound

This paper cites Uniavgen: Unified audio and video generation with asymmetric cross-modal interactions.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Uniavgen: Unified audio and video generation with asymmetric cross-modal interactions

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:16:53.571568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.544882Z digest=sha256:25d7832b870418b149566be25c23fd41129c5f74ebe4e891ed0d6a4199b1fb10

Observation dd650d56-eda4-4d06-80ef-7bbfaa557f56 · outbound

This paper cites UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.548806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.548806Z digest=sha256:fea8224fb92e0876e4317808e87510192f0c3937226641efebcb2718da4a183b

Observation d846b8d5-983e-4fe4-b01d-7996cf6fe358 · outbound

This paper cites InstructAV2AV: Instruction-Guided Audio-Video Joint Editing.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound InstructAV2AV: Instruction-Guided Audio-Video Joint Editing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.552901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.552901Z digest=sha256:ec8c52511d5c5da6b0b0f2d66c43d0c80e0702720082e048f65b7b360264ce9b

Observation 088567b2-ec9d-4d18-bd2c-453312f10e2e · outbound

This paper cites Guibin Chen, Dixuan Lin, Jiangping Yang, Youqiang Zhang, Zhengcong Fei, Debang Li, Sheng Chen, Chaofeng Ao, Nuo Pang, Yiming Wang, et al.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Guibin Chen, Dixuan Lin, Jiangping Yang, Youqiang Zhang, Zhengcong Fei, Debang Li, Sheng Chen, Chaofeng Ao, Nuo Pang, Yiming Wang, et al

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.464969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.464969Z digest=sha256:8de010f62db484227a62bd63716b8595e6ca4ffd8eaf9c5ab30ad52d1f197131

Observation 87f681e6-8702-48ec-9471-5b10076424dd · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Wan: Open and Advanced Large-Scale Video Generative Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.536889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.536889Z digest=sha256:f13ada4a3a73b0d0a867342cf3f44e53b12c3f6caa5573de3d60f6108fdef717

Observation 8454eb56-0efb-475a-93c2-af9c3b5f0718 · outbound

This paper cites Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.517841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.517841Z digest=sha256:b3583e3d567474d1c0f0ad784e6f43708ae0f20a88190fdae884fe588a34972f

Observation d3063727-59cd-4966-925a-49b47f222a7c · outbound

This paper cites AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

Reference 2023

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:52.664679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.525251Z digest=sha256:083154e93bbe3f9c4f1ca2a684aeac1c3f58517d9f053814b631184412467f12

Observation 47bbb723-024c-4a6a-8dc9-451b43a86f7e · outbound

This paper cites CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:16:53.347134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:16:52.469576Z digest=sha256:bd6943fda63f618bdd9b548b4bc43cc066da775dc8deadee27e1b99f527da438

Observation 0fadbedf-3701-49b7-866b-674715952e60 · outbound

This paper cites Speed by simplicity: A single-stream architecture for fast audio-video generative foundation model.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound Speed by simplicity: A single-stream architecture for fast audio-video generative foundation model

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.474289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.474289Z digest=sha256:19a8aceadf3ea42ff781b2402ef7a1026ede3634eb7c8dead33f4a2754e1dd31

Observation ec011e65-83f0-4560-aa54-334be2c026ab · outbound

This paper cites LTX-2: Efficient Joint Audio-Visual Foundation Model.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound LTX-2: Efficient Joint Audio-Visual Foundation Model

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.486493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.486493Z digest=sha256:608b43639fe4833319c4fc55f7f8fa40f916ba8cd522c68561ffad350437cc09

Pith citing papers

No inbound Pith citation observations are available.