Pith. sign in

Paper Citation Record · LEDGER

From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2409.19132.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.19132 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:38:08.575530Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T13:28:18.779237Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5b295371-2952-460f-a006-c9e040a762ad · inbound

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation cites this paper.

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:08.575530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:08.575530Z digest=sha256:a186ecd03a15a632242b4d34e01812dcf6e1d3fa8ca72a6820cd97c58fff6f6c

Observation 76428a5b-8200-43c9-a5fe-e62957cc6005 · inbound

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance cites this paper.

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.763813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T13:10:18.700497Z digest=sha256:f274aec358073114ee8111d54e3e2c1daf0e58aa8a6270af6f0392163b1bee0a

Observation c7ae6a2e-3729-4f90-9057-e3bc1fc0f94c · inbound

ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals cites this paper.

ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:30:50.844186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T18:30:30.431777Z digest=sha256:f92b89d7169f2730c2faec4671876ad3956cdd15344ed4d2da9dbc43bf844dc0

Observation 4c6031e5-0730-446d-a82b-6215795da607 · inbound

Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning cites this paper.

Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:10:59.444346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T17:46:43.584943Z digest=sha256:a2b9f64f4daec5913a6f128149131a764c93ecc57708570b2562e2c415b92d37

Observation 4c469386-d564-4fbb-88f1-64ad2f9fc04e · inbound

Do Joint Audio-Video Generation Models Understand Physics? cites this paper.

Do Joint Audio-Video Generation Models Understand Physics? From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.234269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:262dc40d4aa877babadf6a185eb87263cfb90833b91f7cbc32b93f58fed0a655

Observation 9041fab1-11ca-4481-aad0-9e7d6e411fcc · inbound

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation cites this paper.

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:18.780712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T08:04:48.283908Z digest=sha256:77f06e060022a935294976d63617d1da3f043fe1c269057438c4aefe48cb7452