Pith. sign in

Paper Citation Record · LEDGER

Read, Watch and Scream! Sound Generation from Text and Video

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2407.05551.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.05551 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:26:46.705357Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T13:11:23.797787Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b73e3130-6b2a-48b9-8788-c5136e7682a8 · inbound

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT cites this paper.

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT Read, Watch and Scream! Sound Generation from Text and Video

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T14:26:46.705357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:26:46.705357Z digest=sha256:a6611237d43d0d2d627da4714fa08515f546dc70fbb4f4f5703680d5a9f6c47e

Observation 110c6b42-6d76-4411-a425-12ee6748e1f7 · inbound

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet cites this paper.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet Read, Watch and Scream! Sound Generation from Text and Video

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.831770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.831770Z digest=sha256:dab28ba7f6e0d6c92377fb97b594feb68beadadc2ed285db545ab5af1110cf9e

Observation 4f00a0c2-f96d-4e06-944f-345201604a8e · inbound

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation cites this paper.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Read, Watch and Scream! Sound Generation from Text and Video

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:37:17.308243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:37:17.308243Z digest=sha256:f111b7b8289ee912d3982e6c10ecae855deeae225db14dc2efcedab9c9c36de3

Observation 689c51ec-fe7f-4d5f-a2af-89429960d0d5 · inbound

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance cites this paper.

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance Read, Watch and Scream! Sound Generation from Text and Video

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.800632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T13:10:18.700497Z digest=sha256:4fa5e7487c2c3fbc22142c87706d206752ea5964b373b5d748775730817c53ec

Observation 4573399d-1415-4678-a0fc-f8bc7dbd131b · inbound

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation cites this paper.

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation Read, Watch and Scream! Sound Generation from Text and Video

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:21:06.812467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T08:20:02.986562Z digest=sha256:056d211f86ff040dea0fc82b7de9fc098088f741eb285755c2a50f2e6a17b0c9