Pith. sign in

Paper Citation Record · LEDGER

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation

As of 16 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2507.02271.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02271 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:37:19.640172Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69a048b8-bf54-4b91-a693-0876059b899b · outbound

This paper cites The Foley grail: The art of performing sound for film, games, and animation.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation The Foley grail: The art of performing sound for film, games, and animation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:24.534991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:16.105516Z digest=sha256:e7ee44ba83b637a796d1c29d928af7ef361c5746ece8d4f7bdcf8ed9457a4595

Observation 1b4e5e8d-32a5-4203-9e59-b2e0858f143d · outbound

This paper cites Deep residual learning for image recog- nition.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Deep residual learning for image recog- nition

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:23.444828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:16.695834Z digest=sha256:6dfbfe72267e591894b13a15059dea1f0dc74ebb977f246dea0099fd5f4326b0

Observation b0c46848-49bc-4f87-9f79-5d9cb000a8a2 · outbound

This paper cites Denoising diffusion probabilistic models.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Denoising diffusion probabilistic models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:37:17.017938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:37:17.017938Z digest=sha256:554a9eba4406af7d0a18690a570dafb747898b472049624fec168b2bd0cb726f

Observation ace39fec-2829-4717-88cf-34db85b6c68e · outbound

This paper cites Densely connected convolutional networks.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Densely connected convolutional networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:23.266003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:17.087475Z digest=sha256:ed76323eae3baf23b3b780ff0d5d6d97f45c6a7c5adf3020a0af10e844a94fad

Observation 4f00a0c2-f96d-4e06-944f-345201604a8e · outbound

This paper cites Read, Watch and Scream! Sound Generation from Text and Video.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Read, Watch and Scream! Sound Generation from Text and Video

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:37:17.308243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:37:17.308243Z digest=sha256:50e591181f95676b4076ef7a63e95bd3526a0009be33a26570b6ff3b5a746d00

Observation db35df16-0f1b-47b4-8715-a7105ef1ecd5 · outbound

This paper cites Diverse part dis- covery: Occluded person re-identification with part-aware transformer.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Diverse part dis- covery: Occluded person re-identification with part-aware transformer

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:22.814494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:17.481818Z digest=sha256:3bee999679479eae964937fe8ef13b58d0aada57fc0ece59e282ccef2eb03610

Observation ea159ff9-6ce8-4424-849c-83af5bdab436 · outbound

This paper cites Mitigating and evaluating static bias of ac- tion representations in the background and the foreground.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Mitigating and evaluating static bias of ac- tion representations in the background and the foreground

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:22.678563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:17.568223Z digest=sha256:c97138256fde8c9ab49a753fb9863fc3f99863368ff899604b32eb8271de3b3f

Observation 6dce6e9e-f5dc-432a-b213-05cb950ed00d · outbound

This paper cites AudioLDM: Text-to-audio generation with latent diffusion models.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation AudioLDM: Text-to-audio generation with latent diffusion models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:22.538386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:17.636200Z digest=sha256:8bed99ec07afdf00990970aa1f9fd8e3847c2f46fb7954a43641918d7e37b101

Observation d62a57df-c626-45fd-b39a-a55d007b256a · outbound

This paper cites DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:37:17.798903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:37:17.798903Z digest=sha256:c92038875bb809fe865a5f5f6f042ccab21ebd586753ff23d752c405fac16298

Observation 8f82e688-1ece-4d83-a2b4-dab92dded7ad · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models.Advances in Neural Information Processing Systems, 36,.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models.Advances in Neural Information Processing Systems, 36,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:22.339611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:17.879008Z digest=sha256:43b0afde23672a7d168e9fa916a80c7c44bfbfa624b53db069b44c54fadb8824

Observation 5212535c-5019-435a-9b49-e5dc42b12f9c · outbound

This paper cites The filmmaker’s eye: The language of the lens: The power of lenses and the expressive cinematic image.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation The filmmaker’s eye: The language of the lens: The power of lenses and the expressive cinematic image

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:22.132569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:17.958654Z digest=sha256:d4697fce53070a5ced5479184ff2c49963d311e2a1609db0277fab4a7a4c2615

Observation 25fe33b3-7429-46af-b274-99c798f89081 · outbound

This paper cites Masked generative video-to-audio transformers with enhanced synchronicity.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Masked generative video-to-audio transformers with enhanced synchronicity

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:21.828307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:18.184538Z digest=sha256:3b8e6f6ab83b52113d670163cae8f48386dbd416a4b9eddc092f8363492684d3

Observation e12fde5f-5029-4e60-9cde-02ca65adfc27 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation High-resolution image synthesis with latent diffusion models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:21.590179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:18.403758Z digest=sha256:728b042bfff55576b0cbc8f367b417c519f27221c26ffa9ed2ef7fa733caff92

Observation e6ed0c00-47cb-4112-a669-9ad94db08c51 · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:37:18.510802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:37:18.510802Z digest=sha256:5b7d399d5c8f03cf93c6fcba8b54aee3d1ec76dbfe5a78487097959534ca5654

Observation dbcb4214-e9cb-4123-a6c0-3656e012e064 · outbound

This paper cites Denoising diffusion implicit models.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Denoising diffusion implicit models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:21.420248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:18.589032Z digest=sha256:4c89a2b235ff3957eb37274989d8d71c0ebdc6a9433067c26b1f55578c95befb

Observation a879510a-756e-404b-aee7-f14b9d686219 · outbound

This paper cites Temporally Aligned Audio for Video with Autoregression.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Temporally Aligned Audio for Video with Autoregression

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:37:18.705798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:37:18.705798Z digest=sha256:86a24b2648cff8a50b2720f1cd74308ccab81ed66d363a6089a9e44bec3600ca

Observation 5364271f-4cb6-4cee-88c7-b28e96dcf8dd · outbound

This paper cites Removing the background by adding the background: Towards background robust self- supervised video representation learning.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Removing the background by adding the background: Towards background robust self- supervised video representation learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:21.190004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:18.804143Z digest=sha256:7d1a2c48c162dacbfdd306fa28285615c3278c2bc965b5ca01f492116d7370ba

Observation 61ca4cac-0cdc-4c81-b490-83829effb8c5 · outbound

This paper cites Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:37:18.886087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:37:18.886087Z digest=sha256:362067e9c3735d5321c24c54b8fa137b2eb121a74f149fe9ae2d806e7a375099

Observation dcf0d7b2-3fe4-4f7c-8ca5-6778b45c9083 · outbound

This paper cites Sonicvisionlm: Playing sound with vi- sion language models.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Sonicvisionlm: Playing sound with vi- sion language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:21.041126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:18.967426Z digest=sha256:ba1824dd9bc2747899eef9545b7c8fd38f3a38b119a2ce1cd7d79f7e027f29e4

Observation c58c3f2a-5a27-4d6e-a756-6457b59f5e04 · outbound

This paper cites Data- distortion guided self-distillation for deep neural networks.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Data- distortion guided self-distillation for deep neural networks

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:20.889079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:19.069164Z digest=sha256:53b0f87728f62aceb8b28bf9c287ce297d79f05e9f2cc21ab7b8a353cf6cdf09

Observation 8a4a7f3c-41ba-4ec9-ba7c-be28db54ae05 · outbound

This paper cites Cutmix: Regularization strategy to train strong classifiers with localizable features.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Cutmix: Regularization strategy to train strong classifiers with localizable features

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:20.670457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:19.155163Z digest=sha256:d925e0d8b9e60656850b4ed00635fe1a3d465fff63bba6acadd021a4de3c420e

Observation 13c81013-ec80-41f2-98a5-f4b191318ce7 · outbound

This paper cites Self- supervised scene de-occlusion.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Self- supervised scene de-occlusion

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:20.436510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:19.278087Z digest=sha256:9b08d88c83e6997c021215476bbab49b299c1af2f2832b2d5e47e08e904d06ff

Observation 7a4233b5-94f2-40d4-acc6-21b953d901bb · outbound

This paper cites Be your own teacher: Improve the performance of convolutional neural networks via self distillation.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Be your own teacher: Improve the performance of convolutional neural networks via self distillation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:20.232725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:19.379719Z digest=sha256:fdf754a3c9e6dc32a08979116dd92f471ce57f9c89c21af7553df5de6b564c1c

Observation 1c65f088-3fe7-4bc2-9256-114872a13097 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:37:19.460025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:37:19.460025Z digest=sha256:c0db369a6e99495463cc589c5cb5fa5955018f75b7342e9b64a1fc6df9df38dc

Observation 990972e0-a7ab-4afc-afc2-e81b7f28873e · outbound

This paper cites Human orientation estimation un- der partial observation.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Human orientation estimation un- der partial observation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:20.056340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:19.556189Z digest=sha256:9b9bbbb2f31b07ab9c1c232a3add92079326fd8141d94240140e20b86b2aa182

Observation 085339fe-e831-4b9d-8a2b-553feb5cd84a · outbound

This paper cites Knowledge distillation by on-the-fly native ensemble.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Knowledge distillation by on-the-fly native ensemble

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:19.890122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:19.640172Z digest=sha256:14a93da5d7cf934be9a364afe74c5ba07cefccd2a470f7fef0a576a85dd0cf16

Observation 4e6e21e1-ceb6-42c3-861e-f6b90451c68a · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Vggsound: A large-scale audio-visual dataset

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:24.274677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:16.196262Z digest=sha256:68e0e8a5cc91636f87d021c22b051db054072488cc16e7e744df00ca5beda5d4

Observation ec3772c3-aca5-4cc1-9ee4-a98870fc188b · outbound

This paper cites Classifier-Free Diffusion Guidance.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Classifier-Free Diffusion Guidance

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T20:37:16.946712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:37:16.946712Z digest=sha256:3c6c49ce1a4049670014e4c7ba4c4134f9d869b1db78c70f89cd7fe892fe6400

Observation 60093ecc-6da4-4ccd-9ebf-89a2145e40ac · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Distilling the Knowledge in a Neural Network

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T20:37:16.868550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:37:16.868550Z digest=sha256:ceb12f18c92668174da8c8d9c0ff5a70fa93b23977d46b4c106c6594f9def1a3

Observation e812d4d2-3a71-4312-be93-55647037a939 · outbound

This paper cites Taming visually guided sound generation.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Taming visually guided sound generation

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:23.118610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:17.144015Z digest=sha256:df3b2d5beb3f86abdf21bcc7dd20c0e72944220d772591169b65157b6eaa4a09

Observation 1d3a90e3-5301-4b14-9610-8063bbac5269 · outbound

This paper cites Pose-guided feature alignment for occluded person re-identification.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Pose-guided feature alignment for occluded person re-identification

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:21.977775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:18.081047Z digest=sha256:7b1a901447771bb00cac4679339c962104c705003bc26e97a2a63af74628b802

Observation 73dbc52d-bd2e-4109-872d-1649253d7afb · outbound

This paper cites Diffusion models beat gans on image synthe- sis.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Diffusion models beat gans on image synthe- sis

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:24.073656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:16.299487Z digest=sha256:3f5e443cf6a86cca4e0cc600a977a7b6788e5f9b67af4e31ad99edf4e2188872

Observation 943e0a8e-ed34-4e15-b066-937d0624c7e3 · outbound

This paper cites Motion-aware contrastive video rep- resentation learning via foreground-background merging.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Motion-aware contrastive video rep- resentation learning via foreground-background merging

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:23.904541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:16.385530Z digest=sha256:01f577259d3424cd96c9f035b6a3b82526467aa936c909bf493fd01bba13b564

Observation 58f6c621-ff98-454b-abce-0de6cd318e58 · outbound

This paper cites Conditional gener- ation of audio from video via foley analogies.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Conditional gener- ation of audio from video via foley analogies

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:23.778413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:16.459442Z digest=sha256:8b461a4d71718edce89c5024a0ae2913e2861351b7b63e8c1870bae687007f12

Observation aaedeca7-da18-4389-996e-139a69818b1f · outbound

This paper cites Learn2augment: learning to composite videos for data augmentation in ac- tion recognition.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Learn2augment: learning to composite videos for data augmentation in ac- tion recognition

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:23.603245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:16.563519Z digest=sha256:765a4f0d7680f3202164f4657522044f98744fd27d4c19ba4793033175a58985

Observation 9cfa774b-c41b-4853-96fb-8a535021a89f · outbound

This paper cites Instance-wise occlusion and depth orders in natural scenes.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation Instance-wise occlusion and depth orders in natural scenes

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:37:22.986685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:37:17.406906Z digest=sha256:9bcb7ea9e465dc3befcacf65991eabb50ae1d7d2c6a0fbf7e28186ffc6c4a222

Observation bcfd0b05-4dc7-41cc-801b-84536be8fcc6 · outbound

This paper cites STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T20:37:18.302787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:37:18.302787Z digest=sha256:83b30ba0ae798500ec86167acf3ba4e83e324345fbe2ab9ba36e7af0229903e4

Pith citing papers

No inbound Pith citation observations are available.