Pith. sign in

Paper Citation Record · LEDGER

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance

As of 23 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2412.18157.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18157 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:01:38.747896Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5dfb7121-e594-4dc0-93c3-7e824b842549 · outbound

This paper cites Generating visually aligned sound from videos,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Generating visually aligned sound from videos,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.315954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T05:01:38.603327Z digest=sha256:3689b7cc957ee7addfcd3012dfa4c97ad257a75a2f891e02d3e72000030746d6

Observation 7f37af99-e3dc-4270-8fd2-c3016a9657a5 · outbound

This paper cites Taming visually guided sound generation,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Taming visually guided sound generation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.301234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T05:01:38.608027Z digest=sha256:a538faa4acc96c9f224e75b20c98e3d903feb746172335bcca451688d47b531c

Observation c4581f99-f988-4b81-9220-c5044e0f6b6a · outbound

This paper cites Attention is all you need,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Attention is all you need,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.612589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.612589Z digest=sha256:c77fd444a41b8dd1ac85d251a36172f1f9417647c97a82be13797e2382e16306

Observation 46a1da61-e450-446f-8cfc-a394399df2be · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Learning transferable visual models from natural language supervision,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.618248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.618248Z digest=sha256:a695178949ff52212fe4fda69b5bf85aba6f73bf432e29b3451e4b79e8d0630c

Observation a744e841-2c08-404c-8b66-45e555bc176e · outbound

This paper cites Imagebind: One embedding space to bind them all,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Imagebind: One embedding space to bind them all,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.623065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.623065Z digest=sha256:f793983b19f74b679cc4f06d5819d88db91c3aae053f691c1006975328bda39a

Observation c8e4a828-487c-4ba8-a2c2-779ae22072e6 · outbound

This paper cites Denoising diffusion probabilistic models,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Denoising diffusion probabilistic models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.628058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.628058Z digest=sha256:8434e03c75a24a3128a482f8f3dfeee2994a7ab0d3182b186f3ce310e4703526

Observation f99302b1-5d3d-4c06-ab10-26ca27acd214 · outbound

This paper cites Conditional generation of audio from video via foley analogies,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Conditional generation of audio from video via foley analogies,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.249712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T05:01:38.633488Z digest=sha256:7675f9b1617543ca50df8564a6c8b61697517fdc39698b350c7458d072211158

Observation 0de05aa5-1667-4472-bad8-014d5c258bf7 · outbound

This paper cites Varietysound: Timbre-controllable video to sound generation via unsupervised information disentanglement,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Varietysound: Timbre-controllable video to sound generation via unsupervised information disentanglement,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.235676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T05:01:38.637922Z digest=sha256:56bd7916a0681832734e57d875462b18fe3f4de3c68cef0ad52ac7689127ae34

Observation bb71c8fd-1a60-497d-bd44-2f50c84b26e6 · outbound

This paper cites Diff-foley: Synchronized video- to-audio synthesis with latent diffusion models,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Diff-foley: Synchronized video- to-audio synthesis with latent diffusion models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.221071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T05:01:38.642369Z digest=sha256:33f77d524c7d108ccc8ad0937b59be511ada16918095c037d1a2c5fbf8ed22e3

Observation b0baf4ea-d77a-4717-8936-db8d6dda74ef · outbound

This paper cites I hear your true colors: Image guided audio generation,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance I hear your true colors: Image guided audio generation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.206553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T05:01:38.647528Z digest=sha256:5313966502ed1efeb8c24bfc1455d1fe39edea4aa5d0402adf5fec584d86ea00

Observation 82f2665c-0c2a-40d2-aeef-9889bc041184 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.651937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.651937Z digest=sha256:bb9136cca914c2ceac54b138d78ae3477081f08b03fe53d78b8022b3cfc71d97

Observation 5b036e71-af05-4ca9-8827-5dfeebc4e2ea · outbound

This paper cites Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.656647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.656647Z digest=sha256:28f3ee910077254c96e8011a1668197bf4d546e57f833b68628830b1e3535c63

Observation ebc93496-4ccc-4660-8eb4-33adb2f3926e · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.661167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.661167Z digest=sha256:1de548bc6d451c8e4f7f2314acbc1c45ab3d84bdfec0201121541fd5eb0de03f

Observation 66576f66-c5e7-4e9b-ae71-754e911cb725 · outbound

This paper cites Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.667045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.667045Z digest=sha256:fa2d493deb0e7542457e95c3241c8fd3759c258afc87fa15a6b9a8ed259d636b

Observation dc1c3c51-5603-4238-89e0-28dd1de65d9c · outbound

This paper cites Predicting deep zero-shot con- volutional neural networks using textual descriptions,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Predicting deep zero-shot con- volutional neural networks using textual descriptions,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.183592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T05:01:38.671276Z digest=sha256:a1a31b3d4d3787dfef3c864ed7c535cc4ab8ec6146b28e954319f0fa1909895e

Observation 7f9402c1-e700-4568-98b8-c32d1ff27321 · outbound

This paper cites Integrating language guidance into vision-based deep metric learning,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Integrating language guidance into vision-based deep metric learning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.168740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T05:01:38.675614Z digest=sha256:500da8247e75e7926463f03e7c4d558bc483c3d1e17fb3d60d53d0497ce3e40a

Observation 9f53574b-4bc0-45cc-9953-91bb9bf5360b · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.679692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.679692Z digest=sha256:c8e77c9cd71360bd5f1e0ebdfd5ecfbbfe0005af276e09a4b3e2fb28de080258

Observation ec336e35-c0ad-496c-ac41-f1363c5c8ddb · outbound

This paper cites Adding conditional control to text-to-image diffusion models,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Adding conditional control to text-to-image diffusion models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.684397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.684397Z digest=sha256:48cfc49362315c6bd9f55f798477145fbbea815dad8d75f9d38697ba57f90593

Observation 5c52080e-0a92-40c9-9e0b-7bfaf9c2c1e2 · outbound

This paper cites The benefit of temporally-strong labels in audio event classification,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance The benefit of temporally-strong labels in audio event classification,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.145445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T05:01:38.688822Z digest=sha256:3cd16a89a948a1f581d74d3c5fcda4d97f8943cb2be0736501d87a55d131e31b

Observation 828f7af1-5619-4181-8f10-f5d78f21b61c · outbound

This paper cites Vggsound: A large- scale audio-visual dataset,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Vggsound: A large- scale audio-visual dataset,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.130794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T05:01:38.693180Z digest=sha256:319ddc7f619206d8237ca913bc15287c6a0e85d134f112e1af87ef3d01af81f5

Observation daa08241-334d-499d-8abf-6c34faf50bf4 · outbound

This paper cites Text-to-audio grounding: Build- ing correspondence between captions and sound events,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Text-to-audio grounding: Build- ing correspondence between captions and sound events,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.116072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T05:01:38.697764Z digest=sha256:8be4d901442debe48fe12cb1be66938b58e3386d01206c813e261344df654b2f

Observation 7d0daffe-6166-48ba-a9fb-836f29a9d52f · outbound

This paper cites Towards Weakly Supervised Text-to-Audio Grounding.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Towards Weakly Supervised Text-to-Audio Grounding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.702238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.702238Z digest=sha256:bad9bd5a43fbceb7d171aa3074cd908ca1b4f0af49577aabb7574cd69f34066e

Observation a8e132c8-e7ce-42df-8f5e-1e48c1b3d1a3 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.707509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.707509Z digest=sha256:6576e3e4cc5720db1f0c16f835391f6ece5211b1093e4629bdc811300fd67053

Observation 52cf62e5-1d30-42d6-8671-3b9ea584e164 · outbound

This paper cites Correlation of Fr\'echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Correlation of Fr\'echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.712224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.712224Z digest=sha256:09a2617d09ffe964d304cc752d862030d95842e5661272c7b59749444b7bccd3

Observation 51b0d9ab-a0c1-456f-aa1c-75de199c2702 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.716900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.716900Z digest=sha256:d2e6d00ca5625ecf24e8734c7652cd3fcb53ffea31d3622e502adc17571049ef

Observation dd755cc4-22e3-4c89-81ff-0e3ab5c6b53d · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.091317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T05:01:38.721295Z digest=sha256:1269e50afff1842c68b8f87ed2a6582e30f33b1a7288b00266bf6c33481d9580

Observation 1e89bd65-4bed-4f43-b4ec-6e44e35a5939 · outbound

This paper cites Cnn architectures for large-scale audio classification,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Cnn architectures for large-scale audio classification,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.076628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T05:01:38.725435Z digest=sha256:5ca845a642b3cc522ff768f1f85b2404d3f5c80b49f5d784dd457ee3446e9a4b

Observation 13e56c13-8629-482d-85ad-4487522ced56 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.061764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T05:01:38.729633Z digest=sha256:ebf4cb773e5e2c4a50c359dcbbcc75cb59ce1e167751eba7ae5bab0d422583f7

Observation abc35f88-6c8d-4329-ac81-183c089e950c · outbound

This paper cites Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.734637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.734637Z digest=sha256:b798595fa51acca9e0a5b8782f352933b3616881da0e5c4eeb93244432853fa7

Observation 4df78d47-57ae-4bef-acd9-d64e21b250ed · outbound

This paper cites Video-foley: Two-stage video-to- sound generation via temporal event condition for foley sound,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Video-foley: Two-stage video-to- sound generation via temporal event condition for foley sound,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.739388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.739388Z digest=sha256:08f6f1508af776ce9a380234c814fafa6cb1fd4eef295189a75170ff4111e977

Observation 78d8003f-9395-4bc1-90d8-f339f896c700 · outbound

This paper cites Wav2clip: Learning robust audio representations from clip,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Wav2clip: Learning robust audio representations from clip,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.047018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T05:01:38.743732Z digest=sha256:4fbf8b5802b158788ea0b47dbecf74a954854231bc7eef09c57c5ff88a620321

Observation fb97627d-ccf9-465d-8054-401f6636873f · outbound

This paper cites Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.031776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T05:01:38.747896Z digest=sha256:924fc9eafb9fbda663a5dd75e53d0fecf2d456d94a2be26903e8f426e5e80f6f

Pith citing papers

No inbound Pith citation observations are available.