Pith. sign in

Paper Citation Record · LEDGER

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance

As of 23 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2412.18157.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18157 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:01:38.747896Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5dfb7121-e594-4dc0-93c3-7e824b842549 · outbound

This paper cites Generating visually aligned sound from videos,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Generating visually aligned sound from videos,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.315954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:01:38.603327Z digest=sha256:09c94308ccccfe189a29d54db6bf63ff0230509e043fda99199cc0e26d8c9ea0

Observation 7f37af99-e3dc-4270-8fd2-c3016a9657a5 · outbound

This paper cites Taming visually guided sound generation,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Taming visually guided sound generation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.301234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:01:38.608027Z digest=sha256:4500d03ad87c7d6cc7679e5d6436cdf29151c3e236c089cf3090b3cef93b58b3

Observation c4581f99-f988-4b81-9220-c5044e0f6b6a · outbound

This paper cites Attention is all you need,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Attention is all you need,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.612589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.612589Z digest=sha256:c77fd444a41b8dd1ac85d251a36172f1f9417647c97a82be13797e2382e16306

Observation 46a1da61-e450-446f-8cfc-a394399df2be · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Learning transferable visual models from natural language supervision,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.618248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.618248Z digest=sha256:a695178949ff52212fe4fda69b5bf85aba6f73bf432e29b3451e4b79e8d0630c

Observation a744e841-2c08-404c-8b66-45e555bc176e · outbound

This paper cites Imagebind: One embedding space to bind them all,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Imagebind: One embedding space to bind them all,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.623065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.623065Z digest=sha256:f793983b19f74b679cc4f06d5819d88db91c3aae053f691c1006975328bda39a

Observation c8e4a828-487c-4ba8-a2c2-779ae22072e6 · outbound

This paper cites Denoising diffusion probabilistic models,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Denoising diffusion probabilistic models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.628058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.628058Z digest=sha256:8434e03c75a24a3128a482f8f3dfeee2994a7ab0d3182b186f3ce310e4703526

Observation f99302b1-5d3d-4c06-ab10-26ca27acd214 · outbound

This paper cites Conditional generation of audio from video via foley analogies,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Conditional generation of audio from video via foley analogies,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.249712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:01:38.633488Z digest=sha256:bd07d3b4454304ab5f81dd1bb4f4559d4e0f1f25643a542c2f5cdf1b110a9b58

Observation 0de05aa5-1667-4472-bad8-014d5c258bf7 · outbound

This paper cites Varietysound: Timbre-controllable video to sound generation via unsupervised information disentanglement,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Varietysound: Timbre-controllable video to sound generation via unsupervised information disentanglement,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.235676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:01:38.637922Z digest=sha256:5f8423202e91d2e14a663ab09c008c4fb05804c58ecd922d751fbea9d3474997

Observation bb71c8fd-1a60-497d-bd44-2f50c84b26e6 · outbound

This paper cites Diff-foley: Synchronized video- to-audio synthesis with latent diffusion models,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Diff-foley: Synchronized video- to-audio synthesis with latent diffusion models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.221071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:01:38.642369Z digest=sha256:f1b7a30d3b5b39c3b8f347a8e95846449b4a2314a853d865ecf559834a61b13c

Observation b0baf4ea-d77a-4717-8936-db8d6dda74ef · outbound

This paper cites I hear your true colors: Image guided audio generation,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance I hear your true colors: Image guided audio generation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.206553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:01:38.647528Z digest=sha256:544e71a9580fce52d99e2c0ed2f5551cbd1a7b4e55ef37587b9307ed58d24652

Observation 82f2665c-0c2a-40d2-aeef-9889bc041184 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.651937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.651937Z digest=sha256:bb9136cca914c2ceac54b138d78ae3477081f08b03fe53d78b8022b3cfc71d97

Observation 5b036e71-af05-4ca9-8827-5dfeebc4e2ea · outbound

This paper cites Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.656647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.656647Z digest=sha256:28f3ee910077254c96e8011a1668197bf4d546e57f833b68628830b1e3535c63

Observation ebc93496-4ccc-4660-8eb4-33adb2f3926e · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.661167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.661167Z digest=sha256:1de548bc6d451c8e4f7f2314acbc1c45ab3d84bdfec0201121541fd5eb0de03f

Observation 66576f66-c5e7-4e9b-ae71-754e911cb725 · outbound

This paper cites Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.667045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.667045Z digest=sha256:fa2d493deb0e7542457e95c3241c8fd3759c258afc87fa15a6b9a8ed259d636b

Observation dc1c3c51-5603-4238-89e0-28dd1de65d9c · outbound

This paper cites Predicting deep zero-shot con- volutional neural networks using textual descriptions,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Predicting deep zero-shot con- volutional neural networks using textual descriptions,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.183592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:01:38.671276Z digest=sha256:c05f254f4c2987193236967201864a661b6ada52d11c600f156f91a84a561a34

Observation 7f9402c1-e700-4568-98b8-c32d1ff27321 · outbound

This paper cites Integrating language guidance into vision-based deep metric learning,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Integrating language guidance into vision-based deep metric learning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.168740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:01:38.675614Z digest=sha256:13454baaf14dbf121c9b0f290bf544802f8771b38af1d78eb24d695b756d7aea

Observation 9f53574b-4bc0-45cc-9953-91bb9bf5360b · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.679692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.679692Z digest=sha256:c8e77c9cd71360bd5f1e0ebdfd5ecfbbfe0005af276e09a4b3e2fb28de080258

Observation ec336e35-c0ad-496c-ac41-f1363c5c8ddb · outbound

This paper cites Adding conditional control to text-to-image diffusion models,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Adding conditional control to text-to-image diffusion models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.684397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.684397Z digest=sha256:48cfc49362315c6bd9f55f798477145fbbea815dad8d75f9d38697ba57f90593

Observation 5c52080e-0a92-40c9-9e0b-7bfaf9c2c1e2 · outbound

This paper cites The benefit of temporally-strong labels in audio event classification,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance The benefit of temporally-strong labels in audio event classification,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.145445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:01:38.688822Z digest=sha256:c44f5a52ae26ac2cf45448386031d403998e0559053b93d2e0935053a7f0bad8

Observation 828f7af1-5619-4181-8f10-f5d78f21b61c · outbound

This paper cites Vggsound: A large- scale audio-visual dataset,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Vggsound: A large- scale audio-visual dataset,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.130794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:01:38.693180Z digest=sha256:3b741e111b8467866b81a0a31921ee308d53b0fb7f1c882c5d6e4d43cc369339

Observation daa08241-334d-499d-8abf-6c34faf50bf4 · outbound

This paper cites Text-to-audio grounding: Build- ing correspondence between captions and sound events,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Text-to-audio grounding: Build- ing correspondence between captions and sound events,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.116072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:01:38.697764Z digest=sha256:4dca31d84b9882477fddab5f8b85a869627471b084f6f3ee073d587f508ba30f

Observation 7d0daffe-6166-48ba-a9fb-836f29a9d52f · outbound

This paper cites Towards Weakly Supervised Text-to-Audio Grounding.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Towards Weakly Supervised Text-to-Audio Grounding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.702238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.702238Z digest=sha256:bad9bd5a43fbceb7d171aa3074cd908ca1b4f0af49577aabb7574cd69f34066e

Observation a8e132c8-e7ce-42df-8f5e-1e48c1b3d1a3 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.707509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.707509Z digest=sha256:6576e3e4cc5720db1f0c16f835391f6ece5211b1093e4629bdc811300fd67053

Observation 52cf62e5-1d30-42d6-8671-3b9ea584e164 · outbound

This paper cites Correlation of Fr\'echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Correlation of Fr\'echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.712224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.712224Z digest=sha256:09a2617d09ffe964d304cc752d862030d95842e5661272c7b59749444b7bccd3

Observation 51b0d9ab-a0c1-456f-aa1c-75de199c2702 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.716900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.716900Z digest=sha256:d2e6d00ca5625ecf24e8734c7652cd3fcb53ffea31d3622e502adc17571049ef

Observation dd755cc4-22e3-4c89-81ff-0e3ab5c6b53d · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.091317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:01:38.721295Z digest=sha256:530635921a2d0247b1dde1cdf1e7f844abc8496e48cc5fc53d8fed96b28404f6

Observation 1e89bd65-4bed-4f43-b4ec-6e44e35a5939 · outbound

This paper cites Cnn architectures for large-scale audio classification,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Cnn architectures for large-scale audio classification,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.076628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:01:38.725435Z digest=sha256:9405571188ace725157d090bb10eb2157f5562f2453b6cf44746b0fb091ac3cd

Observation 13e56c13-8629-482d-85ad-4487522ced56 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.061764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:01:38.729633Z digest=sha256:a12b6a98932185e3c03d2653d7ebc5976d20100d152dd7bef9f8c8ec0b3b6263

Observation abc35f88-6c8d-4329-ac81-183c089e950c · outbound

This paper cites Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.734637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.734637Z digest=sha256:b798595fa51acca9e0a5b8782f352933b3616881da0e5c4eeb93244432853fa7

Observation 4df78d47-57ae-4bef-acd9-d64e21b250ed · outbound

This paper cites Video-foley: Two-stage video-to- sound generation via temporal event condition for foley sound,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Video-foley: Two-stage video-to- sound generation via temporal event condition for foley sound,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.739388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.739388Z digest=sha256:08f6f1508af776ce9a380234c814fafa6cb1fd4eef295189a75170ff4111e977

Observation 78d8003f-9395-4bc1-90d8-f339f896c700 · outbound

This paper cites Wav2clip: Learning robust audio representations from clip,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Wav2clip: Learning robust audio representations from clip,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.047018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:01:38.743732Z digest=sha256:c1e3490b789519cacbdc8e4b2cdda8255577c6bd77750357a465b9cea5a4a81a

Observation fb97627d-ccf9-465d-8054-401f6636873f · outbound

This paper cites Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms,.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:01:39.031776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T05:01:38.747896Z digest=sha256:b75e17391f5c260a0454c639aede6de5ecf88d579de202cd865d0af71eb8e226

Pith citing papers

No inbound Pith citation observations are available.