Pith. sign in

Paper Citation Record · LEDGER

Sounding that Object: Interactive Object-Aware Image to Audio Generation

As of 18 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2506.04214.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04214 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:53:18.476414Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e79acf52-0658-4e85-9d0d-4c7491b25a18 · outbound

This paper cites S., and Zisserman, A.

Sounding that Object: Interactive Object-Aware Image to Audio Generation S., and Zisserman, A

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:22.979644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:14.166587Z digest=sha256:e877cedb74251abff2fb3868f180a46ddb8a1c7d0efda805718f53a570e640d6

Observation 690029ba-d601-4347-8455-815102b30820 · outbound

This paper cites Effect of positional encoding.We assess the impact of positional encoding (PE) on our model’s performance.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Effect of positional encoding.We assess the impact of positional encoding (PE) on our model’s performance

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:19.495273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:18.261352Z digest=sha256:607f2ffe68eb30cc3dcd2a59ec75a32a30b146818497d82bbdd4c5f82dcaa91e

Observation f34a7d41-b1ed-4f67-8430-9330dc8265c9 · outbound

This paper cites L., Wu, H.-H., Salamon, J., and Bello, J.

Sounding that Object: Interactive Object-Aware Image to Audio Generation L., Wu, H.-H., Salamon, J., and Bello, J

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:15.135183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:15.135183Z digest=sha256:fd415a2bc3ab547306ba0a8a940fcd1c8c42371f41b8e4395bbea34c1cd40a93

Observation 24645505-7d44-4ab5-ab66-924f51f1f9f0 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Sounding that Object: Interactive Object-Aware Image to Audio Generation BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:15.258122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:15.258122Z digest=sha256:a256f0ff19a17f33e3d42b6a229ba0f6bd17e2e0da2d9a75029e796f5f2ed6fe

Observation 2ab79958-5864-400b-b386-d8ed413de65b · outbound

This paper cites CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos.

Sounding that Object: Interactive Object-Aware Image to Audio Generation CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:15.412519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:15.412519Z digest=sha256:0ff9ecab407506751d9c1fcd019cf1749f9c18287045b9ead81dd83a80675c39

Observation b84b5353-d282-4f1d-b3e2-4fb45c42dbfe · outbound

This paper cites We randomly selected 100 samples for evaluation, each rated by 50 unique participants to ensure reliability.

Sounding that Object: Interactive Object-Aware Image to Audio Generation We randomly selected 100 samples for evaluation, each rated by 50 unique participants to ensure reliability

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:19.678359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:18.190175Z digest=sha256:9c162c10b1e7fbba97bf99ece1da402b656e22c55daf9e721b83d72c78f2497c

Observation 306f4733-c98d-4a9a-a73c-6f422635d477 · outbound

This paper cites B., and Tor- ralba, A.

Sounding that Object: Interactive Object-Aware Image to Audio Generation B., and Tor- ralba, A

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:22.041748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:15.903051Z digest=sha256:d582c37be5b46958e11b3ef4804ac581851f9268ed21e5500f5d82927d8f7658

Observation 03ba7b75-97c6-4683-9c90-352d36f6924d · outbound

This paper cites Gotta Hear Them All: Towards Sound Source Aware Audio Generation.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Gotta Hear Them All: Towards Sound Source Aware Audio Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.014440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.014440Z digest=sha256:ca2ddfbfc943d5ddba7768cc30f8442d8b534cae88dc46b277a3078c9adc0722

Observation 15194fdf-c0ab-4eaf-b1f4-3b133068b7b8 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Classifier-Free Diffusion Guidance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.085528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.085528Z digest=sha256:57ba88a8ab25d0399991ca2255dae768fab15a683228a23d3e1c67453009effe

Observation 352ceb62-d23f-43df-85a7-88a5ea815f5b · outbound

This paper cites D., Kim, B., Lee, H., and Kim, G.

Sounding that Object: Interactive Object-Aware Image to Audio Generation D., Kim, B., Lee, H., and Kim, G

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.285900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.285900Z digest=sha256:c38fc887c64a44e7cfb530d08edaa35289d7b0e9da5352c2a1f2d244939d07b1

Observation 96613bab-25e7-4572-a390-b5f75b4c23af · outbound

This paper cites Hifi-gan: Generative ad- versarial networks for efficient and high fidelity speech synthesis.Advances in Neural Information Processing Systems, 33:17022–17033, 2020a.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Hifi-gan: Generative ad- versarial networks for efficient and high fidelity speech synthesis.Advances in Neural Information Processing Systems, 33:17022–17033, 2020a

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:21.367734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:16.567043Z digest=sha256:0dfe6d877575105da813e08a4a9a0f0881fe8f897b1e4253ffc9102ba40531fb

Observation 1ae440de-da3b-4066-a714-4e99af60fb63 · outbound

This paper cites Soundini: Sound-Guided Diffusion for Natural Video Editing.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Soundini: Sound-Guided Diffusion for Natural Video Editing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.636613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.636613Z digest=sha256:4387e2b095e5b6e3463f762fdad07e81fcd79ff34af4c247649d85b528707754

Observation a6df55b8-a4a8-4bb8-9fb1-8b679d97c704 · outbound

This paper cites Decoupled Weight Decay Regularization.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Decoupled Weight Decay Regularization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.702062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.702062Z digest=sha256:9e4f37ee65f598f6090a44bab1ff4c77a89d028f70f65866df5d9cbca46276b7

Observation e044c547-1dba-4001-add8-0f90575d7d5b · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Representation Learning with Contrastive Predictive Coding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.890076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.890076Z digest=sha256:d66ebe459cc67e41d86f4ec53ae8a7ac996bcce8cbf96b38cdbec6d654d0b985

Observation fab05a56-1416-46ff-9cb6-7079b351df7b · outbound

This paper cites A., Zhang, R., and Zhu, J.-Y.

Sounding that Object: Interactive Object-Aware Image to Audio Generation A., Zhang, R., and Zhu, J.-Y

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:21.187264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:17.029469Z digest=sha256:2e0dc7348f6e0f94f6539e337b8bf9053064e104667e442dc993f0a0e0646a90

Observation d66c4616-7dd0-47b8-ab14-93dde8116703 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Sounding that Object: Interactive Object-Aware Image to Audio Generation SAM 2: Segment Anything in Images and Videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.123505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.123505Z digest=sha256:1d13245904ce331616b763bd113c269fe4604380461149a0ed2d3425b3e30265

Observation d1fc9ea9-2aa6-43f3-a4b8-fe17c98098bb · outbound

This paper cites Self-supervised audio-visual co- segmentation.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Self-supervised audio-visual co- segmentation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:20.993648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:17.231123Z digest=sha256:99a4750ffc072bd63b3d6c693964cb2269cd9e6421efecf182b27872b8779c25

Observation 88f551ae-306b-4a73-9f74-e5bdc202b4ad · outbound

This paper cites and Adi, Y.

Sounding that Object: Interactive Object-Aware Image to Audio Generation and Adi, Y

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:20.773082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:17.324515Z digest=sha256:f037bcb1946970c68353c2a1da104cacb08e23db54f9b8b02277acf0144df2a3

Observation f18cc79b-1cf5-483d-8215-d94d3a14a6e3 · outbound

This paper cites Denoising Diffusion Implicit Models.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Denoising Diffusion Implicit Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.375318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.375318Z digest=sha256:e9dcdd0158461818c47b49ce237923a3606aacbd3a9e7ce3ea1d8c50c6c5235e

Observation 924826e6-0172-4273-b1ae-d6e2d83287d7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.445595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.445595Z digest=sha256:3875aacbe5a21cc08ee38b8036ec79ce583636b2d3fdb2ebe91714a49803b2cb

Observation 59579319-00f7-4000-b8af-ca7706245cb1 · outbound

This paper cites Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.538154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.538154Z digest=sha256:c7aae0b6ac61ada1c0cae0ba98ad5f36c530fcde9ae9959d74f89e0a3330790b

Observation eb6f9f6b-836a-4797-b06a-4a3e7a96640c · outbound

This paper cites P., and Salamon, J.

Sounding that Object: Interactive Object-Aware Image to Audio Generation P., and Salamon, J

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:20.586856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:17.603114Z digest=sha256:2bff316cb9ee513e7c9c25b140d5c26fd8f8e8e3b90a0e13633cdd42b0f65791

Observation 428750bf-740f-4dbf-b041-10c95fbbbbac · outbound

This paper cites Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.684335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.684335Z digest=sha256:aecbbc263fd90bfde813e536b0d97742768e5fcea20a8756948b3e0d26e29595

Observation f491f5a3-11f0-4614-94de-1a71b8fca4f4 · outbound

This paper cites Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.755251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.755251Z digest=sha256:62ca0203e53071c1575480048f6b84ceed94a2d1aa61503e85f4ff8cebf59005

Observation d226c09d-d252-4134-a793-847f7f69ca39 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

Sounding that Object: Interactive Object-Aware Image to Audio Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.851996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.851996Z digest=sha256:ed246cb26e1290f3e3a91223442936f88531ecf8ab5e27efbb11a7ea1032a1be

Observation 014a0f98-9611-4d99-9212-b584e52053f6 · outbound

This paper cites Results Video We provide a results video on the project webpage, which showcases our model’s ability to generate sounds based on the masked object prompts.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Results Video We provide a results video on the project webpage, which showcases our model’s ability to generate sounds based on the masked object prompts

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:20.349641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:17.911045Z digest=sha256:54146579bb09569a3996d75d389c4de046d557f6f3c741798c2eb75d702e8ec4

Observation f2e512bb-446e-45db-9b36-ab6e5a2b0c89 · outbound

This paper cites The original dataset comprises 4,616 hours of video clips, each paired with corresponding labels and captions.

Sounding that Object: Interactive Object-Aware Image to Audio Generation The original dataset comprises 4,616 hours of video clips, each paired with corresponding labels and captions

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:20.122666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:17.979108Z digest=sha256:8196f2c0e85358ca7c8f1a165145354a6f5a287108e73a0867eb4379f9660795

Observation 89c7e839-6d02-4d91-b3e5-4d8946b2b60b · outbound

This paper cites Speech" and “Music.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Speech" and “Music

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:19.887206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:18.061160Z digest=sha256:c637c3a34ef5030abbb7ba84c9f6ad7cd598045cb878d35052e48c17c4bfda89

Observation c4b27ee5-ee27-4ffb-95ef-e47c2b6e99f3 · outbound

This paper cites By extracting features from both modalities and computing cosine similarity, we show in Table 9 that our method consistently outperforms baselines on this metric.

Sounding that Object: Interactive Object-Aware Image to Audio Generation By extracting features from both modalities and computing cosine similarity, we show in Table 9 that our method consistently outperforms baselines on this metric

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:19.281192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:18.389315Z digest=sha256:1ad728c16fca5c4daabfd616e80237617a70d0af507abff811405f948f0615b0

Observation eb734ed4-bcef-4f39-9e5b-731ed4ad6714 · outbound

This paper cites video clips with better audio-visual synchronization, for test- ing.

Sounding that Object: Interactive Object-Aware Image to Audio Generation video clips with better audio-visual synchronization, for test- ing

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:19.118786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:18.476414Z digest=sha256:cdb85c673896866074f6b21c47644d0eb3d43c10f29f969164f82d7a12239b22

Observation 05235672-3c85-44af-928a-a5ffbf6dde8a · outbound

This paper cites MONet: Unsupervised Scene Decomposition and Representation.

Sounding that Object: Interactive Object-Aware Image to Audio Generation MONet: Unsupervised Scene Decomposition and Representation

Reference 1994

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:14.410618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:14.410618Z digest=sha256:d51e1e4cfcb6133994da1266e76cbef4b5ed4966a7be9364977ed5eb2150dc78

Observation 36292d4d-164e-47c1-a47d-b26706e60b98 · outbound

This paper cites Auto-Encoding Variational Bayes.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Auto-Encoding Variational Bayes

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.344543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.344543Z digest=sha256:baf065936bdf72bdd72fbf08db385c595ced0529af3bd5eb91173880339e9c11

Observation 10f7d14b-2f4b-4782-bd2e-6c614d482d2d · outbound

This paper cites D., Carr, C., Zukowski, Z., Taylor, J., and Pons, J.

Sounding that Object: Interactive Object-Aware Image to Audio Generation D., Carr, C., Zukowski, Z., Taylor, J., and Pons, J

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:15.739919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:15.739919Z digest=sha256:7e57e1becde0cacf805438c51cea83a922378eb47d201e09bba9ae333ef23766

Observation a35530dd-e4d0-41b2-9fe0-2674cf50081d · outbound

This paper cites Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.772565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.772565Z digest=sha256:777f783a38ec63f01796af19137f729e7052f90bd83140cac51c8db683696040

Observation 9b5802cf-c42f-4ba2-a2ef-a519db155eee · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Neural Machine Translation by Jointly Learning to Align and Translate

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:14.283844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:14.283844Z digest=sha256:ae16547343f2b061e184e3c6b34b31266cd29ac0491c9652c573e25971822c42

Observation 75e97090-91e3-411d-88d0-3d00fd96cdb5 · outbound

This paper cites Visual acoustic matching.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Visual acoustic matching

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:22.719167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:14.542529Z digest=sha256:2aa68d724f371cc04498a2c9f32cff9ac33ab0d792e81329b0c0177d4ca1d946

Observation 8ac5f492-375f-4c66-a64e-f9ad9a2bc2a7 · outbound

This paper cites Audio-Visual Synchronisation in the wild.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Audio-Visual Synchronisation in the wild

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:14.750169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:14.750169Z digest=sha256:4d57401c9a26127f3bacef6574c4ced1df103916593635438d8f86bc24937697

Observation 92982670-0cc1-4f22-9e56-f5e5bfa7c2e8 · outbound

This paper cites Synch- former: Efficient synchronization from sparse cues.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Synch- former: Efficient synchronization from sparse cues

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:21.806357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:16.157720Z digest=sha256:8c5c0eb4d9d8690b62beea2c39caed287cb867e033e98897dae21146ca3cdb88

Observation ba9961c9-9d3c-42a4-9dfd-ff2d9a87981d · outbound

This paper cites On uni-modal feature learning in supervised multi-modal learning.

Sounding that Object: Interactive Object-Aware Image to Audio Generation On uni-modal feature learning in supervised multi-modal learning

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:22.306427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:15.611374Z digest=sha256:8bf5d81a3e1cde6935cb5db2e20afc7139bbd95a6887b2f12a816ef0fdf6e8ae

Observation f344cc3e-447f-4383-8639-dbe578e98bd8 · outbound

This paper cites S., Wiles, O., Moses, Y ., and Zisserman, A.

Sounding that Object: Interactive Object-Aware Image to Audio Generation S., Wiles, O., Moses, Y ., and Zisserman, A

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:21.536182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:53:16.461969Z digest=sha256:b34c9c6647e0e039795e21da5e866c889d8c46726ff82b9972cc6ad70b1e6251

Observation 14e806fc-87e0-48fa-9c31-79ec9b6b96c0 · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

Sounding that Object: Interactive Object-Aware Image to Audio Generation MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:14.895706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:14.895706Z digest=sha256:a6bf09fd1df7efd81846edf29a50a013507820434a77a14ef93f8e7cd3c4f1d4

Pith citing papers

No inbound Pith citation observations are available.