Pith. sign in

Paper Citation Record · LEDGER

Sounding that Object: Interactive Object-Aware Image to Audio Generation

As of 8 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2506.04214.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04214 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:53:18.476414Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e79acf52-0658-4e85-9d0d-4c7491b25a18 · outbound

This paper cites S., and Zisserman, A.

Sounding that Object: Interactive Object-Aware Image to Audio Generation S., and Zisserman, A

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:22.979644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:53:14.166587Z digest=sha256:4007efcad29d704a82798be63a6ef400272cf121c68af8edfad2f15129ed0043

Observation 690029ba-d601-4347-8455-815102b30820 · outbound

This paper cites Effect of positional encoding.We assess the impact of positional encoding (PE) on our model’s performance.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Effect of positional encoding.We assess the impact of positional encoding (PE) on our model’s performance

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:19.495273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:53:18.261352Z digest=sha256:704dbc89023afd95129f1721dcf9b7774b22d36715fab6dec283d04434b4e6b5

Observation f34a7d41-b1ed-4f67-8430-9330dc8265c9 · outbound

This paper cites L., Wu, H.-H., Salamon, J., and Bello, J.

Sounding that Object: Interactive Object-Aware Image to Audio Generation L., Wu, H.-H., Salamon, J., and Bello, J

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:15.135183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:15.135183Z digest=sha256:46af5ffb149d2a3a2bbda4772f509b18a60df46bdedd1580741912d22db941d4

Observation 24645505-7d44-4ab5-ab66-924f51f1f9f0 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Sounding that Object: Interactive Object-Aware Image to Audio Generation BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:15.258122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:15.258122Z digest=sha256:ff7d0af2086adaf09c842e3139130edacd377719dd95e89d830f7a6942520b82

Observation 2ab79958-5864-400b-b386-d8ed413de65b · outbound

This paper cites CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos.

Sounding that Object: Interactive Object-Aware Image to Audio Generation CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:15.412519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:15.412519Z digest=sha256:570436f07d70dfd7b5611b6a3162b6c37153fcd77ce639e2b36420996ab456d6

Observation b84b5353-d282-4f1d-b3e2-4fb45c42dbfe · outbound

This paper cites We randomly selected 100 samples for evaluation, each rated by 50 unique participants to ensure reliability.

Sounding that Object: Interactive Object-Aware Image to Audio Generation We randomly selected 100 samples for evaluation, each rated by 50 unique participants to ensure reliability

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:19.678359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:53:18.190175Z digest=sha256:4ee204fd49de2d7329ff2d93030808c452aa4b4f65c14a61d8d817c5feab16f6

Observation 306f4733-c98d-4a9a-a73c-6f422635d477 · outbound

This paper cites B., and Tor- ralba, A.

Sounding that Object: Interactive Object-Aware Image to Audio Generation B., and Tor- ralba, A

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:22.041748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:53:15.903051Z digest=sha256:51f26795732c5473751000adc53122f6d72ddf1d7cec51df109019954f6e7d1e

Observation 03ba7b75-97c6-4683-9c90-352d36f6924d · outbound

This paper cites Gotta Hear Them All: Towards Sound Source Aware Audio Generation.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Gotta Hear Them All: Towards Sound Source Aware Audio Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.014440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.014440Z digest=sha256:2b4017aa88458f844da58998eff84cd31afd1874e0a0d53410bd58fcd66585d9

Observation 15194fdf-c0ab-4eaf-b1f4-3b133068b7b8 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Classifier-Free Diffusion Guidance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.085528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.085528Z digest=sha256:a2e841d94079d0f5d217c0bfe34d5521ade5a3a3d2675f8164a285c04cc2a949

Observation 352ceb62-d23f-43df-85a7-88a5ea815f5b · outbound

This paper cites D., Kim, B., Lee, H., and Kim, G.

Sounding that Object: Interactive Object-Aware Image to Audio Generation D., Kim, B., Lee, H., and Kim, G

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.285900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.285900Z digest=sha256:b8351fb3d4049d9d793e45553a16b3f2944f302c3336bf3d97878d54a3aaeebb

Observation 96613bab-25e7-4572-a390-b5f75b4c23af · outbound

This paper cites Hifi-gan: Generative ad- versarial networks for efficient and high fidelity speech synthesis.Advances in Neural Information Processing Systems, 33:17022–17033, 2020a.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Hifi-gan: Generative ad- versarial networks for efficient and high fidelity speech synthesis.Advances in Neural Information Processing Systems, 33:17022–17033, 2020a

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:21.367734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:53:16.567043Z digest=sha256:e210b61260b24a14f7dcb9dd72180e690a958afe712d6a783f4029b2e1fa438d

Observation 1ae440de-da3b-4066-a714-4e99af60fb63 · outbound

This paper cites Soundini: Sound-Guided Diffusion for Natural Video Editing.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Soundini: Sound-Guided Diffusion for Natural Video Editing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.636613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.636613Z digest=sha256:8fe411347368405e976915c31053f326b07e2b5ccef50221d128bfca26a38a42

Observation a6df55b8-a4a8-4bb8-9fb1-8b679d97c704 · outbound

This paper cites Decoupled Weight Decay Regularization.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Decoupled Weight Decay Regularization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.702062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.702062Z digest=sha256:3b807b930464acdd4d846545e0ef1c60488e1b2f6abf8b3f6d22a93f0b412963

Observation e044c547-1dba-4001-add8-0f90575d7d5b · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Representation Learning with Contrastive Predictive Coding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.890076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.890076Z digest=sha256:4d0e0ca79afba516315ec41a7db8f7c6152dd54d83e6b6f5bf9d215c8d1c540c

Observation fab05a56-1416-46ff-9cb6-7079b351df7b · outbound

This paper cites A., Zhang, R., and Zhu, J.-Y.

Sounding that Object: Interactive Object-Aware Image to Audio Generation A., Zhang, R., and Zhu, J.-Y

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:21.187264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:53:17.029469Z digest=sha256:6d6bbc75d0ccb23801f7d6d925db77cadf1fe44201adeb44ca67f062c74373e4

Observation d66c4616-7dd0-47b8-ab14-93dde8116703 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Sounding that Object: Interactive Object-Aware Image to Audio Generation SAM 2: Segment Anything in Images and Videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.123505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.123505Z digest=sha256:4612eb888e5f3d1861d08293193cac6d82b423e81d0b3b15c16114526ea701cb

Observation d1fc9ea9-2aa6-43f3-a4b8-fe17c98098bb · outbound

This paper cites Self-supervised audio-visual co- segmentation.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Self-supervised audio-visual co- segmentation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:20.993648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:53:17.231123Z digest=sha256:e686ca02e3410315b47184f1fc3027581d3c77eb737041baf37d472f13cbbbb2

Observation 88f551ae-306b-4a73-9f74-e5bdc202b4ad · outbound

This paper cites and Adi, Y.

Sounding that Object: Interactive Object-Aware Image to Audio Generation and Adi, Y

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:20.773082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:53:17.324515Z digest=sha256:26db8346e9c08f90087fbbc9d438ba76aa953debe13926ba96c7b8a608437def

Observation f18cc79b-1cf5-483d-8215-d94d3a14a6e3 · outbound

This paper cites Denoising Diffusion Implicit Models.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Denoising Diffusion Implicit Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.375318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.375318Z digest=sha256:80bf4640a03db152b6fb997dde8c2f99c6c817a7777ec87ea58511f0fe5954cf

Observation 924826e6-0172-4273-b1ae-d6e2d83287d7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.445595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.445595Z digest=sha256:4faea1a9d2be1ace0bde7c93b98362213d00e395e9196c26b10c01541a4c3245

Observation 59579319-00f7-4000-b8af-ca7706245cb1 · outbound

This paper cites Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.538154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.538154Z digest=sha256:fd310081f32c54b1c466d95bc7df181635c921bf84d4716d36de4b381c92bc78

Observation eb6f9f6b-836a-4797-b06a-4a3e7a96640c · outbound

This paper cites P., and Salamon, J.

Sounding that Object: Interactive Object-Aware Image to Audio Generation P., and Salamon, J

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:20.586856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:53:17.603114Z digest=sha256:69905c1fcddd403ea89cd51e9fc91d6290ec989c638bb31e650fe0c4049c4a69

Observation 428750bf-740f-4dbf-b041-10c95fbbbbac · outbound

This paper cites Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.684335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.684335Z digest=sha256:f77ccebddad810cc0b3f55ba1d135c99c914f0f3a61a5b1f2cb7a2def5ef8ced

Observation f491f5a3-11f0-4614-94de-1a71b8fca4f4 · outbound

This paper cites Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.755251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.755251Z digest=sha256:946d325cc89dae9771928476e4d9cb7544e29284617bfac06f85cfd7b3827f74

Observation d226c09d-d252-4134-a793-847f7f69ca39 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

Sounding that Object: Interactive Object-Aware Image to Audio Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.851996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.851996Z digest=sha256:e78beb81af7f47563ca7696d9892fa36d355d58c4a27b4642a76f16b7830d7e1

Observation 014a0f98-9611-4d99-9212-b584e52053f6 · outbound

This paper cites Results Video We provide a results video on the project webpage, which showcases our model’s ability to generate sounds based on the masked object prompts.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Results Video We provide a results video on the project webpage, which showcases our model’s ability to generate sounds based on the masked object prompts

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:20.349641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:53:17.911045Z digest=sha256:db7c3490bfd62f3fe7837cd10eb2a9d87c24fd3d76a6b5a9c6bb3e7b08921fcb

Observation f2e512bb-446e-45db-9b36-ab6e5a2b0c89 · outbound

This paper cites The original dataset comprises 4,616 hours of video clips, each paired with corresponding labels and captions.

Sounding that Object: Interactive Object-Aware Image to Audio Generation The original dataset comprises 4,616 hours of video clips, each paired with corresponding labels and captions

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:20.122666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:53:17.979108Z digest=sha256:092dbefd8f33969575addf6dc4446b275c46ad9b8dddfd6904e94bfde0b145bd

Observation 89c7e839-6d02-4d91-b3e5-4d8946b2b60b · outbound

This paper cites Speech" and “Music.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Speech" and “Music

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:19.887206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:53:18.061160Z digest=sha256:345ad1d36783fb055b484dacac8ed77627df2cd8b7104190a4ef5b48862def0c

Observation c4b27ee5-ee27-4ffb-95ef-e47c2b6e99f3 · outbound

This paper cites By extracting features from both modalities and computing cosine similarity, we show in Table 9 that our method consistently outperforms baselines on this metric.

Sounding that Object: Interactive Object-Aware Image to Audio Generation By extracting features from both modalities and computing cosine similarity, we show in Table 9 that our method consistently outperforms baselines on this metric

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:19.281192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:53:18.389315Z digest=sha256:41e3454ffc48b116ec981cfbb39b386efe14b5c031772a09dff3ad592a5640c3

Observation eb734ed4-bcef-4f39-9e5b-731ed4ad6714 · outbound

This paper cites video clips with better audio-visual synchronization, for test- ing.

Sounding that Object: Interactive Object-Aware Image to Audio Generation video clips with better audio-visual synchronization, for test- ing

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:19.118786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:53:18.476414Z digest=sha256:d2b71f433daaddf803f6b6887abcbfa0b3dae9078bbbbc7557856ea292f4889c

Observation 05235672-3c85-44af-928a-a5ffbf6dde8a · outbound

This paper cites MONet: Unsupervised Scene Decomposition and Representation.

Sounding that Object: Interactive Object-Aware Image to Audio Generation MONet: Unsupervised Scene Decomposition and Representation

Reference 1994

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:14.410618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:14.410618Z digest=sha256:62bfb0843b4dd6309b09ded90968fc324aba5ba507f751495c8ea97d37bab160

Observation 36292d4d-164e-47c1-a47d-b26706e60b98 · outbound

This paper cites Auto-Encoding Variational Bayes.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Auto-Encoding Variational Bayes

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.344543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.344543Z digest=sha256:a940e1061f0af597595e58fecff5af48c4d56be9debbbd72a37d830192c4e304

Observation 10f7d14b-2f4b-4782-bd2e-6c614d482d2d · outbound

This paper cites D., Carr, C., Zukowski, Z., Taylor, J., and Pons, J.

Sounding that Object: Interactive Object-Aware Image to Audio Generation D., Carr, C., Zukowski, Z., Taylor, J., and Pons, J

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:15.739919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:15.739919Z digest=sha256:4b212160c4aad35c99575fc2939876dd1682575ea022a3df94327439a2796be2

Observation a35530dd-e4d0-41b2-9fe0-2674cf50081d · outbound

This paper cites Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.772565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.772565Z digest=sha256:da504d9a85980b329ee8ae351e7a4fac99ed9b8e5a964875a62fe728733e40eb

Observation 9b5802cf-c42f-4ba2-a2ef-a519db155eee · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Neural Machine Translation by Jointly Learning to Align and Translate

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:14.283844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:14.283844Z digest=sha256:8c72b0a2f65df50761d8944e6dc1b476233d1ffe50b4f5e0c21d31e9ca121105

Observation 75e97090-91e3-411d-88d0-3d00fd96cdb5 · outbound

This paper cites Visual acoustic matching.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Visual acoustic matching

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:22.719167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:53:14.542529Z digest=sha256:9cc2e8da6bbc99155dd34c3642f26cf96dedc13006b773b7e4e14061cc99c3ee

Observation 8ac5f492-375f-4c66-a64e-f9ad9a2bc2a7 · outbound

This paper cites Audio-Visual Synchronisation in the wild.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Audio-Visual Synchronisation in the wild

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:14.750169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:14.750169Z digest=sha256:6fd28371b80e8f8c7da16eedc80773cbcc7227178f4f7da0cab3ef0f9a75bd1f

Observation 92982670-0cc1-4f22-9e56-f5e5bfa7c2e8 · outbound

This paper cites Synch- former: Efficient synchronization from sparse cues.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Synch- former: Efficient synchronization from sparse cues

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:21.806357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:53:16.157720Z digest=sha256:adbcaa55141a0f93b899d85ee362ad27194ab614a4880ea662e94f51ce636c55

Observation ba9961c9-9d3c-42a4-9dfd-ff2d9a87981d · outbound

This paper cites On uni-modal feature learning in supervised multi-modal learning.

Sounding that Object: Interactive Object-Aware Image to Audio Generation On uni-modal feature learning in supervised multi-modal learning

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:22.306427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:53:15.611374Z digest=sha256:bceba79de677318a0cb3b90ae77e74c84c09d026d4021845670ccd92669b4384

Observation f344cc3e-447f-4383-8639-dbe578e98bd8 · outbound

This paper cites S., Wiles, O., Moses, Y ., and Zisserman, A.

Sounding that Object: Interactive Object-Aware Image to Audio Generation S., Wiles, O., Moses, Y ., and Zisserman, A

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:21.536182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:53:16.461969Z digest=sha256:ad6cfd4c10bde0f2e06cdf767d465ed43f9500c3893f8c9a3056e3d4d6bfcb27

Observation 14e806fc-87e0-48fa-9c31-79ec9b6b96c0 · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

Sounding that Object: Interactive Object-Aware Image to Audio Generation MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:14.895706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:14.895706Z digest=sha256:b0352aa9d01ec8f9916fd831285f7b0039cfd20bebb5ecf269d045f607a2ab25

Pith citing papers

No inbound Pith citation observations are available.