Pith. sign in

Paper Citation Record · LEDGER

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation

As of 19 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2508.12918.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.12918 v2

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:21:43.256404Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:41:34.359873Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T15:41:34.547908Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact2
  • verified fuzzy24
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 634fd715-de61-4ec3-9222-fc0b42e08cb8 · outbound

This paper cites Visual to sound: Generating natural sound for videos in the wild,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Visual to sound: Generating natural sound for videos in the wild,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:44.200768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.027493Z digest=sha256:c27eb3fc0fbc1ac15ad951913590c47dfb26abecc8e255fa964a961e444f6cb5

Observation 1c25f619-bcc7-4fe7-9d59-63485da9ea69 · outbound

This paper cites Generating visually aligned sound from videos,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Generating visually aligned sound from videos,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:44.183376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.034677Z digest=sha256:14a79132cfdb04af9f3774d9d0be896994ba75e180468382c7a7e754b8de8042

Observation 14d8d485-3a71-43f2-a02b-a7a4b0c4b893 · outbound

This paper cites Taming Visually Guided Sound Generation.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Taming Visually Guided Sound Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:43.040880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:43.040880Z digest=sha256:e834b8c970d5808d2994a09cabdb0d9c825c7d46bb44a68ef6b8cbc34351099b

Observation e43d8064-4de2-473f-be73-47c4f09ba2af · outbound

This paper cites Diff-foley: Synchro- nized video-to-audio synthesis with latent diffusion models,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Diff-foley: Synchro- nized video-to-audio synthesis with latent diffusion models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:44.169148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.046097Z digest=sha256:7536342eb0117a6c225374ded9fc2fc52a40f22b7e653819755316b3d0a52a82

Observation 37cc7545-c9ea-4137-b28d-e8dfa059c323 · outbound

This paper cites Frieren: Efficient video-to-audio generation network with rectified flow matching,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Frieren: Efficient video-to-audio generation network with rectified flow matching,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:43.051962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:43.051962Z digest=sha256:2dba910bd3ec3942495ed78fd61dd644a997cbc34769dce74fc109725791fab8

Observation 2ce1d1e6-2bc4-4162-87a2-de85026c3551 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:43.057249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:43.057249Z digest=sha256:9b1b67eea56016588a1f3f9cc4a9c6d13d139f8d0618d7157e0a7b2878c4ef9e

Observation dba47347-d8b9-448c-aa1d-10d3e86ba961 · outbound

This paper cites V2a- mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation V2a- mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:44.139203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.063835Z digest=sha256:ee118b0bcda7a3e64dfa4c1e59ed7660aabd4810cb886cfdef04faa8e67bd194

Observation ab791b33-c29c-48dd-a3d7-c98694753ba5 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Learning transferable visual models from natural language supervision,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:43.069268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:43.069268Z digest=sha256:764a5853f7c4e181ec8e1fe1f6bbd1452ef138bb972b1694cca2ed94edd838b4

Observation b392579d-4bd6-4231-af63-b8a24d5553d8 · outbound

This paper cites Clap learning audio concepts from natural language supervision,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Clap learning audio concepts from natural language supervision,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:43.073200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:43.073200Z digest=sha256:f0f75abf81e0fd544f9ca6f288bc9dcdd3b34e6623bec8a86d3aaf61d30fa89d

Observation 4145e155-a408-496d-9784-c893c8cccf14 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:43.078213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:43.078213Z digest=sha256:33169b6911156a78bdd554b77b94dc83c6944d54abe75909d5eec4c1625ec2d8

Observation a70c76f8-6cba-4d0a-9a7f-ef673a9e7e92 · outbound

This paper cites Mmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Mmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:44.095240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.082858Z digest=sha256:f37d9813224dd7df0e25a0149918ee4e8ef6643b272998bf286fc1471f4e637d

Observation 1e157850-3b2c-46bd-9672-6859671aeb04 · outbound

This paper cites AudioX: A Unified Framework for Anything-to-Audio Generation.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation AudioX: A Unified Framework for Anything-to-Audio Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:43.088608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:43.088608Z digest=sha256:5f2799e09a285f0460c7be50f1e66feaafdf42986fec9c5d1c40fa3bf03d4889

Observation 3fa8e323-1186-40fd-8084-214010a83296 · outbound

This paper cites Thinksound: Chain-of-thought reasoning in multi- modal large language models for audio generation and editing,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Thinksound: Chain-of-thought reasoning in multi- modal large language models for audio generation and editing,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:43.093884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:43.093884Z digest=sha256:791dd4345b66db1998a3fe84d007191d55320c96292419f80bdf0c23a41a0d45

Observation b04b8288-e6b9-4f07-ae4b-ab4e6a986df2 · outbound

This paper cites On our perception of the direotion of a source of sound,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation On our perception of the direotion of a source of sound,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:44.081235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.098196Z digest=sha256:e52824c373f2dae006e9d82a39c691e0d622274b5cd22612ce8a681c5f8c102f

Observation 45540763-bc2c-483f-bda9-cedcd0e7f423 · outbound

This paper cites Binauralgrad: A two-stage conditional diffusion probabilistic model for binaural audio synthesis,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Binauralgrad: A two-stage conditional diffusion probabilistic model for binaural audio synthesis,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:44.066106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.102600Z digest=sha256:ed0483aaaad90412af730ecef6d1d700674e3cdf6b6ff26aca55c97f4295e235

Observation 71feed8d-ad26-431e-baa4-a6dbff4ce751 · outbound

This paper cites SEE-2-SOUND: Zero-Shot Spatial Environment-to-Spatial Sound.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation SEE-2-SOUND: Zero-Shot Spatial Environment-to-Spatial Sound

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:43.107637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:43.107637Z digest=sha256:a50e33742999ad08db96f1bcb42d6d26c377734fc5093acaae2d67a380fae57f

Observation b4c46e2d-4fcc-4609-8663-e2949880deb4 · outbound

This paper cites 2.5 d visual sound,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation 2.5 d visual sound,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:44.040861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.112550Z digest=sha256:05b936b87260232e51f86a1942ce3282cdde46fb2fc7ca8f063ccabb6c2df886

Observation b6e8dc6a-a3a3-4b41-9e87-03707602a641 · outbound

This paper cites Visually-guided audio spa- tialization in video with geometry-aware multi-task learning,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Visually-guided audio spa- tialization in video with geometry-aware multi-task learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:44.021922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.117453Z digest=sha256:14cf08026378b39d684b5d2567a1d1e580dcf357597bb4534f8e2bf2c74d1982

Observation 64cf3ee9-d638-4209-92ed-ad462cbe637f · outbound

This paper cites Visually guided binaural audio generation with cross-modal consistency,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Visually guided binaural audio generation with cross-modal consistency,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:44.001856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.122580Z digest=sha256:e929d8b84d229a275be7d572f5537b81e5a5265816c301a4540e94b4a6528e65

Observation 52c7a403-3bb6-4a9e-b095-6757b1f884e8 · outbound

This paper cites Cross-modal generative model for visual-guided binaural stereo generation,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Cross-modal generative model for visual-guided binaural stereo generation,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:43.983707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.127212Z digest=sha256:e98a601d48581e41d287fdad40bd820cca5273a42b9a57415e8bb321b0e0ed17

Observation ac538c02-559f-4f20-9ff7-ce1d3cf88416 · outbound

This paper cites ImmerseDiffusion: A Generative Spatial Audio Latent Diffusion Model.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation ImmerseDiffusion: A Generative Spatial Audio Latent Diffusion Model

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:21:43.477011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.131398Z digest=sha256:9977ef1ef07e59d39b94debbcb541fd6ea45b7ca5deb03be66844eea44afe2a8

Observation f461fcbf-7193-46a3-a738-86a7273a1a5c · outbound

This paper cites Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:43.136060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:43.136060Z digest=sha256:26f51963b3ac7fadaadb9233ebbbdb9a25330bda8adabb83d5802ddd7dbbbd9a

Observation 216e6c86-42ad-481f-9892-a745f735b33d · outbound

This paper cites Visage: Video-to-spatial audio generation,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Visage: Video-to-spatial audio generation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:43.965132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.141631Z digest=sha256:f5e4bcdb602411c68abfb0701037b19620d8f66ff0b1aa9345d523cfa7b2b268

Observation bfd2d03f-b1c0-4a9c-a520-33339dbfbc99 · outbound

This paper cites OmniAudio: Generating Spatial Audio from 360-Degree Video.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation OmniAudio: Generating Spatial Audio from 360-Degree Video

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:43.146268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:43.146268Z digest=sha256:2ddf65df1387bc60c72477df74850a4a9939e2e2eb795568981d87684db7d132

Observation 7e1036c8-21d4-4e82-b6e6-48612867b93b · outbound

This paper cites Yolo-world: Real-time open-vocabulary object detection,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Yolo-world: Real-time open-vocabulary object detection,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:43.943830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.151063Z digest=sha256:5d52025abbdcb887e1adad1cfd38238c6badbc05df66afa7b84f7cb36c5b1f07

Observation 947ee8dd-1e71-463a-80cc-dd061c17d3a6 · outbound

This paper cites DepthMaster: Taming Diffusion Models for Monocular Depth Estimation.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation DepthMaster: Taming Diffusion Models for Monocular Depth Estimation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:43.156711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:43.156711Z digest=sha256:700573721b4d581cf7f4758df9736a546d20ff082d2bd2cbf481505757584ff0

Observation 8df860bf-3109-4351-8426-bead4cdece57 · outbound

This paper cites DiffWave: A Versatile Diffusion Model for Audio Synthesis.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation DiffWave: A Versatile Diffusion Model for Audio Synthesis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:43.162088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:43.162088Z digest=sha256:6b84d3d543a73360e66dc1682dc08ee48ebced30ab26ccc9452c1d2b2a6c36c4

Observation 51d1486a-3044-4d1c-b003-7a020efbe367 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Vggsound: A large-scale audio-visual dataset,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:43.924828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.167653Z digest=sha256:01441be02270cb4662f3b2ece096ba9c80c1c0075696fa184819330fc8c9d870

Observation 4ee85f91-8f27-44bd-8f61-6365124042fb · outbound

This paper cites The hutubs hrtf database,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation The hutubs hrtf database,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:43.902520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.173339Z digest=sha256:76cf941a3ffd801826963e11df7bfc470b68686ab385d4c2260541c072750f78

Observation 327ebba4-9998-4acf-9dc1-d825a850cfb0 · outbound

This paper cites Interpolation of head-related transfer functions,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Interpolation of head-related transfer functions,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:43.887411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.177132Z digest=sha256:1b258bd272cb61e4cb73a7aa194a16e2fca133eb179a54d84ed8a793083790d6

Observation 09cb53e3-85a0-4d9f-a60f-ff754d0196b3 · outbound

This paper cites gpurir: A python library for room impulse response simulation with gpu acceleration,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation gpurir: A python library for room impulse response simulation with gpu acceleration,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:43.868453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.183072Z digest=sha256:1a73c9a365180f489518c3f2e4f3d90259cc58f16720765cb6a0487db6c4de5f

Observation 4879aecf-5d9b-4f95-a1d6-6c4ae0b976fc · outbound

This paper cites VGGSound-Solo,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation VGGSound-Solo,

Reference 32

Resolution
verified exact
doi, observed 2026-08-15T17:21:43.306330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.187312Z digest=sha256:c19775e4b8c239f83f6dd76692e9c96f6dd5046c7c8887a7883c98a5a351e19d

Observation 317afbdb-f592-4f26-9ec8-00a0f368314f · outbound

This paper cites Image method for efficiently simulating small-room acoustics,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Image method for efficiently simulating small-room acoustics,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:43.841792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.194308Z digest=sha256:5a286e98ab093c70c857e3a069b5c0f65a41191de34203f43b63d0b34364a571

Observation 5075d5b9-e663-4e14-a36b-a1b86f80e065 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:43.818158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.199702Z digest=sha256:436b5478a88f88e93d56d0fcef336101c715e62bbed97ef404229d9b50cccb77

Observation 1889291e-32fe-4df7-899f-b4d126060da3 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Audio set: An ontology and human-labeled dataset for audio events,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:43.207546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:43.207546Z digest=sha256:398c44994ff5c24c28ad09ff8a5d8e7988da95b13e2ed4f1aee77e821001e721

Observation 22a32682-3267-40d1-bff4-a0110f0ea2cd · outbound

This paper cites Effi- cient training of audio transformers with patchout,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Effi- cient training of audio transformers with patchout,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:43.783979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.218905Z digest=sha256:8c6eff9c138f60c8afa799a27df86a4d6cd085103d0d53fca281e94beec964f5

Observation c47ae119-6ff3-4ed3-a697-ccd8b6d11da3 · outbound

This paper cites Improved techniques for training gans,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Improved techniques for training gans,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:43.766574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.224501Z digest=sha256:f17e348567fbe0a07d8a4d8b6281905d07652a103f3ad5103041a72a7fcfb815

Observation 95bfd8e5-77ba-4b93-bdf7-71476b7b07c5 · outbound

This paper cites Temporally aligned audio for video with autoregression,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Temporally aligned audio for video with autoregression,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:43.745587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.229533Z digest=sha256:29abcd4a5122e2ab04d42d85b542ee90a84d4458ddf797ca67079666c141a99c

Observation 2f482fc7-5f0b-402b-ae2f-3db2fb57042e · outbound

This paper cites DualSpec: Text-to-spatial-audio Generation via Dual-Spectrogram Guided Diffusion Model.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation DualSpec: Text-to-spatial-audio Generation via Dual-Spectrogram Guided Diffusion Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:43.238079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:43.238079Z digest=sha256:964897302ad3b41f1fa45406b079d8de13aae9d8f898ae0ae4ecdc6b619279db

Observation 075c5672-fb03-446c-b4e6-e7db40dacc66 · outbound

This paper cites Microsoft coco: Common objects in context,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Microsoft coco: Common objects in context,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:43.722940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.243812Z digest=sha256:1a1cf7ce5e5f288ef5f0dfc1f9d1bf9f91d54e3679dd7f1c96d00cc15b40a4ba

Observation b9a46742-bfed-40b6-b0e0-84c4c9f2be2c · outbound

This paper cites Imagenet: A large-scale hierarchical image database,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Imagenet: A large-scale hierarchical image database,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:43.251406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:43.251406Z digest=sha256:ef6bb0fd0a55dbbe2773b5ff494469c3d67feb4f672d3b808dafd1b913813209

Observation 68a4f987-16da-4c66-b57d-f0b12f3fa4ca · outbound

This paper cites Pyroomacoustics: A python package for audio room simulation and array pro- cessing algorithms,.

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation Pyroomacoustics: A python package for audio room simulation and array pro- cessing algorithms,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:21:43.692892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:21:43.256404Z digest=sha256:32567a5d1db4cc84109101c9077a2206ce6c291950f7d15b91021047f1ea751d

Pith citing papers

Observation 6173af98-3b2e-4a96-9062-41097d0cc30c · inbound

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework cites this paper.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework FoleySpace: Vision-Aligned Binaural Spatial Audio Generation

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:41:34.554124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.359873Z digest=sha256:735c71c54131e60635405252563820ceb9f4aaf67bb092668d70f19cad95f739