Pith. sign in

Paper Citation Record · LEDGER

Visual-based spatial audio generation system for multi-speaker environments

As of 9 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2502.07538.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07538 v2

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:30:26.197293Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:04:00.782536Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T22:54:55.849359Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 01c1e241-7246-4d53-9b27-b2b1450eed7a · outbound

This paper cites Exploring audio- visual information fusion for sound event localization and detection in low-resource realistic scenarios,.

Visual-based spatial audio generation system for multi-speaker environments Exploring audio- visual information fusion for sound event localization and detection in low-resource realistic scenarios,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.536832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:30:26.103584Z digest=sha256:64e1a7214fb8794cfea32bdf816fe93482397fec4cc2cb2d3ec911ddeae737ad

Observation 047e4701-6346-4cf6-8f4d-9ca1b7958006 · outbound

This paper cites Aligning audiovisual features for audiovisual speech recognition,.

Visual-based spatial audio generation system for multi-speaker environments Aligning audiovisual features for audiovisual speech recognition,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.522970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:30:26.108687Z digest=sha256:1dc9dbc4a14411af576312956a93d30dcc96e80ee54e2c6960b740ed7b5df474

Observation c3174224-52d0-4dab-9458-417abf7c78be · outbound

This paper cites Immersive spatial audio reproduction for vr/ar using room acoustic modelling from 360° images,.

Visual-based spatial audio generation system for multi-speaker environments Immersive spatial audio reproduction for vr/ar using room acoustic modelling from 360° images,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.508310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:30:26.113528Z digest=sha256:83aaf01ab07e7964bc77dc911865bd7e27e5ff7e3809a39d4ef2221edfb3ba50

Observation 0a28dc3b-e7de-4625-87b6-e6a8ac9ff269 · outbound

This paper cites Scene-aware audio for 360° videos,.

Visual-based spatial audio generation system for multi-speaker environments Scene-aware audio for 360° videos,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.494322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:30:26.118296Z digest=sha256:708f4a413136d2e2aea5ae10982c82df899de6789835510bac6e1f9f70f8a300

Observation f115ea40-7d76-4016-b312-4845ec80fcd7 · outbound

This paper cites Analysis of a distributed processing model for spatialized audio conferences,.

Visual-based spatial audio generation system for multi-speaker environments Analysis of a distributed processing model for spatialized audio conferences,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.480615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:30:26.122985Z digest=sha256:5daeb7c057cc04e4c428061227beefa18e5132a2423e72be65415d08dc0cd014

Observation 78e8b1d3-6199-452e-a843-782c4e2ff069 · outbound

This paper cites Realistic audio in immersive video conferencing,.

Visual-based spatial audio generation system for multi-speaker environments Realistic audio in immersive video conferencing,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.465535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:30:26.127638Z digest=sha256:abe4e872a6b51aedd3329586bfd9eab27e2667a3a7003b880b314861ca31be48

Observation 69c65982-9250-4479-9afb-2e3243dc5188 · outbound

This paper cites Audio-visual sensing from a quadcopter: dataset and baselines for source localization and sound enhancement,.

Visual-based spatial audio generation system for multi-speaker environments Audio-visual sensing from a quadcopter: dataset and baselines for source localization and sound enhancement,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.451213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:30:26.132615Z digest=sha256:fcd55eed68b7373990721f0a64da5d746f2c845d11b23c06b78b3c71c5c4309c

Observation 55096907-3dba-4bce-b6ba-1c792c2ef5a4 · outbound

This paper cites 2.5d visual sound,.

Visual-based spatial audio generation system for multi-speaker environments 2.5d visual sound,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.434867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:30:26.136964Z digest=sha256:00af4f9fcac1b66b7d458b0f3e4fddf26d424a4835c9e2ff1a05aba0c14fe9ca

Observation 118c667c-4c36-4529-bc12-a5fddc806369 · outbound

This paper cites Visually informed binaural audio generation without binaural audios,.

Visual-based spatial audio generation system for multi-speaker environments Visually informed binaural audio generation without binaural audios,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.418915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:30:26.141461Z digest=sha256:47c379be6fb8be868203407fe5928b6941a587ca2155ca1b8d71a2f1ea777171

Observation ed4ccc89-e858-4317-8a1e-85e3381ececf · outbound

This paper cites A review on yolov8 and its advancements,.

Visual-based spatial audio generation system for multi-speaker environments A review on yolov8 and its advancements,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.403157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:30:26.145960Z digest=sha256:921ee8e7c7c72c12e03c96b18b10592c7cee7663ea22893ceb450df82f3877f3

Observation 7687ebb2-2f03-4d28-adb5-1085d7bacbe0 · outbound

This paper cites Wider face: A face detection benchmark,.

Visual-based spatial audio generation system for multi-speaker environments Wider face: A face detection benchmark,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.388110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:30:26.150357Z digest=sha256:cbc422ab8c98c456c52c3b0de72a43366427277fe226f48231c77036ab721d53

Observation b017bd20-49bb-4854-bd06-289fa64f87f5 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data,.

Visual-based spatial audio generation system for multi-speaker environments Depth anything: Unleashing the power of large-scale unlabeled data,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.372085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:30:26.154803Z digest=sha256:037c5c89e1e17609dc75cb5430b91bacc9fccb9d5961db1947cf5c7c81339382

Observation 535446c5-1b96-47b9-aefd-f809af2ef142 · outbound

This paper cites You only look once: Unified, real-time object detection,.

Visual-based spatial audio generation system for multi-speaker environments You only look once: Unified, real-time object detection,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.356023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:30:26.160648Z digest=sha256:ac47bf8be15e760b24acf7c9d9a762eea424d0d8a3d3089abe7a3009c8e2a439

Observation e8396a1d-7eb7-4c8d-aedc-f42dbcdd999c · outbound

This paper cites Conv-tasnet: Surpassing ideal time– frequency magnitude masking for speech separation,.

Visual-based spatial audio generation system for multi-speaker environments Conv-tasnet: Surpassing ideal time– frequency magnitude masking for speech separation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.339161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:30:26.165241Z digest=sha256:0bfc86e396a0246882cf3d2c9f8ee53784264485274da3d30802a909a0c3aac3

Observation f09ed45e-7d93-432f-83d1-ddcd5ac4e1e3 · outbound

This paper cites Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed.

Visual-based spatial audio generation system for multi-speaker environments Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T12:30:26.169424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:30:26.169424Z digest=sha256:d1599f254eb45f2debb9069732d768c5fe9203a746fb1532d520f8b19c6eabca

Observation f87a8c9c-c216-4bbd-b62f-33cba63d9bda · outbound

This paper cites A perceptual evaluation of individual and non-individual hrtfs: A case study of the sadie ii database,.

Visual-based spatial audio generation system for multi-speaker environments A perceptual evaluation of individual and non-individual hrtfs: A case study of the sadie ii database,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.322485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:30:26.174917Z digest=sha256:c8b0da86dc19585cd66fcf67a1f56178cbd0f3dedc8b81e7e70f6e3e386ec638

Observation df9dfaea-879f-4ace-9f8d-7a3361eb6daa · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Visual-based spatial audio generation system for multi-speaker environments LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T12:30:26.180184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:30:26.180184Z digest=sha256:9ff375b7bd20fc0a661bd59fcb3fa0aee427ed52909996d45eb96c0b74edb47d

Observation ca24f598-e728-466c-8ccb-b99e6fe79ba8 · outbound

This paper cites Self- supervised generation of spatial audio for 360 video,.

Visual-based spatial audio generation system for multi-speaker environments Self- supervised generation of spatial audio for 360 video,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.308132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:30:26.184894Z digest=sha256:e1120d9fc39ed9d675c45165fa27bca6af0a95c2b1cd9e742ae1d70dde9e7fda

Observation 37c35292-c3e0-4301-a298-4f30aa1ebff7 · outbound

This paper cites Peaq-the itu standard for objective measurement of perceived audio quality,.

Visual-based spatial audio generation system for multi-speaker environments Peaq-the itu standard for objective measurement of perceived audio quality,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.293031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:30:26.188982Z digest=sha256:0b34b02d009986108fc22d1c2ab1ec7b756be70e01c3076dc9ef6b7eccb8f266

Observation 196a87c3-a672-415b-8523-196aa0eb2830 · outbound

This paper cites An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,.

Visual-based spatial audio generation system for multi-speaker environments An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.278058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:30:26.193102Z digest=sha256:b0a1b629e27dfe15ddc3e86e1c1cd221a2bca619ed0c50c6df4ef149a92eb2a0

Observation 35500a2e-48ca-4ba1-8b71-f261a0836319 · outbound

This paper cites MOSNet: Deep Learning based Objective Assessment for Voice Conversion.

Visual-based spatial audio generation system for multi-speaker environments MOSNet: Deep Learning based Objective Assessment for Voice Conversion

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T12:30:26.197293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:30:26.197293Z digest=sha256:b0659eae2f129b7d6e810a6ae6620d8514b64d6a29ce39ff6a090dc918ef3f29

Pith citing papers

Observation e32e88ea-5594-4674-8822-0c9f4b73b63c · inbound

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations cites this paper.

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations Visual-based spatial audio generation system for multi-speaker environments

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:04:00.782536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:04:00.782536Z digest=sha256:6450455e662e9edd3bb7cd6dbdcd5aea1343e73e04c2a4748865f3d11f27729a

Observation 2eafc3f1-4726-4f15-9c63-e70ead68af53 · inbound

ASAudio: A Survey of Advanced Spatial Audio Research cites this paper.

ASAudio: A Survey of Advanced Spatial Audio Research Visual-based spatial audio generation system for multi-speaker environments

Reference 102

Resolution
verified exact
local_arxiv, observed 2026-08-05T22:54:55.853160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T22:54:55.101714Z digest=sha256:4748964b502f7c52b213a419b8cb1b7d917f30047e62f6a4e11b23adf5f33fcb