Pith. sign in

Paper Citation Record · LEDGER

Visual-based spatial audio generation system for multi-speaker environments

As of 9 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2502.07538.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07538 v2

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:30:26.197293Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:04:00.782536Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T22:54:55.849359Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 01c1e241-7246-4d53-9b27-b2b1450eed7a · outbound

This paper cites Exploring audio- visual information fusion for sound event localization and detection in low-resource realistic scenarios,.

Visual-based spatial audio generation system for multi-speaker environments Exploring audio- visual information fusion for sound event localization and detection in low-resource realistic scenarios,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.536832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:30:26.103584Z digest=sha256:a2fb0fccfd14707da0bca2bb1c666bbdca29f520aa8d578c64669ce499e47ae5

Observation 047e4701-6346-4cf6-8f4d-9ca1b7958006 · outbound

This paper cites Aligning audiovisual features for audiovisual speech recognition,.

Visual-based spatial audio generation system for multi-speaker environments Aligning audiovisual features for audiovisual speech recognition,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.522970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:30:26.108687Z digest=sha256:f0de985357b87df067e2da4e61f30ffb6888adc0b55a4304a2c312d66880ae82

Observation c3174224-52d0-4dab-9458-417abf7c78be · outbound

This paper cites Immersive spatial audio reproduction for vr/ar using room acoustic modelling from 360° images,.

Visual-based spatial audio generation system for multi-speaker environments Immersive spatial audio reproduction for vr/ar using room acoustic modelling from 360° images,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.508310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:30:26.113528Z digest=sha256:1d226d557842f18a7d0876a322fb1db21acc08cda456abdb2899eb4729499ee2

Observation 0a28dc3b-e7de-4625-87b6-e6a8ac9ff269 · outbound

This paper cites Scene-aware audio for 360° videos,.

Visual-based spatial audio generation system for multi-speaker environments Scene-aware audio for 360° videos,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.494322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:30:26.118296Z digest=sha256:945ccc52ca59ebb5f627c03c2ad30043d1da5663339bbce8f598545f8927568d

Observation f115ea40-7d76-4016-b312-4845ec80fcd7 · outbound

This paper cites Analysis of a distributed processing model for spatialized audio conferences,.

Visual-based spatial audio generation system for multi-speaker environments Analysis of a distributed processing model for spatialized audio conferences,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.480615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:30:26.122985Z digest=sha256:d2bed49513b0b62a3e0539006a878b801991ba14c173ffa996b7f56bb77d9785

Observation 78e8b1d3-6199-452e-a843-782c4e2ff069 · outbound

This paper cites Realistic audio in immersive video conferencing,.

Visual-based spatial audio generation system for multi-speaker environments Realistic audio in immersive video conferencing,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.465535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:30:26.127638Z digest=sha256:c8b2b37ef5781f937ad65a5cd6cc97db733fcc1d993c0fe747a2a06219a46b39

Observation 69c65982-9250-4479-9afb-2e3243dc5188 · outbound

This paper cites Audio-visual sensing from a quadcopter: dataset and baselines for source localization and sound enhancement,.

Visual-based spatial audio generation system for multi-speaker environments Audio-visual sensing from a quadcopter: dataset and baselines for source localization and sound enhancement,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.451213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:30:26.132615Z digest=sha256:0414a33e846b70214103a5e557daf4de17f5491e2bf78f2509f6bf5874339cc8

Observation 55096907-3dba-4bce-b6ba-1c792c2ef5a4 · outbound

This paper cites 2.5d visual sound,.

Visual-based spatial audio generation system for multi-speaker environments 2.5d visual sound,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.434867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:30:26.136964Z digest=sha256:32a8489acd8bac37afb74f824ce90c855516223b3caecab01c8eb1aabe79d0ac

Observation 118c667c-4c36-4529-bc12-a5fddc806369 · outbound

This paper cites Visually informed binaural audio generation without binaural audios,.

Visual-based spatial audio generation system for multi-speaker environments Visually informed binaural audio generation without binaural audios,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.418915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:30:26.141461Z digest=sha256:c38269b3360adc5a19c0f040f52a160f2380989ff891e29a87199657efa897de

Observation ed4ccc89-e858-4317-8a1e-85e3381ececf · outbound

This paper cites A review on yolov8 and its advancements,.

Visual-based spatial audio generation system for multi-speaker environments A review on yolov8 and its advancements,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.403157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:30:26.145960Z digest=sha256:d0b23cb9a0ec30ac247aed2b5feaddbe95067ce9a034289d8f456e81fdfe127e

Observation 7687ebb2-2f03-4d28-adb5-1085d7bacbe0 · outbound

This paper cites Wider face: A face detection benchmark,.

Visual-based spatial audio generation system for multi-speaker environments Wider face: A face detection benchmark,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.388110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:30:26.150357Z digest=sha256:bd74f2bab457a64ea11cc421619d6bfe794e16dba38a7d614e6abc1868907072

Observation b017bd20-49bb-4854-bd06-289fa64f87f5 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data,.

Visual-based spatial audio generation system for multi-speaker environments Depth anything: Unleashing the power of large-scale unlabeled data,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.372085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:30:26.154803Z digest=sha256:b2bad45e2b9fd836f6ef507de8f44d9e78bf9e6f698753dc051ba28000844d05

Observation 535446c5-1b96-47b9-aefd-f809af2ef142 · outbound

This paper cites You only look once: Unified, real-time object detection,.

Visual-based spatial audio generation system for multi-speaker environments You only look once: Unified, real-time object detection,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.356023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:30:26.160648Z digest=sha256:c39570b5cb23676b1e9f5a26188e62407ec32a63526d3213e36c6f10ffa171f2

Observation e8396a1d-7eb7-4c8d-aedc-f42dbcdd999c · outbound

This paper cites Conv-tasnet: Surpassing ideal time– frequency magnitude masking for speech separation,.

Visual-based spatial audio generation system for multi-speaker environments Conv-tasnet: Surpassing ideal time– frequency magnitude masking for speech separation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.339161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:30:26.165241Z digest=sha256:25f8038d1611f84c0a1db6ad511c1db3d8e51591d9326ab2e968da0e52ed07cc

Observation f09ed45e-7d93-432f-83d1-ddcd5ac4e1e3 · outbound

This paper cites Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed.

Visual-based spatial audio generation system for multi-speaker environments Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T12:30:26.169424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:30:26.169424Z digest=sha256:d1599f254eb45f2debb9069732d768c5fe9203a746fb1532d520f8b19c6eabca

Observation f87a8c9c-c216-4bbd-b62f-33cba63d9bda · outbound

This paper cites A perceptual evaluation of individual and non-individual hrtfs: A case study of the sadie ii database,.

Visual-based spatial audio generation system for multi-speaker environments A perceptual evaluation of individual and non-individual hrtfs: A case study of the sadie ii database,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.322485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:30:26.174917Z digest=sha256:98bc6e68894af9a6ae1ac75d93a47d21c5204dff913cd5edb6a55df963ca7660

Observation df9dfaea-879f-4ace-9f8d-7a3361eb6daa · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Visual-based spatial audio generation system for multi-speaker environments LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T12:30:26.180184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:30:26.180184Z digest=sha256:9ff375b7bd20fc0a661bd59fcb3fa0aee427ed52909996d45eb96c0b74edb47d

Observation ca24f598-e728-466c-8ccb-b99e6fe79ba8 · outbound

This paper cites Self- supervised generation of spatial audio for 360 video,.

Visual-based spatial audio generation system for multi-speaker environments Self- supervised generation of spatial audio for 360 video,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.308132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:30:26.184894Z digest=sha256:ef52842254697f1fb5ae37b89e767f2b63251f80ec1ae83aee0fb032a0d59ca6

Observation 37c35292-c3e0-4301-a298-4f30aa1ebff7 · outbound

This paper cites Peaq-the itu standard for objective measurement of perceived audio quality,.

Visual-based spatial audio generation system for multi-speaker environments Peaq-the itu standard for objective measurement of perceived audio quality,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.293031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:30:26.188982Z digest=sha256:65e2edc262eceb08fe260e8109320d65ada5794125c52def2e534843740de98e

Observation 196a87c3-a672-415b-8523-196aa0eb2830 · outbound

This paper cites An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,.

Visual-based spatial audio generation system for multi-speaker environments An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:30:26.278058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:30:26.193102Z digest=sha256:5f1c37f7e7337dede7eb0394e5886fa3b370407dd4b01c4f52bda9547911ba17

Observation 35500a2e-48ca-4ba1-8b71-f261a0836319 · outbound

This paper cites MOSNet: Deep Learning based Objective Assessment for Voice Conversion.

Visual-based spatial audio generation system for multi-speaker environments MOSNet: Deep Learning based Objective Assessment for Voice Conversion

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T12:30:26.197293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:30:26.197293Z digest=sha256:b0659eae2f129b7d6e810a6ae6620d8514b64d6a29ce39ff6a090dc918ef3f29

Pith citing papers

Observation e32e88ea-5594-4674-8822-0c9f4b73b63c · inbound

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations cites this paper.

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations Visual-based spatial audio generation system for multi-speaker environments

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:04:00.782536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:04:00.782536Z digest=sha256:6450455e662e9edd3bb7cd6dbdcd5aea1343e73e04c2a4748865f3d11f27729a

Observation 2eafc3f1-4726-4f15-9c63-e70ead68af53 · inbound

ASAudio: A Survey of Advanced Spatial Audio Research cites this paper.

ASAudio: A Survey of Advanced Spatial Audio Research Visual-based spatial audio generation system for multi-speaker environments

Reference 102

Resolution
verified exact
local_arxiv, observed 2026-08-05T22:54:55.853160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T22:54:55.101714Z digest=sha256:a8c14883da0aadcca73f92252740c6ccb67d82f741bd102fd3c2ffe9c3e31109