Pith. sign in

Paper Citation Record · LEDGER

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues

As of 10 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 1 inbound Pith citation observation for arXiv:2606.24512.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.24512 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-25T22:43:45.086693Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-25T22:43:45.086693Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T18:40:02.573284Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b51a36c-36ce-439d-889c-bb5c5480f1d2 · outbound

This paper cites A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues

Reference 1

Resolution
malformed identifier
local_arxiv, observed 2026-07-04T18:40:02.574504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:f635fbea48a42668d8bb9ebb6c42d8284bbb5579d5bf395a368693cc49344786

Observation 36320534-72e9-4769-b552-880e8da5062f · outbound

This paper cites Framework Overview Figure 1 illustrates our proposed three-stage self-guided framework for joint separation and classification.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues Framework Overview Figure 1 illustrates our proposed three-stage self-guided framework for joint separation and classification

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-25T22:43:45.086693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:11018397690cbaee17131727636da52331be48be2172ed2044087f60af2801dd

Observation 7baf151a-7cd6-46ec-be48-9837f94b8bec · outbound

This paper cites Datasets and Augmentation To train our models, we dynamically generated 4-channel mixtures online for each sample using the SpatialScaper simulator [9].

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues Datasets and Augmentation To train our models, we dynamically generated 4-channel mixtures online for each sample using the SpatialScaper simulator [9]

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-25T22:43:45.086693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:9e185e62e8c2aba019739cd8cc156bf860476287491316198eb1b5ae86029f4d

Observation 03ca6ca4-f405-4a04-b5af-3b9e3cdf54b1 · outbound

This paper cites Ablation Study To validate the effectiveness of our proposed methods across the multi-stage framework, we conducted a comprehensive ablation study on the development test set.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues Ablation Study To validate the effectiveness of our proposed methods across the multi-stage framework, we conducted a comprehensive ablation study on the development test set

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-25T22:43:45.086693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:c488fe01733b82ab345031d3411ed636dd4212032c0b3cee14a7caaaf3381180

Observation a0c4ed46-cc6a-481c-af4f-385e3da5f3b0 · outbound

This paper cites an unresolved cited work.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-25T22:43:45.086693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:b20e516f87a421192a842aef591e86442a70e5871d23a6a175303cae5e5ce6bd

Observation 8b806eea-be57-4f6c-b30a-5d56d2606625 · outbound

This paper cites RS-2024-00337945), STEAM research grant (No.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues RS-2024-00337945), STEAM research grant (No

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-25T22:43:45.086693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:7c3188d1d45e7fa448c49b89985dc53ec9de7263d3c65c256b0bdde8ca36b5db

Observation ea01d0d6-0fc7-4db7-a10a-ea923d19f93b · outbound

This paper cites Description and Discussion on DCASE 2025 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues Description and Discussion on DCASE 2025 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:40:02.569358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:43f504409bb68e0cccb04f1ebf549b963647afd4c26ff8b88fd830676e9aa56c

Observation 788a391b-6e39-4bb6-aab2-5028afe64a4e · outbound

This paper cites Description and Discussion on DCASE 2026 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues Description and Discussion on DCASE 2026 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-04T18:40:02.572018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:d90772008b27734a7894666be074d88d402789007c5b30981bde6cdf441b2452

Observation a3e495b0-c3b6-4ab8-b66e-0f24b7d68dd1 · outbound

This paper cites DeepASA: An object- oriented multi-purpose network for auditory scene analysis,.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues DeepASA: An object- oriented multi-purpose network for auditory scene analysis,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-25T22:43:45.086693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:19f9363c8fa9057ecd4cb9b562e9f8e36574e0c379a32e9f8513355c63f54000

Observation fcdf3d59-deab-4d6c-b55a-2308184f07e3 · outbound

This paper cites Self-guided target sound extraction and classification through universal sound separation model and multiple clues,.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues Self-guided target sound extraction and classification through universal sound separation model and multiple clues,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-25T22:43:45.086693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:b1e2d046d80c77d225b49c67aa7bbe9292990122a4b20c8e3055c15b30029713

Observation 03d474a7-7621-4ad7-82b5-b1389bb51a93 · outbound

This paper cites Sound separation and classifi- cation with object and semantic guidance,.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues Sound separation and classifi- cation with object and semantic guidance,

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:40:02.566433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:ed446c4e3bfe2c2621c8aa9a33c719716c5775123297b2e1761308125d9a5761

Observation ed02b3d9-abb5-4b43-96b9-2726999a4214 · outbound

This paper cites Au- dio flamingo 3: Advancing audio intelligence with fully open large audio language models,.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues Au- dio flamingo 3: Advancing audio intelligence with fully open large audio language models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-25T22:43:45.086693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:68a58ee0cd27b5b808a5beb6e03d6f49f534901ef4adecc292217c4391e6b09e

Observation 09421c7e-9eb6-42a4-90ab-2f086a862667 · outbound

This paper cites Temporal film: Capturing long-range sequence depen- dencies with feature-wise modulations.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues Temporal film: Capturing long-range sequence depen- dencies with feature-wise modulations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-25T22:43:45.086693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:32783823b23cc4832306087751ad1c4e8dc051bfe9c938f40d22b42b0352f28d

Observation 68551387-7597-47ad-aafe-b3e89b2a4622 · outbound

This paper cites Film: Visual reasoning with a general condi- tioning layer,.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues Film: Visual reasoning with a general condi- tioning layer,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-25T22:43:45.086693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:27b39ab3d345aee0cfa11ca58439f631c8744cf84d63952a51bdcd6f20fecf2e

Observation 28e1d384-1534-4c98-99a3-9a19901a24b0 · outbound

This paper cites Spatial scaper: a library to simulate and aug- ment soundscapes for sound event localization and detec- tion in realistic rooms,.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues Spatial scaper: a library to simulate and aug- ment soundscapes for sound event localization and detec- tion in realistic rooms,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-25T22:43:45.086693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:5affa0c67dbe8df4450bea779346eea8d20707db740d6cbddf117f7a1da570c9

Observation 98b72138-d199-4b5c-9a3b-5fbfced7c314 · outbound

This paper cites The voice bank cor- pus: Design, collection and data analysis of a large re- gional accent speech database,.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues The voice bank cor- pus: Design, collection and data analysis of a large re- gional accent speech database,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-25T22:43:45.086693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:2b167c4fa866ad012aff5d1da6a832bc4713dc0a6e4e9f2f3bd7a03e4ed47d6f

Observation 358bedae-16c1-409b-a8b9-6614560f5202 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues Audio set: An ontology and human-labeled dataset for audio events,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-25T22:43:45.086693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:b10ef8960faf09f5b8ba00202fe8cca9c654edda9b5adcb16124e34f158e22ab

Observation dc0dd560-28f3-43d5-9ee0-446668ffb8cb · outbound

This paper cites Sa-sdr: A novel loss function for sepa- ration of meeting style data,.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues Sa-sdr: A novel loss function for sepa- ration of meeting style data,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-25T22:43:45.086693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:4fcc560011024bc5a4a125adb0cc891fd86ebe53785644949a4fdc030e8454e9

Observation 466dd0ab-473d-424d-ad45-b2998ce069e7 · outbound

This paper cites Arcface: Ad- ditive angular margin loss for deep face recognition,.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues Arcface: Ad- ditive angular margin loss for deep face recognition,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-25T22:43:45.086693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:3c0ea190808e1ce4960d655973bff3d052b37cd5b2182eb0acd599c151b1e263

Observation c1214f66-2143-49af-ba76-77486bcda045 · outbound

This paper cites Class-aware permutation-invariant signal-to- distortion ratio for semantic segmentation of sound scene with same-class sources,.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues Class-aware permutation-invariant signal-to- distortion ratio for semantic segmentation of sound scene with same-class sources,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-25T22:43:45.086693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:812d37127246df18d6819284ffd134b9bd42fd382d657d98126a42ecd29cd6b8

Pith citing papers

Observation 8b51a36c-36ce-439d-889c-bb5c5480f1d2 · inbound

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues cites this paper.

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues

Reference 1

Resolution
malformed identifier
local_arxiv, observed 2026-07-04T18:40:02.574504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-25T22:43:45.086693Z digest=sha256:f635fbea48a42668d8bb9ebb6c42d8284bbb5579d5bf395a368693cc49344786