Pith. sign in

Paper Citation Record · LEDGER

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization

As of 19 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2412.01488.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01488 v3

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:23:41.635307Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy45
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e6e7773a-7197-48d7-94fb-c29b0d7da8fb · outbound

This paper cites Emerging properties in self-supervised vision transformers.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Emerging properties in self-supervised vision transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.485201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.365557Z digest=sha256:a4aa125c0008172c7efa2068433397f1d6926bfe5cc8ba73e5efe274ccb9dc6e

Observation 2da26d9d-7167-467b-a525-caf0e041a8ed · outbound

This paper cites Localizing visual sounds the hard way.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Localizing visual sounds the hard way

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.468994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.371489Z digest=sha256:b7b8b4e1f6f25aaa6338fb660990112f1829023b1b47e909a97d34fc54d34668

Observation 67e6dbe0-b1d5-4eee-be42-36d42a8b93b6 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Vggsound: A large-scale audio-visual dataset

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.451434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.378005Z digest=sha256:cafb59eba55f1fa028c3b13ed82f03296b9567300a7424b115bb19d92e9ad2fc

Observation 8f790430-b373-4c93-a839-ff6924b20b6d · outbound

This paper cites Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset, 2023.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset, 2023

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.434573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.383816Z digest=sha256:c1973bd69ade8ee32c5465b651701e99211be89ca35aa0e9d219c1a15a92f986

Observation be7389c0-8488-4d44-91d3-bbaa887a0bb1 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Reproducible scaling laws for contrastive language-image learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.418177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.389017Z digest=sha256:160eda6e15cfdac0c1c499028b2047962ca9457dab04ce3a31b94d8dc7045ab6

Observation 76edd49a-a693-4939-9fcc-f3d76e00b591 · outbound

This paper cites Meerkat: Audio-visual large language model for grounding in space and time.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Meerkat: Audio-visual large language model for grounding in space and time

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.397715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.393789Z digest=sha256:5a8f4c2d91bb5170c50d1233552f3c852f113412bb4cb59ba64c57be75cc5ace

Observation ac52431f-cea6-4b2b-85a6-6cc1bcd83f86 · outbound

This paper cites Multilayer nonnegative matrix factorisation.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Multilayer nonnegative matrix factorisation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.379997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.399495Z digest=sha256:55d4aa93fbf2f6b62d70d0a448a3a2d440776dd7be466048d41186f7139829e5

Observation 9bfa13ba-454b-4d59-a8d4-e989254fe9c9 · outbound

This paper cites Deep feature factorization for concept discovery.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Deep feature factorization for concept discovery

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.363505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.404938Z digest=sha256:94e377930db99058b45740f54b30886c87fb495dd0de7c8ff84f30e7755c473e

Observation 43db80ca-88b3-45ec-8b0e-0852439267b8 · outbound

This paper cites Neural Network Matrix Factorization.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Neural Network Matrix Factorization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T04:23:41.411051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:23:41.411051Z digest=sha256:01cc2ff6820741a144091413da650a467c641ee04c56a6f4eee3eb991b787986

Observation 2066cef2-5c37-433d-8131-9aae3db6beb2 · outbound

This paper cites Clap learning audio concepts from natural language supervision.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Clap learning audio concepts from natural language supervision

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.344663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.416086Z digest=sha256:1f7612dc08823c3a631a9ac3301ae02e594b5485dbefa1d5cb8ce03ee7aa4e4b

Observation e77553b9-a849-4cd7-99f6-c81de089b320 · outbound

This paper cites Avsegformer: Audio-visual segmentation with transformer.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Avsegformer: Audio-visual segmentation with transformer

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.329719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.420719Z digest=sha256:281eabeea521959694ea14a7a548b9394c6020c1472298f2cbef843aec9b8e06

Observation f3cbbf35-8fa8-4d06-bc28-55fa4df9cf14 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Audio set: An ontology and human-labeled dataset for audio events

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.314823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.425640Z digest=sha256:62f20967903390652f3f8f7a5b8f83d1bfadccc26515a82abb1ff3900fe086bf

Observation 394d892c-586e-4ea6-92fa-e73e57718423 · outbound

This paper cites The why and how of nonnegative matrix factorization.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization The why and how of nonnegative matrix factorization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.297240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.431194Z digest=sha256:69a9a4fc5bd63dfd1145b3815352348d162e58727f9b1be5feedeef3ad6c877b

Observation ab4b0125-4112-4438-8aa5-9e52508fe2c7 · outbound

This paper cites Imagebind: One embedding space to bind them all.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Imagebind: One embedding space to bind them all

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:23:41.435751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:23:41.435751Z digest=sha256:4bd6b2524c014d5e24f66146e4cc92bd2ff8fdb0deb4bbd630b5a73c39d4a253

Observation a983b572-ac08-48bc-a575-5ef48fcf5c18 · outbound

This paper cites Contrastive Audio-Visual Masked Autoencoder.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Contrastive Audio-Visual Masked Autoencoder

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:23:41.440713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:23:41.440713Z digest=sha256:c74611cfd487b12d5a56242182a21fcbeb95256b49ae848ca04f74ee94463e29

Observation 09c971e0-ce71-45dd-9998-ccf2ff70be83 · outbound

This paper cites Single channel speech music separation using nonnegative matrix factorization and spectral masks.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Single channel speech music separation using nonnegative matrix factorization and spectral masks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.267080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.445696Z digest=sha256:a904dd718af10f3704033b93b848c2637b93cd191a90ff43552c4c8f0dc515f6

Observation a70468c1-e73e-4d8a-b01c-0f6802c3f190 · outbound

This paper cites Non-negative matrix factorization for face recognition.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Non-negative matrix factorization for face recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.250748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.450214Z digest=sha256:3801deec276589dc9e98825764b009f07bc8459820bbaac65fd4620da6563117

Observation 07f012fd-eed4-4236-9d3d-2f0d2baff75c · outbound

This paper cites chirp" from the.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization chirp" from the

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.233219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.455041Z digest=sha256:21e404e9fea12607503c46ab128c6729f9530e7b808093fc9f8e8a996af43d63

Observation 7e272c82-839e-4054-ac26-32fcd7b74da5 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.216372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.460373Z digest=sha256:ad25dd1759639a127cb89143e05e6137470987786f977dc6053c08993ae2c8a9

Observation f537bdda-6134-4bdd-8480-0dd85db28d83 · outbound

This paper cites Transfer Learning from Audio-Visual Grounding to Speech Recognition.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Transfer Learning from Audio-Visual Grounding to Speech Recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T04:23:41.465895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:23:41.465895Z digest=sha256:2d1b8b94cb2f5206cbef3eaf0e3b7e754261b838a12b1d1961936a814253c744

Observation cba18e4f-a2b0-4137-b677-8a65b7e2ad6a · outbound

This paper cites A Critical Assessment of Visual Sound Source Localization Models Including Negative Audio.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization A Critical Assessment of Visual Sound Source Localization Models Including Negative Audio

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T04:23:41.473842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:23:41.473842Z digest=sha256:28ba421f85270ea3fa6ad93f5d1fcd88cfc03f96fe8a2d7c6d2d8ea28dd5d096

Observation fd5de8c9-5f6d-4680-9409-9d29238e6539 · outbound

This paper cites Algorithms for non-negative matrix factorization.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Algorithms for non-negative matrix factorization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.199422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.478984Z digest=sha256:5c0ce561c489b5ed28929b47939307a4b9ad71f660f82b15247037c737442425

Observation 55ab5fa2-739a-4b4c-ba38-ce50bc4d74fc · outbound

This paper cites Unsupervised sound localization via iterative contrastive learning.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Unsupervised sound localization via iterative contrastive learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.184269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.483918Z digest=sha256:52b92bbc334fa570b4a6effbc430305370d1c78990f3d701009103e8bad7f2e8

Observation c174fe3b-9ac8-49ea-9e27-4352b71c6c08 · outbound

This paper cites Audio-visual segmentation by exploring cross-modal mutual semantics.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Audio-visual segmentation by exploring cross-modal mutual semantics

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.168896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.489158Z digest=sha256:6fadc54aec27453eecbdd84eea96fded45f2258cfe32ce061cd153373c1e5d3e

Observation adea6297-dde1-4f77-ae3c-3ed2af783d06 · outbound

This paper cites Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T04:23:41.494188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:23:41.494188Z digest=sha256:29af73e6315cfe2b3d1160ef653e2f8f44d44669056f4c5fa457ed5d0ed43020

Observation 03102f43-292a-420b-836e-129cfb75311a · outbound

This paper cites Exploiting transformation invariance and equivariance for self-supervised sound localisation.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Exploiting transformation invariance and equivariance for self-supervised sound localisation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.154728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.499270Z digest=sha256:541a60898a1454b56f0e66299db44b3d20382a944079dd0ec9d5b40fd58b9638

Observation dea00248-ac9b-4fdf-a823-d51478db0ec6 · outbound

This paper cites Image segmentation using text and image prompts.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Image segmentation using text and image prompts

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.141024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.504042Z digest=sha256:03e875fb23caf0486ae00e0f7925ab46688ad014df3b2561be2556e7dc3b130d

Observation a8cd910e-e30e-4e4e-a3c8-d8d16a8f0f6b · outbound

This paper cites An eye for an ear: zero-shot audio description leveraging an image captioner with audio-visual token distribution matching.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization An eye for an ear: zero-shot audio description leveraging an image captioner with audio-visual token distribution matching

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.127218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.509314Z digest=sha256:5f149b30eb6d5d2d57d9127321057997b9aff8d45427aad4ca5b27d3dca1b54e

Observation 4eea906a-1d55-4d2a-9cb1-ffbf7d517a40 · outbound

This paper cites A closer look at weakly-supervised audio-visual source localization.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization A closer look at weakly-supervised audio-visual source localization

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.112948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.514798Z digest=sha256:f98476426c21e2afbf6fadc75ee78aa4f782089cce8dc023a018853d8e49237c

Observation 0e90062f-60ed-447b-97a6-977b09e920cf · outbound

This paper cites A Concept-Based Explainability Framework for Large Multimodal Models.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization A Concept-Based Explainability Framework for Large Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:23:41.520270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:23:41.520270Z digest=sha256:228c84baf7d473aa8632cd2e13ee9f0345a77ee8982ebd921c381dccb3e24d40

Observation aac9c43f-6969-48e9-89ad-4defc724afdc · outbound

This paper cites Listen to interpret: Post-hoc interpretability for audio networks with nmf.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Listen to interpret: Post-hoc interpretability for audio networks with nmf

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.096290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.526646Z digest=sha256:cfb91d7a0db14726a8012ca6a8cd553b4eda45f03c2f0eed19091a37e893cf13

Observation b38c4eb1-38f8-459f-ba47-2a794b9aad4d · outbound

This paper cites Guiding audio source separation by video object information.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Guiding audio source separation by video object information

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.081241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.532777Z digest=sha256:fcf8eb79043b5ce5e032cc01fa55bae94e898c6abfd073ff9890501a699ebd6f

Observation 31bf4e73-ac6d-4ed4-b01a-842b831e4de3 · outbound

This paper cites Marginnce: Robust sound localization with a negative margin.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Marginnce: Robust sound localization with a negative margin

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.065183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.538478Z digest=sha256:2dc5309218218f8cfaf02e49ebc52799a0076d4b7cac233a4621643d8272b2a6

Observation 049cf943-4e67-4b62-b51c-603333264557 · outbound

This paper cites Can clip help sound source localization? In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), January 2024.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Can clip help sound source localization? In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), January 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.049396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.543730Z digest=sha256:d0d130beedf4dde1f7bd8ab772b2c4f7b05d996bca00dcd9b6618fbd95112fc6

Observation e1bd8688-db34-4e0d-b068-cd945dadc870 · outbound

This paper cites Learning transferable visual models from natural language supervision.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Learning transferable visual models from natural language supervision

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.035083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.548694Z digest=sha256:62c3ba27606984cf5fdc613d4764462695ee89bff6c6e3421ff8eeae5289a434

Observation b96c824a-9c57-4100-ad82-ff70fedabe54 · outbound

This paper cites Laion- 5b: An open large-scale dataset for training next generation image-text models.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Laion- 5b: An open large-scale dataset for training next generation image-text models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.020302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.553671Z digest=sha256:c089d298c90b55cee334b68d01c6d538e38ef2e6216aa3b47799caa67214cd74

Observation d2d773af-542d-4b50-8b11-ddbe30127d4e · outbound

This paper cites Soft nonnegative matrix co-factorization.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Soft nonnegative matrix co-factorization

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:42.002528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.558095Z digest=sha256:5fa5598d0ab4941d051c0c18f8ff2fa2e498ed6fb441e8c756fec050fbd752dd

Observation ade26ffd-c0f6-48ad-9fb0-7d20d6860cfa · outbound

This paper cites Learning to localize sound source in visual scenes.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Learning to localize sound source in visual scenes

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:41.987034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.563297Z digest=sha256:a997f679dcef7967b78a5ce793decac85ce37247f2871a35b2132106b60b44ae

Observation ba10f421-b590-41a8-a293-dea8d5040411 · outbound

This paper cites Learning to localize sound sources in visual scenes: Analysis and applications.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Learning to localize sound sources in visual scenes: Analysis and applications

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:41.971619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.568327Z digest=sha256:b023a04c8f192de16fe57c7de2923d5d1f7744293233b72f8010916de1ba6cc8

Observation 724eaf77-0934-4d99-847f-4bbdb6051a6e · outbound

This paper cites Learning sound localization better from semantically similar samples.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Learning sound localization better from semantically similar samples

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:41.953907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.573269Z digest=sha256:c54453dba5664fbc56bffe614dd3f00dd0b1ac7c0b51c41c131aaa7ad569542a

Observation 6d4a439d-eaff-4ac3-b460-6904964ce2af · outbound

This paper cites Sound source localization is all about cross-modal alignment.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Sound source localization is all about cross-modal alignment

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:41.940508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.577960Z digest=sha256:149a61567641663a99826dea65fc37866967f9bbf7594e40f5b67c94208ab17d

Observation 6e811ac2-eba3-4bb7-a1c1-1576ac0e8b8a · outbound

This paper cites Increasing importance of joint analysis of audio and video in computer vision: A survey.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Increasing importance of joint analysis of audio and video in computer vision: A survey

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:41.926479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.582017Z digest=sha256:3c7f59f6235d5e309c617ec6e1c258e6e310caf6067cc0a8105cbecb917e3ce0

Observation bd27ecd2-f5e5-4248-a694-64ffe7e5c8c2 · outbound

This paper cites Learning audio-visual source localization via false negative aware contrastive learning.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Learning audio-visual source localization via false negative aware contrastive learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:41.911290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.591779Z digest=sha256:a969762a2f0892dc70a3226581e47b9bb2e9cc6879a02b69a221edfe53d58c25

Observation affd8a84-9597-4234-8128-c7bf6f43aea3 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Audio-visual event localization in unconstrained videos

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:41.894808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.596649Z digest=sha256:04924b7f5564ff56eb276652e0b3725bc4c717f87d8b6851c065ddaf6b42d7e2

Observation d5c4fbac-c93b-4f84-92c5-798058cb14da · outbound

This paper cites Combining non-negative matrix factoriza- tion and deep neural networks for speech enhancement and automatic speech recognition.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Combining non-negative matrix factoriza- tion and deep neural networks for speech enhancement and automatic speech recognition

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:41.878827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.602241Z digest=sha256:1c28b742b526c27bc8ac5ba662d68d435033564342a7a553c71611094c43bb9f

Observation d38866b8-2cd0-4dce-a59b-2038c55374e8 · outbound

This paper cites Document clustering based on non-negative matrix factorization.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Document clustering based on non-negative matrix factorization

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:41.862798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.606898Z digest=sha256:a2961ec68a0627c094e19e1d8878af0c57cbe21a8802f025e21f498f0e112276

Observation 0bf2f4e7-82f8-4ff4-9c6c-21629b283753 · outbound

This paper cites Coupled nonnegative matrix factorization unmixing for hyperspectral and multispectral data fusion.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Coupled nonnegative matrix factorization unmixing for hyperspectral and multispectral data fusion

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:41.846293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.612291Z digest=sha256:3a6f8b4b7f9d2122fb9ddb937b46b790b32b0b4bc33b41964ffc7aabcc8a53ff

Observation cea3375a-ad04-4c25-8af0-163da9acb579 · outbound

This paper cites Matrix co-factorization on compressed sensing.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Matrix co-factorization on compressed sensing

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:41.830156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.618257Z digest=sha256:bdfa8e33e18c87e97890cf714d7aee40fb3b18a9a73c5ca164c37f9e8d477cb9

Observation fad3c349-078c-4638-9a83-0e7c9a13e8d8 · outbound

This paper cites Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:41.815704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.622563Z digest=sha256:8677b7fe9211b7ccbe9d2e46f05716a31fdafce010255954e23f423f06792c71

Observation d70059dc-7c51-4dee-8833-81b812226705 · outbound

This paper cites Semantic understanding of scenes through the ade20k dataset.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Semantic understanding of scenes through the ade20k dataset

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:41.799884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.627328Z digest=sha256:b60c71add56e94348c5cec16d39163a0e663ebb0e96198b076a4a94515227c14

Observation 212c30af-bc19-468b-8a4b-0497ba2bbe45 · outbound

This paper cites Audio-visual segmentation with semantics.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Audio-visual segmentation with semantics

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:41.784862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.631230Z digest=sha256:ca523f7fba0cecf150257ccc9854e46f1b5f33f1f387fd8cb74d648b68debad8

Observation eadfaf95-d7b5-4b70-a9d6-2ed3f29471da · outbound

This paper cites Audio–visual segmentation.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Audio–visual segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:23:41.768870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:23:41.635307Z digest=sha256:7a7976a4ce94633eb22c5639aba4102b1ccaaa9187744b43dec5727ee406eca2

Pith citing papers

No inbound Pith citation observations are available.