Pith. sign in

Paper Citation Record · LEDGER

Implicit Counterfactual Learning for Audio-Visual Segmentation

As of 9 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 0 inbound Pith citation observations for arXiv:2507.20740.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20740 v1

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:24:31.180804Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

86 of 86 outbound references displayed

  • verified exact6
  • verified fuzzy58
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 92232e28-b925-4df0-98d6-3cd58f3e6a75 · outbound

This paper cites Remembering the past and imagining the future: Common and distinct neural substrates during event construction and elaboration.

Implicit Counterfactual Learning for Audio-Visual Segmentation Remembering the past and imagining the future: Common and distinct neural substrates during event construction and elaboration

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:20.667712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:20.667712Z digest=sha256:557f4653d03efb8099c41be423eedf6946b01ffdef0bc4be37d19c7c7365e408

Observation 9eb1ad31-b9f1-4f74-96f2-1c505fd5a6b4 · outbound

This paper cites Unsupervised Audio-Visual Segmentation with Modality Alignment.

Implicit Counterfactual Learning for Audio-Visual Segmentation Unsupervised Audio-Visual Segmentation with Modality Alignment

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:32.277437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:20.791836Z digest=sha256:6f5175bce617f913ca65ec31831ab3808490d1e11844abe3c9739a8d88f334b1

Observation a94ba4dc-11e1-4712-acd0-31967afc369f · outbound

This paper cites Numerics of gram-schmidt orthogonalization.

Implicit Counterfactual Learning for Audio-Visual Segmentation Numerics of gram-schmidt orthogonalization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:20.951174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:20.951174Z digest=sha256:15b4ebfec832a682643a094a7891aae9d3b79b8230605e4fdd61fc2707d237e2

Observation a04bf18d-2e62-4197-8c4a-71e71f82acf2 · outbound

This paper cites Self-projection and the brain.

Implicit Counterfactual Learning for Audio-Visual Segmentation Self-projection and the brain

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:21.092219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:21.092219Z digest=sha256:ab1b1ff62bde6a75f92534275b3b7432e29ced4304df11bd5da5a9102cb6888a

Observation d1420b7c-6981-4053-8105-cf2018875008 · outbound

This paper cites Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues.

Implicit Counterfactual Learning for Audio-Visual Segmentation Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:32.180551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:21.302100Z digest=sha256:d917e5c22607c23cfee7ffb68c861aec14bec1e9c1fc0292d43dc05f1a6f3d7a

Observation f0f5345d-472c-470c-9d4f-604e5b30339e · outbound

This paper cites Unraveling in- stance associations: A closer look for audio-visual segmenta- tion.

Implicit Counterfactual Learning for Audio-Visual Segmentation Unraveling in- stance associations: A closer look for audio-visual segmenta- tion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:21.405095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:21.405095Z digest=sha256:7e67c4874101546064fcf9c165425b316a89b74ef8795bf37ccf13c0b2476578

Observation 030edca1-1e34-4e8e-beaa-9300ac0f420a · outbound

This paper cites Cpm: Class-conditional prompting ma- chine for audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cpm: Class-conditional prompting ma- chine for audio-visual segmentation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:21.553958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:21.553958Z digest=sha256:94446f0a334087c902ffe5aa8d067ac6177191d17e33c268fffdad50f7c1fff6

Observation a6ea029d-b907-43e9-9046-4109dee5f372 · outbound

This paper cites Cpm: Class-conditional prompting ma- chine for audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cpm: Class-conditional prompting ma- chine for audio-visual segmentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:21.724840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:21.724840Z digest=sha256:29a7cc2a5ee6a04b54992e18b0f1a1edd673e8e297c7bbd9c54989359ac0bc57

Observation 1ce8f95b-8eaa-45d5-aea5-7067c2022032 · outbound

This paper cites C-cam: Causal cam for weakly supervised seman- tic segmentation on medical image.

Implicit Counterfactual Learning for Audio-Visual Segmentation C-cam: Causal cam for weakly supervised seman- tic segmentation on medical image

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.508793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:21.852776Z digest=sha256:ed832e6ad965332906fea5b5092e230a4e4804c42a0255bce6a7df83596487d0

Observation c6142001-4757-44a9-bc47-8aa03bb6c9c2 · outbound

This paper cites Asi- seg: Audio-driven surgical instrument segmentation with surgeon intention understanding.

Implicit Counterfactual Learning for Audio-Visual Segmentation Asi- seg: Audio-driven surgical instrument segmentation with surgeon intention understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.498445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:21.943556Z digest=sha256:b3d13841b4a48dd86cc8e37fa9a99013c187a4737043daaff95bdb586e844120

Observation fb8e7385-9045-4591-820f-a702f4b0b486 · outbound

This paper cites Regret and its avoidance: a neuroimaging study of choice behavior.

Implicit Counterfactual Learning for Audio-Visual Segmentation Regret and its avoidance: a neuroimaging study of choice behavior

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.487167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:22.148549Z digest=sha256:b6a1acea3caec704680abfbb9aeeb12d22592080ac485969b5ae893322ca6dc4

Observation fbd22937-7218-4a6c-b92f-ef4c741c36b6 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

Implicit Counterfactual Learning for Audio-Visual Segmentation An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:22.313670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:22.313670Z digest=sha256:86505649a87263176bcf1a4a8823620111380b0da0ffb7ecd4f39393927cc192

Observation a1df29dc-3460-4ad4-b47e-84b0ecfdb377 · outbound

This paper cites Avsegformer: Audio-visual segmentation with trans- former.

Implicit Counterfactual Learning for Audio-Visual Segmentation Avsegformer: Audio-visual segmentation with trans- former

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.477009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:22.499286Z digest=sha256:d6be26ca7b21046134d5f9cd426b276133664c59c0673388487eb42ba7339834

Observation 55ba5b81-6af3-40a3-896e-5c9ab45c8ab2 · outbound

This paper cites Open- vocabulary audio-visual semantic segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Open- vocabulary audio-visual semantic segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.467530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:22.665424Z digest=sha256:168c476ff188b24033d1398f6e05312443d35696d91f35497617a31d7dcc5f9e

Observation a40a676a-cf8a-4e3e-af99-e85fb9a9efcf · outbound

This paper cites Embodied intelligence via learning and evolution.

Implicit Counterfactual Learning for Audio-Visual Segmentation Embodied intelligence via learning and evolution

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.458399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:22.804976Z digest=sha256:f7a34e16ee950e3ad323de8eb5fbab665f6649467e7fc3e0770a931226b57985

Observation 7d8ebeb7-ef7a-4eca-9289-3da185fc510d · outbound

This paper cites Improving audio-visual segmenta- tion with bidirectional generation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Improving audio-visual segmenta- tion with bidirectional generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.449069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:22.948652Z digest=sha256:d19f34bfe782bdf6a857aeed07969ca8d486b561812d81b4fe2ca7f234d556f0

Observation 7db11bdb-d39e-4059-89d5-b0291245b940 · outbound

This paper cites Deep residual learning for image recognition.

Implicit Counterfactual Learning for Audio-Visual Segmentation Deep residual learning for image recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:23.073062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:23.073062Z digest=sha256:061465415d4d0c2fa9d740c9a6188d7974da0d285abccd7bcf6d4a07a7595191

Observation 7db395b5-bece-416d-a433-19a801035327 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Momentum contrast for unsupervised visual rep- resentation learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.433031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:23.261042Z digest=sha256:1bc557e174fa5fbb45a9361d2119c0077eefc3496b109fea1794e1e7ee2f4177

Observation 11788e0e-e091-4466-80f6-be4069bbd29f · outbound

This paper cites Decoupling static and hier- archical motion perception for referring video segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Decoupling static and hier- archical motion perception for referring video segmentation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.424334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:23.491221Z digest=sha256:1539581b9cc39b0aac5212200d5c65cb7a75829dfcd4b8e742339f07fee0d6b4

Observation a084624f-8829-4e51-9103-40308de46c4f · outbound

This paper cites A general mechanism for perceptual decision-making in the human brain.

Implicit Counterfactual Learning for Audio-Visual Segmentation A general mechanism for perceptual decision-making in the human brain

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.415636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:24.452927Z digest=sha256:1c99ee9a8a399b55f4fd2af47df3b536738005e8e30b08f117a187df6297d3a7

Observation 25d03ac9-fee1-4fbe-b4e1-10adddf4bc64 · outbound

This paper cites Cnn archi- tectures for large-scale audio classification.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cnn archi- tectures for large-scale audio classification

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.406901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:24.614224Z digest=sha256:7a7c54aaa5a9e84e15c8e27ee82be8ad2a5047a359823470b7a332045cb1600d

Observation 0439d45d-3416-427a-a8f1-2f5fcc1fe83e · outbound

This paper cites Aleatory and epistemic uncertainty in prob- ability elicitation with an example from hazardous waste management.

Implicit Counterfactual Learning for Audio-Visual Segmentation Aleatory and epistemic uncertainty in prob- ability elicitation with an example from hazardous waste management

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.397744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:24.755532Z digest=sha256:65f930d232ea146b9232fe0f9e7bed86a5da39f94ff3584cc44cfa4b69e236eb

Observation ad0cc980-5875-4b85-8a50-691e04dbeb66 · outbound

This paper cites Mix and local- ize: Localizing sound sources in mixtures.

Implicit Counterfactual Learning for Audio-Visual Segmentation Mix and local- ize: Localizing sound sources in mixtures

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.389427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:24.860562Z digest=sha256:f3c919ae34315e3d8df8481959a1e3ce434f505e4ba364796b3c67f509a69ed8

Observation 4bd25916-ba48-4257-b022-57a94eebb045 · outbound

This paper cites Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models.

Implicit Counterfactual Learning for Audio-Visual Segmentation Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.380217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:25.010580Z digest=sha256:9f99a9c53bec81c5b4386f367836320ac041aedd81bd0b7f6487ca2009afb746

Observation 22163450-fa2b-4b2c-b5fa-5ccb7264a524 · outbound

This paper cites Discovering Sounding Objects by Audio Queries for Audio Visual Segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Discovering Sounding Objects by Audio Queries for Audio Visual Segmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:25.169098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:25.169098Z digest=sha256:8109a4bee08ab341d4df5247a97c4c9e99fead23060c411a34875504510dd9ec

Observation 68326339-9d84-4988-b9c3-9c53df2ff659 · outbound

This paper cites Adaptive mixtures of local experts.Neu- ral computation, 3(1):79–87, 1991.

Implicit Counterfactual Learning for Audio-Visual Segmentation Adaptive mixtures of local experts.Neu- ral computation, 3(1):79–87, 1991

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.371190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:25.330912Z digest=sha256:63b02386293fd580436be71db2275fc4cc31691ff0552e1c141a73dc5cf36711

Observation 6c4c9f0d-b966-4443-b3cb-2633c9986a55 · outbound

This paper cites Counterfactually augmented event matching for de-biased temporal sentence grounding.

Implicit Counterfactual Learning for Audio-Visual Segmentation Counterfactually augmented event matching for de-biased temporal sentence grounding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.362368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:25.559673Z digest=sha256:bb7b9bb526e77d889402ee1a2ac77c2b7044a3fd72035534809ceac8573dc214

Observation 993faee2-f328-4068-9126-e24260f76fb1 · outbound

This paper cites Learning to visually localize sound sources from mix- tures without prior source knowledge.

Implicit Counterfactual Learning for Audio-Visual Segmentation Learning to visually localize sound sources from mix- tures without prior source knowledge

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.353458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:25.728411Z digest=sha256:ff5a3fbb7e692da1ad72ebeaab3ffde384c9c4d5737cddc0c09a9a6752244d7a

Observation dd417650-4dd9-4d51-b7d5-2064dcdfbd7e · outbound

This paper cites Segment Anything.

Implicit Counterfactual Learning for Audio-Visual Segmentation Segment Anything

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:25.802594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:25.802594Z digest=sha256:17661c50a5b0f7dcec65ed5772b35d2443059e66b230a789ba7f43241890af7c

Observation e63916f1-06da-44b1-9a0a-9fb875251c68 · outbound

This paper cites The singular value decompo- sition: Its computation and some applications.

Implicit Counterfactual Learning for Audio-Visual Segmentation The singular value decompo- sition: Its computation and some applications

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.344825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:25.806450Z digest=sha256:e9eee67f25bdd7360e6667c5134ed989cd15d26b627703baa26a06931fd5b6eb

Observation 474a0f74-f627-4521-b5f7-23d07f5141c8 · outbound

This paper cites Improving vision and language concepts understanding with multimodal counterfactual samples.

Implicit Counterfactual Learning for Audio-Visual Segmentation Improving vision and language concepts understanding with multimodal counterfactual samples

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.335417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:25.892511Z digest=sha256:0aab6f382f9aadf709f73d6718edca7a921854b49a827c44a54f28424e3ceb17

Observation 408135f8-3618-45b7-bf4c-cdddf4533d7d · outbound

This paper cites Selm: Selective mechanism based audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Selm: Selective mechanism based audio-visual segmentation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.325244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:26.058249Z digest=sha256:d2a0ea3b2bad712fa7a4da13511f37d2e57e7fa2673937a2aac47c4de5b2a2e0

Observation cb089aa7-6c78-41ad-8705-7df5c63882fe · outbound

This paper cites Catr: Combinatorial-dependence audio-queried transformer for audio-visual video segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Catr: Combinatorial-dependence audio-queried transformer for audio-visual video segmentation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.315333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:26.185331Z digest=sha256:b4caf59339a457783bae0e369dbd0b56088749cc487d45ec991a044dc69079af

Observation 4bda86bf-c11d-4172-9739-69e6e3a052d8 · outbound

This paper cites Dice Loss for Data-imbalanced NLP Tasks.

Implicit Counterfactual Learning for Audio-Visual Segmentation Dice Loss for Data-imbalanced NLP Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:26.309876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:26.309876Z digest=sha256:5a4b00099a4a983677e46f59c4f4befc2e0806167855fecf2cfb1003f836203a

Observation dc6ce626-8fa9-44c3-8a01-d4ab5140b521 · outbound

This paper cites Qdformer: Towards ro- bust audiovisual segmentation in complex environments with quantization-based semantic decomposition.

Implicit Counterfactual Learning for Audio-Visual Segmentation Qdformer: Towards ro- bust audiovisual segmentation in complex environments with quantization-based semantic decomposition

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.305293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:26.423838Z digest=sha256:f41d388ba49481bcafdc261d2e2eb7f54baaa786809082db91edfe87a87d126b

Observation 07b5e024-bb4d-4dfc-8e01-9cb7c3ac55b4 · outbound

This paper cites Bavs: bootstrapping audio- visual segmentation by integrating foundation knowledge.

Implicit Counterfactual Learning for Audio-Visual Segmentation Bavs: bootstrapping audio- visual segmentation by integrating foundation knowledge

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.084022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:26.526751Z digest=sha256:a1433e772c2d31b0d5699b1ee209164a9deb45964a564f0b15203e1678b465ce

Observation cf977619-d618-4910-b7d3-7d9da6e76caf · outbound

This paper cites Benchmarking au- dio visual segmentation for long-untrimmed videos.

Implicit Counterfactual Learning for Audio-Visual Segmentation Benchmarking au- dio visual segmentation for long-untrimmed videos

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.893219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:26.586204Z digest=sha256:9d721fc1b31b29811e1ac71de3ffd97b408315cf5a2390f0428a8372a42e2d71

Observation b31a92b5-bc64-4781-84c0-0900693af95b · outbound

This paper cites Pay attention to mlps.

Implicit Counterfactual Learning for Audio-Visual Segmentation Pay attention to mlps

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.688252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:26.693841Z digest=sha256:22093b35d3f58ad728f9261c48dde8cc384e4d0e16c8246aa74e760a6c72d3dd

Observation 83406f5f-7bb9-4ab9-8444-f8fed48f887d · outbound

This paper cites Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:26.828928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:26.828928Z digest=sha256:36920f1a5e2b0a86c0bb8a7d6c99ff60a1fffdb15abf6d0597c922e7c684ec59

Observation 2c8f30a5-6a29-4d72-bdfd-1edacc700458 · outbound

This paper cites Annotation-free Audio-Visual Segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Annotation-free Audio-Visual Segmentation

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:31.926622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:26.936793Z digest=sha256:567a092d4933ce23fc4dae5b3defc4229959609df25154bd7de94dcc23ddf901

Observation 625ce05b-28ab-495a-b273-91ed366358e5 · outbound

This paper cites Audio-visual segmentation via unlabeled frame exploitation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio-visual segmentation via unlabeled frame exploitation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.518602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:27.047524Z digest=sha256:778faef15f6a6bd6c93b52e64c150fbbd285d35348f4019ab8546b8b706c7905

Observation c7579be3-fa9c-49a3-be3a-10f786024ce9 · outbound

This paper cites Cross-modal causal relational reasoning for event-level visual question answer- ing.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cross-modal causal relational reasoning for event-level visual question answer- ing

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.282649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:27.109504Z digest=sha256:0c245aec9dcc4c5dcb3ae764a8ec6a6fa315b933c1f6f6ff9bb4edce9b6cb60c

Observation a06da19b-3e21-4ffb-b1d3-ac841e088740 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Implicit Counterfactual Learning for Audio-Visual Segmentation Swin transformer: Hierarchical vision transformer using shifted windows

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:27.197551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:27.197551Z digest=sha256:07003cfa1dc2c43fab3804325841a1d624b9540210660d4be0565688a49a55a9

Observation 13a86cc8-f036-4e73-9a36-d75679152aaa · outbound

This paper cites Step- ping stones: a progressive training strategy for audio-visual semantic segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Step- ping stones: a progressive training strategy for audio-visual semantic segmentation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.131875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:27.320757Z digest=sha256:795798d000300155e96017b771fc8d6ec00f3c5d3b5ef86bceb78ccc81da1da1

Observation b64d31e0-fc4b-46aa-a4d2-719d938986bf · outbound

This paper cites T-vsl: Text-guided visual sound source localization in mixtures.

Implicit Counterfactual Learning for Audio-Visual Segmentation T-vsl: Text-guided visual sound source localization in mixtures

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.993803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:27.386920Z digest=sha256:ef144e5ed7c449d650b90fabcd3c1d30d0c32408d53610eb12492cedf37ebb7d

Observation 076f56ba-c832-4b60-b184-fb3d02046c3e · outbound

This paper cites Contrastive Conditional Latent Diffusion for Audio-visual Segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Contrastive Conditional Latent Diffusion for Audio-visual Segmentation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:31.766124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:27.507201Z digest=sha256:61015cb7f0442ad9a1071c0775d167b6db25621985aa5279dae91a8126f1262b

Observation 055e248f-983a-45a3-8192-908b2a38d400 · outbound

This paper cites Multimodal variational auto-encoder based audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Multimodal variational auto-encoder based audio-visual segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.818435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:27.623102Z digest=sha256:04a92dbfb3bba38570155b0c488488e75e9bddd33aacafdf78147ac86eb1f469

Observation 6bd68199-fb71-4bcd-95f9-262304aef98b · outbound

This paper cites Weakly-supervised audio- visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Weakly-supervised audio- visual segmentation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:27.707306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:27.707306Z digest=sha256:fc2427d60222b7728dbf20ee97aa83949e2960b3ac9347fa01ac296f961e2654

Observation 679f4081-6369-4b62-8236-66eb37155659 · outbound

This paper cites Audio-visual grouping net- work for sound localization from mixtures.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio-visual grouping net- work for sound localization from mixtures

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.647904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:27.831618Z digest=sha256:a928f1451a6ab2ba4ba5cf37a7aa72b33889da96f6eb946a493fcfda017dc5e9

Observation b1a01a9e-85cd-4fce-8f44-da6ae9d201d1 · outbound

This paper cites The book of why: the new science of cause and effect.

Implicit Counterfactual Learning for Audio-Visual Segmentation The book of why: the new science of cause and effect

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.501190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:27.932310Z digest=sha256:a90667afb50f0f2a979a28d5607b2c4334f86912b8c3758679d3b874c58a5b36

Observation bcabcb0c-4af6-402f-ac10-adf0ba432cfd · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Implicit Counterfactual Learning for Audio-Visual Segmentation Learning transferable visual models from natural language supervi- sion

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:28.047650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:28.047650Z digest=sha256:9a5b883e5898afaede56d9bf04c2ef1040b1a69e7dd979ecb54632d1a493de0b

Observation 550e563d-bb74-48bb-8da8-98f2af648330 · outbound

This paper cites Zero-shot text-to-image generation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Zero-shot text-to-image generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:28.155249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:28.155249Z digest=sha256:0636565bc9ae20109df0007ed0c06c55d6d74116a7b176db11bb6a1bc3fea6a8

Observation a0617c11-9af7-4ca2-aed7-99a53b509b4b · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Implicit Counterfactual Learning for Audio-Visual Segmentation High-resolution image synthesis with latent diffusion models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.393079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:28.258287Z digest=sha256:734a54f318e9c7525511dc8f4eaed0efaf1e0fb6399bc50dd15641aae6d805f2

Observation 5e9251e1-5e22-4c0c-87d5-841fd30864bb · outbound

This paper cites Focal loss for dense ob- ject detection.

Implicit Counterfactual Learning for Audio-Visual Segmentation Focal loss for dense ob- ject detection

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:28.331897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:28.331897Z digest=sha256:5eb1ff7a454127e5e5a9e0cb085442f62a52b49351efb0d5f899ea319c209649

Observation af45f266-bf12-4f4b-8d4c-88b9a904e7be · outbound

This paper cites The wasserstein distance and approxi- mation theorems.

Implicit Counterfactual Learning for Audio-Visual Segmentation The wasserstein distance and approxi- mation theorems

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.249885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:28.439113Z digest=sha256:4c4dbec4f35aa7e1a58cc1380aab5241555fda7700eb359ed10fce4d50b1ee41

Observation f7ab43a7-705a-400e-92ea-0d19b46e369c · outbound

This paper cites Better aggregation in test-time augmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Better aggregation in test-time augmentation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.097862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:28.519677Z digest=sha256:f04a4960a40a65034ab49b3050778f59e9533ccac6fb8b821c42544c8670f447

Observation ffb19ff6-108c-4007-954d-feb8189ad7d3 · outbound

This paper cites A mathematical theory of commu- nication.

Implicit Counterfactual Learning for Audio-Visual Segmentation A mathematical theory of commu- nication

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.925906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:28.634597Z digest=sha256:fd6c4598d011b16f188e136b8c2e0579411cc4ea9dbbf5f2538f9fc6d5c80991

Observation 99535a89-fd96-4c7c-bd56-9077b0df0bef · outbound

This paper cites Looking similar sounding different: Leveraging counterfactual cross-modal pairs for audiovisual representa- tion learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Looking similar sounding different: Leveraging counterfactual cross-modal pairs for audiovisual representa- tion learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.804483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:28.704904Z digest=sha256:27d0c6d0e2cee381373fcb813f45041ca23d08088288e0000eabc3c17eeec3d7

Observation 23110967-391c-4226-986c-db595b6da29e · outbound

This paper cites 3d audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation 3d audio-visual segmentation

Reference 59

Resolution
verified exact
raw_fallback, observed 2026-08-06T13:24:31.614704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:28.777772Z digest=sha256:da438d14d2788adaabaa32c632ebb043b462e451bc62c65354abfdd0881ed185

Observation e2a02759-95dd-43e7-ac93-50bfdd89d177 · outbound

This paper cites Unveiling and mitigating bias in audio visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Unveiling and mitigating bias in audio visual segmentation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.702149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:28.858392Z digest=sha256:cb2ea0cfd9b276589725e7de3b8c2ed9024dc2be697dcb1e3ce92c5662a3fa9e

Observation 15377690-589c-47c1-8b04-87f7639ab53a · outbound

This paper cites Learning audio-visual source localization via false negative aware contrastive learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Learning audio-visual source localization via false negative aware contrastive learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.572731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:28.979464Z digest=sha256:e713ee32543153335551487e32aa82029f4e9f6ce03638e1219e584bc310312c

Observation 2515ccf1-998f-4f4d-ac57-f3a328a0e817 · outbound

This paper cites Language-guided audio-visual source separation via trimodal consistency.

Implicit Counterfactual Learning for Audio-Visual Segmentation Language-guided audio-visual source separation via trimodal consistency

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.417838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:29.061286Z digest=sha256:de7b50717d9a6950e4878125722e770f2ed6147e36d5c32e64af6096eeb99100

Observation 0c2e86cd-6ec9-4654-a9a2-c8867e76ca0a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Implicit Counterfactual Learning for Audio-Visual Segmentation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:29.167740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:29.167740Z digest=sha256:03ff9f1a17e14fe7ff7a71de0f2ce7cc864aec50da4fa2aa5ab67a8e634c9b63

Observation 8a2ca42f-3d2f-4d93-a5d3-fb2d02c7faac · outbound

This paper cites Neural discrete representation learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Neural discrete representation learning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.232776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:29.269444Z digest=sha256:3d1c70b3282f1ce8fc4c197fdbf40d23d1cef5b359be4c8bc278d4cdca93c3f1

Observation f2efaf7a-148a-4976-ac2b-eb2b10eaeb84 · outbound

This paper cites Counterfactual cycle-consistent learn- ing for instruction following and generation in vision- language navigation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Counterfactual cycle-consistent learn- ing for instruction following and generation in vision- language navigation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.985658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:29.340354Z digest=sha256:00b7a1565298451e09b714c0ef8ead3d54aea9e80bca6d8033c474e17522420e

Observation dfe75355-433d-4e55-8086-5c7bb61ea5b4 · outbound

This paper cites Vision-and-language naviga- tion via causal learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Vision-and-language naviga- tion via causal learning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.817968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:29.450485Z digest=sha256:423c31c7817e27f86473db1996f033aa0e8d8f386e997bcb7bb068e963b115e8

Observation 6b0363dd-d240-4a07-b11c-99257ee012d2 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Implicit Counterfactual Learning for Audio-Visual Segmentation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:29.575827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:29.575827Z digest=sha256:f4c693c152df88b26848ee9777fbf47f213643fb592acca63e0d4b02391036bf

Observation 57326a31-7fed-419a-9e03-f04f782c88ca · outbound

This paper cites Pvt v2: Improved baselines with pyramid vision transformer.

Implicit Counterfactual Learning for Audio-Visual Segmentation Pvt v2: Improved baselines with pyramid vision transformer

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.591109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:29.651671Z digest=sha256:c44abfc6fd06e4b02168a7a026c15f6f916969321146f96ca42c2c6463464e3c

Observation 2d8e8a11-5f77-4b10-8d4c-a21c080864b3 · outbound

This paper cites Drivedreamer: Towards real-world- drive world models for autonomous driving.

Implicit Counterfactual Learning for Audio-Visual Segmentation Drivedreamer: Towards real-world- drive world models for autonomous driving

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:29.738949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:29.738949Z digest=sha256:292d56573f9c208393d6c4ac791d2d5867a2bb836f6549cd78b2182e08c5cf73

Observation b00d8a91-7e04-44c4-be8e-414d4a62e7f1 · outbound

This paper cites Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer.

Implicit Counterfactual Learning for Audio-Visual Segmentation Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:31.372411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:29.839727Z digest=sha256:98166c218cf3db3cfa1f90d4579ad4a739c44db9e1494a75543ed2be805fe52e

Observation d84413af-577e-40c0-bc13-1a6e731270b3 · outbound

This paper cites Prompting segmentation with sound is gen- eralizable audio-visual source localizer.

Implicit Counterfactual Learning for Audio-Visual Segmentation Prompting segmentation with sound is gen- eralizable audio-visual source localizer

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.366851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:29.911767Z digest=sha256:95416279444c6048f44e370b522ebaca557ebc13f9c58ba21bd79b0620f1e028

Observation 2e1162cf-9d17-4e68-b0cc-8e8d71258c23 · outbound

This paper cites Can textual semantics mitigate sounding object segmentation preference? In European Conference on Com- puter Vision, pages 340–356.

Implicit Counterfactual Learning for Audio-Visual Segmentation Can textual semantics mitigate sounding object segmentation preference? In European Conference on Com- puter Vision, pages 340–356

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.206530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:30.002229Z digest=sha256:f040bbc02a4bfc8e54ec81dcea16e624b88836e8970ecda9610cc6616fa1c398

Observation 145afdec-8d79-44d4-a233-1db959720e1c · outbound

This paper cites Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.033283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:30.078314Z digest=sha256:f9cb204af3de0a12a0c943bdeb23a629116d06c1e7c6647e4b5c68f612c26d44

Observation 2f93e96f-c19c-41b4-8489-c070a09d9865 · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

Implicit Counterfactual Learning for Audio-Visual Segmentation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:30.193705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:30.193705Z digest=sha256:1ccf4d4c51320280ec58fe7b1a7950e2083ee65febd70fce31a539df0628e0d8

Observation 25e4326e-a60c-4e3f-b7ee-99d4bb643f4a · outbound

This paper cites Referred by multi-modality: A unified tem- poral transformer for video object segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Referred by multi-modality: A unified tem- poral transformer for video object segmentation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.919685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:30.264676Z digest=sha256:6d98e077fff0bf050e3cda701aa2c07186d992ec8682e6091693be7afc6459c4

Observation 5434c471-de08-4491-804c-0091f009c7c1 · outbound

This paper cites Cooperation does matter: Exploring multi-order bilateral relations for audio- visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cooperation does matter: Exploring multi-order bilateral relations for audio- visual segmentation

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.757091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:30.395792Z digest=sha256:9518262cc8754b9998347ad6885c3195fc3f7dca76f84eb3966fd5c366967c1a

Observation c1f36da8-5497-4e1f-9186-b2268aa953a2 · outbound

This paper cites Revisiting counterfactual prob- lems in referring expression comprehension.

Implicit Counterfactual Learning for Audio-Visual Segmentation Revisiting counterfactual prob- lems in referring expression comprehension

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.617230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:30.482824Z digest=sha256:55aa1b8b809d59aa358c5be925a8e578e15b393b458d3434d4ca59543105d325

Observation fc20ea18-d27a-4797-a966-e1b8bd72dcee · outbound

This paper cites Losh: Long-short text joint prediction network for referring video object segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Losh: Long-short text joint prediction network for referring video object segmentation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.459830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:30.591093Z digest=sha256:8a257f942198778279a3ad4db02882f10b558ade601ca3d73aedf0c5cdd3aca8

Observation 9c325e08-f56f-4645-b594-f7a34b8e575a · outbound

This paper cites Discovering the real association: Multimodal causal rea- soning in video question answering.

Implicit Counterfactual Learning for Audio-Visual Segmentation Discovering the real association: Multimodal causal rea- soning in video question answering

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.321674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:30.648645Z digest=sha256:39e681870a3c3a356520a48f3153c8e5a8749b1ffa0651fcd90bb7a0ad1c4446

Observation de25b491-8657-4d1f-8b76-893d79779aad · outbound

This paper cites Weakly- supervised mirror detection via scribble annotations.

Implicit Counterfactual Learning for Audio-Visual Segmentation Weakly- supervised mirror detection via scribble annotations

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.147692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:30.711366Z digest=sha256:cc101479bb35ea8d976930321633e38dc27650b67d0972421a2231a3f0827ede

Observation 7410b16a-d8ab-4b01-b633-cbee76be50da · outbound

This paper cites Heterogeneous experts and hierarchical perception for un- derwater salient object detection.

Implicit Counterfactual Learning for Audio-Visual Segmentation Heterogeneous experts and hierarchical perception for un- derwater salient object detection

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.969158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:30.787086Z digest=sha256:cc663e06859a681760cb103fea06ad64e583d195957f1186a43d7db93ae1248e

Observation 5310ff46-d21b-4fc5-836b-00592f7203dc · outbound

This paper cites Causal intervention for weakly- supervised semantic segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Causal intervention for weakly- supervised semantic segmentation

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.874213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:30.860023Z digest=sha256:6e81f5cae78dee387936bc74f4ce9f59e016f74dce46d3d1d381e486cbce28db

Observation 919094cc-387d-44d6-9e58-d36771852dd4 · outbound

This paper cites Audio–visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio–visual segmentation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:30.918766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:30.918766Z digest=sha256:12651ac90450750ed715857c92a61407b5fc5b6239e09a72488c7dc505c009d4

Observation ed7a1761-fd70-4dab-bcb5-a3e28308dedc · outbound

This paper cites Audio-visual segmentation with semantics.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio-visual segmentation with semantics

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.591154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:31.109915Z digest=sha256:3ae89059ca503c4a9289e4699a32155d5bddb5258984f5caece67f05d53465e0

Observation cc711852-876c-4aff-83cc-ffd8923a2ae3 · outbound

This paper cites Exploring pre-trained text- to-video diffusion models for referring video object segmen- tation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Exploring pre-trained text- to-video diffusion models for referring video object segmen- tation

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.424693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:31.180804Z digest=sha256:4e30056f17844db497d8312d39dc1909cbc79caecf7ee9fdf2708dfb4f1b66db

Observation 9abe49c0-d912-40e2-9f47-04af5b5e656a · outbound

This paper cites 2, 5, 6, 8.

Implicit Counterfactual Learning for Audio-Visual Segmentation 2, 5, 6, 8

Reference 403

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.709159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:24:31.029697Z digest=sha256:b09589d9cc6723ec6bdab8460a7d2f309c9641ffccba4f384e570d0bfbc90001

Pith citing papers

No inbound Pith citation observations are available.