Pith. sign in

Paper Citation Record · LEDGER

Implicit Counterfactual Learning for Audio-Visual Segmentation

As of 10 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 0 inbound Pith citation observations for arXiv:2507.20740.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20740 v1

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:24:31.180804Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

86 of 86 outbound references displayed

  • verified exact6
  • verified fuzzy58
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 92232e28-b925-4df0-98d6-3cd58f3e6a75 · outbound

This paper cites Remembering the past and imagining the future: Common and distinct neural substrates during event construction and elaboration.

Implicit Counterfactual Learning for Audio-Visual Segmentation Remembering the past and imagining the future: Common and distinct neural substrates during event construction and elaboration

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:20.667712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:20.667712Z digest=sha256:557f4653d03efb8099c41be423eedf6946b01ffdef0bc4be37d19c7c7365e408

Observation 9eb1ad31-b9f1-4f74-96f2-1c505fd5a6b4 · outbound

This paper cites Unsupervised Audio-Visual Segmentation with Modality Alignment.

Implicit Counterfactual Learning for Audio-Visual Segmentation Unsupervised Audio-Visual Segmentation with Modality Alignment

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:32.277437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:20.791836Z digest=sha256:43af08d59aaca7526d5811f13b1e8f3c9fd2262e69829c30fc688ffbc778e744

Observation a94ba4dc-11e1-4712-acd0-31967afc369f · outbound

This paper cites Numerics of gram-schmidt orthogonalization.

Implicit Counterfactual Learning for Audio-Visual Segmentation Numerics of gram-schmidt orthogonalization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:20.951174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:20.951174Z digest=sha256:15b4ebfec832a682643a094a7891aae9d3b79b8230605e4fdd61fc2707d237e2

Observation a04bf18d-2e62-4197-8c4a-71e71f82acf2 · outbound

This paper cites Self-projection and the brain.

Implicit Counterfactual Learning for Audio-Visual Segmentation Self-projection and the brain

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:21.092219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:21.092219Z digest=sha256:ab1b1ff62bde6a75f92534275b3b7432e29ced4304df11bd5da5a9102cb6888a

Observation d1420b7c-6981-4053-8105-cf2018875008 · outbound

This paper cites Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues.

Implicit Counterfactual Learning for Audio-Visual Segmentation Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:32.180551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:21.302100Z digest=sha256:5831a460b7cca23d3c4c0a892b3919989ac83baf978e4fe70fc2696abeef21f6

Observation f0f5345d-472c-470c-9d4f-604e5b30339e · outbound

This paper cites Unraveling in- stance associations: A closer look for audio-visual segmenta- tion.

Implicit Counterfactual Learning for Audio-Visual Segmentation Unraveling in- stance associations: A closer look for audio-visual segmenta- tion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:21.405095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:21.405095Z digest=sha256:7e67c4874101546064fcf9c165425b316a89b74ef8795bf37ccf13c0b2476578

Observation 030edca1-1e34-4e8e-beaa-9300ac0f420a · outbound

This paper cites Cpm: Class-conditional prompting ma- chine for audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cpm: Class-conditional prompting ma- chine for audio-visual segmentation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:21.553958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:21.553958Z digest=sha256:94446f0a334087c902ffe5aa8d067ac6177191d17e33c268fffdad50f7c1fff6

Observation a6ea029d-b907-43e9-9046-4109dee5f372 · outbound

This paper cites Cpm: Class-conditional prompting ma- chine for audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cpm: Class-conditional prompting ma- chine for audio-visual segmentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:21.724840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:21.724840Z digest=sha256:29a7cc2a5ee6a04b54992e18b0f1a1edd673e8e297c7bbd9c54989359ac0bc57

Observation 1ce8f95b-8eaa-45d5-aea5-7067c2022032 · outbound

This paper cites C-cam: Causal cam for weakly supervised seman- tic segmentation on medical image.

Implicit Counterfactual Learning for Audio-Visual Segmentation C-cam: Causal cam for weakly supervised seman- tic segmentation on medical image

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.508793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:21.852776Z digest=sha256:376bb59798e73aa6d8d7b3db5c2cb0092705c9ccb4378498b9e2c2a245a4b00b

Observation c6142001-4757-44a9-bc47-8aa03bb6c9c2 · outbound

This paper cites Asi- seg: Audio-driven surgical instrument segmentation with surgeon intention understanding.

Implicit Counterfactual Learning for Audio-Visual Segmentation Asi- seg: Audio-driven surgical instrument segmentation with surgeon intention understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.498445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:21.943556Z digest=sha256:d3b5d1cb218f7415748f51a912e45ea71133b2f09ecd1f9f67c9bc781b96e685

Observation fb8e7385-9045-4591-820f-a702f4b0b486 · outbound

This paper cites Regret and its avoidance: a neuroimaging study of choice behavior.

Implicit Counterfactual Learning for Audio-Visual Segmentation Regret and its avoidance: a neuroimaging study of choice behavior

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.487167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:22.148549Z digest=sha256:845a17b7a7ba29357b35ac13e4a8c169e93ae1a5c1e58430e5428d61ac057612

Observation fbd22937-7218-4a6c-b92f-ef4c741c36b6 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

Implicit Counterfactual Learning for Audio-Visual Segmentation An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:22.313670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:22.313670Z digest=sha256:86505649a87263176bcf1a4a8823620111380b0da0ffb7ecd4f39393927cc192

Observation a1df29dc-3460-4ad4-b47e-84b0ecfdb377 · outbound

This paper cites Avsegformer: Audio-visual segmentation with trans- former.

Implicit Counterfactual Learning for Audio-Visual Segmentation Avsegformer: Audio-visual segmentation with trans- former

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.477009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:22.499286Z digest=sha256:a974bfe99f517880261b554e01a96419c000b1053a22aa02b7651d941b784aad

Observation 55ba5b81-6af3-40a3-896e-5c9ab45c8ab2 · outbound

This paper cites Open- vocabulary audio-visual semantic segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Open- vocabulary audio-visual semantic segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.467530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:22.665424Z digest=sha256:545b12e51a0e16b4e31c80e7d6458387ebc6e51a4d065946db6788d1cd6df025

Observation a40a676a-cf8a-4e3e-af99-e85fb9a9efcf · outbound

This paper cites Embodied intelligence via learning and evolution.

Implicit Counterfactual Learning for Audio-Visual Segmentation Embodied intelligence via learning and evolution

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.458399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:22.804976Z digest=sha256:3c3d5af90c8d017d0f6f86187f2453a56571b51051518eb25e1c2dbad1211b21

Observation 7d8ebeb7-ef7a-4eca-9289-3da185fc510d · outbound

This paper cites Improving audio-visual segmenta- tion with bidirectional generation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Improving audio-visual segmenta- tion with bidirectional generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.449069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:22.948652Z digest=sha256:4e12bcc2398316575a96dd7f93b1275f5652e7cec67d9d46b7e47f547978bd7e

Observation 7db11bdb-d39e-4059-89d5-b0291245b940 · outbound

This paper cites Deep residual learning for image recognition.

Implicit Counterfactual Learning for Audio-Visual Segmentation Deep residual learning for image recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:23.073062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:23.073062Z digest=sha256:061465415d4d0c2fa9d740c9a6188d7974da0d285abccd7bcf6d4a07a7595191

Observation 7db395b5-bece-416d-a433-19a801035327 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Momentum contrast for unsupervised visual rep- resentation learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.433031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:23.261042Z digest=sha256:bdd62ae16ddffaca62f1002e9feee2b2e06e681b433ed38a24f76cd22c4038cf

Observation 11788e0e-e091-4466-80f6-be4069bbd29f · outbound

This paper cites Decoupling static and hier- archical motion perception for referring video segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Decoupling static and hier- archical motion perception for referring video segmentation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.424334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:23.491221Z digest=sha256:cd13b4ae609779b48925219737e3128d5bd3bfbd21962fab9347be3f3d77ead9

Observation a084624f-8829-4e51-9103-40308de46c4f · outbound

This paper cites A general mechanism for perceptual decision-making in the human brain.

Implicit Counterfactual Learning for Audio-Visual Segmentation A general mechanism for perceptual decision-making in the human brain

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.415636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:24.452927Z digest=sha256:9b7e17c7adef83a01c7ee106af171f9ac5f0e5bca4cd858df009595adcacad9a

Observation 25d03ac9-fee1-4fbe-b4e1-10adddf4bc64 · outbound

This paper cites Cnn archi- tectures for large-scale audio classification.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cnn archi- tectures for large-scale audio classification

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.406901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:24.614224Z digest=sha256:384746a83ee73124649b5432f6e37d9927c7b782c68702c8a9ba68854baa872b

Observation 0439d45d-3416-427a-a8f1-2f5fcc1fe83e · outbound

This paper cites Aleatory and epistemic uncertainty in prob- ability elicitation with an example from hazardous waste management.

Implicit Counterfactual Learning for Audio-Visual Segmentation Aleatory and epistemic uncertainty in prob- ability elicitation with an example from hazardous waste management

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.397744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:24.755532Z digest=sha256:807ce024ccf02bc3280ab0d64d1f698e8c51180b9f641db057a595c755cd5075

Observation ad0cc980-5875-4b85-8a50-691e04dbeb66 · outbound

This paper cites Mix and local- ize: Localizing sound sources in mixtures.

Implicit Counterfactual Learning for Audio-Visual Segmentation Mix and local- ize: Localizing sound sources in mixtures

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.389427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:24.860562Z digest=sha256:637bc916e895c40c8dadae574256b740a05d1a4e8843efbc67e8074d1f2fc7d2

Observation 4bd25916-ba48-4257-b022-57a94eebb045 · outbound

This paper cites Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models.

Implicit Counterfactual Learning for Audio-Visual Segmentation Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.380217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:25.010580Z digest=sha256:e947f6bc0dff45eed0129d671d23bcc4945a6a0e1d3b0de7cdf9d995db212c16

Observation 22163450-fa2b-4b2c-b5fa-5ccb7264a524 · outbound

This paper cites Discovering Sounding Objects by Audio Queries for Audio Visual Segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Discovering Sounding Objects by Audio Queries for Audio Visual Segmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:25.169098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:25.169098Z digest=sha256:11ca3039f9c4fa9572394224418338f6e034ca0c45aecd1048deb50b945dcf73

Observation 68326339-9d84-4988-b9c3-9c53df2ff659 · outbound

This paper cites Adaptive mixtures of local experts.Neu- ral computation, 3(1):79–87, 1991.

Implicit Counterfactual Learning for Audio-Visual Segmentation Adaptive mixtures of local experts.Neu- ral computation, 3(1):79–87, 1991

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.371190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:25.330912Z digest=sha256:6e23b6ecf10259a9b257035935f7ccd9334e4d5135b27331478d8465f3470fc3

Observation 6c4c9f0d-b966-4443-b3cb-2633c9986a55 · outbound

This paper cites Counterfactually augmented event matching for de-biased temporal sentence grounding.

Implicit Counterfactual Learning for Audio-Visual Segmentation Counterfactually augmented event matching for de-biased temporal sentence grounding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.362368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:25.559673Z digest=sha256:88d89f6c5bbe8757768159dd7aef9af6dc3a396c7504d9ba0cc6b567e0e4122d

Observation 993faee2-f328-4068-9126-e24260f76fb1 · outbound

This paper cites Learning to visually localize sound sources from mix- tures without prior source knowledge.

Implicit Counterfactual Learning for Audio-Visual Segmentation Learning to visually localize sound sources from mix- tures without prior source knowledge

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.353458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:25.728411Z digest=sha256:faff8259293f8b7dde5bc1fd20b3767371b0c278f05113497c1d9752946bfa4b

Observation dd417650-4dd9-4d51-b7d5-2064dcdfbd7e · outbound

This paper cites Segment Anything.

Implicit Counterfactual Learning for Audio-Visual Segmentation Segment Anything

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:25.802594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:25.802594Z digest=sha256:17661c50a5b0f7dcec65ed5772b35d2443059e66b230a789ba7f43241890af7c

Observation e63916f1-06da-44b1-9a0a-9fb875251c68 · outbound

This paper cites The singular value decompo- sition: Its computation and some applications.

Implicit Counterfactual Learning for Audio-Visual Segmentation The singular value decompo- sition: Its computation and some applications

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.344825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:25.806450Z digest=sha256:e89fe1e3bd670d97a2c3b2640c7287df0093dc9709ce5a1e7282e8362e4e23d6

Observation 474a0f74-f627-4521-b5f7-23d07f5141c8 · outbound

This paper cites Improving vision and language concepts understanding with multimodal counterfactual samples.

Implicit Counterfactual Learning for Audio-Visual Segmentation Improving vision and language concepts understanding with multimodal counterfactual samples

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.335417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:25.892511Z digest=sha256:a9af550c8ea0ed31288d1987f54253e4be8b4a81d2381c3bdefffbd77338fbae

Observation 408135f8-3618-45b7-bf4c-cdddf4533d7d · outbound

This paper cites Selm: Selective mechanism based audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Selm: Selective mechanism based audio-visual segmentation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.325244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:26.058249Z digest=sha256:7616f02c5e8315143528e802dbaa63290c839570d922d3669f9c34b1a43b3f9e

Observation cb089aa7-6c78-41ad-8705-7df5c63882fe · outbound

This paper cites Catr: Combinatorial-dependence audio-queried transformer for audio-visual video segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Catr: Combinatorial-dependence audio-queried transformer for audio-visual video segmentation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.315333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:26.185331Z digest=sha256:acd15dd6a6e51d36d34279f3cef0325b4d3cf2f95e4a2c8d7839f92c0d8309f0

Observation 4bda86bf-c11d-4172-9739-69e6e3a052d8 · outbound

This paper cites Dice Loss for Data-imbalanced NLP Tasks.

Implicit Counterfactual Learning for Audio-Visual Segmentation Dice Loss for Data-imbalanced NLP Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:26.309876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:26.309876Z digest=sha256:5a4b00099a4a983677e46f59c4f4befc2e0806167855fecf2cfb1003f836203a

Observation dc6ce626-8fa9-44c3-8a01-d4ab5140b521 · outbound

This paper cites Qdformer: Towards ro- bust audiovisual segmentation in complex environments with quantization-based semantic decomposition.

Implicit Counterfactual Learning for Audio-Visual Segmentation Qdformer: Towards ro- bust audiovisual segmentation in complex environments with quantization-based semantic decomposition

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.305293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:26.423838Z digest=sha256:78fb0a2dfcca9ae980ea0835ef730f8f57bee02a3b3cdf32dd98c77c4e10279d

Observation 07b5e024-bb4d-4dfc-8e01-9cb7c3ac55b4 · outbound

This paper cites Bavs: bootstrapping audio- visual segmentation by integrating foundation knowledge.

Implicit Counterfactual Learning for Audio-Visual Segmentation Bavs: bootstrapping audio- visual segmentation by integrating foundation knowledge

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.084022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:26.526751Z digest=sha256:7a81b1bd3ad43dfc1e715578f97db1520b0421b11693368c18372803ef4fdeb5

Observation cf977619-d618-4910-b7d3-7d9da6e76caf · outbound

This paper cites Benchmarking au- dio visual segmentation for long-untrimmed videos.

Implicit Counterfactual Learning for Audio-Visual Segmentation Benchmarking au- dio visual segmentation for long-untrimmed videos

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.893219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:26.586204Z digest=sha256:ef86c9f8c88382589d0fc5b381bb9f02e682a4e0a7ef198b0dfbe8bea98c0216

Observation b31a92b5-bc64-4781-84c0-0900693af95b · outbound

This paper cites Pay attention to mlps.

Implicit Counterfactual Learning for Audio-Visual Segmentation Pay attention to mlps

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.688252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:26.693841Z digest=sha256:37397920f9a5c2833c863b283731b5dd3037dccb44c0827d35b5f67cf3dc3231

Observation 83406f5f-7bb9-4ab9-8444-f8fed48f887d · outbound

This paper cites Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:26.828928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:26.828928Z digest=sha256:36920f1a5e2b0a86c0bb8a7d6c99ff60a1fffdb15abf6d0597c922e7c684ec59

Observation 2c8f30a5-6a29-4d72-bdfd-1edacc700458 · outbound

This paper cites Annotation-free Audio-Visual Segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Annotation-free Audio-Visual Segmentation

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:31.926622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:26.936793Z digest=sha256:5f64953fe699e96e9f3b8d95ab87d8558a972d1611bb6b1c850dfd0d713a1e24

Observation 625ce05b-28ab-495a-b273-91ed366358e5 · outbound

This paper cites Audio-visual segmentation via unlabeled frame exploitation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio-visual segmentation via unlabeled frame exploitation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.518602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:27.047524Z digest=sha256:4180f02833a458e31f1ea8c28faf4d3cc0b9b15a5dfdc955c0b98be087abcdee

Observation c7579be3-fa9c-49a3-be3a-10f786024ce9 · outbound

This paper cites Cross-modal causal relational reasoning for event-level visual question answer- ing.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cross-modal causal relational reasoning for event-level visual question answer- ing

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.282649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:27.109504Z digest=sha256:3a45716e1e1b5bda5128b0c8422f6382bd8813075d0dd09f7662cd6358c823d2

Observation a06da19b-3e21-4ffb-b1d3-ac841e088740 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Implicit Counterfactual Learning for Audio-Visual Segmentation Swin transformer: Hierarchical vision transformer using shifted windows

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:27.197551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:27.197551Z digest=sha256:07003cfa1dc2c43fab3804325841a1d624b9540210660d4be0565688a49a55a9

Observation 13a86cc8-f036-4e73-9a36-d75679152aaa · outbound

This paper cites Step- ping stones: a progressive training strategy for audio-visual semantic segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Step- ping stones: a progressive training strategy for audio-visual semantic segmentation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.131875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:27.320757Z digest=sha256:fefb642abc6410bb088407c6a653bdf7c88ffb5ca4c5d929408320de7dab94ff

Observation b64d31e0-fc4b-46aa-a4d2-719d938986bf · outbound

This paper cites T-vsl: Text-guided visual sound source localization in mixtures.

Implicit Counterfactual Learning for Audio-Visual Segmentation T-vsl: Text-guided visual sound source localization in mixtures

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.993803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:27.386920Z digest=sha256:8bcbe51ca7c37b23366e1a0c1b6e45f0c56852563da5f0e181e9692e92b61f33

Observation 076f56ba-c832-4b60-b184-fb3d02046c3e · outbound

This paper cites Contrastive Conditional Latent Diffusion for Audio-visual Segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Contrastive Conditional Latent Diffusion for Audio-visual Segmentation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:31.766124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:27.507201Z digest=sha256:7728c5f4bf7b709d604b78f1ff9b0cf82ac76f2afcb7669b7d34fb074d4a3d25

Observation 055e248f-983a-45a3-8192-908b2a38d400 · outbound

This paper cites Multimodal variational auto-encoder based audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Multimodal variational auto-encoder based audio-visual segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.818435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:27.623102Z digest=sha256:a8949912e74b215601e289d354065ed231f3bd43c880c7dd2059d76d0fcf8572

Observation 6bd68199-fb71-4bcd-95f9-262304aef98b · outbound

This paper cites Weakly-supervised audio- visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Weakly-supervised audio- visual segmentation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:27.707306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:27.707306Z digest=sha256:fc2427d60222b7728dbf20ee97aa83949e2960b3ac9347fa01ac296f961e2654

Observation 679f4081-6369-4b62-8236-66eb37155659 · outbound

This paper cites Audio-visual grouping net- work for sound localization from mixtures.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio-visual grouping net- work for sound localization from mixtures

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.647904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:27.831618Z digest=sha256:555c42912e199f5f9cf017046e98cab412b5d402733da4f1b6a01630bf0f0fea

Observation b1a01a9e-85cd-4fce-8f44-da6ae9d201d1 · outbound

This paper cites The book of why: the new science of cause and effect.

Implicit Counterfactual Learning for Audio-Visual Segmentation The book of why: the new science of cause and effect

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.501190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:27.932310Z digest=sha256:c510c1f2fff024b7695da0002646201dbb1363f627a7dbdffa0a40cd9daad967

Observation bcabcb0c-4af6-402f-ac10-adf0ba432cfd · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Implicit Counterfactual Learning for Audio-Visual Segmentation Learning transferable visual models from natural language supervi- sion

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:28.047650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:28.047650Z digest=sha256:9a5b883e5898afaede56d9bf04c2ef1040b1a69e7dd979ecb54632d1a493de0b

Observation 550e563d-bb74-48bb-8da8-98f2af648330 · outbound

This paper cites Zero-shot text-to-image generation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Zero-shot text-to-image generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:28.155249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:28.155249Z digest=sha256:0636565bc9ae20109df0007ed0c06c55d6d74116a7b176db11bb6a1bc3fea6a8

Observation a0617c11-9af7-4ca2-aed7-99a53b509b4b · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Implicit Counterfactual Learning for Audio-Visual Segmentation High-resolution image synthesis with latent diffusion models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.393079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:28.258287Z digest=sha256:ff4c2b170570fe97ce68f17c39d898e623ed643f992767e1a4e4bf3a5b041f2d

Observation 5e9251e1-5e22-4c0c-87d5-841fd30864bb · outbound

This paper cites Focal loss for dense ob- ject detection.

Implicit Counterfactual Learning for Audio-Visual Segmentation Focal loss for dense ob- ject detection

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:28.331897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:28.331897Z digest=sha256:5eb1ff7a454127e5e5a9e0cb085442f62a52b49351efb0d5f899ea319c209649

Observation af45f266-bf12-4f4b-8d4c-88b9a904e7be · outbound

This paper cites The wasserstein distance and approxi- mation theorems.

Implicit Counterfactual Learning for Audio-Visual Segmentation The wasserstein distance and approxi- mation theorems

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.249885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:28.439113Z digest=sha256:1bc189f71bcc2f0c85d4ccc4c1ef1e414950cbd047af122a6d834a20581c9dcc

Observation f7ab43a7-705a-400e-92ea-0d19b46e369c · outbound

This paper cites Better aggregation in test-time augmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Better aggregation in test-time augmentation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.097862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:28.519677Z digest=sha256:fc392fe43bb8f6ae2ee012e6fb875cbce5011f04d223937a0523c7435efae46a

Observation ffb19ff6-108c-4007-954d-feb8189ad7d3 · outbound

This paper cites A mathematical theory of commu- nication.

Implicit Counterfactual Learning for Audio-Visual Segmentation A mathematical theory of commu- nication

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.925906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:28.634597Z digest=sha256:524081768ee81bbb93357c0b4b8bf956f93ceab9f221bee6a2c7e2efe1a1f868

Observation 99535a89-fd96-4c7c-bd56-9077b0df0bef · outbound

This paper cites Looking similar sounding different: Leveraging counterfactual cross-modal pairs for audiovisual representa- tion learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Looking similar sounding different: Leveraging counterfactual cross-modal pairs for audiovisual representa- tion learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.804483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:28.704904Z digest=sha256:852ba0f8a9113a4a930c2db502a84d59b4b73b61f0e6f10d63461cd970f9ebbc

Observation 23110967-391c-4226-986c-db595b6da29e · outbound

This paper cites 3d audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation 3d audio-visual segmentation

Reference 59

Resolution
verified exact
raw_fallback, observed 2026-08-06T13:24:31.614704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:28.777772Z digest=sha256:15120b18af2e6ef60c5c39b83c5b2d040f040fda80845df669a0fe956f3cd161

Observation e2a02759-95dd-43e7-ac93-50bfdd89d177 · outbound

This paper cites Unveiling and mitigating bias in audio visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Unveiling and mitigating bias in audio visual segmentation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.702149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:28.858392Z digest=sha256:187a8adf30b483b0c92110d03a6fb14555efa6da4fd5afb5b73ebde35919ffff

Observation 15377690-589c-47c1-8b04-87f7639ab53a · outbound

This paper cites Learning audio-visual source localization via false negative aware contrastive learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Learning audio-visual source localization via false negative aware contrastive learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.572731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:28.979464Z digest=sha256:5f2ca65a89e37f45762b0f7d9eeed587163f8d618625601cf18f98f252f58cba

Observation 2515ccf1-998f-4f4d-ac57-f3a328a0e817 · outbound

This paper cites Language-guided audio-visual source separation via trimodal consistency.

Implicit Counterfactual Learning for Audio-Visual Segmentation Language-guided audio-visual source separation via trimodal consistency

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.417838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:29.061286Z digest=sha256:929f612950716ca3dc9f0309d98187142ffd7805f78ead3ecb28f87776f2eba8

Observation 0c2e86cd-6ec9-4654-a9a2-c8867e76ca0a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Implicit Counterfactual Learning for Audio-Visual Segmentation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:29.167740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:29.167740Z digest=sha256:03ff9f1a17e14fe7ff7a71de0f2ce7cc864aec50da4fa2aa5ab67a8e634c9b63

Observation 8a2ca42f-3d2f-4d93-a5d3-fb2d02c7faac · outbound

This paper cites Neural discrete representation learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Neural discrete representation learning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.232776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:29.269444Z digest=sha256:73c1b8338bc844fdd762a067913eefd1dccc0d3753f369761aa79eef1017522f

Observation f2efaf7a-148a-4976-ac2b-eb2b10eaeb84 · outbound

This paper cites Counterfactual cycle-consistent learn- ing for instruction following and generation in vision- language navigation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Counterfactual cycle-consistent learn- ing for instruction following and generation in vision- language navigation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.985658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:29.340354Z digest=sha256:19b1a85b1a954d7bbc3774f25a86ae729e774a82f97c9ba42c79ea913c98935c

Observation dfe75355-433d-4e55-8086-5c7bb61ea5b4 · outbound

This paper cites Vision-and-language naviga- tion via causal learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Vision-and-language naviga- tion via causal learning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.817968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:29.450485Z digest=sha256:dfc544cfae3891ba6c42ab31571bdeb3ccea0629e37c9cefea6170f3314b6fcc

Observation 6b0363dd-d240-4a07-b11c-99257ee012d2 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Implicit Counterfactual Learning for Audio-Visual Segmentation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:29.575827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:29.575827Z digest=sha256:f4c693c152df88b26848ee9777fbf47f213643fb592acca63e0d4b02391036bf

Observation 57326a31-7fed-419a-9e03-f04f782c88ca · outbound

This paper cites Pvt v2: Improved baselines with pyramid vision transformer.

Implicit Counterfactual Learning for Audio-Visual Segmentation Pvt v2: Improved baselines with pyramid vision transformer

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.591109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:29.651671Z digest=sha256:762961798c9726d0e4b573085be3105a5c76494b05f314cba498a98ef2f00100

Observation 2d8e8a11-5f77-4b10-8d4c-a21c080864b3 · outbound

This paper cites Drivedreamer: Towards real-world- drive world models for autonomous driving.

Implicit Counterfactual Learning for Audio-Visual Segmentation Drivedreamer: Towards real-world- drive world models for autonomous driving

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:29.738949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:29.738949Z digest=sha256:292d56573f9c208393d6c4ac791d2d5867a2bb836f6549cd78b2182e08c5cf73

Observation b00d8a91-7e04-44c4-be8e-414d4a62e7f1 · outbound

This paper cites Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer.

Implicit Counterfactual Learning for Audio-Visual Segmentation Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:31.372411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:29.839727Z digest=sha256:55b7ab21417e2c3f6f6e4f9cf494748bb3c2d2c51d7b897e8ab53cabaaadd753

Observation d84413af-577e-40c0-bc13-1a6e731270b3 · outbound

This paper cites Prompting segmentation with sound is gen- eralizable audio-visual source localizer.

Implicit Counterfactual Learning for Audio-Visual Segmentation Prompting segmentation with sound is gen- eralizable audio-visual source localizer

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.366851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:29.911767Z digest=sha256:93aabdc26adc148ece14f106fe69902a4438c5d19158a069d6c5224165202157

Observation 2e1162cf-9d17-4e68-b0cc-8e8d71258c23 · outbound

This paper cites Can textual semantics mitigate sounding object segmentation preference? In European Conference on Com- puter Vision, pages 340–356.

Implicit Counterfactual Learning for Audio-Visual Segmentation Can textual semantics mitigate sounding object segmentation preference? In European Conference on Com- puter Vision, pages 340–356

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.206530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:30.002229Z digest=sha256:31da2c19796706b0ff7115805886e0d49373fefdf8fbe4f90b05408f9492cd79

Observation 145afdec-8d79-44d4-a233-1db959720e1c · outbound

This paper cites Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.033283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:30.078314Z digest=sha256:43a6767f87813ba718d6e59e562c9210882573c7aa5e4a85c66003321fdcb170

Observation 2f93e96f-c19c-41b4-8489-c070a09d9865 · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

Implicit Counterfactual Learning for Audio-Visual Segmentation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:30.193705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:30.193705Z digest=sha256:1ccf4d4c51320280ec58fe7b1a7950e2083ee65febd70fce31a539df0628e0d8

Observation 25e4326e-a60c-4e3f-b7ee-99d4bb643f4a · outbound

This paper cites Referred by multi-modality: A unified tem- poral transformer for video object segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Referred by multi-modality: A unified tem- poral transformer for video object segmentation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.919685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:30.264676Z digest=sha256:7f14f23e1a8008fe49c715af3a970f6d6f4e37b2e690539027911c5a045e307f

Observation 5434c471-de08-4491-804c-0091f009c7c1 · outbound

This paper cites Cooperation does matter: Exploring multi-order bilateral relations for audio- visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cooperation does matter: Exploring multi-order bilateral relations for audio- visual segmentation

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.757091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:30.395792Z digest=sha256:9c140b6bcad687e718465871488ddd46a0d8f17604130a0814b863d971f59b0a

Observation c1f36da8-5497-4e1f-9186-b2268aa953a2 · outbound

This paper cites Revisiting counterfactual prob- lems in referring expression comprehension.

Implicit Counterfactual Learning for Audio-Visual Segmentation Revisiting counterfactual prob- lems in referring expression comprehension

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.617230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:30.482824Z digest=sha256:b5b0dc472e323d6020cc82c2ee69ef52df3e6aba1f70b22ebe18ee6d2ecad5d7

Observation fc20ea18-d27a-4797-a966-e1b8bd72dcee · outbound

This paper cites Losh: Long-short text joint prediction network for referring video object segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Losh: Long-short text joint prediction network for referring video object segmentation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.459830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:30.591093Z digest=sha256:ba74c53a2ee702464aad2f5803948802da6428f6d82128c52604e3554a8b103f

Observation 9c325e08-f56f-4645-b594-f7a34b8e575a · outbound

This paper cites Discovering the real association: Multimodal causal rea- soning in video question answering.

Implicit Counterfactual Learning for Audio-Visual Segmentation Discovering the real association: Multimodal causal rea- soning in video question answering

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.321674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:30.648645Z digest=sha256:2385a85ee2ecc2ea0aaf2b52a88c86ae642dad8ab0c12783bda85ddd8f2a240a

Observation de25b491-8657-4d1f-8b76-893d79779aad · outbound

This paper cites Weakly- supervised mirror detection via scribble annotations.

Implicit Counterfactual Learning for Audio-Visual Segmentation Weakly- supervised mirror detection via scribble annotations

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.147692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:30.711366Z digest=sha256:19051cabd747110a1cf9131a9d90b07550b27fd848be728b0c4c21f549460055

Observation 7410b16a-d8ab-4b01-b633-cbee76be50da · outbound

This paper cites Heterogeneous experts and hierarchical perception for un- derwater salient object detection.

Implicit Counterfactual Learning for Audio-Visual Segmentation Heterogeneous experts and hierarchical perception for un- derwater salient object detection

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.969158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:30.787086Z digest=sha256:d57dda5bb7735e52fe5a2ff2aa454f0cb1b50c846b7637416849f51e7b254d69

Observation 5310ff46-d21b-4fc5-836b-00592f7203dc · outbound

This paper cites Causal intervention for weakly- supervised semantic segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Causal intervention for weakly- supervised semantic segmentation

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.874213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:30.860023Z digest=sha256:deba6af22d2adb04748b8b0594e129da244baf0a77c53d984f4a084ec951cd57

Observation 919094cc-387d-44d6-9e58-d36771852dd4 · outbound

This paper cites Audio–visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio–visual segmentation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:30.918766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:30.918766Z digest=sha256:12651ac90450750ed715857c92a61407b5fc5b6239e09a72488c7dc505c009d4

Observation ed7a1761-fd70-4dab-bcb5-a3e28308dedc · outbound

This paper cites Audio-visual segmentation with semantics.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio-visual segmentation with semantics

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.591154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:31.109915Z digest=sha256:b408993832fc45027539988be5a48ada40dbd1cb08f1415e673e964681ff5402

Observation cc711852-876c-4aff-83cc-ffd8923a2ae3 · outbound

This paper cites Exploring pre-trained text- to-video diffusion models for referring video object segmen- tation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Exploring pre-trained text- to-video diffusion models for referring video object segmen- tation

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.424693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:31.180804Z digest=sha256:090f8233120204a81b1cb1462f825b3376de478a95434dcb459678dca26e0d53

Observation 9abe49c0-d912-40e2-9f47-04af5b5e656a · outbound

This paper cites 2, 5, 6, 8.

Implicit Counterfactual Learning for Audio-Visual Segmentation 2, 5, 6, 8

Reference 403

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.709159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:24:31.029697Z digest=sha256:c415a4d70527a0746b43e7598ea363d8246e55993bc52be0e75fb6587f12d6ad

Pith citing papers

No inbound Pith citation observations are available.