Pith. sign in

Paper Citation Record · LEDGER

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes

As of 15 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2506.01558.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01558 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:44:12.200891Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5f657d51-1ebb-49b7-a1c1-a31055a06513 · outbound

This paper cites Self-calibrated clip for training-free open-vocabulary segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Self-calibrated clip for training-free open-vocabulary segmentation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.322434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.322434Z digest=sha256:4fc35f4bff02282a445f9e10af6ff1c765110dde2f584bb52dde38a02d57078a

Observation 4106bdea-8ebf-42b7-b522-168cf3635064 · outbound

This paper cites Unraveling in- stance associations: A closer look for audio-visual segmenta- tion.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Unraveling in- stance associations: A closer look for audio-visual segmenta- tion

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.348892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.348892Z digest=sha256:eedae5d44f9be0ab4336b7d0ecdc976bf1389a9a6b9a8e65d14df486068d3c4b

Observation 28ddbb4c-1743-43ca-a927-991312d9307f · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.369535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.369535Z digest=sha256:d237b16e1fb2084b95bf156638822f02a6a2993c51e007124bf1f9fac60d43bc

Observation 9dbd80ee-966b-44cf-9bd8-158540a1080f · outbound

This paper cites Avsegformer: Audio-visual segmentation with trans- former.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Avsegformer: Audio-visual segmentation with trans- former

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.736426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:11.393623Z digest=sha256:9d03d14715d7b25a6378de7afdd6b07ea753703cab5fd6973b824d5c906d08a6

Observation 20ec1905-41b2-4500-bd1c-255311445b5a · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Audio set: An ontology and human- labeled dataset for audio events

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.699455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:11.422743Z digest=sha256:069556514d41858b3c234fe6d98a215391e242ab06f57d0d3cb49bedf2fbf3bf

Observation 36ef3bfc-767c-4f91-9234-e25c73b26324 · outbound

This paper cites Cnn archi- tectures for large-scale audio classification.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Cnn archi- tectures for large-scale audio classification

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.667367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:11.443369Z digest=sha256:915e39c8e1b1d05586e143f4a77099578eb84f9362526a8955033298536c6419

Observation 0bd5aab6-3c94-4e93-9d15-3be7573625d0 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.631580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:11.467929Z digest=sha256:6c9a98192e935b33be96ecea2d1fede0f2da664b3e9427af059c6c1f343a8c6e

Observation 112943e8-e328-4570-85bf-0f711cd5c1a1 · outbound

This paper cites Segment any- thing.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Segment any- thing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.487126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.487126Z digest=sha256:0e682d11528fc62b752547de927c1c864e520f1d387317ae41e33fb0ce77c4ef

Observation 44aa1544-d441-432f-a99f-8a2339efc465 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Lisa: Reasoning segmentation via large language model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.507306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.507306Z digest=sha256:22055fc0c159e234ba4386ced636f0cb9787a5bb36a84d2a13a792e26794768a

Observation a8e998dc-9bfe-4169-a895-9b7f0e2d6e10 · outbound

This paper cites Learning to answer questions in dynamic audio-visual scenarios.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Learning to answer questions in dynamic audio-visual scenarios

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.580940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:11.527309Z digest=sha256:240dd24b7a0d6b683b33afa59beed4463f83ee28b180b96d07e2392e1aba3a97

Observation dd7a614f-769c-40c2-a826-1465021b808b · outbound

This paper cites Robust referring video object segmentation with cyclic structural consensus.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Robust referring video object segmentation with cyclic structural consensus

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.551397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.551397Z digest=sha256:6127de74d0ced5f32f8da292db517823bf90b3f1cdbbec480f7607a94217cb0d

Observation 68ac84f6-97a6-4d9e-9619-a4d26330f4de · outbound

This paper cites Libero: Benchmarking knowl- edge transfer for lifelong robot learning.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Libero: Benchmarking knowl- edge transfer for lifelong robot learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.532561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:11.568454Z digest=sha256:e169f15aa697d6d57958aa2e1f538edac658bd89a7dfb03ff46f13471ab8fd6f

Observation b7696e5c-1259-4a90-a136-bc97a26f253b · outbound

This paper cites GRES: Gen- eralized referring expression segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes GRES: Gen- eralized referring expression segmentation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.498870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:11.584902Z digest=sha256:1a8ec96147be2f1533ddb58854d7dfa3b1de3827a2e77e471a7d85a504f7f7f9

Observation abedcc9d-ea78-471c-b193-3feea0248855 · outbound

This paper cites Improved baselines with visual instruction tuning.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Improved baselines with visual instruction tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.601043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.601043Z digest=sha256:5f5c379df9390d7fcf485f04659568d54ae414036f916696b5ad2e4ed84c9677

Observation e108d5a6-4276-4322-8943-95722e94cc8f · outbound

This paper cites Visual instruction tuning.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Visual instruction tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.621005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.621005Z digest=sha256:0b90096efd2c771c69ddfa6e2df7635bd8315bd40b158c02467c7eba658aeae8

Observation a2c1f8d2-72ed-4b08-804b-9a28108aa5a3 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.639745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.639745Z digest=sha256:d8c0b52fb9550ac3f5dd9bce060eb015adfe002d233cfc22d3f7ecd8b87066bb

Observation e1580b59-e6a2-4b27-adb5-7382a9c7bf06 · outbound

This paper cites Open-vocabulary segmentation with semantic-assisted calibration.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Open-vocabulary segmentation with semantic-assisted calibration

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.446323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:11.660057Z digest=sha256:1bb6217e203c055b5a21c429a7600e18caf812504c6b0de29a0a212208dd08fd

Observation d3be685d-3edd-4cb2-a54e-cce2bfe266c4 · outbound

This paper cites Universal segmentation at arbi- trary granularity with language instruction.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Universal segmentation at arbi- trary granularity with language instruction

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.399232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:11.677555Z digest=sha256:4ab474f289bf2ff21f6de990a62abc21566d1f204bd063a9f4ee5bbfc6a270c1

Observation f94e348d-ca8a-47fc-ace8-4688e635702d · outbound

This paper cites Decoupled weight de- cay regularization.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Decoupled weight de- cay regularization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.704295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.704295Z digest=sha256:34a4e4ef211c0428a799575b9b953e57a62ec40c53fbaa6dee80bbe6afbde797

Observation 09e0cbf6-df4a-49e9-94f7-25215f04a837 · outbound

This paper cites Soc: Semantic-assisted object cluster for referring video object segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Soc: Semantic-assisted object cluster for referring video object segmentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.354324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:11.721300Z digest=sha256:6055e0442e4b595194c519761094db132838b4cdbb0786517dcca2c2eaba5a5a

Observation 16f510cc-11dc-4c5a-a067-d6c1aef562ba · outbound

This paper cites Mod- eling context between objects for referring expression un- derstanding.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Mod- eling context between objects for referring expression un- derstanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.313045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:11.738165Z digest=sha256:070ee02ecd7845dd65278546fb79a81addbd4721211beb5a93f1909c059223cb

Observation b645e32b-6c58-4403-94c1-1563de57e8aa · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes SAM 2: Segment Anything in Images and Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.756990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.756990Z digest=sha256:1ac5097651624f8625a6ec4e83aeb3171978a5be6a1094e06e3611c95208eb38

Observation 247f7548-6d8e-430a-9382-b699f0dee7a6 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Pixellm: Pixel reasoning with large multimodal model

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.269637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:11.778749Z digest=sha256:c3d7b5f592881d34c00c954ad398f5bdff2d0bca016323e37d68d72e985dccd6

Observation e51c9301-f233-417f-95a4-7afa4d1929a7 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.802462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.802462Z digest=sha256:48fce0e5e679b447e67d30a2106915f471852fc6916efc56a882a5b8b28a5e2f

Observation 3028eb4f-1c46-4812-92eb-c73560c00a93 · outbound

This paper cites Efficient attention: Attention with lin- ear complexities.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Efficient attention: Attention with lin- ear complexities

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.225927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:11.832827Z digest=sha256:e0fac4aa7165bb584605a7167c7921338e9a382672d32c8880321561f27db7d0

Observation 5878c5f0-72c0-41ed-a211-55eac3e2713c · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Roformer: Enhanced transformer with rotary position embedding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.848405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.848405Z digest=sha256:f480b15db9134300fa3f74201f9ed5a6a0f261f445e89bafad0b166c8469ddc0

Observation bc52a43d-4679-4847-ad56-7721d435448c · outbound

This paper cites Auto- acd: A large-scale dataset for audio-language representation learning.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Auto- acd: A large-scale dataset for audio-language representation learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.182438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:11.871678Z digest=sha256:ecf748fc1b20c1e1f49aedfbde12c05098a585d9885c774c7b9174f787e6163d

Observation f31b91df-7213-47e0-aca7-bd32ddac1e79 · outbound

This paper cites Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.892307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.892307Z digest=sha256:c19802ef4c4378c721b1e624b81632ba138780cdfb573d7173657ede5c29b926

Observation 64f558e1-147b-468f-878c-483ff6fa7bbe · outbound

This paper cites Efficient remote sensing transformer for coastline detection with sentinel-2 satellite imagery.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Efficient remote sensing transformer for coastline detection with sentinel-2 satellite imagery

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.151753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:11.907122Z digest=sha256:b63b260d02b67929347631f5adeaaafc2534a6ee74d05566037eaece7d9ebf39

Observation 58ab5cbc-bd89-45e4-b97f-f47dc87a431e · outbound

This paper cites Prompting segmentation with sound is gen- eralizable audio-visual source localizer.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Prompting segmentation with sound is gen- eralizable audio-visual source localizer

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.117916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:11.924209Z digest=sha256:91d1d3ce3c8d884526a7573ee9baafb1e340217b1be1877cd5aa77b983143512

Observation 2e111730-5219-4b15-8ebf-ad71dd6e9371 · outbound

This paper cites Ref-avs: Refer and segment objects in audio-visual scenes.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Ref-avs: Refer and segment objects in audio-visual scenes

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.060729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:11.945322Z digest=sha256:e81fd7ed35f9a2f501cc40b4993d54cd62f72743f21aeca8545a8b6ecdbea524

Observation 493dc4ad-5486-4df8-804b-0238b2ed74bc · outbound

This paper cites Convolution meets trans- former: Efficient hybrid transformer for semantic segmenta- tion with very high resolution imagery.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Convolution meets trans- former: Efficient hybrid transformer for semantic segmenta- tion with very high resolution imagery

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.026308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:11.959188Z digest=sha256:8b936f5c1727f05638d1c200592ad9e72b56c9a7b8f8b15b02171e7d72fe7414

Observation 5cb92975-e83f-4199-a288-d262b10ca69c · outbound

This paper cites IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word Emphasis.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word Emphasis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.980254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.980254Z digest=sha256:b791da66ed2803b6a1b1b9ccc2cc6018015e597fdef6b4ecd517f60594ca5ac9

Observation 4138baf4-d4f3-423a-bb44-b3b7c22d4c7c · outbound

This paper cites Onlinerefer: A simple online baseline for referring video object segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Onlinerefer: A simple online baseline for referring video object segmentation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.997831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.997831Z digest=sha256:5d60d6b50e9c400c590cb7692c20bcd71119375b2a80b2fbfb6b6202d6457911

Observation b4aace26-db3c-4949-8b24-73b9f780121b · outbound

This paper cites Language as queries for referring video object seg- mentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Language as queries for referring video object seg- mentation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.014816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.014816Z digest=sha256:36d155068b10ee5687f85b889859dc5464ce595227eea081f57e956702fec6a3

Observation 0882748c-073e-47fa-aa91-48331e4454a4 · outbound

This paper cites Language as queries for referring video object seg- mentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Language as queries for referring video object seg- mentation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:13.972399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:12.032542Z digest=sha256:a6d676685dc9759f141a2dd6647ef93eb57580f6a28a4371c17156938d979c50

Observation 4d8b0770-911d-4830-955e-49fbd5ba5748 · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Gsva: Generalized segmentation via multimodal large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:13.935682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:12.045116Z digest=sha256:7f624fd5820da1fdde091966a9e89f9d11fb19c49ef470b6cd66580f48076016

Observation d7575162-9125-4f72-a778-1645e4c6b5a8 · outbound

This paper cites Efficientsam: Leveraged masked image pretraining for efficient segment anything.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Efficientsam: Leveraged masked image pretraining for efficient segment anything

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.061070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.061070Z digest=sha256:195c9cb7bc1f21d770c4efa10979a5b0ea0b9d61c78fab9792c36bb9a9a35d33

Observation 338baf62-16e1-439a-bd25-27965cef8bea · outbound

This paper cites Avqa: A dataset for audio- visual question answering on videos.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Avqa: A dataset for audio- visual question answering on videos

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:13.878090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:12.078141Z digest=sha256:652f37a4b5038dfc93f03117f772ed0051b94e51a6b2460ef0e481a6d4b5c9cf

Observation 9556304f-7048-4819-9f84-f36361c06f50 · outbound

This paper cites Lavt: Language-aware vision transformer for referring image segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Lavt: Language-aware vision transformer for referring image segmentation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:12.578490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:12.092128Z digest=sha256:0dfc762f398fa0196e43c6f238023ceb1841cb2feb70011ba9c9277a3677c785

Observation 16a50437-db4c-45f7-9e15-127705cd219f · outbound

This paper cites Language- aware vision transformer for referring segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Language- aware vision transformer for referring segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:12.541709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:12.110003Z digest=sha256:647ba191cb8b8f42313a4697507d596c55775e357af7563a9f93bd9a3f05ebb5

Observation fd632782-eadc-4606-8d53-4ee658adf5c9 · outbound

This paper cites Faster Segment Anything: Towards Lightweight SAM for Mobile Applications.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Faster Segment Anything: Towards Lightweight SAM for Mobile Applications

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.122600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.122600Z digest=sha256:54bd893767730443fcdc40477f3f09e2973c422c083e1dd6358c75d70953d698

Observation d7f2f6e5-a0fd-4d05-91d9-d1a3c3424bb8 · outbound

This paper cites EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.137620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.137620Z digest=sha256:91cae8eab09553e363cd15d1744c38e372e35f2ca2c20da4789cb925dc7ec7fe

Observation 6d85d712-5c57-410d-8660-822dd4315a3c · outbound

This paper cites Fast Segment Anything.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Fast Segment Anything

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.151305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.151305Z digest=sha256:4c0164825fbdaf2f5c0673e892fdb7dca363db1b11936a9775886e54b8c8e413

Observation 98aa6ac5-3a2e-4830-9d09-ed2199079c42 · outbound

This paper cites Audio-visual segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Audio-visual segmentation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:12.500281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:44:12.170613Z digest=sha256:175f211db4234df6ed4c8dffe8a77de16bb95671dc16ea55bcc8985441e3658a

Observation 3321ace5-26c3-49e0-8a45-7ab64cb4f444 · outbound

This paper cites Audio-Visual Segmentation with Semantics.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Audio-Visual Segmentation with Semantics

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.186738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.186738Z digest=sha256:c61cfc352376cbd6c564c8b1e7cd435febd32a65e6e95f996831df60418f78c4

Observation 518e8f9d-885c-4021-83d6-3e19bb0a794f · outbound

This paper cites Generalized decoding for pixel, image, and lan- guage.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Generalized decoding for pixel, image, and lan- guage

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.200891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.200891Z digest=sha256:e9d1ea4880e0ecd14d986b9e8d57265649751ea72953afc68903b7554508e9c6

Pith citing papers

No inbound Pith citation observations are available.