Pith. sign in

Paper Citation Record · LEDGER

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference

As of 17 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2411.15851.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15851 v2

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:53:33.608671Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact3
  • verified fuzzy37
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b298baa7-e8e4-4466-8c33-215c03636890 · outbound

This paper cites Self- supervised multimodal versatile networks.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Self- supervised multimodal versatile networks

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:36.259398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:32.563359Z digest=sha256:2a665503b611578efa8a1f9178671fc51491e07b59e4bc2068f736f9331ca337

Observation f2b0ec18-3fad-4177-aa0e-b39aae75fd0c · outbound

This paper cites Vqa: Visual question answering.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Vqa: Visual question answering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:36.192976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:32.740385Z digest=sha256:149404ef79353ec3915604e6cf9d42aaeee5abc754657eeccee93bcc31c10a97

Observation bcfe492a-4536-4a33-9dc6-109cb470c38c · outbound

This paper cites Single-stage semantic segmentation from image labels.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Single-stage semantic segmentation from image labels

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:36.085633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:32.809661Z digest=sha256:c202f628ba26a0abebbc420a6576c816723a58af78ee3870b096d0bd9ed31023

Observation 78d05ed0-9048-44e8-b94c-765f9f6bf9f4 · outbound

This paper cites Fossil: Free open-vocabulary semantic seg- mentation through synthetic references retrieval.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Fossil: Free open-vocabulary semantic seg- mentation through synthetic references retrieval

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:36.070007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:32.817434Z digest=sha256:632754816eb7e530ba9af00d8765dcb4332046e19e3985f506e09a9d0ab5ff56

Observation b64f6d17-a166-48fc-9519-e760de86b75a · outbound

This paper cites Grounding everything: Emerging localiza- tion properties in vision-language transformers.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Grounding everything: Emerging localiza- tion properties in vision-language transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:36.052570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:32.826968Z digest=sha256:fd949d37d5cbaa1b36796c7a2d08e4424dc841b66bce65493313efd57704726b

Observation 1d9a4e0e-d4e1-4ef4-b8f2-087d251ad34b · outbound

This paper cites Language Models are Few-Shot Learners.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Language Models are Few-Shot Learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:32.834829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:32.834829Z digest=sha256:f27e5b33d419a633b9e77fa81ca32bd6cd11516ba333532ae9b5e4550be33721

Observation 39fcef99-4425-417e-a316-34af1ae24a62 · outbound

This paper cites Coco- stuff: Thing and stuff classes in context.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Coco- stuff: Thing and stuff classes in context

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:36.037620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:32.839966Z digest=sha256:2849096c52fad8db9a02e6dce4b936843ea2a6c6648204193f4242d8cfedb3d4

Observation e6cb987b-2747-43c4-b9fc-f252f3e88bd2 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Emerg- ing properties in self-supervised vision transformers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:36.022720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:32.844364Z digest=sha256:5a7aff94d2234fecb07cf8774f90c71e189b8a448943a6023f9c71c46318a492

Observation 720e2a45-2f21-4802-bd70-94c76f7817f6 · outbound

This paper cites Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.909851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:32.849187Z digest=sha256:303d4ca08b341cc3e908e2901897e5baf9e04186ec9447ea1e8c0c0fdd3fdd92

Observation 59add72b-cbed-41f5-9e12-e55cbe5d4cfc · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Reproducible scal- ing laws for contrastive language-image learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.893379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:32.854083Z digest=sha256:d043de1158494b98b7773d5c9b8d625b30aedc81f86fc25afb0951438dfad910

Observation 6bcddcd9-c84b-4a20-b53c-8ccdaa8da91f · outbound

This paper cites Mmsegmentation: Open- mmlab semantic segmentation toolbox and benchmark,.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Mmsegmentation: Open- mmlab semantic segmentation toolbox and benchmark,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:32.858803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:32.858803Z digest=sha256:5bfa228d5982b9b0a17f43e4b648c51725c2a2be918b70d8d622382bb0268788

Observation 05dd4e8e-6feb-4753-a520-e088b664eade · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference The cityscapes dataset for semantic urban scene understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.867348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:32.863655Z digest=sha256:e9ac5699a16ea0942b3ca8a583bec96275118a88b548be828e6f5dee4ebb177a

Observation 65821c0f-2527-4333-a5a6-b022c59676b1 · outbound

This paper cites Vision Transformers Need Registers.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Vision Transformers Need Registers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:32.868403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:32.868403Z digest=sha256:07411b38fbcbe6ee6328bbd6659c78e3e2c7e72571d0057ab0f505c19c6a986f

Observation e2e783d8-34df-405e-9b0f-a595dc122f2f · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:32.873421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:32.873421Z digest=sha256:6c12593839a3272c5a8b6ecc83b75768d88fe2d6bc21b4c486389ed413a2c202

Observation 5ce91ed8-829d-4a69-ae6b-be0bb16edddf · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:32.878221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:32.878221Z digest=sha256:39a7301cc92f70830d067d303f795e7beaa7e023c5e9b66c536c137cd0bfd862

Observation f4a721d3-574e-4cfa-98fa-d8452bc59ba2 · outbound

This paper cites The pascal visual object classes challenge: A retrospective.IJCV, 111:98–136, 2015.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference The pascal visual object classes challenge: A retrospective.IJCV, 111:98–136, 2015

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.847718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:32.882936Z digest=sha256:ca1cc1c7a2d29641a17249f66c7db1c5923427acadb95214b93cb7b50656b516

Observation 9d6152d3-944f-4766-ba2c-87ff0c6d2b23 · outbound

This paper cites Improved baselines for vision-language pre-training.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Improved baselines for vision-language pre-training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:32.887371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:32.887371Z digest=sha256:2b703d7be370a53c415511261c28743c375900afb1406fe7dfa74e7b637d3ae4

Observation 8f14eeba-163f-4c60-9780-699a1d138dcb · outbound

This paper cites Pay attention to your neighbours: Training-free open-vocabulary semantic segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Pay attention to your neighbours: Training-free open-vocabulary semantic segmentation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.635478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:32.893255Z digest=sha256:05c892bf00359a6343f4a4788ba77da804c0aa8a5a58427d9fca003b8d2beb7a

Observation 6121bcd2-b035-490e-94e8-23a5962e7320 · outbound

This paper cites Open-vocabulary semantic segmentation with decou- pled one-pass network.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Open-vocabulary semantic segmentation with decou- pled one-pass network

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.575648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:32.898117Z digest=sha256:06bbc99170c7818f2d27e19ad8c3ac28f5c01f72bc01b8a46d55ed01b5b95547

Observation e3a8ae2b-cd0f-4067-b4fe-e29cfeb11fa7 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Scaling up visual and vision-language representation learning with noisy text supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:32.953042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:32.953042Z digest=sha256:2f1a851012cb5d2a6854d1bb1f7588af8ecf94db5573f9d41ed617484c3940c3

Observation 5331fbc8-5e34-4f7d-83ea-1e3387bc8fe5 · outbound

This paper cites Learning mask-aware clip representations for zero-shot segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Learning mask-aware clip representations for zero-shot segmentation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.013103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.013103Z digest=sha256:123c7a5c8cbf164a12a87dba481eeb24809a1560d0eb9c17f057a0cece21e82e

Observation 82495263-bf02-403d-aded-aefa128c51e0 · outbound

This paper cites In Defense of Lazy Visual Grounding for Open-Vocabulary Semantic Segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference In Defense of Lazy Visual Grounding for Open-Vocabulary Semantic Segmentation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-12T13:53:33.969268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.113821Z digest=sha256:816ed2d0221e4c2b0fd7e97c22b21a77fc2eb190594daf3c701ab2e9af1436eb

Observation a035c38d-50ba-4429-b6a1-7e054f3c9525 · outbound

This paper cites Weakly supervised ground- ing for vqa in vision-language transformers.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Weakly supervised ground- ing for vqa in vision-language transformers

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.537853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.211579Z digest=sha256:e3bd3fb068c283f21f79f042cee50d7283d600dc3a916f64aa822e13eea4c721

Observation 05a8c5c2-5c51-425b-a53f-b0a8f8c728fe · outbound

This paper cites Vilt: Vision- and-language transformer without convolution or region su- pervision.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Vilt: Vision- and-language transformer without convolution or region su- pervision

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.522030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.219038Z digest=sha256:44ce19a001b9cf9039113c6d9ed0b77850caee3b2271e6d87dddbda891dc5f47

Observation 3c09e933-e46d-4822-bb77-c8a69a614e8c · outbound

This paper cites Segment any- thing.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Segment any- thing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.503755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.223688Z digest=sha256:4b580b35b1d946c1693f3ce1482c1fe7e2bc3b5bbda1a384e46595a18215e2ea

Observation f277c252-8644-46ae-903d-4dcefa7a4774 · outbound

This paper cites Efficient infer- ence in fully connected crfs with gaussian edge potentials.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Efficient infer- ence in fully connected crfs with gaussian edge potentials

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.449911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.228363Z digest=sha256:b9fe63e74255adfa0f968f738bf2915350b452f4710d618ef434d768b301e359

Observation dc6e9316-0db8-46c5-9b96-4ecbdc0d28b7 · outbound

This paper cites Clearclip: Decom- posing clip representations for dense vision-language infer- ence.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Clearclip: Decom- posing clip representations for dense vision-language infer- ence

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.235150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.232981Z digest=sha256:4d510d30635ae3103d5a91f189495a872c28ee5023e6df559931c248aa749475

Observation 2be1f5cc-aa44-4f78-8645-ace1b37d7181 · outbound

This paper cites ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.237441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.237441Z digest=sha256:01be67cf661d929888947603452dfb1ba8e9bde3d15eb967d50d9ea12973b964

Observation ab0901a4-81c4-4ba2-9f83-76a579a32092 · outbound

This paper cites Align before fuse: Vision and language representation learn- ing with momentum distillation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Align before fuse: Vision and language representation learn- ing with momentum distillation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.242396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.242396Z digest=sha256:6bb998da71985a22e68517bfb18ade7d156ea50056b6b5b66ef06157fc36e026

Observation 23839c05-77b6-4ef0-bb42-87eacf6e4b2f · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.247066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.247066Z digest=sha256:37ce44fce3d425bab15b3da676d7b4ad8a257e89ab931494b9359e5c8e8a6741

Observation 052452b5-df35-4041-887e-eb0f2d1f471c · outbound

This paper cites A Closer Look at the Explainability of Contrastive Language-Image Pre-training.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference A Closer Look at the Explainability of Contrastive Language-Image Pre-training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.251700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.251700Z digest=sha256:bd56aa05ad5bc67c8246faa37a829b5a7924bbbaac12abb6f1e2e51083d0a564

Observation 99d91e53-7b94-45bf-ab0a-1f3a77f3cb69 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Open-vocabulary semantic segmentation with mask-adapted clip

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.256226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.256226Z digest=sha256:ddaf68c19b93bcc0b0737efe4e8e8da7aeb720d1a9470e4ca7dd4b0d4bb68d27

Observation 9dc0fa3e-36b6-492c-b91d-952f29390bb3 · outbound

This paper cites Segclip: Patch aggregation with learn- able centers for open-vocabulary semantic segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Segclip: Patch aggregation with learn- able centers for open-vocabulary semantic segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.976103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.261076Z digest=sha256:6d4666e897a81a3a8e5a4c17b306c65ec44cd0fe3592b5b29b92b76c463352fc

Observation a080340c-8dfb-4982-916d-f4c1bb5675d5 · outbound

This paper cites End-to-end learning of visual representations from uncurated instruc- tional videos.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference End-to-end learning of visual representations from uncurated instruc- tional videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.266020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.266020Z digest=sha256:d9563fdb3fdd35c673df7232fd3e6c6fde6555e8429cb3ad80cc169173dfb61e

Observation 07a32e02-d329-4720-bf45-1be978ca6b42 · outbound

This paper cites The role of context for object detection and se- mantic segmentation in the wild.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference The role of context for object detection and se- mantic segmentation in the wild

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.271239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.271239Z digest=sha256:13db43ac652d1c5ac6e0efe181495422b700f2db4acb1beca4ed39d9d1549149

Observation c78933b3-c3bd-4d27-9c5b-a2b6b168bd99 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference DINOv2: Learning Robust Visual Features without Supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.276041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.276041Z digest=sha256:95020dba268c191ad036ebb485dcbb129a48d2203ed131f354d50b3c7f5f3f67

Observation 2ee3dd2a-8a54-48d9-92b7-9b8123489458 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Learn- ing transferable visual models from natural language super- vision

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.851135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.280808Z digest=sha256:750caa4a31e7473459138f74bab40cd6fea9662ae9c0f59c68944b7d14af8454

Observation d61ff42f-ccf5-4c68-9b8c-c3c114d007eb · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.703300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.285175Z digest=sha256:a376a6976b8f17fee9c9f9a840103889877b52fac37dc000b67c4eb7bd6ca1fc

Observation 0a9cb8d4-7eba-4a1b-86c3-5e25ecd96a1c · outbound

This paper cites Denseclip: Language-guided dense prediction with context- aware prompting.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Denseclip: Language-guided dense prediction with context- aware prompting

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.562702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.289597Z digest=sha256:4057a50a9cb56f2fe6e74be18ae936f6f32ecca6f1016891a5c7674d3f916318

Observation 4940a514-f9f0-47a2-933e-1f529ca59c0d · outbound

This paper cites ViewCo: Discovering Text-Supervised Segmentation Masks via Multi-View Semantic Consistency.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference ViewCo: Discovering Text-Supervised Segmentation Masks via Multi-View Semantic Consistency

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.294063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.294063Z digest=sha256:6503340385f521d23f2e4ef6e0888c174c4274972cb483729628bb5598ff7d51

Observation 1fe42a55-416f-4c26-b78c-4bd0eca49955 · outbound

This paper cites Ex- plore the potential of clip for training-free open vocabulary semantic segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Ex- plore the potential of clip for training-free open vocabulary semantic segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.547330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.319556Z digest=sha256:7923c0b652f903176f12d87522f494551e39d2e044892ea13b1ed52d4bccba81

Observation b25e8ae9-8542-4ee1-b4a2-370b2ea732c9 · outbound

This paper cites Reco: Re- trieve and co-segment for zero-shot transfer.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Reco: Re- trieve and co-segment for zero-shot transfer

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.533047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.354188Z digest=sha256:d1a206d7385e3bc63de0d4cbaf11f8f56fb9feddcba6455ad0006269b805ff10

Observation 3ed885df-0962-4438-9671-6fc7714cc06f · outbound

This paper cites Clip as rnn: Segment countless visual concepts without training endeavor.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Clip as rnn: Segment countless visual concepts without training endeavor

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.517041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.404055Z digest=sha256:e9c75bc0cb01c4d15dffa77e3b15b489f76879adbb6369e725ff31b92049b8c6

Observation bde8b7f7-9db7-424d-adab-d4f7a237a95d · outbound

This paper cites Learning to Decompose Visual Features with Latent Textual Prompts.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Learning to Decompose Visual Features with Latent Textual Prompts

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-12T13:53:33.862640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.444207Z digest=sha256:f421a8e3eb1b3fde1d1037288c6a184e2b0e68c48ddbe0f5e7d5c03cbfefa71a

Observation 220c7c8f-da1e-47f2-8467-dfd54cb65217 · outbound

This paper cites Sclip: Rethinking self-attention for dense vision-language inference.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Sclip: Rethinking self-attention for dense vision-language inference

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.502054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.524250Z digest=sha256:85aaeafb21ba066f85ad304a39bcc765e43100f7cbe46b903f1b6c493b687ab6

Observation 6bde215f-8dd1-4f6f-90a9-d397ab646c45 · outbound

This paper cites Medklip: Medical knowledge enhanced language-image pre-training for x-ray diagnosis.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Medklip: Medical knowledge enhanced language-image pre-training for x-ray diagnosis

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.461061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.529413Z digest=sha256:8af67b92de962beeb7e74cb8f7242978996df04f9025b85703adff73a81adf2e

Observation 15686629-73c2-46e6-8c92-3aebaa9b02bd · outbound

This paper cites Rewrite caption semantics: Bridging se- mantic gaps for language-supervised semantic segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Rewrite caption semantics: Bridging se- mantic gaps for language-supervised semantic segmentation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.391322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.534204Z digest=sha256:c1bbaeb12ad8190d5edd153cc237f230b718f6c6d99f38d0e84d4c28c205becd

Observation 577d6387-f93b-470b-b003-397677aea621 · outbound

This paper cites Demystifying CLIP Data.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Demystifying CLIP Data

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.538321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.538321Z digest=sha256:56e2c3c02c0d9c7cf01fe7aa69d95ae0072ba5163fbc8b515c993ab8ff440057

Observation ee7fde15-8078-4e60-a3d6-3542e81401fa · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Groupvit: Semantic segmentation emerges from text supervision

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.298390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.542913Z digest=sha256:4485633882fb5e33049028dd529493a94b2f9aab41dcb66843888d313c8465c2

Observation 033d3e8f-faff-4b8b-8e78-1df6548ae8f3 · outbound

This paper cites Learning open-vocabulary semantic segmentation models from natural language supervision.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Learning open-vocabulary semantic segmentation models from natural language supervision

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.230098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.547629Z digest=sha256:52bd8cb798cea0155b61b9d7d4ebd20f90a5dbbbb1ed3d46154803bd24a74dfb

Observation dfb67b0f-d00c-425c-b1bf-513888743bcd · outbound

This paper cites Side adapter network for open-vocabulary semantic segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Side adapter network for open-vocabulary semantic segmentation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.162876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.552385Z digest=sha256:b3358e00f74516db26e3044423f929ccc38d0defff2e6f8110271c13f207a34d

Observation 2ae95653-c54e-4ad2-b70c-6256400a508f · outbound

This paper cites Unified contrastive learning in image-text-label space.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Unified contrastive learning in image-text-label space

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.148172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.557229Z digest=sha256:586eedfe0b0855d36106f0be923117c14e84f3d4be180ce1105be8357c557104

Observation 5bcb9ca7-fc93-449f-96aa-e5003c30bcec · outbound

This paper cites Tuning-free Universally-Supervised Semantic Segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Tuning-free Universally-Supervised Semantic Segmentation

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-12T13:53:33.773603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.561686Z digest=sha256:9000c9921738565f8bf5f5f3e5f46bf7d00982914410503d51053a7ea470553c

Observation 3964313e-b3b3-41aa-a326-138ad3d039dd · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.566440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.566440Z digest=sha256:75686d2e5f648b1fc4df508fd8d61e08e0a5c231a3d75b02694062c0dbb4f82a

Observation 3d48020e-1367-4386-8a20-804609d647b4 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.571272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.571272Z digest=sha256:5bd76dc5ff7874f340eb0812df25478d56586ca4c9e69f591df734bae70df054

Observation effb5a11-af05-47e4-95f2-dfb9f63b7af0 · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Florence: A New Foundation Model for Computer Vision

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.575798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.575798Z digest=sha256:fa0ea253f85392bd9f8af53ce80b10ef552bb11410578a8c3313737cead35dd7

Observation 0c08ad36-e9e0-4cb8-bd6f-3ba52b74a209 · outbound

This paper cites Uncovering prototypical knowledge for weakly open- vocabulary semantic segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Uncovering prototypical knowledge for weakly open- vocabulary semantic segmentation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.133288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.580214Z digest=sha256:529e4b0b1db3e934ac8f664794850a7484707aab46db24429354f10581e6144b

Observation 074d1995-6a2f-406a-80fd-9932ce131e50 · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Semantic under- standing of scenes through the ade20k dataset

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.117946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.584897Z digest=sha256:22be13084ddc74505704732a8485851610c1cc06def1d75fe9d6a3f209eae4dd

Observation 76a4c42b-17b4-408e-a6ce-2c1a25984d72 · outbound

This paper cites Extract free dense labels from clip.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Extract free dense labels from clip

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.589293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.589293Z digest=sha256:f8ea9635c700992db6cf9c0f22602c169beae5213f5e93618d9d6612ae1ef0e9

Observation cd9440bd-40a5-4666-becc-c6a063c32c23 · outbound

This paper cites Image Segmentation in Foundation Model Era: A Survey.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Image Segmentation in Foundation Model Era: A Survey

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.593862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.593862Z digest=sha256:20e37e27aeb851dc42c60b02e2bb50fd52fd5a0ceb8ca319352a43117f608887

Observation 1ce81bfe-7811-4a4c-9db3-99703333e1cf · outbound

This paper cites Zegclip: Towards adapting clip for zero-shot se- mantic segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Zegclip: Towards adapting clip for zero-shot se- mantic segmentation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.092536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.598810Z digest=sha256:3f67b7ba988044b19ad1d71e6978f2cbf0eb0d7db8e63a52d5f48ee15a614ffe

Observation 355102fb-2c94-4bb2-9db8-6887d5fa32f5 · outbound

This paper cites Segment everything everywhere all at once.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Segment everything everywhere all at once

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.076309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.603563Z digest=sha256:610070079a7cb802b8a77e69329dfd7289abd5646a90a53cde77a9c4d51d09f7

Observation 5238c073-aadf-411b-b77b-57a069dc7835 · outbound

This paper cites For example, in the COCO Ob- ject dataset (see Fig.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference For example, in the COCO Ob- ject dataset (see Fig

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.061494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T13:53:33.608671Z digest=sha256:66c1edc4f0cf6eb8e502f32a765c95661d4a9179532692fa3aba6091088ea18f

Pith citing papers

No inbound Pith citation observations are available.