Pith. sign in

Paper Citation Record · LEDGER

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference

As of 16 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2411.15851.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15851 v2

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:53:33.608671Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact3
  • verified fuzzy37
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b298baa7-e8e4-4466-8c33-215c03636890 · outbound

This paper cites Self- supervised multimodal versatile networks.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Self- supervised multimodal versatile networks

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:36.259398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:32.563359Z digest=sha256:b8c662bbb8a4efeabe5bbe20f39587cd5e6d778f3ca95cca28377f99f9243540

Observation f2b0ec18-3fad-4177-aa0e-b39aae75fd0c · outbound

This paper cites Vqa: Visual question answering.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Vqa: Visual question answering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:36.192976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:32.740385Z digest=sha256:f17c9fac665376834d929b0a14a38bc90bdebb7b2fb09fb563704dcd6b3f565f

Observation bcfe492a-4536-4a33-9dc6-109cb470c38c · outbound

This paper cites Single-stage semantic segmentation from image labels.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Single-stage semantic segmentation from image labels

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:36.085633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:32.809661Z digest=sha256:b71003fb9768612ef05e579032528df57382b2cb882719504b4efb387c5a357b

Observation 78d05ed0-9048-44e8-b94c-765f9f6bf9f4 · outbound

This paper cites Fossil: Free open-vocabulary semantic seg- mentation through synthetic references retrieval.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Fossil: Free open-vocabulary semantic seg- mentation through synthetic references retrieval

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:36.070007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:32.817434Z digest=sha256:4e44583486b3f6cd9dc96565cb2ba507af67bcaceb761b03ca7311a61ecb3835

Observation b64f6d17-a166-48fc-9519-e760de86b75a · outbound

This paper cites Grounding everything: Emerging localiza- tion properties in vision-language transformers.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Grounding everything: Emerging localiza- tion properties in vision-language transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:36.052570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:32.826968Z digest=sha256:22adac25d963cae6a519481860969acd923a97ee87acf472527681ba2f193dc3

Observation 1d9a4e0e-d4e1-4ef4-b8f2-087d251ad34b · outbound

This paper cites Language Models are Few-Shot Learners.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Language Models are Few-Shot Learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:32.834829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:32.834829Z digest=sha256:f27e5b33d419a633b9e77fa81ca32bd6cd11516ba333532ae9b5e4550be33721

Observation 39fcef99-4425-417e-a316-34af1ae24a62 · outbound

This paper cites Coco- stuff: Thing and stuff classes in context.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Coco- stuff: Thing and stuff classes in context

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:36.037620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:32.839966Z digest=sha256:b934910f2bb3e2bbe40544810bd4fe45034413a3645ff85ead921924be7c462f

Observation e6cb987b-2747-43c4-b9fc-f252f3e88bd2 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Emerg- ing properties in self-supervised vision transformers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:36.022720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:32.844364Z digest=sha256:5f730abe20e9e783ed1fb993e3f05a268cb44b3eb97d4a3a74d609b1f6b70753

Observation 720e2a45-2f21-4802-bd70-94c76f7817f6 · outbound

This paper cites Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.909851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:32.849187Z digest=sha256:8ba132ff7121ea3c55ecbb5b7491689a98afe372b5399cf6b288a07828193a41

Observation 59add72b-cbed-41f5-9e12-e55cbe5d4cfc · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Reproducible scal- ing laws for contrastive language-image learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.893379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:32.854083Z digest=sha256:16c77e86a52b33030f9e7018306fc7ce8f718a3970b6fb530617181f24631f0e

Observation 6bcddcd9-c84b-4a20-b53c-8ccdaa8da91f · outbound

This paper cites Mmsegmentation: Open- mmlab semantic segmentation toolbox and benchmark,.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Mmsegmentation: Open- mmlab semantic segmentation toolbox and benchmark,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:32.858803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:32.858803Z digest=sha256:5bfa228d5982b9b0a17f43e4b648c51725c2a2be918b70d8d622382bb0268788

Observation 05dd4e8e-6feb-4753-a520-e088b664eade · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference The cityscapes dataset for semantic urban scene understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.867348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:32.863655Z digest=sha256:39b57eb5ca9e8f2778176b4926836e8bd990c0f3c5c8e640c6eb3b498db205b7

Observation 65821c0f-2527-4333-a5a6-b022c59676b1 · outbound

This paper cites Vision Transformers Need Registers.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Vision Transformers Need Registers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:32.868403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:32.868403Z digest=sha256:07411b38fbcbe6ee6328bbd6659c78e3e2c7e72571d0057ab0f505c19c6a986f

Observation e2e783d8-34df-405e-9b0f-a595dc122f2f · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:32.873421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:32.873421Z digest=sha256:6c12593839a3272c5a8b6ecc83b75768d88fe2d6bc21b4c486389ed413a2c202

Observation 5ce91ed8-829d-4a69-ae6b-be0bb16edddf · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:32.878221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:32.878221Z digest=sha256:26429453da508e92974f41399cd4d3160d6e2545ac989dcde8fdfae5bf835eaf

Observation f4a721d3-574e-4cfa-98fa-d8452bc59ba2 · outbound

This paper cites The pascal visual object classes challenge: A retrospective.IJCV, 111:98–136, 2015.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference The pascal visual object classes challenge: A retrospective.IJCV, 111:98–136, 2015

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.847718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:32.882936Z digest=sha256:ece46dbf54d2cef1874b7f6b40e04e74ba30387d636e8a68ea27056e3c6f9080

Observation 9d6152d3-944f-4766-ba2c-87ff0c6d2b23 · outbound

This paper cites Improved baselines for vision-language pre-training.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Improved baselines for vision-language pre-training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:32.887371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:32.887371Z digest=sha256:dc908a8225332fefa40b09fd8c910db57a60265b64161c445001a803da4ecc87

Observation 8f14eeba-163f-4c60-9780-699a1d138dcb · outbound

This paper cites Pay attention to your neighbours: Training-free open-vocabulary semantic segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Pay attention to your neighbours: Training-free open-vocabulary semantic segmentation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.635478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:32.893255Z digest=sha256:2c2109f1ef8bac0d8fd8c932496bfe221a11227774aa7fc6600c34945b1ce6bd

Observation 6121bcd2-b035-490e-94e8-23a5962e7320 · outbound

This paper cites Open-vocabulary semantic segmentation with decou- pled one-pass network.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Open-vocabulary semantic segmentation with decou- pled one-pass network

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.575648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:32.898117Z digest=sha256:6703bfb442838683b461c45e968c284e891994561ef3a94cb1be16d439eebde4

Observation e3a8ae2b-cd0f-4067-b4fe-e29cfeb11fa7 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Scaling up visual and vision-language representation learning with noisy text supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:32.953042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:32.953042Z digest=sha256:2f1a851012cb5d2a6854d1bb1f7588af8ecf94db5573f9d41ed617484c3940c3

Observation 5331fbc8-5e34-4f7d-83ea-1e3387bc8fe5 · outbound

This paper cites Learning mask-aware clip representations for zero-shot segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Learning mask-aware clip representations for zero-shot segmentation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.013103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.013103Z digest=sha256:123c7a5c8cbf164a12a87dba481eeb24809a1560d0eb9c17f057a0cece21e82e

Observation 82495263-bf02-403d-aded-aefa128c51e0 · outbound

This paper cites In Defense of Lazy Visual Grounding for Open-Vocabulary Semantic Segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference In Defense of Lazy Visual Grounding for Open-Vocabulary Semantic Segmentation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-12T13:53:33.969268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.113821Z digest=sha256:e3918e0926d2fcc0c1f892f1b679178abbe45f89398656e06b718663211394e5

Observation a035c38d-50ba-4429-b6a1-7e054f3c9525 · outbound

This paper cites Weakly supervised ground- ing for vqa in vision-language transformers.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Weakly supervised ground- ing for vqa in vision-language transformers

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.537853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.211579Z digest=sha256:8ec8cedd9ad35cd47d90f92f50f97c72b2aac6f47b628e0c9aded7e15d23d438

Observation 05a8c5c2-5c51-425b-a53f-b0a8f8c728fe · outbound

This paper cites Vilt: Vision- and-language transformer without convolution or region su- pervision.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Vilt: Vision- and-language transformer without convolution or region su- pervision

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.522030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.219038Z digest=sha256:91eaa532e684b4c72c63bec75b8c29cefd02688a9cc93ecd78293963790210ff

Observation 3c09e933-e46d-4822-bb77-c8a69a614e8c · outbound

This paper cites Segment any- thing.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Segment any- thing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.503755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.223688Z digest=sha256:2164a34461763b3aac6502c94081019e0ab004f1516e5bb6fa24faa6eb2bd797

Observation f277c252-8644-46ae-903d-4dcefa7a4774 · outbound

This paper cites Efficient infer- ence in fully connected crfs with gaussian edge potentials.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Efficient infer- ence in fully connected crfs with gaussian edge potentials

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.449911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.228363Z digest=sha256:dbb5f986b6fe468acbbf84a07aadad2ea308a536cbf49880177c8f57e01edb3f

Observation dc6e9316-0db8-46c5-9b96-4ecbdc0d28b7 · outbound

This paper cites Clearclip: Decom- posing clip representations for dense vision-language infer- ence.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Clearclip: Decom- posing clip representations for dense vision-language infer- ence

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:35.235150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.232981Z digest=sha256:f918f072c337df609dff855a76ab9390300d7164b41ad0bf187b0fb436bf184f

Observation 2be1f5cc-aa44-4f78-8645-ace1b37d7181 · outbound

This paper cites ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.237441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.237441Z digest=sha256:dff12bfbdd98d2d19d3d5b334ff431aee8ea8a7f0bfe352ca5fc3962b8afa240

Observation ab0901a4-81c4-4ba2-9f83-76a579a32092 · outbound

This paper cites Align before fuse: Vision and language representation learn- ing with momentum distillation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Align before fuse: Vision and language representation learn- ing with momentum distillation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.242396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.242396Z digest=sha256:6bb998da71985a22e68517bfb18ade7d156ea50056b6b5b66ef06157fc36e026

Observation 23839c05-77b6-4ef0-bb42-87eacf6e4b2f · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.247066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.247066Z digest=sha256:37ce44fce3d425bab15b3da676d7b4ad8a257e89ab931494b9359e5c8e8a6741

Observation 052452b5-df35-4041-887e-eb0f2d1f471c · outbound

This paper cites A Closer Look at the Explainability of Contrastive Language-Image Pre-training.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference A Closer Look at the Explainability of Contrastive Language-Image Pre-training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.251700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.251700Z digest=sha256:d3beb988c9e0969bfc09cef26bbeebba944871664b79bb7288e44aa95fc1e3d3

Observation 99d91e53-7b94-45bf-ab0a-1f3a77f3cb69 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Open-vocabulary semantic segmentation with mask-adapted clip

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.256226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.256226Z digest=sha256:ddaf68c19b93bcc0b0737efe4e8e8da7aeb720d1a9470e4ca7dd4b0d4bb68d27

Observation 9dc0fa3e-36b6-492c-b91d-952f29390bb3 · outbound

This paper cites Segclip: Patch aggregation with learn- able centers for open-vocabulary semantic segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Segclip: Patch aggregation with learn- able centers for open-vocabulary semantic segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.976103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.261076Z digest=sha256:353df40be752e0f75d29caa947ce5f7c766bbf0e4176fed9734a27eca772c18a

Observation a080340c-8dfb-4982-916d-f4c1bb5675d5 · outbound

This paper cites End-to-end learning of visual representations from uncurated instruc- tional videos.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference End-to-end learning of visual representations from uncurated instruc- tional videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.266020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.266020Z digest=sha256:d9563fdb3fdd35c673df7232fd3e6c6fde6555e8429cb3ad80cc169173dfb61e

Observation 07a32e02-d329-4720-bf45-1be978ca6b42 · outbound

This paper cites The role of context for object detection and se- mantic segmentation in the wild.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference The role of context for object detection and se- mantic segmentation in the wild

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.271239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.271239Z digest=sha256:13db43ac652d1c5ac6e0efe181495422b700f2db4acb1beca4ed39d9d1549149

Observation c78933b3-c3bd-4d27-9c5b-a2b6b168bd99 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference DINOv2: Learning Robust Visual Features without Supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.276041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.276041Z digest=sha256:95020dba268c191ad036ebb485dcbb129a48d2203ed131f354d50b3c7f5f3f67

Observation 2ee3dd2a-8a54-48d9-92b7-9b8123489458 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Learn- ing transferable visual models from natural language super- vision

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.851135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.280808Z digest=sha256:fd0b85ca7cd1c14b6ba2ceff9cf4e0d78178baed1818e4d8a2e4b2d54a3964b2

Observation d61ff42f-ccf5-4c68-9b8c-c3c114d007eb · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.703300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.285175Z digest=sha256:2187f950417c5f39a18364355aa23407967e70ed3cf137257c7370b03814a790

Observation 0a9cb8d4-7eba-4a1b-86c3-5e25ecd96a1c · outbound

This paper cites Denseclip: Language-guided dense prediction with context- aware prompting.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Denseclip: Language-guided dense prediction with context- aware prompting

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.562702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.289597Z digest=sha256:f0b63586da9532d8f489eb0b187e55c751172254d1a4f09f6aef74ccc5f633f1

Observation 4940a514-f9f0-47a2-933e-1f529ca59c0d · outbound

This paper cites ViewCo: Discovering Text-Supervised Segmentation Masks via Multi-View Semantic Consistency.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference ViewCo: Discovering Text-Supervised Segmentation Masks via Multi-View Semantic Consistency

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.294063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.294063Z digest=sha256:fd5f4c67520240c2206cac4e0e38dea63d44253a7258a84c642aff03f15da892

Observation 1fe42a55-416f-4c26-b78c-4bd0eca49955 · outbound

This paper cites Ex- plore the potential of clip for training-free open vocabulary semantic segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Ex- plore the potential of clip for training-free open vocabulary semantic segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.547330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.319556Z digest=sha256:d82fa3ac5cb3096b4e08f3026e484404d49ade98b779f646cfa3a21c4acecf15

Observation b25e8ae9-8542-4ee1-b4a2-370b2ea732c9 · outbound

This paper cites Reco: Re- trieve and co-segment for zero-shot transfer.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Reco: Re- trieve and co-segment for zero-shot transfer

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.533047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.354188Z digest=sha256:46277f9827eecfce247dc5041dfe576bee31d717684ac0a2cd833cc4c9b710b2

Observation 3ed885df-0962-4438-9671-6fc7714cc06f · outbound

This paper cites Clip as rnn: Segment countless visual concepts without training endeavor.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Clip as rnn: Segment countless visual concepts without training endeavor

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.517041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.404055Z digest=sha256:d30b5191cbe166f6759417e360353ac9d443b87eba4a1168b798c44c4163b827

Observation bde8b7f7-9db7-424d-adab-d4f7a237a95d · outbound

This paper cites Learning to Decompose Visual Features with Latent Textual Prompts.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Learning to Decompose Visual Features with Latent Textual Prompts

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-12T13:53:33.862640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.444207Z digest=sha256:3ab7da4c066ca48268947d422a6c68387709278fb69abbd80278b9a7bfcd441d

Observation 220c7c8f-da1e-47f2-8467-dfd54cb65217 · outbound

This paper cites Sclip: Rethinking self-attention for dense vision-language inference.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Sclip: Rethinking self-attention for dense vision-language inference

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.502054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.524250Z digest=sha256:7def91e646febc40952b88a9fc93f403faddafe30e4e91a95302a328fc7c2c29

Observation 6bde215f-8dd1-4f6f-90a9-d397ab646c45 · outbound

This paper cites Medklip: Medical knowledge enhanced language-image pre-training for x-ray diagnosis.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Medklip: Medical knowledge enhanced language-image pre-training for x-ray diagnosis

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.461061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.529413Z digest=sha256:dd99833d2cbc5115a2d66c02f0e1486058cec35f110b06263052d70535a192db

Observation 15686629-73c2-46e6-8c92-3aebaa9b02bd · outbound

This paper cites Rewrite caption semantics: Bridging se- mantic gaps for language-supervised semantic segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Rewrite caption semantics: Bridging se- mantic gaps for language-supervised semantic segmentation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.391322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.534204Z digest=sha256:ad1a9f348e28102478d1a26dab075f3516c10a3f52838c52c358d7b6bc8484d5

Observation 577d6387-f93b-470b-b003-397677aea621 · outbound

This paper cites Demystifying CLIP Data.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Demystifying CLIP Data

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.538321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.538321Z digest=sha256:56e2c3c02c0d9c7cf01fe7aa69d95ae0072ba5163fbc8b515c993ab8ff440057

Observation ee7fde15-8078-4e60-a3d6-3542e81401fa · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Groupvit: Semantic segmentation emerges from text supervision

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.298390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.542913Z digest=sha256:229b39dd93bf0734ab263fc129d6a1635ba4692ca5302ff531ec5e3ef57c4ff0

Observation 033d3e8f-faff-4b8b-8e78-1df6548ae8f3 · outbound

This paper cites Learning open-vocabulary semantic segmentation models from natural language supervision.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Learning open-vocabulary semantic segmentation models from natural language supervision

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.230098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.547629Z digest=sha256:bba0d8b4a25ec8a0354b76f3006c2ffdf74af74b05bbb973581ff89ee951d573

Observation dfb67b0f-d00c-425c-b1bf-513888743bcd · outbound

This paper cites Side adapter network for open-vocabulary semantic segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Side adapter network for open-vocabulary semantic segmentation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.162876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.552385Z digest=sha256:fcaf0e468967f98775666b6c39dc532d8de040d78e7a9111579309b474cb938d

Observation 2ae95653-c54e-4ad2-b70c-6256400a508f · outbound

This paper cites Unified contrastive learning in image-text-label space.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Unified contrastive learning in image-text-label space

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.148172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.557229Z digest=sha256:b8855ffae369001aca20536950ea33f0195daa9c16e73139da6fa7ceac073fb1

Observation 5bcb9ca7-fc93-449f-96aa-e5003c30bcec · outbound

This paper cites Tuning-free Universally-Supervised Semantic Segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Tuning-free Universally-Supervised Semantic Segmentation

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-12T13:53:33.773603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.561686Z digest=sha256:fc96a30737fd79f4df7fe260d5e0d408b391a2704bd0124fcd4b1c9176096ccf

Observation 3964313e-b3b3-41aa-a326-138ad3d039dd · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.566440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.566440Z digest=sha256:a79bacaa354ca73e2720297df2efb46e787e5e6d445048f38d5e6b96e34855d6

Observation 3d48020e-1367-4386-8a20-804609d647b4 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.571272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.571272Z digest=sha256:5bd76dc5ff7874f340eb0812df25478d56586ca4c9e69f591df734bae70df054

Observation effb5a11-af05-47e4-95f2-dfb9f63b7af0 · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Florence: A New Foundation Model for Computer Vision

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.575798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.575798Z digest=sha256:fa0ea253f85392bd9f8af53ce80b10ef552bb11410578a8c3313737cead35dd7

Observation 0c08ad36-e9e0-4cb8-bd6f-3ba52b74a209 · outbound

This paper cites Uncovering prototypical knowledge for weakly open- vocabulary semantic segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Uncovering prototypical knowledge for weakly open- vocabulary semantic segmentation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.133288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.580214Z digest=sha256:ddc3f0cbc45c85836a1947e4b94e858f193f4444dbbab58937a1e5ca3ce94d72

Observation 074d1995-6a2f-406a-80fd-9932ce131e50 · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Semantic under- standing of scenes through the ade20k dataset

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.117946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.584897Z digest=sha256:7388d80189992506b3cb9e16b3292a1b7dcf7a3681350361e8bf35ebb1f9fcb7

Observation 76a4c42b-17b4-408e-a6ce-2c1a25984d72 · outbound

This paper cites Extract free dense labels from clip.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Extract free dense labels from clip

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.589293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.589293Z digest=sha256:f8ea9635c700992db6cf9c0f22602c169beae5213f5e93618d9d6612ae1ef0e9

Observation cd9440bd-40a5-4666-becc-c6a063c32c23 · outbound

This paper cites Image Segmentation in Foundation Model Era: A Survey.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Image Segmentation in Foundation Model Era: A Survey

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.593862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.593862Z digest=sha256:0e5df1a1f45d9307e27e1907c5a9baef3acceb37b8666cdf05b5658b8d4639d0

Observation 1ce81bfe-7811-4a4c-9db3-99703333e1cf · outbound

This paper cites Zegclip: Towards adapting clip for zero-shot se- mantic segmentation.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Zegclip: Towards adapting clip for zero-shot se- mantic segmentation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.092536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.598810Z digest=sha256:6f453d73bf950da3618cd22c010dd07bf6447fbe4f0c4324e650c9582b3af127

Observation 355102fb-2c94-4bb2-9db8-6887d5fa32f5 · outbound

This paper cites Segment everything everywhere all at once.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Segment everything everywhere all at once

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.076309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.603563Z digest=sha256:9309363ccdfd554c8c6cc596e023a8b571c4bfc2831ae69bfbc3f767b748d7e9

Observation 5238c073-aadf-411b-b77b-57a069dc7835 · outbound

This paper cites For example, in the COCO Ob- ject dataset (see Fig.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference For example, in the COCO Ob- ject dataset (see Fig

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:53:34.061494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:53:33.608671Z digest=sha256:864a0c3bb60c39b83e127f4eb214675532de66c878b7125fcbcc2b56275e67b3

Pith citing papers

No inbound Pith citation observations are available.