Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T13:53:33.608671Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2411.15851.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T13:53:33.608671Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
63 of 63 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b298baa7-e8e4-4466-8c33-215c03636890 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Self- supervised multimodal versatile networks
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f2b0ec18-3fad-4177-aa0e-b39aae75fd0c · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Vqa: Visual question answering
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bcfe492a-4536-4a33-9dc6-109cb470c38c · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Single-stage semantic segmentation from image labels
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 78d05ed0-9048-44e8-b94c-765f9f6bf9f4 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Fossil: Free open-vocabulary semantic seg- mentation through synthetic references retrieval
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b64f6d17-a166-48fc-9519-e760de86b75a · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Grounding everything: Emerging localiza- tion properties in vision-language transformers
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1d9a4e0e-d4e1-4ef4-b8f2-087d251ad34b · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Language Models are Few-Shot Learners
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39fcef99-4425-417e-a316-34af1ae24a62 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Coco- stuff: Thing and stuff classes in context
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e6cb987b-2747-43c4-b9fc-f252f3e88bd2 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Emerg- ing properties in self-supervised vision transformers
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 720e2a45-2f21-4802-bd70-94c76f7817f6 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 59add72b-cbed-41f5-9e12-e55cbe5d4cfc · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Reproducible scal- ing laws for contrastive language-image learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6bcddcd9-c84b-4a20-b53c-8ccdaa8da91f · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Mmsegmentation: Open- mmlab semantic segmentation toolbox and benchmark,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05dd4e8e-6feb-4753-a520-e088b664eade · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference The cityscapes dataset for semantic urban scene understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 65821c0f-2527-4333-a5a6-b022c59676b1 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Vision Transformers Need Registers
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2e783d8-34df-405e-9b0f-a595dc122f2f · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ce91ed8-829d-4a69-ae6b-be0bb16edddf · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4a721d3-574e-4cfa-98fa-d8452bc59ba2 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference The pascal visual object classes challenge: A retrospective.IJCV, 111:98–136, 2015
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9d6152d3-944f-4766-ba2c-87ff0c6d2b23 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Improved baselines for vision-language pre-training
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f14eeba-163f-4c60-9780-699a1d138dcb · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Pay attention to your neighbours: Training-free open-vocabulary semantic segmentation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6121bcd2-b035-490e-94e8-23a5962e7320 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Open-vocabulary semantic segmentation with decou- pled one-pass network
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e3a8ae2b-cd0f-4067-b4fe-e29cfeb11fa7 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Scaling up visual and vision-language representation learning with noisy text supervision
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5331fbc8-5e34-4f7d-83ea-1e3387bc8fe5 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Learning mask-aware clip representations for zero-shot segmentation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82495263-bf02-403d-aded-aefa128c51e0 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference In Defense of Lazy Visual Grounding for Open-Vocabulary Semantic Segmentation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a035c38d-50ba-4429-b6a1-7e054f3c9525 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Weakly supervised ground- ing for vqa in vision-language transformers
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 05a8c5c2-5c51-425b-a53f-b0a8f8c728fe · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Vilt: Vision- and-language transformer without convolution or region su- pervision
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3c09e933-e46d-4822-bb77-c8a69a614e8c · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Segment any- thing
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f277c252-8644-46ae-903d-4dcefa7a4774 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Efficient infer- ence in fully connected crfs with gaussian edge potentials
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation dc6e9316-0db8-46c5-9b96-4ecbdc0d28b7 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Clearclip: Decom- posing clip representations for dense vision-language infer- ence
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2be1f5cc-aa44-4f78-8645-ace1b37d7181 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab0901a4-81c4-4ba2-9f83-76a579a32092 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Align before fuse: Vision and language representation learn- ing with momentum distillation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23839c05-77b6-4ef0-bb42-87eacf6e4b2f · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 052452b5-df35-4041-887e-eb0f2d1f471c · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference A Closer Look at the Explainability of Contrastive Language-Image Pre-training
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99d91e53-7b94-45bf-ab0a-1f3a77f3cb69 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Open-vocabulary semantic segmentation with mask-adapted clip
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dc0fa3e-36b6-492c-b91d-952f29390bb3 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Segclip: Patch aggregation with learn- able centers for open-vocabulary semantic segmentation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a080340c-8dfb-4982-916d-f4c1bb5675d5 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference End-to-end learning of visual representations from uncurated instruc- tional videos
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07a32e02-d329-4720-bf45-1be978ca6b42 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference The role of context for object detection and se- mantic segmentation in the wild
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c78933b3-c3bd-4d27-9c5b-a2b6b168bd99 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference DINOv2: Learning Robust Visual Features without Supervision
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ee3dd2a-8a54-48d9-92b7-9b8123489458 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Learn- ing transferable visual models from natural language super- vision
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d61ff42f-ccf5-4c68-9b8c-c3c114d007eb · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0a9cb8d4-7eba-4a1b-86c3-5e25ecd96a1c · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Denseclip: Language-guided dense prediction with context- aware prompting
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4940a514-f9f0-47a2-933e-1f529ca59c0d · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference ViewCo: Discovering Text-Supervised Segmentation Masks via Multi-View Semantic Consistency
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fe42a55-416f-4c26-b78c-4bd0eca49955 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Ex- plore the potential of clip for training-free open vocabulary semantic segmentation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b25e8ae9-8542-4ee1-b4a2-370b2ea732c9 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Reco: Re- trieve and co-segment for zero-shot transfer
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3ed885df-0962-4438-9671-6fc7714cc06f · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Clip as rnn: Segment countless visual concepts without training endeavor
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bde8b7f7-9db7-424d-adab-d4f7a237a95d · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Learning to Decompose Visual Features with Latent Textual Prompts
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 220c7c8f-da1e-47f2-8467-dfd54cb65217 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Sclip: Rethinking self-attention for dense vision-language inference
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6bde215f-8dd1-4f6f-90a9-d397ab646c45 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Medklip: Medical knowledge enhanced language-image pre-training for x-ray diagnosis
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 15686629-73c2-46e6-8c92-3aebaa9b02bd · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Rewrite caption semantics: Bridging se- mantic gaps for language-supervised semantic segmentation
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 577d6387-f93b-470b-b003-397677aea621 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Demystifying CLIP Data
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee7fde15-8078-4e60-a3d6-3542e81401fa · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Groupvit: Semantic segmentation emerges from text supervision
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 033d3e8f-faff-4b8b-8e78-1df6548ae8f3 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Learning open-vocabulary semantic segmentation models from natural language supervision
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation dfb67b0f-d00c-425c-b1bf-513888743bcd · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Side adapter network for open-vocabulary semantic segmentation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2ae95653-c54e-4ad2-b70c-6256400a508f · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Unified contrastive learning in image-text-label space
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5bcb9ca7-fc93-449f-96aa-e5003c30bcec · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Tuning-free Universally-Supervised Semantic Segmentation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3964313e-b3b3-41aa-a326-138ad3d039dd · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d48020e-1367-4386-8a20-804609d647b4 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference CoCa: Contrastive Captioners are Image-Text Foundation Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation effb5a11-af05-47e4-95f2-dfb9f63b7af0 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Florence: A New Foundation Model for Computer Vision
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c08ad36-e9e0-4cb8-bd6f-3ba52b74a209 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Uncovering prototypical knowledge for weakly open- vocabulary semantic segmentation
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 074d1995-6a2f-406a-80fd-9932ce131e50 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Semantic under- standing of scenes through the ade20k dataset
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 76a4c42b-17b4-408e-a6ce-2c1a25984d72 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Extract free dense labels from clip
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd9440bd-40a5-4666-becc-c6a063c32c23 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Image Segmentation in Foundation Model Era: A Survey
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ce81bfe-7811-4a4c-9db3-99703333e1cf · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Zegclip: Towards adapting clip for zero-shot se- mantic segmentation
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 355102fb-2c94-4bb2-9db8-6887d5fa32f5 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference Segment everything everywhere all at once
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5238c073-aadf-411b-b77b-57a069dc7835 · outbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference For example, in the COCO Ob- ject dataset (see Fig
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
No inbound Pith citation observations are available.