Pith. sign in

Paper Citation Record · LEDGER

SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2312.01597.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.01597 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:34:51.075189Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T19:36:29.309759Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 95d6b0c8-dd8e-4220-8118-9c4fa2e4ccbb · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.295493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:72a1e75f9430dfb5e7c612f1bc42d06a290cdfbfa4b3a61446779e2ea1f9f31a

Observation d9606a73-0f19-4eea-9f67-e31315b3625e · inbound

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements cites this paper.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.166483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.166483Z digest=sha256:f78f76cbd00b6e3651f69b4ad330a282b012328141aa623cd3a4a10870022fcc

Observation e83e9750-a805-4292-b29e-2252dc0bbc73 · inbound

FLAIR: VLM with Fine-grained Language-informed Image Representations cites this paper.

FLAIR: VLM with Fine-grained Language-informed Image Representations SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:06.859418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:06.859418Z digest=sha256:0adc2b70744807ac158eb72d1739187382c8db705b08af413fca2431874a7211

Observation fa374632-36a5-41e3-ac66-f08ac7cf0f30 · inbound

TeD-Loc: Text Distillation for Weakly Supervised Object Localization cites this paper.

TeD-Loc: Text Distillation for Weakly Supervised Object Localization SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:12:34.882674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-23T05:11:17.213575Z digest=sha256:160fb8d3b86d89040ead5de8e6c15d5a1f9ad8c2abd4dd2622c3c60bc29c364a

Observation 45b0c1f1-b919-43be-83aa-6433dd16fcd7 · inbound

DynAlign: Unsupervised Dynamic Taxonomy Alignment for Cross-Domain Segmentation cites this paper.

DynAlign: Unsupervised Dynamic Taxonomy Alignment for Cross-Domain Segmentation SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T13:35:59.711061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:35:59.711061Z digest=sha256:74145323e718a7e96c892add0045f73ba338222f82a62eb6ab5a1fbe98b42cc9

Observation 0ae76e9d-c700-4f74-812a-449e6060ed27 · inbound

Laser: Efficient Language-Guided Segmentation in Neural Radiance Fields cites this paper.

Laser: Efficient Language-Guided Segmentation in Neural Radiance Fields SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T21:29:34.973117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:29:34.973117Z digest=sha256:c8e8070422d5804d3d967d0bf2f1763d3980c13ce8ba8ea9d67c80f982dbc8e5

Observation ccdf75bc-5191-4cec-915c-76003b377e57 · inbound

DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception cites this paper.

DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T23:34:51.075189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:34:51.075189Z digest=sha256:db17d7a796acfb4b31fcae2cb844bac071ac069c37ecb20ed3c022a9b36b4e6b

Observation dea6da46-5bb6-46f9-afb3-3b67f67e0f14 · inbound

RemoteSAM: Towards Segment Anything for Earth Observation cites this paper.

RemoteSAM: Towards Segment Anything for Earth Observation SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:39:37.747635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:39:37.747635Z digest=sha256:d69035aab5d9ad5b6dee94303de068ac9e6de18dc64d557770e602d7f6c53b45

Observation 09dbf5f1-110e-488e-9ac2-1ae488ea66a6 · inbound

G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models cites this paper.

G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:46:58.263286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:46:58.263286Z digest=sha256:0d7ea4aa3af1c847a7aabf8a2b6764dfb660c7a9a4b027829a928fe181a56004

Observation 32c87e97-1ff2-4d3e-9086-afb29d14dacf · inbound

Annotation-Free Open-Vocabulary Segmentation for Remote-Sensing Images cites this paper.

Annotation-Free Open-Vocabulary Segmentation for Remote-Sensing Images SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T16:41:41.714250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:41:41.714250Z digest=sha256:8885d8dc0d783ad72442bc08a4f0a9999ef308ee83961c8de2a268848eddb754

Observation e2bf5ff8-3e61-47e1-87bc-2b01717e3ca2 · inbound

Plug-in Feedback Self-adaptive Attention in CLIP for Training-free Open-Vocabulary Segmentation cites this paper.

Plug-in Feedback Self-adaptive Attention in CLIP for Training-free Open-Vocabulary Segmentation SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:45.993153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:45.993153Z digest=sha256:7f9b55423c914bab924b78763ca4f7133f9bc6e8578c145f0cf801cde0f6d918

Observation 5ada7ad0-ce1b-44ae-a086-e5c6120bff8c · inbound

SegEarth-OV3: Exploring SAM 3 for Open-Vocabulary Semantic Segmentation in Remote Sensing Images cites this paper.

SegEarth-OV3: Exploring SAM 3 for Open-Vocabulary Semantic Segmentation in Remote Sensing Images SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:53:42.275004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T23:53:27.350335Z digest=sha256:6b67373b08f0776e57f49634f45456ad954e957385c2c326849632cf6b0bf751

Observation 904df424-5af3-42ff-a8ee-23f9b138e10f · inbound

Best Segmentation Buddies for Image-Shape Correspondence cites this paper.

Best Segmentation Buddies for Image-Shape Correspondence SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:13:13.524486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T11:10:29.360224Z digest=sha256:ac05d72a31869f29163d306494ff988bbe131c8ed0196ab9939eba45e37aeb43

Observation d3127534-8c85-4768-9f13-e70fa482d691 · inbound

Vision Harnessing Agent for Open Ad-hoc Segmentation cites this paper.

Vision Harnessing Agent for Open Ad-hoc Segmentation SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.512876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:52:40.429412Z digest=sha256:a2bbe5e664a6ae76ed9f3af7b4da3b76d893e583c629a3515d6ecff5894c0975

Observation efff4043-dec5-4d7f-8b6c-0bd95051a59a · inbound

LARE: Low-Attention Region Encoding for Text-Image Retrieval cites this paper.

LARE: Low-Attention Region Encoding for Text-Image Retrieval SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:09:15.262035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-26T21:25:12.373068Z digest=sha256:e67403b0e8aa5302d003ff1c8ae4bf96bcaee6d2fdef784845a00889ca17ba99

Observation 1a10c2ce-6aeb-45f7-b0a5-da57f1b67a4b · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 241

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.544202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:77c9b6623cf1ae31ae4aff2bfd0c2b1d601d067699b42ab187fa5baa2e682f17

Observation f47636fc-376e-4604-be5c-99e36a71e536 · inbound

Sparse Attention for Dense Open-Vocabulary Prediction in CLIP cites this paper.

Sparse Attention for Dense Open-Vocabulary Prediction in CLIP SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T19:36:29.311029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-09T19:26:49.805499Z digest=sha256:c275225716fda529b4618704934c48e8e4563f9d667351b6ab8e14b351de982b

Observation ecf2435e-9775-47b6-8ef4-e92d53dad321 · inbound

Sparse Attention for Dense Open-Vocabulary Prediction in CLIP cites this paper.

Sparse Attention for Dense Open-Vocabulary Prediction in CLIP SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T15:51:57.429693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:51:57.429693Z digest=sha256:2a301eeeb4eddd3a4e21044a8fcc0cbe96c0d30546c5a28147346a3f1b4fa9af

Observation 28551c3f-0a01-4fbb-9289-d09829405572 · inbound

Standalone DINOv3 for Training-Free Open-Vocabulary Semantic Segmentation in Remote Sensing cites this paper.

Standalone DINOv3 for Training-Free Open-Vocabulary Semantic Segmentation in Remote Sensing SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:20.323156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T04:20:20.323156Z digest=sha256:4a277569d44e6dcc5a346260a50985ef87d7f189bc86124df7e4d1fe861f02f4