Pith. sign in

Paper Citation Record · LEDGER

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation

As of 17 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2411.13836.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13836 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:53:31.889301Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:04:44.106636Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T22:21:15.797283Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact2
  • verified fuzzy42
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e8d4222c-c7a9-42dd-8f17-90c0bb26a74d · outbound

This paper cites Weakly su- pervised learning of instance segmentation with inter-pixel relations.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Weakly su- pervised learning of instance segmentation with inter-pixel relations

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:35.104096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:29.953476Z digest=sha256:06311f860b9ce79a5f672828a59c171fa8e8475f33cdea5488905c6f82faf7d9

Observation 6c6c5a6a-84a4-4045-aa37-4967a6e6d61d · outbound

This paper cites Zero-shot semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Zero-shot semantic segmentation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:35.009572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:29.960642Z digest=sha256:a22e4593feb079e2515c09938733601b387ee796e9d962b0665e4eda9bd070bd

Observation d18a48e3-a586-4cd5-b360-d17f0e54e0a9 · outbound

This paper cites Coco- stuff: Thing and stuff classes in context.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Coco- stuff: Thing and stuff classes in context

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.916255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:29.967487Z digest=sha256:be7ce87feb5232ebde0bc62249ca5e1d8f5d70b7ccaa1670d777586441d28c15

Observation 197c4d1a-34e0-4973-abff-865a2fdd9132 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Emerg- ing properties in self-supervised vision transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.852705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:29.973143Z digest=sha256:d74d377e94e53a4487126584be222123bcd37d8ef388af24f1baf3ae80ce4398

Observation 2fb29648-8d04-4c9b-b931-1e7a56e978bd · outbound

This paper cites Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.806179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:29.980180Z digest=sha256:e1cc6092586ab006504c5fdef6b8a4ec8959adf4da32047e387894ebc434ad61

Observation dfd3e05b-4488-4f30-9402-a96375d05028 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Masked-attention mask transformer for universal image segmentation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.742847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:29.992618Z digest=sha256:a433921d5392ba2dd20161dc49ddcfafd892409bcbf282b695a25c8bb546cb27

Observation bfa636b9-7eb7-4292-983e-64c112e2cc83 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Reproducible scaling laws for contrastive language-image learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.704693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:30.005895Z digest=sha256:7ea852079ade669e6805151a4da3b6b99ee771a4abd18db9d27b5ae20c089088

Observation 5272cc8f-b55b-45cd-92b1-ba3b62e3fe37 · outbound

This paper cites CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:30.079121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:30.079121Z digest=sha256:1a3e0d417829de25f2647e2b5c0590d406d2263668f0d564c3965aae406aa4d2

Observation 5c3e7584-fa90-43a4-8185-0a4fe387b6ed · outbound

This paper cites Open-Vocabulary Universal Image Segmentation with MaskCLIP.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Open-Vocabulary Universal Image Segmentation with MaskCLIP

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:30.129614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:30.129614Z digest=sha256:745b9b58b035cba9199470280b8443451357ee7abf5af7b01609a0310912e1ba

Observation d277024b-19b1-4fc2-b1f8-077edea43f61 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:30.187878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:30.187878Z digest=sha256:ea130b23e72a114b1071001d559df765302466451a80bc10c21dbd262e1a99dd

Observation 78f240fe-ec8b-4796-aadf-3b4b65160abe · outbound

This paper cites The pascal visual object classes (voc) challenge.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation The pascal visual object classes (voc) challenge

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.669246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:30.241754Z digest=sha256:43b50c1c7b5907069d78d69e0463f10013da7aacf53006dabff25fbf5b5f6fb2

Observation 0881ea7d-4ff2-4390-8134-9d8081342a30 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.636723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:30.253622Z digest=sha256:c1879326d57f2bd4409997ee03484e15533b8f7c72a8e51990ab252ddd71411b

Observation 2274c32d-a5ad-4847-8272-79a291cfca52 · outbound

This paper cites Diffusion models for zero-shot open-vocabulary segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Diffusion models for zero-shot open-vocabulary segmentation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.594676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:30.262474Z digest=sha256:e6a25a581e6ce1e5edaaf5b852147cc765399663135e91433964b51449f38a70

Observation 6b827282-485e-49ef-88eb-515e84b63026 · outbound

This paper cites Segment anything.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Segment anything

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.561979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:30.323897Z digest=sha256:9c3493007548b9f406d666175c1655164dbac536021fcb274ac308a45aaf0384

Observation 423ec143-ef68-46a0-8dd1-e6d917960360 · outbound

This paper cites ProxyCLIP: Proxy attention improves clip for open-vocabulary segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation ProxyCLIP: Proxy attention improves clip for open-vocabulary segmentation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.527327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:30.434077Z digest=sha256:57d7990fa6f35bc84c4ac8b75d6dde8f2419c9ffd11b07740dfa0a9198e08ad4

Observation 6dc9b584-b0b6-4919-87f8-cb622f46ef15 · outbound

This paper cites ClearCLIP: Decom- posing clip representations for dense vision-language infer- ence.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation ClearCLIP: Decom- posing clip representations for dense vision-language infer- ence

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.469805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:30.511695Z digest=sha256:9eef5a57e51a5aaec9cc3e7f208c18c135dd960ae995775de352dd299a858749

Observation 0e733d73-93f9-4fc4-b743-8420134f2b27 · outbound

This paper cites Anti- adversarially manipulated attributions for weakly and semi- supervised semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Anti- adversarially manipulated attributions for weakly and semi- supervised semantic segmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.352659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:30.524335Z digest=sha256:aeb464e16e297e853e2ff63a93a8750dcad0c054ac8b7f1a9c8d3b8bdcd3ae17

Observation 9fa5cd2c-22c7-4835-9522-776468f09def · outbound

This paper cites A Closer Look at the Explainability of Contrastive Language-Image Pre-training.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation A Closer Look at the Explainability of Contrastive Language-Image Pre-training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:30.578203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:30.578203Z digest=sha256:3c0a4162ba7ce6fc7685412f730619e42fc5912dda00c66327493125556781fa

Observation 29a748cd-ff89-4e14-8ada-6c5fa9bea170 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Open-vocabulary semantic segmentation with mask-adapted clip

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.308320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:30.672403Z digest=sha256:20325ba161aa8b7f588e82f5875dad795d985ad1822f04ad65a0327a53e6bff3

Observation cafa38ba-9d28-4b8a-a449-f375d9a8c235 · outbound

This paper cites CLIP is also an efficient segmenter: A text-driven approach for weakly su- pervised semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation CLIP is also an efficient segmenter: A text-driven approach for weakly su- pervised semantic segmentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.274869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:30.744405Z digest=sha256:d9e4b37151195e5b1d6914e41abc27e2eea87301d043a191e82099b0d4110228

Observation 449e98ac-1963-48a5-babc-1bd3c2baeffb · outbound

This paper cites TagCLIP: A local-to-global framework to enhance open-vocabulary multi-label classification of clip without training.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation TagCLIP: A local-to-global framework to enhance open-vocabulary multi-label classification of clip without training

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.237480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:30.756293Z digest=sha256:38aa4e74a906cb1f35aaac5393a92fdffd916ffe79df401814bde9432cc9f76f

Observation 9744087a-2167-4d00-b9a8-cbc741bbe050 · outbound

This paper cites Fully convolutional networks for semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Fully convolutional networks for semantic segmentation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.201120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:30.764526Z digest=sha256:fce4eb0f772050975efa4528a6be5dff7f9287a7bee0f0bc6ea8b0f713477c1c

Observation 66449b3c-c91c-4a61-8ca8-18a167b6a519 · outbound

This paper cites SegCLIP: Patch aggregation with learnable centers for open-vocabulary semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation SegCLIP: Patch aggregation with learnable centers for open-vocabulary semantic segmentation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.169300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:30.854185Z digest=sha256:d4cfbab16145915ebe6787768e97b5be0e3a73498cfda21536b44ca4b40ff0a3

Observation 4cbf94d9-3b2a-4eb4-9e34-5ee0d946aa24 · outbound

This paper cites The role of context for object detection and semantic segmentation in the wild.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation The role of context for object detection and semantic segmentation in the wild

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.127614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:30.956732Z digest=sha256:423f445e0ccac8c3347fe01bd03274882f17fb67addc129ea68f4181e86df22b

Observation 82a288b3-01c7-41c6-9ee1-579ecf5ed70b · outbound

This paper cites Learning transferable visual models from natural language supervision.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Learning transferable visual models from natural language supervision

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.983972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:30.980568Z digest=sha256:79ab7e34c0d35108e26ff380ad8cf3f3ec5320310cdc909e4fd2f5cc74a08211

Observation d3175b13-b7f6-4075-a543-9417b6f3b717 · outbound

This paper cites Per- ceptual grouping in contrastive vision-language models.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Per- ceptual grouping in contrastive vision-language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.861006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.004171Z digest=sha256:4c1b4e10ea309fa991e3c77a1506cbfa83a6dcd94be68d3aa1cffd2248ff42e6

Observation e7ea2c76-8fe9-4204-97ac-b02c22479857 · outbound

This paper cites ViewCo: Discovering text-supervised segmentation masks via multi-view semantic consistency.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation ViewCo: Discovering text-supervised segmentation masks via multi-view semantic consistency

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.818391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.016261Z digest=sha256:9869f262b43afd4dfeba71f81d5df21ca1b3a2da63d5698d5628e6d05218d319

Observation dc4bb987-e0bd-48da-8858-9c13f054bca6 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation High-resolution image syn- thesis with latent diffusion models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.785884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.091640Z digest=sha256:13b0648c86e40dd06687ce2ca61e2eab9a1d53eb081e4f14b959800caabe6759

Observation 8db690b3-cad4-42db-9946-f1bf9512de3a · outbound

This paper cites To- ken contrast for weakly-supervised semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation To- ken contrast for weakly-supervised semantic segmentation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.749028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.205206Z digest=sha256:4903397eb2f8386bd029061472c5adeb7ac1d2d6cb528bdb456ee24bf6d28fef

Observation ff6fa4f0-1142-43e1-ac19-a7faab042712 · outbound

This paper cites Laion-5B: An open large-scale dataset for training next generation image-text models.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Laion-5B: An open large-scale dataset for training next generation image-text models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.713015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.270985Z digest=sha256:4a5b84420098f805e161cb9e51300595bd56461974ba76e9620d77072c341b11

Observation 3ef15bd6-675a-46e3-ae5b-d66e8b41b6db · outbound

This paper cites ReCo: Re- trieve and co-segment for zero-shot transfer.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation ReCo: Re- trieve and co-segment for zero-shot transfer

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.672270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.288916Z digest=sha256:6105e7bf93d587c39234af061ac10f9132b81f87fd5bdafd04082db32b7263b6

Observation 76af3956-dcb7-43e3-82b4-dade04414c75 · outbound

This paper cites iSeg: An Iterative Refinement-based Framework for Training-free Segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation iSeg: An Iterative Refinement-based Framework for Training-free Segmentation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:53:32.128129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.299753Z digest=sha256:5df110a79dbd293b6a48e4c253bed5ce8198f84acab3f9bf13131158411e4eb6

Observation ee87535e-4d4b-41ad-8387-f8c20bdb7c73 · outbound

This paper cites CLIP as RNN: Segment countless visual concepts with- out training endeavor.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation CLIP as RNN: Segment countless visual concepts with- out training endeavor

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.634456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.324636Z digest=sha256:ec5581a1b535f017dbe69def1bb074a272fb25f5c85510e65b53a1848718d372

Observation 0e62b1ef-f7ee-410a-9bee-51decc852a1f · outbound

This paper cites Sclip: Rethinking self-attention for dense vision-language inference.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Sclip: Rethinking self-attention for dense vision-language inference

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.574412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.421744Z digest=sha256:bbd7681bc08449c106a4d33c76b85eb02f212bfc5bb71d56c3bffe4a4e926b04

Observation 537c9fe7-9180-47fc-beb5-44f643700f30 · outbound

This paper cites Sam-clip: Merging vision foundation models towards semantic and spatial understanding.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Sam-clip: Merging vision foundation models towards semantic and spatial understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.398863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.466108Z digest=sha256:d5c14f220406b87f4d221f5e37fe98cf891071a2db83f687748372f6ced0b240

Observation f1e187ce-2dfd-4150-8125-3768a669df13 · outbound

This paper cites Diffusion Model is Secretly a Training-free Open Vocabulary Semantic Segmenter.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Diffusion Model is Secretly a Training-free Open Vocabulary Semantic Segmenter

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:31.546345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:31.546345Z digest=sha256:30f771aca6ca0edfec81425ec6f246da72c56dcdd4ff5c67296708c9b6ac9a12

Observation e021f2c3-832d-4647-886f-e9733f0fac57 · outbound

This paper cites Clip-dinoiser: Teaching clip a few dino tricks for open- vocabulary semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Clip-dinoiser: Teaching clip a few dino tricks for open- vocabulary semantic segmentation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.361545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.678102Z digest=sha256:d67fc28ee489cabf731107a5b425a6d8c74244d91285ceb3c21a290d168e75a7

Observation 5268430a-a04e-4686-bde8-0acc42f79d99 · outbound

This paper cites From Text to Mask: Localizing Entities Using the Attention of Text-to-Image Diffusion Models.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation From Text to Mask: Localizing Entities Using the Attention of Text-to-Image Diffusion Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:53:32.041705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.718299Z digest=sha256:793c3ed99f9e2626b4c46ef93ef80cdeaa3a4003394726f66fbfd71a281b4d27

Observation 0fd812d0-79ab-49e5-8126-9f64443fd613 · outbound

This paper cites Sed: A simple encoder-decoder for open- vocabulary semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Sed: A simple encoder-decoder for open- vocabulary semantic segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.203098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.738935Z digest=sha256:669a1eff873773062550ffc59d8bb3a68ec008eff88a2d4bf2b2dd1318a3023b

Observation 6818b0d7-7fbb-4083-b73b-4807879449f2 · outbound

This paper cites Alvarez, and Ping Luo.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Alvarez, and Ping Luo

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.169597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.752197Z digest=sha256:ab1d0e98140f05c88b937e2b47493ea5574d5139a77045d014e93c7e57067f53

Observation 82e486ea-48b1-475c-bb02-2c9994ba1305 · outbound

This paper cites Clims: Cross language image matching for weakly supervised se- mantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Clims: Cross language image matching for weakly supervised se- mantic segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.990468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.765289Z digest=sha256:0da06631e11bb176d836a6ebe27c77da819a9f58c4bc085e3a30e0d7a67847fb

Observation 179b678d-4304-41da-b23a-b38123aa2586 · outbound

This paper cites Rewrite caption semantics: Bridging seman- tic gaps for language-supervised semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Rewrite caption semantics: Bridging seman- tic gaps for language-supervised semantic segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.915948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.782187Z digest=sha256:41448ff405b30eca7176264403a7540d9f680d97c1dcde67e7c71a626b3bc8e7

Observation aa91ddfa-5545-4106-83cb-1dbe4aaf9647 · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Groupvit: Semantic segmentation emerges from text supervision

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.850323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.798853Z digest=sha256:046d8661b9ccd0b91915455f0de5de07470d07012de0123f764a7c717b52cad7

Observation 34810fe9-a51d-4920-83db-c970a2641733 · outbound

This paper cites Learning open-vocabulary seman- tic segmentation models from natural language supervision.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Learning open-vocabulary seman- tic segmentation models from natural language supervision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.809151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.808098Z digest=sha256:2d971f7afa32af6e0af809051bbee7e99d8ee99f93e39b4a789a317c0a2b2c43

Observation 2f00c189-f6a3-4241-a794-bda1d061773d · outbound

This paper cites Open-vocabulary panoptic segmentation with text-to-image diffusion models.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Open-vocabulary panoptic segmentation with text-to-image diffusion models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.773430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.817712Z digest=sha256:2b49c0fcac78d5ae386bd016749a22c6b678993b5c4f7d420089e29c93fdefab

Observation 16223171-313d-4031-a611-a09d431217d6 · outbound

This paper cites Multi-class token transformer for weakly supervised se- mantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Multi-class token transformer for weakly supervised se- mantic segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.726561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.824494Z digest=sha256:ad66e20cde9f61a5038abb0efe49c00293769c78518050e3f35f7fed0e67d241

Observation 0732b187-27e6-4bd9-9ecd-79a5528df144 · outbound

This paper cites Side adapter network for open-vocabulary semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Side adapter network for open-vocabulary semantic segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.660574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.838226Z digest=sha256:b79d0b35b664527229d58656d06165bf913cfddcf729c9631a14edfcb6d5fd12

Observation dedbb98a-65d2-424e-b811-32470ae6b714 · outbound

This paper cites Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIP.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIP

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:31.853023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:31.853023Z digest=sha256:83cf666cdff0cdb6aede4470af5bcf7f16f6a7db6f1d7364777094573f759003

Observation 91ee6cf4-fc84-47f8-bc5f-272acd54846b · outbound

This paper cites Open vocabulary scene parsing.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Open vocabulary scene parsing

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.616043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.865201Z digest=sha256:6c81b2c9b8a702d89e2e294b0eae29eec953197f4a8d3146dd13a5dc437f7e1c

Observation c5655f20-8b1e-4db0-8adb-12423fbf6252 · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Semantic under- standing of scenes through the ade20k dataset

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:31.878658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:31.878658Z digest=sha256:505f7975869a41132522a972eabd6f96e978911d1f6838aa1e2d930ba649383d

Observation 3241dc5e-106a-4fec-82df-4def85422460 · outbound

This paper cites Extract free dense labels from clip.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Extract free dense labels from clip

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.461065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T15:53:31.889301Z digest=sha256:2e653beda7125a1639a7b6fc4f1135bca3f30a88a72e68bd3941cee3ce750956

Pith citing papers

Observation a7d5046b-d83a-4772-9b7f-346ca2037fe4 · inbound

CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic Segmentation cites this paper.

CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic Segmentation CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T20:04:44.106636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:04:44.106636Z digest=sha256:be35a92335b01dfdbfd7a9bdde675c70356475f9519cf99d9da51210e8108823

Observation 1fee426a-3576-4ef6-b5f8-473170d116dc · inbound

Perception Encoder: The best visual embeddings are not at the output of the network cites this paper.

Perception Encoder: The best visual embeddings are not at the output of the network CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:21:15.798787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T22:21:15.681336Z digest=sha256:884990dc9bbbe261d5650fe6c50c6a6c9e868a1513ce8a6e5b9841cf724744de