Pith. sign in

Paper Citation Record · LEDGER

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation

As of 17 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2411.13836.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13836 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:53:31.889301Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:04:44.106636Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T22:21:15.797283Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact2
  • verified fuzzy42
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e8d4222c-c7a9-42dd-8f17-90c0bb26a74d · outbound

This paper cites Weakly su- pervised learning of instance segmentation with inter-pixel relations.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Weakly su- pervised learning of instance segmentation with inter-pixel relations

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:35.104096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:29.953476Z digest=sha256:76b7e13683b656dbfaf3939c8c5ef8ab48c99800218a2e8bc55b4bf25e0021b2

Observation 6c6c5a6a-84a4-4045-aa37-4967a6e6d61d · outbound

This paper cites Zero-shot semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Zero-shot semantic segmentation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:35.009572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:29.960642Z digest=sha256:20910d79af19f78aa382ad34765c2513b9002eb00166b086e145087d0f02bae5

Observation d18a48e3-a586-4cd5-b360-d17f0e54e0a9 · outbound

This paper cites Coco- stuff: Thing and stuff classes in context.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Coco- stuff: Thing and stuff classes in context

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.916255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:29.967487Z digest=sha256:63f2cda66f6f2de1bf30369fa7f605503cb3235823b9c74420f8cf02b598a30b

Observation 197c4d1a-34e0-4973-abff-865a2fdd9132 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Emerg- ing properties in self-supervised vision transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.852705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:29.973143Z digest=sha256:bcf89f66eda7ed07d3ed5572ba055ad0e8211be922b9fc395a103eab5abf7849

Observation 2fb29648-8d04-4c9b-b931-1e7a56e978bd · outbound

This paper cites Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.806179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:29.980180Z digest=sha256:ca0b670702f5e8cada506013ff60ce44d4ac490d205067771a429db558030989

Observation dfd3e05b-4488-4f30-9402-a96375d05028 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Masked-attention mask transformer for universal image segmentation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.742847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:29.992618Z digest=sha256:7acd2b5e19e1a9efd1fe24097f51ca6d3c79b8ff5c16d5ececadefe5ffc3a214

Observation bfa636b9-7eb7-4292-983e-64c112e2cc83 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Reproducible scaling laws for contrastive language-image learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.704693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:30.005895Z digest=sha256:2d2d43f2021ae8c43bd074647cffaac4da0b675a68b7fc71cab7fd24109182c5

Observation 5272cc8f-b55b-45cd-92b1-ba3b62e3fe37 · outbound

This paper cites CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:30.079121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:30.079121Z digest=sha256:1a3e0d417829de25f2647e2b5c0590d406d2263668f0d564c3965aae406aa4d2

Observation 5c3e7584-fa90-43a4-8185-0a4fe387b6ed · outbound

This paper cites Open-Vocabulary Universal Image Segmentation with MaskCLIP.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Open-Vocabulary Universal Image Segmentation with MaskCLIP

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:30.129614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:30.129614Z digest=sha256:745b9b58b035cba9199470280b8443451357ee7abf5af7b01609a0310912e1ba

Observation d277024b-19b1-4fc2-b1f8-077edea43f61 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:30.187878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:30.187878Z digest=sha256:ea130b23e72a114b1071001d559df765302466451a80bc10c21dbd262e1a99dd

Observation 78f240fe-ec8b-4796-aadf-3b4b65160abe · outbound

This paper cites The pascal visual object classes (voc) challenge.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation The pascal visual object classes (voc) challenge

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.669246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:30.241754Z digest=sha256:5c91504829e1beab81c62f5ac15fd3140f7021643c611d4ab2124b8080539104

Observation 0881ea7d-4ff2-4390-8134-9d8081342a30 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.636723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:30.253622Z digest=sha256:3cc73e1eddd437c668c2df7eed677af156b74805c41fdf14c63b44ee0bee4642

Observation 2274c32d-a5ad-4847-8272-79a291cfca52 · outbound

This paper cites Diffusion models for zero-shot open-vocabulary segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Diffusion models for zero-shot open-vocabulary segmentation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.594676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:30.262474Z digest=sha256:d2e2eac41e04313bffd5db02c8e43c5867772850d597e333f7ce73f1b07a28f2

Observation 6b827282-485e-49ef-88eb-515e84b63026 · outbound

This paper cites Segment anything.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Segment anything

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.561979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:30.323897Z digest=sha256:7aa1fc51ebfee504b2e8c7bed467b657cb0a1e4b358ac41db51fd330f4d8f485

Observation 423ec143-ef68-46a0-8dd1-e6d917960360 · outbound

This paper cites ProxyCLIP: Proxy attention improves clip for open-vocabulary segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation ProxyCLIP: Proxy attention improves clip for open-vocabulary segmentation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.527327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:30.434077Z digest=sha256:0c7f0b14d0bca54afe8fca767d23e80f69f0b693a148e25769aefd5a14fbc322

Observation 6dc9b584-b0b6-4919-87f8-cb622f46ef15 · outbound

This paper cites ClearCLIP: Decom- posing clip representations for dense vision-language infer- ence.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation ClearCLIP: Decom- posing clip representations for dense vision-language infer- ence

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.469805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:30.511695Z digest=sha256:79b1534eb44d83c37ca0bf98752ff48df22f9e07fb642a7cdd67e54744d368e3

Observation 0e733d73-93f9-4fc4-b743-8420134f2b27 · outbound

This paper cites Anti- adversarially manipulated attributions for weakly and semi- supervised semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Anti- adversarially manipulated attributions for weakly and semi- supervised semantic segmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.352659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:30.524335Z digest=sha256:f60a4e0bcedad4a9ed4530f1ea83f1804f74ec901e1f87d13ab0deb6ccffd003

Observation 9fa5cd2c-22c7-4835-9522-776468f09def · outbound

This paper cites A Closer Look at the Explainability of Contrastive Language-Image Pre-training.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation A Closer Look at the Explainability of Contrastive Language-Image Pre-training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:30.578203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:30.578203Z digest=sha256:3c0a4162ba7ce6fc7685412f730619e42fc5912dda00c66327493125556781fa

Observation 29a748cd-ff89-4e14-8ada-6c5fa9bea170 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Open-vocabulary semantic segmentation with mask-adapted clip

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.308320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:30.672403Z digest=sha256:4c3c95ef9de23f6f0f5f1ce9578ad8be126c3f985f9b51428228252417b25163

Observation cafa38ba-9d28-4b8a-a449-f375d9a8c235 · outbound

This paper cites CLIP is also an efficient segmenter: A text-driven approach for weakly su- pervised semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation CLIP is also an efficient segmenter: A text-driven approach for weakly su- pervised semantic segmentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.274869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:30.744405Z digest=sha256:0e33390e81c3caa32e8736999ccb68b956ae69b3c23f431b9f181a9a00c807fd

Observation 449e98ac-1963-48a5-babc-1bd3c2baeffb · outbound

This paper cites TagCLIP: A local-to-global framework to enhance open-vocabulary multi-label classification of clip without training.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation TagCLIP: A local-to-global framework to enhance open-vocabulary multi-label classification of clip without training

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.237480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:30.756293Z digest=sha256:561703f8d21a84ad42be229f33ae2b8c594afc412e3a0f31ed0b5ac2e082fd94

Observation 9744087a-2167-4d00-b9a8-cbc741bbe050 · outbound

This paper cites Fully convolutional networks for semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Fully convolutional networks for semantic segmentation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.201120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:30.764526Z digest=sha256:72a223b7675aa0a9cc6449b8a00d03c7f6840226677b3c71a9c6a8dd795753f9

Observation 66449b3c-c91c-4a61-8ca8-18a167b6a519 · outbound

This paper cites SegCLIP: Patch aggregation with learnable centers for open-vocabulary semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation SegCLIP: Patch aggregation with learnable centers for open-vocabulary semantic segmentation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.169300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:30.854185Z digest=sha256:81dcf5d7383989ee9ce12ea48951d2cf25f69fdf114c0ad4338a2701b58a8f38

Observation 4cbf94d9-3b2a-4eb4-9e34-5ee0d946aa24 · outbound

This paper cites The role of context for object detection and semantic segmentation in the wild.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation The role of context for object detection and semantic segmentation in the wild

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.127614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:30.956732Z digest=sha256:0e0685385af75b41140e1a0141ddf9b1e13481b45ef914a0d3e228ff11d63e3d

Observation 82a288b3-01c7-41c6-9ee1-579ecf5ed70b · outbound

This paper cites Learning transferable visual models from natural language supervision.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Learning transferable visual models from natural language supervision

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.983972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:30.980568Z digest=sha256:8b9a47f2468b2d81bc2b99b0a9285b54cfc91c4a6ea97ab4a26aed7d3411f2c7

Observation d3175b13-b7f6-4075-a543-9417b6f3b717 · outbound

This paper cites Per- ceptual grouping in contrastive vision-language models.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Per- ceptual grouping in contrastive vision-language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.861006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.004171Z digest=sha256:ce6c724f72a8ea9c00f822b618410b7e6f214651b70483436257ed82cdfecdaf

Observation e7ea2c76-8fe9-4204-97ac-b02c22479857 · outbound

This paper cites ViewCo: Discovering text-supervised segmentation masks via multi-view semantic consistency.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation ViewCo: Discovering text-supervised segmentation masks via multi-view semantic consistency

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.818391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.016261Z digest=sha256:07115bab5ac7bb17b434ae709075856fa1373efa0b57ee94dfbc76b8f41c11f7

Observation dc4bb987-e0bd-48da-8858-9c13f054bca6 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation High-resolution image syn- thesis with latent diffusion models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.785884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.091640Z digest=sha256:c5825e75ef441496eda4ea26365cc9d952a7424127a32d311574fdd8a2eb8be4

Observation 8db690b3-cad4-42db-9946-f1bf9512de3a · outbound

This paper cites To- ken contrast for weakly-supervised semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation To- ken contrast for weakly-supervised semantic segmentation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.749028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.205206Z digest=sha256:76f063a8c5982d5010289cd5ac8fc5d9b2a8f3f55d70f10565ac87852e37e015

Observation ff6fa4f0-1142-43e1-ac19-a7faab042712 · outbound

This paper cites Laion-5B: An open large-scale dataset for training next generation image-text models.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Laion-5B: An open large-scale dataset for training next generation image-text models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.713015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.270985Z digest=sha256:7c9685812be313302177b4592a363a2067065e2e6155129b3166969a2ef9c675

Observation 3ef15bd6-675a-46e3-ae5b-d66e8b41b6db · outbound

This paper cites ReCo: Re- trieve and co-segment for zero-shot transfer.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation ReCo: Re- trieve and co-segment for zero-shot transfer

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.672270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.288916Z digest=sha256:de927dc2b8475039cf2cfa6365781b6c6654d556e541b23e8fedb07166cb15e3

Observation 76af3956-dcb7-43e3-82b4-dade04414c75 · outbound

This paper cites iSeg: An Iterative Refinement-based Framework for Training-free Segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation iSeg: An Iterative Refinement-based Framework for Training-free Segmentation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:53:32.128129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.299753Z digest=sha256:4b1eccf756bb97ecebb8f3012a73194c558fd57c4e1226520e3fbe2713213e71

Observation ee87535e-4d4b-41ad-8387-f8c20bdb7c73 · outbound

This paper cites CLIP as RNN: Segment countless visual concepts with- out training endeavor.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation CLIP as RNN: Segment countless visual concepts with- out training endeavor

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.634456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.324636Z digest=sha256:84b69843177bdffb708ca5f53ac65b0fe21116348f9bbe999a8b9873fdd2d851

Observation 0e62b1ef-f7ee-410a-9bee-51decc852a1f · outbound

This paper cites Sclip: Rethinking self-attention for dense vision-language inference.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Sclip: Rethinking self-attention for dense vision-language inference

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.574412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.421744Z digest=sha256:b441837d7018684dbea54d8f895ff46437a691354f3ae8400ba0c34adf83067b

Observation 537c9fe7-9180-47fc-beb5-44f643700f30 · outbound

This paper cites Sam-clip: Merging vision foundation models towards semantic and spatial understanding.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Sam-clip: Merging vision foundation models towards semantic and spatial understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.398863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.466108Z digest=sha256:dd5e9e755c2b1fcd34ede54a48b8169268cc2c85e68727d86bbf08336b4011dc

Observation f1e187ce-2dfd-4150-8125-3768a669df13 · outbound

This paper cites Diffusion Model is Secretly a Training-free Open Vocabulary Semantic Segmenter.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Diffusion Model is Secretly a Training-free Open Vocabulary Semantic Segmenter

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:31.546345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:31.546345Z digest=sha256:30f771aca6ca0edfec81425ec6f246da72c56dcdd4ff5c67296708c9b6ac9a12

Observation e021f2c3-832d-4647-886f-e9733f0fac57 · outbound

This paper cites Clip-dinoiser: Teaching clip a few dino tricks for open- vocabulary semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Clip-dinoiser: Teaching clip a few dino tricks for open- vocabulary semantic segmentation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.361545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.678102Z digest=sha256:c2ade24935dfa4b23f375dd24057e13facc124bea4874f5f9c0227fbf29791c5

Observation 5268430a-a04e-4686-bde8-0acc42f79d99 · outbound

This paper cites From Text to Mask: Localizing Entities Using the Attention of Text-to-Image Diffusion Models.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation From Text to Mask: Localizing Entities Using the Attention of Text-to-Image Diffusion Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:53:32.041705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.718299Z digest=sha256:e8c811334dfe2b4ba49d0f3a7e9af7270a3effdc6c745eee1b29d687fcfbba40

Observation 0fd812d0-79ab-49e5-8126-9f64443fd613 · outbound

This paper cites Sed: A simple encoder-decoder for open- vocabulary semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Sed: A simple encoder-decoder for open- vocabulary semantic segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.203098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.738935Z digest=sha256:1d4dc418fa1d16f7ad9077c7746c61d634ed3891a29924cbf84cb80991fcf386

Observation 6818b0d7-7fbb-4083-b73b-4807879449f2 · outbound

This paper cites Alvarez, and Ping Luo.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Alvarez, and Ping Luo

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.169597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.752197Z digest=sha256:cac52ce83e60f3f6b105549266b5348f435945549ca73b18d5fa84bd2f266522

Observation 82e486ea-48b1-475c-bb02-2c9994ba1305 · outbound

This paper cites Clims: Cross language image matching for weakly supervised se- mantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Clims: Cross language image matching for weakly supervised se- mantic segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.990468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.765289Z digest=sha256:e7344d1dc5bd305d5c90dc817d852ce2a55ef99a25ff252b0111c996ace32816

Observation 179b678d-4304-41da-b23a-b38123aa2586 · outbound

This paper cites Rewrite caption semantics: Bridging seman- tic gaps for language-supervised semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Rewrite caption semantics: Bridging seman- tic gaps for language-supervised semantic segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.915948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.782187Z digest=sha256:49278bc6b8cb9d92ccf3bdd8998d6fa1c2952e5131403256777e8c0f461abfc8

Observation aa91ddfa-5545-4106-83cb-1dbe4aaf9647 · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Groupvit: Semantic segmentation emerges from text supervision

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.850323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.798853Z digest=sha256:5a2341909c82b63abe78a574eeaca8f007982e3855fd6fe1187ba34b213393ef

Observation 34810fe9-a51d-4920-83db-c970a2641733 · outbound

This paper cites Learning open-vocabulary seman- tic segmentation models from natural language supervision.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Learning open-vocabulary seman- tic segmentation models from natural language supervision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.809151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.808098Z digest=sha256:0da28aa9961fb3fc805974fd7cd8759e87de7266517f950715e86e7d870b8757

Observation 2f00c189-f6a3-4241-a794-bda1d061773d · outbound

This paper cites Open-vocabulary panoptic segmentation with text-to-image diffusion models.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Open-vocabulary panoptic segmentation with text-to-image diffusion models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.773430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.817712Z digest=sha256:c5dc266b7f8491aee2d3ea286fff46e282f884e9e925121ab7b5a79d38c52ba4

Observation 16223171-313d-4031-a611-a09d431217d6 · outbound

This paper cites Multi-class token transformer for weakly supervised se- mantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Multi-class token transformer for weakly supervised se- mantic segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.726561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.824494Z digest=sha256:2ec1c69c079ea51ea0bea226437c745c3383522b1676b160add10f0d07246397

Observation 0732b187-27e6-4bd9-9ecd-79a5528df144 · outbound

This paper cites Side adapter network for open-vocabulary semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Side adapter network for open-vocabulary semantic segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.660574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.838226Z digest=sha256:a500ccb840d139b40878eb167b018b6746bdb3e87b695d3471d1d605e31d9021

Observation dedbb98a-65d2-424e-b811-32470ae6b714 · outbound

This paper cites Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIP.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIP

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:31.853023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:31.853023Z digest=sha256:83cf666cdff0cdb6aede4470af5bcf7f16f6a7db6f1d7364777094573f759003

Observation 91ee6cf4-fc84-47f8-bc5f-272acd54846b · outbound

This paper cites Open vocabulary scene parsing.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Open vocabulary scene parsing

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.616043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.865201Z digest=sha256:78858f7f2edeb836881e2660d4e7f9a3b72b7b87d5bf737c8330d38018898272

Observation c5655f20-8b1e-4db0-8adb-12423fbf6252 · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Semantic under- standing of scenes through the ade20k dataset

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:31.878658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:31.878658Z digest=sha256:505f7975869a41132522a972eabd6f96e978911d1f6838aa1e2d930ba649383d

Observation 3241dc5e-106a-4fec-82df-4def85422460 · outbound

This paper cites Extract free dense labels from clip.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Extract free dense labels from clip

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.461065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T15:53:31.889301Z digest=sha256:6ccac07776d6eb2bfe9bc90b23af2d03a08d13d24eb11af3c40e884250d8d04d

Pith citing papers

Observation a7d5046b-d83a-4772-9b7f-346ca2037fe4 · inbound

CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic Segmentation cites this paper.

CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic Segmentation CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T20:04:44.106636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:04:44.106636Z digest=sha256:be35a92335b01dfdbfd7a9bdde675c70356475f9519cf99d9da51210e8108823

Observation 1fee426a-3576-4ef6-b5f8-473170d116dc · inbound

Perception Encoder: The best visual embeddings are not at the output of the network cites this paper.

Perception Encoder: The best visual embeddings are not at the output of the network CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:21:15.798787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T22:21:15.681336Z digest=sha256:8566fac06f7feeddc7f32c1a6f9563c86e61211627f5a839ef5c6576d3a3de5f