Pith. sign in

Paper Citation Record · LEDGER

Fine-grained CLIP fine-tuning with self-annotated region alignment

As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2607.13661.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13661 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T04:34:24.068813Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8bd193dc-b30d-4078-8177-d7083e37a48e · outbound

This paper cites Learning transferable visual models from natural language supervision.

Fine-grained CLIP fine-tuning with self-annotated region alignment Learning transferable visual models from natural language supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:19.732075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:19.732075Z digest=sha256:854e0db86a3746fe80801e4c14504d850ebf988fb84589e6a9c0e42f2f46de56

Observation 63d26381-afdc-4909-bda2-548169a26959 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Fine-grained CLIP fine-tuning with self-annotated region alignment Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:19.783657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:19.783657Z digest=sha256:004be4dd18983241b78c751d717c1edf5b2e8ca450d96e1b109b8a40e43e967d

Observation ddb9cc4e-543c-423d-991f-478e2ee00029 · outbound

This paper cites Domain generalization by mutual-information regularization with pre-trained models.

Fine-grained CLIP fine-tuning with self-annotated region alignment Domain generalization by mutual-information regularization with pre-trained models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:19.911115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:19.911115Z digest=sha256:d70076dab69a4c67d990f1afcadcb24ddd5ac7b498d18f5c139f75692205ca94

Observation 3c0642e0-1d44-4817-adae-3bde31c03eab · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.Neurocomputing, 508:293–304, 2022.

Fine-grained CLIP fine-tuning with self-annotated region alignment Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.Neurocomputing, 508:293–304, 2022

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.028038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.028038Z digest=sha256:40715e6145c38a28900a9d1a8049d61a7720bc0238c4c0552f62b31489bf0995

Observation 32ba5634-332f-4950-a23b-8ee151898524 · outbound

This paper cites A simple baseline for zeroshot semantic segmentation with pre-trained vision-language model.ECCV, 2022.

Fine-grained CLIP fine-tuning with self-annotated region alignment A simple baseline for zeroshot semantic segmentation with pre-trained vision-language model.ECCV, 2022

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.137894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.137894Z digest=sha256:45c000eedb0945903a5b4793369416d221dbfb7bb823d14e8ae8dd2bac77929c

Observation 6de4a91c-378c-4c3d-9f61-982631c05f86 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

Fine-grained CLIP fine-tuning with self-annotated region alignment Open-vocabulary semantic segmentation with mask-adapted clip

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.267337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.267337Z digest=sha256:77fa8ef2a24ad354794bed0e7c86fb8f6a8051cd03a832b9f1d8cc2dcd490c67

Observation 11bf04f8-b573-4897-8467-4c55297b54c3 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.ICLR, 2022.

Fine-grained CLIP fine-tuning with self-annotated region alignment Open-vocabulary object detection via vision and language knowledge distillation.ICLR, 2022

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.395029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.395029Z digest=sha256:ba693f72ea84756c2520b614a99ea891ce8d27096f4fb03dc722b4bbee923a4f

Observation 1b1ce2ea-0b29-4882-b758-c8d3ec22c2e1 · outbound

This paper cites Learning to prompt for open-vocabulary object detection with vision-language model.

Fine-grained CLIP fine-tuning with self-annotated region alignment Learning to prompt for open-vocabulary object detection with vision-language model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.522685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.522685Z digest=sha256:f8c9e862b3d6a2bfa82724482cda68e90adc497c787201c4c1ba1715ce80efcd

Observation 565de053-86d3-4a2e-8d37-626adad563dc · outbound

This paper cites F-vlm: Open-vocabulary object detection upon frozen vision and language models.

Fine-grained CLIP fine-tuning with self-annotated region alignment F-vlm: Open-vocabulary object detection upon frozen vision and language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.598985Z digest=sha256:06d8f0f9b1543403fd877153b9d77be85ae4e8826d20b370adfe8d467d3d2bac

Observation b8dcd1f4-4855-4adf-963d-405c14e8163e · outbound

This paper cites Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching.

Fine-grained CLIP fine-tuning with self-annotated region alignment Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.710137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.710137Z digest=sha256:e98732e4cf51990070a746707660e879eff71154827ee1c1c2ec4f2f95ce1e70

Observation d6a36991-b888-41d2-8104-2e835b78156e · outbound

This paper cites Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip.NeurIPS, 36, 2024.

Fine-grained CLIP fine-tuning with self-annotated region alignment Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip.NeurIPS, 36, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.784279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.784279Z digest=sha256:aad9910a019d314dae9216d60309e40269719ef10ae1def70b1e4c5845c6899a

Observation eb9e5c88-b18c-48dd-b5fb-419445fe7bb6 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Fine-grained CLIP fine-tuning with self-annotated region alignment SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.882417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.882417Z digest=sha256:5542c88ffc730039b20ae3f00c73596e569b351753005cac4b685044d011f4ae

Observation 35d417ee-a6ef-4c79-9ebc-15a5b96c56dd · outbound

This paper cites Fg-clip: Fine-grained visual and textual alignment.

Fine-grained CLIP fine-tuning with self-annotated region alignment Fg-clip: Fine-grained visual and textual alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.994264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.994264Z digest=sha256:423a840cea79b180bff72cb210952809a452cf7800ee18037c40b7e38f0cb0c6

Observation 4025cb0e-223d-4a04-9a14-b3b58f99fea0 · outbound

This paper cites Improving fine-grained understanding in image-text pre-training.

Fine-grained CLIP fine-tuning with self-annotated region alignment Improving fine-grained understanding in image-text pre-training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.067883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.067883Z digest=sha256:017094d90a03afbaef2291aaaf5817f62ac0e2ad62949ca98a723a18ec47e57d

Observation 77529e6f-7493-4b61-9f75-161c2dccf433 · outbound

This paper cites Position-guided text prompt for vision-language pre-training.

Fine-grained CLIP fine-tuning with self-annotated region alignment Position-guided text prompt for vision-language pre-training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.144614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.144614Z digest=sha256:8a5603e7869102d8c6ab0acbc5aaac92ace109acbc78a70e8eb13d58b72e3e78

Observation 6a4a53db-8026-4009-8207-31fb8c29c6fe · outbound

This paper cites Regionclip: Region-based language-image pretraining.

Fine-grained CLIP fine-tuning with self-annotated region alignment Regionclip: Region-based language-image pretraining

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.272963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.272963Z digest=sha256:ae3ab23add657ef97f9db1948ebbccbe85a7e8ac58489ce9e3071a021292d318

Observation f74a3a0a-d803-456f-874a-7f819ad08ca3 · outbound

This paper cites Densevlm: A retrieval and decoupled alignment framework for open-vocabulary dense prediction.

Fine-grained CLIP fine-tuning with self-annotated region alignment Densevlm: A retrieval and decoupled alignment framework for open-vocabulary dense prediction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.399489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.399489Z digest=sha256:befb52d5ba36929f9e1c5ed104322ed0f3a9c0bfb942cf2b7b094bc6c16c68e6

Observation 2fc4abb7-f457-4d88-90c9-7a70dce8e874 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.NeurIPS, 28, 2015.

Fine-grained CLIP fine-tuning with self-annotated region alignment Faster r-cnn: Towards real-time object detection with region proposal networks.NeurIPS, 28, 2015

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.508614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.508614Z digest=sha256:dcab42ccd7af6a1399cc105afba9d0b44d09555a08d989b1637486741b8dc99b

Observation 32a894cb-2843-4da8-a076-3556dcf10c95 · outbound

This paper cites Fineclip: Self-distilled region-based clip for better fine-grained understanding.Advances in Neural Information Processing Systems, 37:27896–27918, 2024.

Fine-grained CLIP fine-tuning with self-annotated region alignment Fineclip: Self-distilled region-based clip for better fine-grained understanding.Advances in Neural Information Processing Systems, 37:27896–27918, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.610915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.610915Z digest=sha256:9be2f4d26aeafdc83e93df532fbf8a8e9061224b4902ddde2bea0aff7ce5d4ff

Observation c1be3bf9-5335-4ab8-b12e-af7abd41c793 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Fine-grained CLIP fine-tuning with self-annotated region alignment Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.688429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.688429Z digest=sha256:75334bc97f91bd68e9d3c66e73cbb1b355c7581c84fd643ff80365fc4ce7cee7

Observation f86e41b0-c080-4a67-b931-7a6d6092232a · outbound

This paper cites CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction.

Fine-grained CLIP fine-tuning with self-annotated region alignment CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.757260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.757260Z digest=sha256:1ee4b670ece7bb476c44b8db2cdb00b31d4566e79d06da416e33376bd52baef1

Observation 6fea05e2-1c49-4392-903e-2e2ec146ea25 · outbound

This paper cites Atas: Any-to-any self-distillation for enhanced open-vocabulary dense prediction.

Fine-grained CLIP fine-tuning with self-annotated region alignment Atas: Any-to-any self-distillation for enhanced open-vocabulary dense prediction

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.837936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.837936Z digest=sha256:bd52a890ebb4b8d45c5b29b57a8fe3f4ce0c0d39388796decb0137ca55586306

Observation e2969ff4-6d48-4357-b08b-10d160aea840 · outbound

This paper cites Clim: Contrastive language-image mosaic for region representation.

Fine-grained CLIP fine-tuning with self-annotated region alignment Clim: Contrastive language-image mosaic for region representation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.960049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.960049Z digest=sha256:711de5a072786c0e7e69a3e2de7f67a05e75c410590e44c160542ad55b2ebe0a

Observation 2ec8a351-5b2b-41af-a3d7-29ee243695a9 · outbound

This paper cites Learning to prompt for vision-language models.

Fine-grained CLIP fine-tuning with self-annotated region alignment Learning to prompt for vision-language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.052633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.052633Z digest=sha256:4465816447c72c7f635cf44429e0e1fbaf833632d8d2f625f493fd58c799a61e

Observation f34aa329-6d56-427e-be60-15064d712336 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Fine-grained CLIP fine-tuning with self-annotated region alignment CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.167250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.167250Z digest=sha256:2460e1cac6e27bf71f52ecc9da723fffbf4031705b0d0429c60103a96b83bc67

Observation 0818cb07-8563-4be6-83b6-66b9ae21c98c · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Fine-grained CLIP fine-tuning with self-annotated region alignment Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.240787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.240787Z digest=sha256:328f9e86b0a36bcfe9ab33ba3f4c898914d170b8e1c2f8c7f3749ab3cf22ced8

Observation ea92005a-c474-4d63-8c91-8cf61cc49883 · outbound

This paper cites Region-aware pretraining for open-vocabulary object detection with vision transformers.

Fine-grained CLIP fine-tuning with self-annotated region alignment Region-aware pretraining for open-vocabulary object detection with vision transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.348171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.348171Z digest=sha256:afe6c27a39062715a887d3c591c26776051769872592cf3cca7257e491809fb3

Observation b348c575-dc87-487b-b4c1-837e80beb3cb · outbound

This paper cites Open-vocabulary object detection upon frozen vision and language models.

Fine-grained CLIP fine-tuning with self-annotated region alignment Open-vocabulary object detection upon frozen vision and language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.421378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.421378Z digest=sha256:381df4d36ed48a755d002902fd241d47afa3f6a19f4d42eb2c8980084153ddca

Observation ee5b9da2-0c64-4109-bc6c-6afea2ff4193 · outbound

This paper cites Simple open-vocabulary object detection.

Fine-grained CLIP fine-tuning with self-annotated region alignment Simple open-vocabulary object detection

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.482490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.482490Z digest=sha256:c16bae7981c340b57d91271c829787034896b5c1154962c072f7e0af6cb76b79

Observation 461b57f4-e6e2-4136-8ece-e49634b213ad · outbound

This paper cites Scaling open-vocabulary image segmentation with image-level labels.

Fine-grained CLIP fine-tuning with self-annotated region alignment Scaling open-vocabulary image segmentation with image-level labels

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.551463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.551463Z digest=sha256:8f8c9dc64b0207836e9127667108947e6a7cf66bd2e16c85686da94b412ff89b

Observation 597862d6-c43b-4b4e-8c4d-21297c1af748 · outbound

This paper cites Cascade-clip: Cascaded vision-language embeddings alignment for zero-shot semantic segmentation.

Fine-grained CLIP fine-tuning with self-annotated region alignment Cascade-clip: Cascaded vision-language embeddings alignment for zero-shot semantic segmentation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.605796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.605796Z digest=sha256:36ce19df200f2c291e040028313ab417463faf6f372ea995816cddee39aeb635

Observation 0dddf0e2-4cd3-4e4c-8838-73e619e389cc · outbound

This paper cites Visualizing and understanding convolutional networks.

Fine-grained CLIP fine-tuning with self-annotated region alignment Visualizing and understanding convolutional networks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.666618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.666618Z digest=sha256:8c89870bda2aef10b6d3eba47b2f4579d7f7309f90a881ea1366141b0f50ce58

Observation fe6d806c-2905-435d-8059-6219fb50845b · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

Fine-grained CLIP fine-tuning with self-annotated region alignment Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.716689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.716689Z digest=sha256:b8e474f9680f8e05c0d43d56ac7fb75cbfb17d467f01e0b7347f4e0c368520d1

Observation 30586115-75d7-4000-bae3-65275a1d9758 · outbound

This paper cites why should i trust you?.

Fine-grained CLIP fine-tuning with self-annotated region alignment why should i trust you?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.774483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.774483Z digest=sha256:80c38f230add6a2d7cc1af55f4b01858b4932e49b6c2ff0b219d5a86c57e42c0

Observation 515bce79-e3d9-414a-8667-04882cf09d0a · outbound

This paper cites RISE: Randomized Input Sampling for Explanation of Black-box Models.

Fine-grained CLIP fine-tuning with self-annotated region alignment RISE: Randomized Input Sampling for Explanation of Black-box Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.815449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.815449Z digest=sha256:6d3347c86e4e5f20aacc7666f4483e6dd99d0c8a6aaea5ffd65ee50ed5c5130c

Observation 9a1c5731-fc8d-4aec-9dff-2fcb024282b4 · outbound

This paper cites Odam: Gradient-based instance-specific visual explanations for object detection.

Fine-grained CLIP fine-tuning with self-annotated region alignment Odam: Gradient-based instance-specific visual explanations for object detection

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.875888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.875888Z digest=sha256:d38ccff8152ddc70049077d75f58ee460df6f3fe49fa4c41d602a15dbfd914a0

Observation 600f2382-940b-4e1c-99e6-583c91c474da · outbound

This paper cites Attcat: Explaining transformers via attentive class activation tokens.NeurIPS, 2022.

Fine-grained CLIP fine-tuning with self-annotated region alignment Attcat: Explaining transformers via attentive class activation tokens.NeurIPS, 2022

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.926313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.926313Z digest=sha256:f4b9a9ddc2f40a849b2d5ba2602f222780c98349dd4a9c7e4f4cbd0a0d204650

Observation 2fcd5a22-9bb6-4959-906a-6f43fa45edc7 · outbound

This paper cites Vit-cx: Causal explanation of vision transformers.

Fine-grained CLIP fine-tuning with self-annotated region alignment Vit-cx: Causal explanation of vision transformers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.978761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.978761Z digest=sha256:39457ebcf4d70faaba46465e53c9436ae6b7d7a11271267898ac7616aaf947f5

Observation 426aea57-7706-4963-a6ad-75770ba38d26 · outbound

This paper cites X-pruner: explainable pruning for vision transformers.

Fine-grained CLIP fine-tuning with self-annotated region alignment X-pruner: explainable pruning for vision transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.027755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.027755Z digest=sha256:7291a219a5e384b14020442446f45d5f8982daafe30a279b7b132b8d3cd7030e

Observation 444c21dc-5c62-448f-9e75-d2e059dc324b · outbound

This paper cites Gradient-based visual explanation for transformer-based clip.

Fine-grained CLIP fine-tuning with self-annotated region alignment Gradient-based visual explanation for transformer-based clip

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.086757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.086757Z digest=sha256:6ba80e986da09f804e096081db5fa39054f46481790a8cf9e526a5eb2facfcf2

Observation c9b6e25d-0459-4d2d-a378-0d1fce8d4c8d · outbound

This paper cites Extract free dense labels from clip.

Fine-grained CLIP fine-tuning with self-annotated region alignment Extract free dense labels from clip

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.130648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.130648Z digest=sha256:6e38babb6f56b547fc9f0e6f3459f0f2386768a64527199687a997b279f78d0e

Observation 94584bb7-1049-4471-ae94-be1cd70f3a1b · outbound

This paper cites O’Reilly Media, Inc., 2009.

Fine-grained CLIP fine-tuning with self-annotated region alignment O’Reilly Media, Inc., 2009

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.187146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.187146Z digest=sha256:00a541ec9d9f564230e9ae3d2d04ab8de2ced7d1c8e2779244a4ff0cbfb98851

Observation 99c64763-8baf-433a-992d-4803d6f90527 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Fine-grained CLIP fine-tuning with self-annotated region alignment EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.228619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.228619Z digest=sha256:2cdb51b20b11195648faaca7e06959744c280b2370e4c6f7b05ac11ff6537889

Observation ff33e22f-c38a-4388-a02e-9cb2916b1545 · outbound

This paper cites Microsoft coco: Common objects in context.

Fine-grained CLIP fine-tuning with self-annotated region alignment Microsoft coco: Common objects in context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.320295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.320295Z digest=sha256:89de9c2c3cc0ab1cc3754e93ac33e14cc6d8adc2f312cd04af5ec79134febea7

Observation 1cda5196-ea45-4be9-b5f0-450626689087 · outbound

This paper cites Scene parsing through ade20k dataset.

Fine-grained CLIP fine-tuning with self-annotated region alignment Scene parsing through ade20k dataset

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.367081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.367081Z digest=sha256:ae43b52fa323a93c25497638040341e7324ecea3ed57918ee257a659cf170e4a

Observation 7ec207f7-1166-4fec-b74b-41a90a1cb363 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Fine-grained CLIP fine-tuning with self-annotated region alignment Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.413004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.413004Z digest=sha256:7182ce8b74131f54fc559db5a7efbf989bf73006078d8264437d175deb90606f

Observation 60a84edb-86d0-41fd-ba53-cab7dbf6fd26 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

Fine-grained CLIP fine-tuning with self-annotated region alignment Lvis: A dataset for large vocabulary instance segmentation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.484746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.484746Z digest=sha256:7c749d1c69a79af76715f47830941224f21cf3638d7d1dbe0aef5e6a20fa1d09

Observation 3aedc7db-11ac-459e-878e-0cb683dd086c · outbound

This paper cites Open-vocabulary object detection using captions.

Fine-grained CLIP fine-tuning with self-annotated region alignment Open-vocabulary object detection using captions

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.564582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.564582Z digest=sha256:33c2aa3a7bda7f1edb8d5fb9dae65ed2555264f69a9e78a9201e71d32f4921bf

Observation 37a13a99-87f3-4a96-8828-89617d4ca713 · outbound

This paper cites Detecting twenty-thousand classes using image-level supervision.

Fine-grained CLIP fine-tuning with self-annotated region alignment Detecting twenty-thousand classes using image-level supervision

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.619581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.619581Z digest=sha256:5e8ebcf62ccc9907ded1621c0d12a0aed74897ef4536f766a40c7b4e9e2b0111

Observation 70f1f7f9-eba6-42e5-9a90-bbe7c2f29837 · outbound

This paper cites Learning Object-Language Alignments for Open-Vocabulary Object Detection.

Fine-grained CLIP fine-tuning with self-annotated region alignment Learning Object-Language Alignments for Open-Vocabulary Object Detection

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.697318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.697318Z digest=sha256:fcce69318fddd9355be6687afc3accf50dacdd334de65d4eff65fc283eddbb00

Observation 4015e22e-3554-412d-a9ef-d73385683845 · outbound

This paper cites Everingham, L.

Fine-grained CLIP fine-tuning with self-annotated region alignment Everingham, L

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.775181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.775181Z digest=sha256:3cd4fdb543f45a7039c9c7435fe2cccff1df94d76ac888de977d3294591a4c9a

Observation 5291320a-5457-4b97-80fa-605b2501cc42 · outbound

This paper cites The role of context for object detection and semantic segmentation in the wild.

Fine-grained CLIP fine-tuning with self-annotated region alignment The role of context for object detection and semantic segmentation in the wild

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.855596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.855596Z digest=sha256:26541b76179fb284d396a934620120958c8753d938a5238c0eaaae51cb42558e

Observation d2d19235-5285-41e9-9ae4-545f4796b220 · outbound

This paper cites Cat-seg: Cost aggregation for open-vocabulary semantic segmentation.

Fine-grained CLIP fine-tuning with self-annotated region alignment Cat-seg: Cost aggregation for open-vocabulary semantic segmentation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.912590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.912590Z digest=sha256:59e240b1ef91a403793732cd4063004cff957ac4153988edff6c089b6f955b32

Observation deaab706-55ba-4d15-928f-54ad94fe9b78 · outbound

This paper cites Coco-stuff: Thing and stuff classes in context.

Fine-grained CLIP fine-tuning with self-annotated region alignment Coco-stuff: Thing and stuff classes in context

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.991946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.991946Z digest=sha256:66b4541d16d18ba61b22549ee85f8718a1e04ff5da56433d206148d10a6ee653

Observation dbedda91-2484-4bca-b016-c1f364cb5091 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models.

Fine-grained CLIP fine-tuning with self-annotated region alignment Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:24.068813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:24.068813Z digest=sha256:8c8472559c222a5e5262ef142d67f3cf89ca77f1c8bd7a4c3b28d539a014171b

Pith citing papers

No inbound Pith citation observations are available.