Pith. sign in

Paper Citation Record · LEDGER

Fine-grained CLIP fine-tuning with self-annotated region alignment

As of 14 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2607.13661.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13661 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T04:34:24.068813Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8bd193dc-b30d-4078-8177-d7083e37a48e · outbound

This paper cites Learning transferable visual models from natural language supervision.

Fine-grained CLIP fine-tuning with self-annotated region alignment Learning transferable visual models from natural language supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:19.732075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:19.732075Z digest=sha256:117f885221d0e521a456764dfc1563b5d714f0f9bfb58122e0f578c0f3c9b325

Observation 63d26381-afdc-4909-bda2-548169a26959 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Fine-grained CLIP fine-tuning with self-annotated region alignment Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:19.783657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:19.783657Z digest=sha256:9585862a4e84f5194af87e2025a3d0a92cec75956597c66ccd0d6654f90781ef

Observation ddb9cc4e-543c-423d-991f-478e2ee00029 · outbound

This paper cites Domain generalization by mutual-information regularization with pre-trained models.

Fine-grained CLIP fine-tuning with self-annotated region alignment Domain generalization by mutual-information regularization with pre-trained models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:19.911115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:19.911115Z digest=sha256:734e7357c1bb876901dd5c535524b385de40ba542361f8795b802c04a92e7d8d

Observation 3c0642e0-1d44-4817-adae-3bde31c03eab · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.Neurocomputing, 508:293–304, 2022.

Fine-grained CLIP fine-tuning with self-annotated region alignment Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.Neurocomputing, 508:293–304, 2022

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.028038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.028038Z digest=sha256:e112f520b7ada3a406680d21776e3f78a92847e3a3f10153b0ae842e351eac45

Observation 32ba5634-332f-4950-a23b-8ee151898524 · outbound

This paper cites A simple baseline for zeroshot semantic segmentation with pre-trained vision-language model.ECCV, 2022.

Fine-grained CLIP fine-tuning with self-annotated region alignment A simple baseline for zeroshot semantic segmentation with pre-trained vision-language model.ECCV, 2022

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.137894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.137894Z digest=sha256:959559cac4215c14a7845416b487e61883717b9a573b59956e15e3511b8db779

Observation 6de4a91c-378c-4c3d-9f61-982631c05f86 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

Fine-grained CLIP fine-tuning with self-annotated region alignment Open-vocabulary semantic segmentation with mask-adapted clip

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.267337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.267337Z digest=sha256:dbf37ac0f099679557a94ed5c2add11c09363d7109418760571a05c1dbb1e3c4

Observation 11bf04f8-b573-4897-8467-4c55297b54c3 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.ICLR, 2022.

Fine-grained CLIP fine-tuning with self-annotated region alignment Open-vocabulary object detection via vision and language knowledge distillation.ICLR, 2022

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.395029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.395029Z digest=sha256:6363a2209fb8ce1ca946bdf0df1e1ab1322193d0f46377ac403f590b489500e4

Observation 1b1ce2ea-0b29-4882-b758-c8d3ec22c2e1 · outbound

This paper cites Learning to prompt for open-vocabulary object detection with vision-language model.

Fine-grained CLIP fine-tuning with self-annotated region alignment Learning to prompt for open-vocabulary object detection with vision-language model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.522685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.522685Z digest=sha256:8016352172b2b0f5beddadaf07556d7185f0b408d3fbdfef967d71a8ff110f21

Observation 565de053-86d3-4a2e-8d37-626adad563dc · outbound

This paper cites F-vlm: Open-vocabulary object detection upon frozen vision and language models.

Fine-grained CLIP fine-tuning with self-annotated region alignment F-vlm: Open-vocabulary object detection upon frozen vision and language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.598985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.598985Z digest=sha256:36ed62bd47d76186d777f452c2be9746d7a289d2fc7b2b8c1e6e08bbd83e123b

Observation b8dcd1f4-4855-4adf-963d-405c14e8163e · outbound

This paper cites Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching.

Fine-grained CLIP fine-tuning with self-annotated region alignment Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.710137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.710137Z digest=sha256:409df954a941a81c3d906e80340d950fdd6b722365af4ae1b6b1dae3313ca20d

Observation d6a36991-b888-41d2-8104-2e835b78156e · outbound

This paper cites Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip.NeurIPS, 36, 2024.

Fine-grained CLIP fine-tuning with self-annotated region alignment Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip.NeurIPS, 36, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.784279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.784279Z digest=sha256:910e2c8ea1e5a99b68611776988823c19c4ffbf7dc1f3ad2a5bd19c1c029174f

Observation eb9e5c88-b18c-48dd-b5fb-419445fe7bb6 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Fine-grained CLIP fine-tuning with self-annotated region alignment SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.882417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.882417Z digest=sha256:e5af85ba2acf364a08c871ca3e2486466b895fc6d022923774174c346adbed6e

Observation 35d417ee-a6ef-4c79-9ebc-15a5b96c56dd · outbound

This paper cites Fg-clip: Fine-grained visual and textual alignment.

Fine-grained CLIP fine-tuning with self-annotated region alignment Fg-clip: Fine-grained visual and textual alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:20.994264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:20.994264Z digest=sha256:253d7ed23be1ada71df5b64237929daec0700edd44ff4f4a4d9e54624baf5806

Observation 4025cb0e-223d-4a04-9a14-b3b58f99fea0 · outbound

This paper cites Improving fine-grained understanding in image-text pre-training.

Fine-grained CLIP fine-tuning with self-annotated region alignment Improving fine-grained understanding in image-text pre-training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.067883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.067883Z digest=sha256:dd868bf504027fbc6d6bbf23ceb652464ba6e5144a6cffb150715915fa4b38d6

Observation 77529e6f-7493-4b61-9f75-161c2dccf433 · outbound

This paper cites Position-guided text prompt for vision-language pre-training.

Fine-grained CLIP fine-tuning with self-annotated region alignment Position-guided text prompt for vision-language pre-training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.144614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.144614Z digest=sha256:5cff36e995efc65bb0e4bf4ca5ecf3011e20b80d9db56307cb34b489d4165172

Observation 6a4a53db-8026-4009-8207-31fb8c29c6fe · outbound

This paper cites Regionclip: Region-based language-image pretraining.

Fine-grained CLIP fine-tuning with self-annotated region alignment Regionclip: Region-based language-image pretraining

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.272963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.272963Z digest=sha256:47b792353583506b13d7bf9bc99c115c4fe8fb9490b77d06608edb9497be36f8

Observation f74a3a0a-d803-456f-874a-7f819ad08ca3 · outbound

This paper cites Densevlm: A retrieval and decoupled alignment framework for open-vocabulary dense prediction.

Fine-grained CLIP fine-tuning with self-annotated region alignment Densevlm: A retrieval and decoupled alignment framework for open-vocabulary dense prediction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.399489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.399489Z digest=sha256:73c51341ae2d459974ba0d78b76984784acd16e113f74f22c3c4432dad9ddafc

Observation 2fc4abb7-f457-4d88-90c9-7a70dce8e874 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.NeurIPS, 28, 2015.

Fine-grained CLIP fine-tuning with self-annotated region alignment Faster r-cnn: Towards real-time object detection with region proposal networks.NeurIPS, 28, 2015

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.508614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.508614Z digest=sha256:a3a93049f8193420abf202b72e3229a311fcb136c0a8a960e98afd5d94c5709f

Observation 32a894cb-2843-4da8-a076-3556dcf10c95 · outbound

This paper cites Fineclip: Self-distilled region-based clip for better fine-grained understanding.Advances in Neural Information Processing Systems, 37:27896–27918, 2024.

Fine-grained CLIP fine-tuning with self-annotated region alignment Fineclip: Self-distilled region-based clip for better fine-grained understanding.Advances in Neural Information Processing Systems, 37:27896–27918, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.610915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.610915Z digest=sha256:7de999a4a374ec2e3f610771d2d99ae9a73602df8f7e08636a104ca9fd806ee4

Observation c1be3bf9-5335-4ab8-b12e-af7abd41c793 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Fine-grained CLIP fine-tuning with self-annotated region alignment Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.688429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.688429Z digest=sha256:3a68190b16a50068c132b51a722a3c2a417a78474a72cf3902206962e2ce5434

Observation f86e41b0-c080-4a67-b931-7a6d6092232a · outbound

This paper cites CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction.

Fine-grained CLIP fine-tuning with self-annotated region alignment CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.757260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.757260Z digest=sha256:1ef91d707dc4a4ef7b4d7ceeeddc45506b5080b1e4886ee33b2e3f5b8c5713b2

Observation 6fea05e2-1c49-4392-903e-2e2ec146ea25 · outbound

This paper cites Atas: Any-to-any self-distillation for enhanced open-vocabulary dense prediction.

Fine-grained CLIP fine-tuning with self-annotated region alignment Atas: Any-to-any self-distillation for enhanced open-vocabulary dense prediction

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.837936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.837936Z digest=sha256:6c864c0fef2528c615463331608ddd93469d93af2c5fd04bc8e735fd479872b1

Observation e2969ff4-6d48-4357-b08b-10d160aea840 · outbound

This paper cites Clim: Contrastive language-image mosaic for region representation.

Fine-grained CLIP fine-tuning with self-annotated region alignment Clim: Contrastive language-image mosaic for region representation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:21.960049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:21.960049Z digest=sha256:7e9a2569ba02542cfdea29c5d56a70cbf068d4c53c004c5354858d371398c78b

Observation 2ec8a351-5b2b-41af-a3d7-29ee243695a9 · outbound

This paper cites Learning to prompt for vision-language models.

Fine-grained CLIP fine-tuning with self-annotated region alignment Learning to prompt for vision-language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.052633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.052633Z digest=sha256:3a66c6702feccab4b45679f4a90c96d8b4a541bad71b50be5b70e527317dc33b

Observation f34aa329-6d56-427e-be60-15064d712336 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Fine-grained CLIP fine-tuning with self-annotated region alignment CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.167250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.167250Z digest=sha256:a1626b6a4c91e32280c4fbc8b7dce752895fdfaa59ead41ec1bd7673cd790cbf

Observation 0818cb07-8563-4be6-83b6-66b9ae21c98c · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Fine-grained CLIP fine-tuning with self-annotated region alignment Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.240787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.240787Z digest=sha256:4651515344814592fdb7e2626d03363d796f3ec46ad60fb718ad951559f61534

Observation ea92005a-c474-4d63-8c91-8cf61cc49883 · outbound

This paper cites Region-aware pretraining for open-vocabulary object detection with vision transformers.

Fine-grained CLIP fine-tuning with self-annotated region alignment Region-aware pretraining for open-vocabulary object detection with vision transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.348171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.348171Z digest=sha256:9fad5d4e9f74afaec618d5a0c4d17c9b4e63a8e94132570adb34cb7e99765746

Observation b348c575-dc87-487b-b4c1-837e80beb3cb · outbound

This paper cites Open-vocabulary object detection upon frozen vision and language models.

Fine-grained CLIP fine-tuning with self-annotated region alignment Open-vocabulary object detection upon frozen vision and language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.421378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.421378Z digest=sha256:e9988dd602ae960037d7bd4ac562cbf8b5bc5d522f8b70b312a3385f09df307e

Observation ee5b9da2-0c64-4109-bc6c-6afea2ff4193 · outbound

This paper cites Simple open-vocabulary object detection.

Fine-grained CLIP fine-tuning with self-annotated region alignment Simple open-vocabulary object detection

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.482490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.482490Z digest=sha256:e08628d6924fd42f7823f64ec7cc90dd08e0e5ae5af5a371d6abb81980eb4910

Observation 461b57f4-e6e2-4136-8ece-e49634b213ad · outbound

This paper cites Scaling open-vocabulary image segmentation with image-level labels.

Fine-grained CLIP fine-tuning with self-annotated region alignment Scaling open-vocabulary image segmentation with image-level labels

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.551463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.551463Z digest=sha256:77e4e0d9186494467b23583da000a6abeedec50155386d3c1733d2ee54ad9895

Observation 597862d6-c43b-4b4e-8c4d-21297c1af748 · outbound

This paper cites Cascade-clip: Cascaded vision-language embeddings alignment for zero-shot semantic segmentation.

Fine-grained CLIP fine-tuning with self-annotated region alignment Cascade-clip: Cascaded vision-language embeddings alignment for zero-shot semantic segmentation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.605796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.605796Z digest=sha256:f75d591c99290cbb61e7ef7fff9c834e59b1dc4436ac02d4ec949d17815433ff

Observation 0dddf0e2-4cd3-4e4c-8838-73e619e389cc · outbound

This paper cites Visualizing and understanding convolutional networks.

Fine-grained CLIP fine-tuning with self-annotated region alignment Visualizing and understanding convolutional networks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.666618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.666618Z digest=sha256:c7a45e8a2bec8cc03eb3375241c09d6fbe939924488a5d60576af865c758f137

Observation fe6d806c-2905-435d-8059-6219fb50845b · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

Fine-grained CLIP fine-tuning with self-annotated region alignment Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.716689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.716689Z digest=sha256:74023ccb053534b439546d04fec2358fd427361f6a31d99a03d3b5cf00162b94

Observation 30586115-75d7-4000-bae3-65275a1d9758 · outbound

This paper cites why should i trust you?.

Fine-grained CLIP fine-tuning with self-annotated region alignment why should i trust you?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.774483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.774483Z digest=sha256:1fb327c43179bd1fcb662ee9b8597f8b6fd0bc8b83f0f9b2ff841666ecf7c9d2

Observation 515bce79-e3d9-414a-8667-04882cf09d0a · outbound

This paper cites RISE: Randomized Input Sampling for Explanation of Black-box Models.

Fine-grained CLIP fine-tuning with self-annotated region alignment RISE: Randomized Input Sampling for Explanation of Black-box Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.815449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.815449Z digest=sha256:0317c9604eaf9edfbb27ae7739a49cd30ac6598f88d6691f5c6922f6a0952f86

Observation 9a1c5731-fc8d-4aec-9dff-2fcb024282b4 · outbound

This paper cites Odam: Gradient-based instance-specific visual explanations for object detection.

Fine-grained CLIP fine-tuning with self-annotated region alignment Odam: Gradient-based instance-specific visual explanations for object detection

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.875888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.875888Z digest=sha256:c624ee991a1e2cf5f18e8f53fff2e37746105886284b1047d42db10f9f3941ef

Observation 600f2382-940b-4e1c-99e6-583c91c474da · outbound

This paper cites Attcat: Explaining transformers via attentive class activation tokens.NeurIPS, 2022.

Fine-grained CLIP fine-tuning with self-annotated region alignment Attcat: Explaining transformers via attentive class activation tokens.NeurIPS, 2022

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.926313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.926313Z digest=sha256:898a76299626d6810945487f7038240722793a970de3ce605538116305650d11

Observation 2fcd5a22-9bb6-4959-906a-6f43fa45edc7 · outbound

This paper cites Vit-cx: Causal explanation of vision transformers.

Fine-grained CLIP fine-tuning with self-annotated region alignment Vit-cx: Causal explanation of vision transformers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:22.978761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:22.978761Z digest=sha256:b8048ec54d00a314e71eb97a3611c282ba70e65eeee4ac2a4d9998a78d378ef3

Observation 426aea57-7706-4963-a6ad-75770ba38d26 · outbound

This paper cites X-pruner: explainable pruning for vision transformers.

Fine-grained CLIP fine-tuning with self-annotated region alignment X-pruner: explainable pruning for vision transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.027755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.027755Z digest=sha256:8f9fc6102caa9e2f143f71291ae6353d203c43eb04876d59217c9559d3a42106

Observation 444c21dc-5c62-448f-9e75-d2e059dc324b · outbound

This paper cites Gradient-based visual explanation for transformer-based clip.

Fine-grained CLIP fine-tuning with self-annotated region alignment Gradient-based visual explanation for transformer-based clip

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.086757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.086757Z digest=sha256:465559b13182bf154b00440a1f98eeeef056e8259065e5da39e3eca84841eaf6

Observation c9b6e25d-0459-4d2d-a378-0d1fce8d4c8d · outbound

This paper cites Extract free dense labels from clip.

Fine-grained CLIP fine-tuning with self-annotated region alignment Extract free dense labels from clip

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.130648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.130648Z digest=sha256:3cb9033475378ea40c737264bb9fae392eb20821e2836047ce664834a91c7082

Observation 94584bb7-1049-4471-ae94-be1cd70f3a1b · outbound

This paper cites O’Reilly Media, Inc., 2009.

Fine-grained CLIP fine-tuning with self-annotated region alignment O’Reilly Media, Inc., 2009

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.187146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.187146Z digest=sha256:cf4024d6bfb9bbf3df775757f6ec6eaa51f50fe208dad8ef60669dc7a8a1e9dc

Observation 99c64763-8baf-433a-992d-4803d6f90527 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Fine-grained CLIP fine-tuning with self-annotated region alignment EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.228619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.228619Z digest=sha256:0c73d7559472a9fe137429a6e9cb041cd318ca46a12a5b5471243450296ad827

Observation ff33e22f-c38a-4388-a02e-9cb2916b1545 · outbound

This paper cites Microsoft coco: Common objects in context.

Fine-grained CLIP fine-tuning with self-annotated region alignment Microsoft coco: Common objects in context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.320295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.320295Z digest=sha256:dd1d96dd487f7916402a76c9c7e74e86d18c598147fbc7ccda2334484da545ca

Observation 1cda5196-ea45-4be9-b5f0-450626689087 · outbound

This paper cites Scene parsing through ade20k dataset.

Fine-grained CLIP fine-tuning with self-annotated region alignment Scene parsing through ade20k dataset

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.367081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.367081Z digest=sha256:549f00e8f48673a93718bcea58774fd1b90b73ca3afeca1aebef0af3f521d06e

Observation 7ec207f7-1166-4fec-b74b-41a90a1cb363 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Fine-grained CLIP fine-tuning with self-annotated region alignment Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.413004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.413004Z digest=sha256:4634717021042d77c19e2eaaaa4fa59ef42417fd0d25f3acb99a68b7712ace49

Observation 60a84edb-86d0-41fd-ba53-cab7dbf6fd26 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

Fine-grained CLIP fine-tuning with self-annotated region alignment Lvis: A dataset for large vocabulary instance segmentation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.484746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.484746Z digest=sha256:ff5db831b77c404b8aa51de4dac4c1a968d05fbb9d71ba3f1951862b93f3647a

Observation 3aedc7db-11ac-459e-878e-0cb683dd086c · outbound

This paper cites Open-vocabulary object detection using captions.

Fine-grained CLIP fine-tuning with self-annotated region alignment Open-vocabulary object detection using captions

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.564582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.564582Z digest=sha256:aa081cd59b645a8486fd242595f28775589a366bc0f4572bb2fd7beb8462cbc9

Observation 37a13a99-87f3-4a96-8828-89617d4ca713 · outbound

This paper cites Detecting twenty-thousand classes using image-level supervision.

Fine-grained CLIP fine-tuning with self-annotated region alignment Detecting twenty-thousand classes using image-level supervision

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.619581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.619581Z digest=sha256:b5d91ff633ae9608d44cc7c0c0f839cc75ad662cbd73be4eca840da9be5c19f7

Observation 70f1f7f9-eba6-42e5-9a90-bbe7c2f29837 · outbound

This paper cites Learning Object-Language Alignments for Open-Vocabulary Object Detection.

Fine-grained CLIP fine-tuning with self-annotated region alignment Learning Object-Language Alignments for Open-Vocabulary Object Detection

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.697318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.697318Z digest=sha256:959c4819b73d7e64c3d2c00bf0da16e82eb3ef4ba4fb1f49ea1c1af3788a7f62

Observation 4015e22e-3554-412d-a9ef-d73385683845 · outbound

This paper cites Everingham, L.

Fine-grained CLIP fine-tuning with self-annotated region alignment Everingham, L

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.775181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.775181Z digest=sha256:72d8ab4958c1b01f2df303a776ddda4b6928db435c641c8eeae02293967f2df8

Observation 5291320a-5457-4b97-80fa-605b2501cc42 · outbound

This paper cites The role of context for object detection and semantic segmentation in the wild.

Fine-grained CLIP fine-tuning with self-annotated region alignment The role of context for object detection and semantic segmentation in the wild

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.855596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.855596Z digest=sha256:3df0683cdfe597379a87229720d1b17b815fea2ae4d18cbd584d83ddea957e53

Observation d2d19235-5285-41e9-9ae4-545f4796b220 · outbound

This paper cites Cat-seg: Cost aggregation for open-vocabulary semantic segmentation.

Fine-grained CLIP fine-tuning with self-annotated region alignment Cat-seg: Cost aggregation for open-vocabulary semantic segmentation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.912590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.912590Z digest=sha256:e6a9687d8d9eae8431ba6f2c3016e647c139bf7d959c93ac45525fa12a5527fc

Observation deaab706-55ba-4d15-928f-54ad94fe9b78 · outbound

This paper cites Coco-stuff: Thing and stuff classes in context.

Fine-grained CLIP fine-tuning with self-annotated region alignment Coco-stuff: Thing and stuff classes in context

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:23.991946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:23.991946Z digest=sha256:f303644c9b31c993c375a4e47ab08866b28bb1c3a2d99bd019b876f3619708fb

Observation dbedda91-2484-4bca-b016-c1f364cb5091 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models.

Fine-grained CLIP fine-tuning with self-annotated region alignment Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T04:34:24.068813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:34:24.068813Z digest=sha256:0168d75277014fb3c309bc718b09910831339d8f03578cea41c8e24437176688

Pith citing papers

No inbound Pith citation observations are available.