Pith. sign in

Paper Citation Record · LEDGER

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score

As of 20 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2507.09615.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09615 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:00:31.793739Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy48
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 79e205f0-ecce-4568-890a-16885b5f701d · outbound

This paper cites GPT-4 Technical Report.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:31.578274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:31.578274Z digest=sha256:bc31c5ea7a5e3dd43657dcc2f03786b51f6bba159d908794b42242eff284ea77

Observation 7b1fdebe-9eaa-493a-b054-f104ddbdb6d5 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:31.582097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:31.582097Z digest=sha256:c3465789e8ccd9f4d039bb5180a1228c7c90ae2886fe628969df502c07df1ae8

Observation 6a5415d6-04ea-45ee-8119-866eb49c6b1d · outbound

This paper cites Dpa: Dual prototypes alignment for unsupervised adaptation of vision-language models.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Dpa: Dual prototypes alignment for unsupervised adaptation of vision-language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:48.132976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.585637Z digest=sha256:1ffeb22751500687e5c05d8502e89a9bba95700d6fae714e0efd6c09601c3e79

Observation cbd47111-fbb9-4fdb-8833-42cbeed3568d · outbound

This paper cites Layer Normalization.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Layer Normalization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:31.590153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:31.590153Z digest=sha256:0e967f3c2b065b87e73543a22387ec9d94a6e74bab5844e42793d524d4d388f5

Observation a5512491-c663-4db5-ae11-20dedcfaf51f · outbound

This paper cites Food-101–mining discriminative components with random forests.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Food-101–mining discriminative components with random forests

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:48.114909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.593840Z digest=sha256:c0f075966cd80884646287c52c986dcaf36b4f19610f079ae8618a3b8a8c4309

Observation ace2ce56-843d-440d-b772-84807d0caf1e · outbound

This paper cites Lan- guage models are few-shot learners.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Lan- guage models are few-shot learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:31.597300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:31.597300Z digest=sha256:2f17aee1a8ee0fe767be896ae72eb4de66be700981fc74667de0ca29c8746b15

Observation ed33b855-a438-4791-98d6-ce15d6803ed2 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Emerg- ing properties in self-supervised vision transformers

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:47.971557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.600420Z digest=sha256:03b5f7a079200449bca3291b9b0b30b9ee3945a94c16d359a43f1d040ce20ca3

Observation 485d8ceb-e661-4dd2-bc57-357f0385a646 · outbound

This paper cites Remote sens- ing image scene classification: Benchmark and state of the art.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Remote sens- ing image scene classification: Benchmark and state of the art

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:41.128504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.603507Z digest=sha256:a779054fadbbca27a5c72f47ebfc3610756975d45f409b86683d0b76458100e3

Observation 92de0686-018a-4a4d-84fe-68d94dce94cb · outbound

This paper cites Distribution-aware prompt tuning for vision-language mod- els.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Distribution-aware prompt tuning for vision-language mod- els

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:40.021432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.606726Z digest=sha256:0cc10158372c1696bebddbcfc1cb605efff03fe2f23e6fc04406f41497cdb393

Observation 979bc9d2-d280-441d-82f4-54cb061b8e2a · outbound

This paper cites Describing textures in the wild.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Describing textures in the wild

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:39.076639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.610582Z digest=sha256:d0f4898f9ec74cf2fc60679dc54bcfe63c439f5bdf44d193b59401d45e9e08fc

Observation 394e09b0-237e-4daf-9234-c31d5cbc7a95 · outbound

This paper cites Randaugment: Practical automated data augmen- tation with a reduced search space.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Randaugment: Practical automated data augmen- tation with a reduced search space

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:37.621766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.613881Z digest=sha256:0297677765570b9b3f2b9e9c02ff652429a941a1364cc8a8c2438d4135f52dd1

Observation 656a29f7-b13f-4c54-b155-8ca53831e369 · outbound

This paper cites Improving clip training with language rewrites.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Improving clip training with language rewrites

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:31.617017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:31.617017Z digest=sha256:ce610cd3d1ac1fccb9a2aba2b967aa7e8383bdc50263cc5126595873decea5ff

Observation 0f95cd84-16d6-4d58-8193-fa8802a1c336 · outbound

This paper cites Learning gener- ative visual models from few training examples: An incre- mental bayesian approach tested on 101 object categories.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Learning gener- ative visual models from few training examples: An incre- mental bayesian approach tested on 101 object categories

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:31.620299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:31.620299Z digest=sha256:c50261ce15f5f863f4ebf34d7e202860b0988b3937803791b2a5400b1ec8ee0d

Observation 65913c6e-e78a-4af6-96ef-91f519c03ee8 · outbound

This paper cites Gpt-3: Its nature, scope, limits, and consequences.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Gpt-3: Its nature, scope, limits, and consequences

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:37.162329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.623602Z digest=sha256:cc1022c9e832c67895d2772fd42331fba31f77b2d34af69ac0f00aa34be8bf9e

Observation 70e8545d-9d3f-4954-8284-2763ee765e32 · outbound

This paper cites Perceptron-based learning algo- rithms.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Perceptron-based learning algo- rithms

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:36.922923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.626521Z digest=sha256:7ee88071e471346455c28cfdae16abfefea7d3b480248465c58220016d54acd2

Observation 79fed6b7-45ae-4068-a4ec-3f41f18ce05d · outbound

This paper cites an unresolved cited work.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:00:36.697822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.629523Z digest=sha256:fe669c7a4247b58718b623c9e6901b631274fd3a967d19bd741cb6bb79b92a92

Observation f7c32b89-cd12-43fa-b1d7-32e9e4ea1092 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Open-vocabulary object detection via vision and language knowledge distillation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:36.445314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.632814Z digest=sha256:78b0151806f5f9804383633e3d4e9e649a3053f7d09a23cfc48e59c04f99cc7f

Observation 180fed66-0730-4a74-be02-6688e6159f84 · outbound

This paper cites Parameter-efficient model adaptation for vision transformers.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Parameter-efficient model adaptation for vision transformers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:35.801606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.636225Z digest=sha256:1799d28263240e6911f64c3162c03ff46c90b903f086a53c0d2e38df01c212cc

Observation d048d510-6bd8-4160-81f9-df2b6e03d7f2 · outbound

This paper cites Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:34.909999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.639317Z digest=sha256:9dd7731f9a43545b5133019e3c26c8bfde119e2b5f27d806693f53ea9f3fd02a

Observation 1791a035-7e8b-4c1e-b104-74c4cb178305 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Lora: Low-rank adaptation of large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:34.078815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.642414Z digest=sha256:12e7d320ef4f55c2a3b34c821feb69373f6d982b7354545ed10ce6fdcc7bb33d

Observation 7b719fc3-c289-45c4-b065-2079ba40adf1 · outbound

This paper cites Reclip: Refine contrastive language image pre-training with source free domain adaptation.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Reclip: Refine contrastive language image pre-training with source free domain adaptation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:33.659160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.645861Z digest=sha256:82d0d26f2bf374417bf8c60747880a29adc690c0b21ff34db60fd1bcb42bcd32

Observation fe4deb7d-5b2d-432e-b098-eb343015940b · outbound

This paper cites an unresolved cited work.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:00:33.480601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.649283Z digest=sha256:a9f09a2d192945d4d46997998a56db7a40e9a3e94a40bb945378c4871f8c22e5

Observation 56e7a5d4-43be-4891-b553-0aea7b719fe3 · outbound

This paper cites an unresolved cited work.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:00:33.310928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.652124Z digest=sha256:dda8d9c619b77f41dc4002dd827f30db6c1d4f0df3aecd4ebae27e5a7523c322

Observation 982c8e1c-d73d-4bad-97d6-b0c96d50782a · outbound

This paper cites Adapting visual-language models for generalizable anomaly detection in medical im- ages.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Adapting visual-language models for generalizable anomaly detection in medical im- ages

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:33.150754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.655641Z digest=sha256:005fe3417df9e58892eed520ba0701415e1dfba3ae0aa2e6d1adccf158e1c75a

Observation e5ba3e4d-2e7e-4bf6-bda4-1ead158abbc5 · outbound

This paper cites Unsupervised Prompt Learning for Vision-Language Models.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Unsupervised Prompt Learning for Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:31.658407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:31.658407Z digest=sha256:3809e82ec2366fd1bf782b970912b29eafb28dd7a84ae61d92944b3b1fb740ab

Observation a8602394-960a-46ab-b943-75eab1bb27e4 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.963582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.661594Z digest=sha256:e274bf42257cc1ba35eb5c1aae8522ad2d3295c1201187facd299b2f9159383d

Observation 3bc2c75e-43fc-4222-b5cc-4ff9bf067757 · outbound

This paper cites Learning to prompt with text only supervision for vision- language models.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Learning to prompt with text only supervision for vision- language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.856613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.664429Z digest=sha256:8390f55bd1490fddbac76650443e41c6189a32ed90d848ab92fc4ee2e178c4bc

Observation 74cbbb46-ca1f-49a6-9a47-751acf0b59ee · outbound

This paper cites Maple: Multi-modal prompt learning.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Maple: Multi-modal prompt learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.717699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.668129Z digest=sha256:5e89d3f2367c3e1b4edb0e16b597bb50095d23d4ddb0512a7c840349295b98ef

Observation 41bc28ee-386e-41c1-adfa-de32a484d263 · outbound

This paper cites 3d object representations for fine-grained categorization.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score 3d object representations for fine-grained categorization

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.581788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.671469Z digest=sha256:84bad5a3a9cd66fb471d26488ae521b8ff34e8795f3129e206a79ca0166a1de0

Observation 6e63a22f-ad22-48f0-9b0a-38b2e67e4813 · outbound

This paper cites Learning multiple layers of features from tiny images.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Learning multiple layers of features from tiny images

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:31.674245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:31.674245Z digest=sha256:845c229865e57cac54ba6a4c47ad37270a4a1125566e46d8bfd1e1b04e405db4

Observation d8a1f0da-c252-49c3-b750-fac5561df7fe · outbound

This paper cites Language-driven semantic seg- mentation.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Language-driven semantic seg- mentation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.458270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.677141Z digest=sha256:e76b2db02c695102a136e8096c9d346fa239c0288c036ff3dd9a252694022385

Observation 73579cf1-fa10-4bef-82b2-bd71be44c17f · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:31.679900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:31.679900Z digest=sha256:961819ca3c939c25245c5e9342edce7e8b9d28239befbd5872d893b1c5563766

Observation 8faac6ef-67fe-46d0-beed-8d8329a52c20 · outbound

This paper cites Visual-text cross alignment: Refining the similarity score in vision-language models.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Visual-text cross alignment: Refining the similarity score in vision-language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.419435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.682743Z digest=sha256:877928c8a44a6357107282fe6b9c3c8ec68c51e0439be3605c19ca693a6050e6

Observation ed775e83-1abe-4f8a-bf21-b4af6b0d760e · outbound

This paper cites Masked unsupervised self-training for label-free image classifica- tion.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Masked unsupervised self-training for label-free image classifica- tion

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.383725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.685585Z digest=sha256:25f8b31038a3272382c4c9bdbaafcf1e49da40fd9570c786eeaa630fda844a4b

Observation 7b0b31b6-7410-402c-b8b8-759db0e3a227 · outbound

This paper cites Align before fuse: Vision and language representation learn- ing with momentum distillation.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Align before fuse: Vision and language representation learn- ing with momentum distillation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:31.689196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:31.689196Z digest=sha256:bdd2492ecb7092feb018213bdb28581b43578805c295a5782ba45a70784f94d2

Observation ee1ea375-1198-4e30-b99b-9fb98566ba80 · outbound

This paper cites Promptkd: Unsupervised prompt distillation for vision-language models.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Promptkd: Unsupervised prompt distillation for vision-language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.343106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.692163Z digest=sha256:6833d2de04bbb84c3a683aa49d231055a2d1c0bef9be3491c641067c059645a6

Observation 202df6ea-64c8-43f9-ac34-2dba144e4cbb · outbound

This paper cites Clip is also an efficient segmenter: A text-driven approach for weakly supervised semantic segmentation.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Clip is also an efficient segmenter: A text-driven approach for weakly supervised semantic segmentation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.308495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.695185Z digest=sha256:cd4be790e7fcfbd1443012ecda4fd143b6c00a9ef6e1923fcfb77c052f2a6e4c

Observation 1827d0cd-f261-4640-a897-eb088f709497 · outbound

This paper cites Decoupled weight de- cay regularization.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Decoupled weight de- cay regularization

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.278601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.698095Z digest=sha256:247a4ba40507c6655a940e60bc0b74c94601afe591ae2778ccfd8d44ebe36507

Observation f32ab6fe-2a9f-49e3-a242-aa913b32d224 · outbound

This paper cites Prompt distribution learning.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Prompt distribution learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:31.701788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:31.701788Z digest=sha256:efc92fd3f26472f8d6706069fd8d397260c131dfb0d8de53ba7e41dd4e779a08

Observation 5cc4c2a8-7367-4246-ba86-bc7f4c10fefd · outbound

This paper cites Lafter: Label-free tuning of zero-shot clas- sifier using language and unlabeled image collections.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Lafter: Label-free tuning of zero-shot clas- sifier using language and unlabeled image collections

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.237722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.704735Z digest=sha256:75af7f4ba4fc4ef246a7bad50e740ec092c45e4af560e3174e700d8f1ee5af50

Observation 5cfe96ac-acb1-4f57-901d-2d48338343be · outbound

This paper cites Automated flower classification over a large number of classes.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Automated flower classification over a large number of classes

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.204964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.708317Z digest=sha256:60cbfc262f383ef1ff5f98a6b73eba35dc01b863561bbb297b163b77a8a8eee0

Observation d10282c8-8334-4ede-bb57-6f4e2f9c17f1 · outbound

This paper cites Valse: A task- independent benchmark for vision and language models cen- tered on linguistic phenomena.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Valse: A task- independent benchmark for vision and language models cen- tered on linguistic phenomena

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.169214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.711700Z digest=sha256:6b8c9c9a59b784bee04694651ae8e1c85fd67302c4ab60cbb79eadad8ada40a5

Observation 66825e8a-9afa-4745-999d-c6d0192393ca · outbound

This paper cites Cats and dogs.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Cats and dogs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:31.715225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:31.715225Z digest=sha256:f32fccd901f28579be60f094dfc4e220a1584c99f65aee7117eeb65ea245d94b

Observation 0038cdbd-f8e9-49ad-8985-3fbb38c1a255 · outbound

This paper cites What does a platypus look like? generating customized prompts for zero-shot image classification.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score What does a platypus look like? generating customized prompts for zero-shot image classification

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.134732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.718576Z digest=sha256:72bada773f65815719dee27dd1432d78779eb80c1cb165f5bb9d6e7858c92c56

Observation 42624763-e39f-4487-9b18-9ba7fcc64e89 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Learning transferable visual models from natural language supervi- sion

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.109810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.721668Z digest=sha256:5370b516cef208610bc59d810f02af45b125132ebf81030a6cf3b2e467791396

Observation b7b611f0-c2e0-48f8-90c4-d0129691857a · outbound

This paper cites Flava: A foundational language and vision alignment model.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Flava: A foundational language and vision alignment model

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.078177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.724951Z digest=sha256:af2e5915f0ecb0764549261802a15242d91254016250dd0dc3ed60b3710811de

Observation 56e30ede-9e7a-4366-b262-cc780a661c48 · outbound

This paper cites Ucf101: A dataset of 101 human actions classes from videos in the wild.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Ucf101: A dataset of 101 human actions classes from videos in the wild

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.043348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.727916Z digest=sha256:8f732232437f294029e5a952d3dddaeb375d708f82cf1dd2f3a2a46387aededd

Observation ff89abd8-f48f-4518-8880-349e63f980e7 · outbound

This paper cites Pouf: Prompt-oriented unsupervised fine-tuning for large pre-trained models.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Pouf: Prompt-oriented unsupervised fine-tuning for large pre-trained models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.018115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.730962Z digest=sha256:0d2af81376076a3bff28ab19a700fb61019ad0313a3d3a2a2f0e8e51cb3d4f0f

Observation d0953c9e-28d4-4b8b-b18c-539778e2fc56 · outbound

This paper cites Clip the gap: A single domain generalization approach for object detection.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Clip the gap: A single domain generalization approach for object detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.009428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.734511Z digest=sha256:5b9a2e235eceea75cd9bf1fb4db03a73103b56ec3ee3b021f6fe633ba6723c62

Observation a694403d-ef77-4df9-8599-a5562d97768c · outbound

This paper cites The caltech-ucsd birds-200-2011 dataset.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score The caltech-ucsd birds-200-2011 dataset

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:32.000745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.737504Z digest=sha256:f3e4d7f91368a25f078bddd208dfa576c5dded68663e6c5accfa1e195af2f03e

Observation 9735b591-1b8d-423b-a5d9-d74a508b6b3b · outbound

This paper cites Tent: Fully test-time adaptation by entropy minimization.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Tent: Fully test-time adaptation by entropy minimization

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:31.991474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.740497Z digest=sha256:68c734fa8ba3f9234394487ed4720991a120e4691beb98b0026bd74bc5cb4685

Observation 15261457-3c75-4e0e-9c5a-e5b7f855b58c · outbound

This paper cites Debiased learning from naturally imbalanced pseudo-labels.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Debiased learning from naturally imbalanced pseudo-labels

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:31.983127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.743685Z digest=sha256:66138126310377d65297b3eb802a64ad20ece827be15da19024faad103c0aade

Observation 9598e3ed-8a95-473a-83e3-4fa0267077cc · outbound

This paper cites Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:31.974225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.746809Z digest=sha256:043fd4507cb78c96c892cf2a2553dec6640d75711774ff9040b2f5ca653746f7

Observation e6e5bd0e-d4e7-4196-9e9c-a8abc1cc1428 · outbound

This paper cites Aid: A benchmark data set for performance evaluation of aerial scene classification.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Aid: A benchmark data set for performance evaluation of aerial scene classification

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:31.965642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.749883Z digest=sha256:23f56a2765788093d0f327d6d9e5be1fb6be41d7d5ecf187e8f1fb08ca53880e

Observation 279d2b8f-d5d3-4501-9fa3-1a38e5adba0d · outbound

This paper cites Sun database: Large-scale scene recognition from abbey to zoo.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Sun database: Large-scale scene recognition from abbey to zoo

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:31.955970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.753154Z digest=sha256:494f4f994dc3a62c96a4615002a5c5508f783e0217799ffabac4a7bf635e2ded

Observation 5616dde8-9a62-406a-b7f2-097f2697f444 · outbound

This paper cites Demystifying CLIP Data.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Demystifying CLIP Data

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:31.755954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:31.755954Z digest=sha256:b4ff8486c1ba19afb01c6c1f9e49530d29f9e92f28800c624446d543cf14072d

Observation 98575ec9-0d5f-47d7-9047-1a5703b2179c · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Groupvit: Semantic segmentation emerges from text supervision

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:31.946697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.759535Z digest=sha256:a56c892ab29ccff73f5a408e75c04f60b578a1aa53c10d313997ebc5b27a18cc

Observation 80e44ed2-8cdc-4f7d-aaf8-3bd1b58eb878 · outbound

This paper cites Lever- aging cross-modal neighbor representation for improved clip classification.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Lever- aging cross-modal neighbor representation for improved clip classification

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:31.937415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.762821Z digest=sha256:776b81f0f7b659ef5e6cf0cb4efa0137650e483c58f83fc5d74c2ab0e01e77a0

Observation defbc4fe-7c63-4deb-8414-e981028221de · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Florence: A New Foundation Model for Computer Vision

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:31.766827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:31.766827Z digest=sha256:c0715ce0e095270d3bd768c25e7c3cd79a204a4e503c2f4d5636f88f3c0b137e

Observation 5d1c9f41-27fe-4607-abed-98fd44903b6c · outbound

This paper cites When and why vision- language models behave like bags-of-words, and what to do about it? In International Conference on Learning Repre- sentations, 2023.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score When and why vision- language models behave like bags-of-words, and what to do about it? In International Conference on Learning Repre- sentations, 2023

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:31.927830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.770741Z digest=sha256:34567223fe135c2daadadf230da45b373c5eda2411fa4c06a680eeca8149ae1d

Observation d70ea8bb-a120-44c6-8267-56cfe698725a · outbound

This paper cites Boost- ing vision-language models with transduction.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Boost- ing vision-language models with transduction

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:31.918023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.774760Z digest=sha256:1b555f7281db4641bf2fda3c15d5908f749ec96d1a21bea9642f7ed8f03debac

Observation 07062fc6-0403-4b6c-a1d8-5293407d2482 · outbound

This paper cites Candidate pseudolabel learning: Enhancing vision-language models by prompt tuning with unlabeled data.International Conference on Machine Learning, 2024.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Candidate pseudolabel learning: Enhancing vision-language models by prompt tuning with unlabeled data.International Conference on Machine Learning, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:31.908386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.777795Z digest=sha256:524a7ebd1dd91cb5b0031ab3c921ff451d43e42960a97831b4098a2fd4cce7c9

Observation d1c336b6-34fc-4121-b33c-f8248c34c39e · outbound

This paper cites Prompt, generate, then cache: Cascade of foundation models makes strong few-shot learners.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Prompt, generate, then cache: Cascade of foundation models makes strong few-shot learners

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:31.898969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.780899Z digest=sha256:c3f86c0e6d430468a0f7f5b1a9a96695f88f2936704b82eab61ed1d79a40ecac

Observation 526443bb-fdc9-46f5-b5cb-5ee7bb89c047 · outbound

This paper cites Mediclip: Adapting clip for few-shot medical image anomaly detection.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Mediclip: Adapting clip for few-shot medical image anomaly detection

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:31.889355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.784147Z digest=sha256:37faa8316f9268fade49aee6c2689c224ff277d2f240b2a808e9d9fcc93819ec

Observation 520e91b5-7b14-4a86-93d3-508d886d42b1 · outbound

This paper cites Conditional prompt learning for vision-language mod- els.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Conditional prompt learning for vision-language mod- els

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:31.787539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:31.787539Z digest=sha256:eb54af256f2fbf1e09dc5cbf6bb643ef521e445dd9f4fca04ee90b2a489992d4

Observation 3e64f725-ad24-4d13-b39a-14f9c4e3d653 · outbound

This paper cites Learning to prompt for vision-language models.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Learning to prompt for vision-language models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:31.790759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:31.790759Z digest=sha256:ea98208ba68ba65e0b07a73fc4de6f056cd7b5f23208f3d0bae59772916a2131

Observation d3b6211f-cfaf-4675-b5c0-6eab7820fc98 · outbound

This paper cites Not all features mat- ter: Enhancing few-shot clip with adaptive prior refinement.

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score Not all features mat- ter: Enhancing few-shot clip with adaptive prior refinement

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:00:31.869263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T18:00:31.793739Z digest=sha256:11195f86a402bd80c35dd8008b4e514cc556fd2a224eb9eaf3c136d92bdfc421

Pith citing papers

No inbound Pith citation observations are available.