Pith. sign in

Paper Citation Record · LEDGER

GLAD: Generalizable Tuning for Vision-Language Models

As of 9 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 3 inbound Pith citation observations for arXiv:2507.13089.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13089 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:37:24.063644Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:24:53.437554Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact3
  • verified fuzzy44
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 8c631083-f522-4e07-a307-b4355579d76a · outbound

This paper cites Food-101–mining discriminative components with random forests.

GLAD: Generalizable Tuning for Vision-Language Models Food-101–mining discriminative components with random forests

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:17.647963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:17.647963Z digest=sha256:34a6704059f62a4042f75902236a8f9ad2dec95f6a12d39e8028118a800fbe16

Observation 18c46557-4c87-4e5d-9b2b-23ae3baf775b · outbound

This paper cites Domain prompt learning with quaternion networks.

GLAD: Generalizable Tuning for Vision-Language Models Domain prompt learning with quaternion networks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.642311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:17.697297Z digest=sha256:29b5f73d9f2b22bbd87eee70f4019a5f27b52a4fa4736afffad2511fd4063973

Observation f5650e68-a17e-43f8-802b-123d6870fa79 · outbound

This paper cites Tokenmixup: Efficient attention-guided token-level data augmentation for transformers.

GLAD: Generalizable Tuning for Vision-Language Models Tokenmixup: Efficient attention-guided token-level data augmentation for transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.627929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:17.771183Z digest=sha256:c9c08ad929f6e839b85e1ac4bb15b9ae2ea5f342178037fe7815a82285024c79

Observation 7ef1785a-36f0-40fd-9ea4-f5bb33dd768d · outbound

This paper cites Describing textures in the wild.

GLAD: Generalizable Tuning for Vision-Language Models Describing textures in the wild

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:17.939549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:17.939549Z digest=sha256:765ebaf71386b27e38b65485c02cdd05542f6da59bd15c256c8a7ac03bc12ed1

Observation 0ad76c32-0c90-486f-b8e4-3ea618b7fccf · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

GLAD: Generalizable Tuning for Vision-Language Models Imagenet: A large-scale hierarchical image database

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.603846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:18.028779Z digest=sha256:c43a00ccec78ff9a7accb74bc21af7de1b5267320c891a7e2bb8aec17514df15

Observation 92e68a92-edaa-4485-8bfe-43b4805b65ee · outbound

This paper cites Learning to prompt for open-vocabulary ob- ject detection with vision-language model.

GLAD: Generalizable Tuning for Vision-Language Models Learning to prompt for open-vocabulary ob- ject detection with vision-language model

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.589929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:18.104529Z digest=sha256:e9750a30eba2af988d4a48715e95381677f2f48a2d8da57ffb8b518333505f85

Observation 219ff661-2bad-498b-8dc5-b478deef58ff · outbound

This paper cites Learning gener- ative visual models from few training examples: An incre- mental bayesian approach tested on 101 object categories.

GLAD: Generalizable Tuning for Vision-Language Models Learning gener- ative visual models from few training examples: An incre- mental bayesian approach tested on 101 object categories

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.574755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:18.176667Z digest=sha256:32709d8d722558ee06662c2385d9e8ea9d01348135bde917e905be9960adbc65

Observation 1cdbdb46-701f-4237-be09-1ffb4e5a7ad3 · outbound

This paper cites Prompt- det: Towards open-vocabulary detection using uncurated im- ages.

GLAD: Generalizable Tuning for Vision-Language Models Prompt- det: Towards open-vocabulary detection using uncurated im- ages

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.555382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:18.272795Z digest=sha256:7ec6bec60fe8a50b1421b707efee80b065596aae8cddf9b892152af0105bea10

Observation 3123dc9a-92b4-4b59-a7cf-d4ce526c1dec · outbound

This paper cites Sharpness-Aware Minimization for Efficiently Improving Generalization.

GLAD: Generalizable Tuning for Vision-Language Models Sharpness-Aware Minimization for Efficiently Improving Generalization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:18.401393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:18.401393Z digest=sha256:29994870464088e8590bda1efbfe4fb7d4adaf57c8d92bddc36bfa0c44177174

Observation 062cd90d-afa2-40bf-b4ed-a5ecf1e1fef6 · outbound

This paper cites CLIP-Adapter: Better Vision-Language Models with Feature Adapters.

GLAD: Generalizable Tuning for Vision-Language Models CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:18.530702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:18.530702Z digest=sha256:d5cc8cff0dab72745238f62c66dcb48b6b724a6dc186d4d9748b89773918d0b5

Observation c8cdf682-f7dd-47c4-8bd7-fa1559d24ea6 · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

GLAD: Generalizable Tuning for Vision-Language Models Explaining and Harnessing Adversarial Examples

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:18.616605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:18.616605Z digest=sha256:87bdac6c7e6910c92dbcbbf7ce6779c41af1d6c001a42502a2ea4ea48e8b4b91

Observation c557da71-b43d-4aa4-ad2c-953260e28b21 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

GLAD: Generalizable Tuning for Vision-Language Models Open-vocabulary object detection via vision and language knowledge distillation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.540229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:18.662303Z digest=sha256:bbc1ace7478fc3552171b7c9896933e1d10bf732218143c3f7ce328cbbb5e804

Observation 6d2940be-bf72-4e9f-98a7-d6b6929d039f · outbound

This paper cites Eurosat: A novel dataset and deep learn- ing benchmark for land use and land cover classification.

GLAD: Generalizable Tuning for Vision-Language Models Eurosat: A novel dataset and deep learn- ing benchmark for land use and land cover classification

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.525518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:18.791568Z digest=sha256:bda40aa6b3086ff28383ad73bf0ebb21ba6117739d9b8e849d16fd6155aefe62

Observation f44a092f-0253-48bb-a48b-226d4d514e5d · outbound

This paper cites The many faces of robust- ness: A critical analysis of out-of-distribution generalization.

GLAD: Generalizable Tuning for Vision-Language Models The many faces of robust- ness: A critical analysis of out-of-distribution generalization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.509559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:18.857488Z digest=sha256:239d554b37836eee0ebc7dd54d3cbcad0ad79c436d09f8b28af21b0e977c5c9f

Observation 89a001d4-2771-45ac-b562-633982e981f0 · outbound

This paper cites Natural adversarial examples.

GLAD: Generalizable Tuning for Vision-Language Models Natural adversarial examples

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.495020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:18.963976Z digest=sha256:2708c69b0652e12031770d3202af794651cea0da17f8bd59ddf361fb61f89594

Observation 6b185be9-37b5-46cb-a022-c7019faad37d · outbound

This paper cites Lora: Low-rank adaptation of large language models.

GLAD: Generalizable Tuning for Vision-Language Models Lora: Low-rank adaptation of large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:19.090863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:19.090863Z digest=sha256:0177a213b5bd43b13326fceb5bbebb3e5f314e0b0222aee4b0452a470594e335

Observation 00d50d34-2784-4f23-aa36-adebb894f1a2 · outbound

This paper cites Learning a Better Initialization for Soft Prompts via Meta-Learning.

GLAD: Generalizable Tuning for Vision-Language Models Learning a Better Initialization for Soft Prompts via Meta-Learning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:37:24.853475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:19.243859Z digest=sha256:47a7bf46bb6625e19662d670a9753ed728d5d0fdbb5d31ede2b5f101d61f9127

Observation 7f0920ef-a73e-4379-8f36-d74481668e8a · outbound

This paper cites Patching open-vocabulary models by interpolating weights.

GLAD: Generalizable Tuning for Vision-Language Models Patching open-vocabulary models by interpolating weights

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:19.347269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:19.347269Z digest=sha256:c360b7d3fed06eda64574a43052e409dab044da27b166c7060b7da603cb86054

Observation 5ff13c00-186d-45e9-97e9-fff6c9f2e3fe · outbound

This paper cites Vi- sual prompt tuning.

GLAD: Generalizable Tuning for Vision-Language Models Vi- sual prompt tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:19.448372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:19.448372Z digest=sha256:442be352c6a5803cd2037e1cb761d76ca28526491f711fbe37be6c6813cd3aa5

Observation adf384f3-9ef9-412d-9a1c-73dffa930e01 · outbound

This paper cites Maple: Multi-modal prompt learning.

GLAD: Generalizable Tuning for Vision-Language Models Maple: Multi-modal prompt learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.456352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:19.508974Z digest=sha256:88d20f33ba7242ea5ba7906e8a07fcbc30236310b9b84a2cdc149dbb2d040e30

Observation d58f73ca-6225-4079-8f74-8bc9337af873 · outbound

This paper cites Self-regulating prompts: Foundational model adaptation without forgetting.

GLAD: Generalizable Tuning for Vision-Language Models Self-regulating prompts: Foundational model adaptation without forgetting

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.440487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:19.619850Z digest=sha256:744699c7a22cc6c72ab1f61d260d875e8ed0785524ec63c3d062e8fef5d09de9

Observation ec6763f4-2486-43f2-8d8a-cdec6d4795fd · outbound

This paper cites Co-mixup: Saliency guided joint mixup with super- modular diversity.

GLAD: Generalizable Tuning for Vision-Language Models Co-mixup: Saliency guided joint mixup with super- modular diversity

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.425528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:19.807655Z digest=sha256:718aa8dc4af3226fe4e76aa73f3ff52a1bd264aaaa6e33eae94c7ab4bf4e2755

Observation 5b63c5e1-dc1a-4444-a1e4-dd0c43ae81e3 · outbound

This paper cites 3d object representations for fine-grained categorization.

GLAD: Generalizable Tuning for Vision-Language Models 3d object representations for fine-grained categorization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:19.982942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:19.982942Z digest=sha256:b4bd1dc74f5761e25d50087e9af7fe2b2ad8800c43423a6ef7eead84995bfa91

Observation b57eb404-5fac-42f4-ada8-4ea5bff4d2d0 · outbound

This paper cites Read-only prompt op- timization for vision-language few-shot learning.

GLAD: Generalizable Tuning for Vision-Language Models Read-only prompt op- timization for vision-language few-shot learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.396667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:20.056213Z digest=sha256:89e5d69b1cac45d3b3257a656a8fcc703c63d52d8d9557c755e336355b473799

Observation 8db438d4-ce38-4fc4-a3f7-f481a3744da0 · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

GLAD: Generalizable Tuning for Vision-Language Models The power of scale for parameter-efficient prompt tuning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.380993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:20.165333Z digest=sha256:7b5c4b6ebc99b999e5dce5dacbb0b76e92892ef9fcf4724af3a6aeba1c049527

Observation 6825f01f-c45e-4115-805c-0a6b5cea8238 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

GLAD: Generalizable Tuning for Vision-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:20.278857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:20.278857Z digest=sha256:07a3525350fd47a5825a874d8c8562cd9c5d1e49545a3b0b7b2355c4bf7dcd2d

Observation 8d1fffa2-0021-4fa0-a42f-194812283d3a · outbound

This paper cites Promptkd: Unsupervised prompt distillation for vision-language models.

GLAD: Generalizable Tuning for Vision-Language Models Promptkd: Unsupervised prompt distillation for vision-language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:20.463882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:20.463882Z digest=sha256:7775484addb1bb5ad6a92b6d1c4ace18c787f61acbde2d94d0f10ff7b9630168

Observation 1db82ce7-f022-48a6-abd9-dca872383d73 · outbound

This paper cites Visual instruction tuning.

GLAD: Generalizable Tuning for Vision-Language Models Visual instruction tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:20.549960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:20.549960Z digest=sha256:ebb7e5d907da5681d64f05505ba0dc5579470e49a65426c106ab9ca36760363f

Observation a3d9f104-71f8-4166-86f3-088edfee8a7f · outbound

This paper cites Pre-train, prompt, and predict: A systematic survey of prompting methods in nat- ural language processing.

GLAD: Generalizable Tuning for Vision-Language Models Pre-train, prompt, and predict: A systematic survey of prompting methods in nat- ural language processing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.333469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:20.621651Z digest=sha256:8971531bd2e357667163edfbbe85e4b996ac9b23b5987f05a63b553aba50edae

Observation 9afda434-7a57-409d-b40c-f35ca7b1eb54 · outbound

This paper cites Decoupled weight decay regularization.

GLAD: Generalizable Tuning for Vision-Language Models Decoupled weight decay regularization

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.318282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:20.716825Z digest=sha256:3b3ddd2b47031bde88f9da508c5cacfecca0baaf4fe3f72f45142fbf60b40a0d

Observation ba63cdc1-477a-44e0-b77a-aef4a4a45e5f · outbound

This paper cites Prompt distribution learning.

GLAD: Generalizable Tuning for Vision-Language Models Prompt distribution learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.302878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:20.801162Z digest=sha256:be502276cd499fcfb2e54b703cc78ab48c6f65ebfaf3a068482269eeb1e423df

Observation 24a22a1d-c2eb-4b73-ad51-a3ba60b337d0 · outbound

This paper cites Image segmentation using text and image prompts.

GLAD: Generalizable Tuning for Vision-Language Models Image segmentation using text and image prompts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:20.912376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:20.912376Z digest=sha256:2ca9b694830248be2fb5469818d53741a49085d4db15f412e287f1f004df7db8

Observation a25130ff-2c5c-44c8-8ab4-ba7bad27944f · outbound

This paper cites Fine-Grained Visual Classification of Aircraft.

GLAD: Generalizable Tuning for Vision-Language Models Fine-Grained Visual Classification of Aircraft

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:21.038250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:21.038250Z digest=sha256:d3afcbdb42ac69cc48dfc99e7beaa4b3a16c7980ba95ab3275ed2f75d893250f

Observation 5772af9d-23eb-449e-94e4-c3813760953a · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

GLAD: Generalizable Tuning for Vision-Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:21.189934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:21.189934Z digest=sha256:cdb7bfd2e6bbc4ff5a0654d1efc0274f823f9b4cc30d6b8a8c57ddfdd1b0e884

Observation 715a6811-a62d-4ac5-b41b-be757be0b828 · outbound

This paper cites Lookbehind-SAM: k steps back, 1 step forward.

GLAD: Generalizable Tuning for Vision-Language Models Lookbehind-SAM: k steps back, 1 step forward

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:37:24.559955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:21.263196Z digest=sha256:b33ad7afabb021801b8353a3593e1910c392a3e62b30f1578bb6cb2fecc5517f

Observation cea1fbff-c8c4-4478-bdc4-839e46f9c63a · outbound

This paper cites When does label smoothing help? Advances in neural in- formation processing systems, 32, 2019.

GLAD: Generalizable Tuning for Vision-Language Models When does label smoothing help? Advances in neural in- formation processing systems, 32, 2019

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.275633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:21.345862Z digest=sha256:e67f0f1f4e77363fb32313acab4dbf11acaa2d42244714abaf07804e99af945a

Observation cbe497ab-911b-43df-b53c-d286e638ee63 · outbound

This paper cites Automated flower classification over a large number of classes.

GLAD: Generalizable Tuning for Vision-Language Models Automated flower classification over a large number of classes

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.255053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:21.409595Z digest=sha256:b5c4e5f32a28dc4b525cdf76bd31afaf03aeeb37c4e64cd58383f9d5d8dd323d

Observation 50e15440-8eba-4475-8bd7-813569dbe26d · outbound

This paper cites Metropolis-hastings data augmentation for graph neu- ral networks.

GLAD: Generalizable Tuning for Vision-Language Models Metropolis-hastings data augmentation for graph neu- ral networks

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.239720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:21.560898Z digest=sha256:babbc7b07c000fe2bef2a740b287b6c03380d8cf415f17025911f0f3189fe737

Observation 2dd8802d-63e3-42cc-8c26-ab2ba19371a3 · outbound

This paper cites Cats and dogs.

GLAD: Generalizable Tuning for Vision-Language Models Cats and dogs

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.224586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:21.659196Z digest=sha256:655958732515b5814c342d3f8ff647a7660de70e3de47fb3854844c7ee0da385

Observation 4efa1e0e-858b-4470-91e6-cb0d0e004511 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

GLAD: Generalizable Tuning for Vision-Language Models Learn- ing transferable visual models from natural language super- vision

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.209884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:21.754468Z digest=sha256:cf09524ae867b7450269e25865d8cd081353391ff2bd793a3d7f6e0b979c3efd

Observation f66149b4-eb02-46e6-bea3-8b47a049627a · outbound

This paper cites Do imagenet classifiers generalize to im- agenet? In ICML, pages 5389–5400.

GLAD: Generalizable Tuning for Vision-Language Models Do imagenet classifiers generalize to im- agenet? In ICML, pages 5389–5400

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.193331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:21.848212Z digest=sha256:4c786314cd91ff0d064750c6baca1c6cc87a007fec3308bce5e40816c66d31f8

Observation 8ab4ee32-8f35-4cfc-b375-66f9e12e4cd6 · outbound

This paper cites Multimodal Instruction Tuning with Conditional Mixture of LoRA.

GLAD: Generalizable Tuning for Vision-Language Models Multimodal Instruction Tuning with Conditional Mixture of LoRA

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:37:24.375379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:21.969328Z digest=sha256:54dc7c4ac60d4571b62512960a57a5ab007d18dad70dd6a2db4fdea8e2b941df

Observation 15bc974f-d69b-4c6c-bbac-a30160cdad02 · outbound

This paper cites Flava: A foundational language and vision alignment model.

GLAD: Generalizable Tuning for Vision-Language Models Flava: A foundational language and vision alignment model

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.175590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:22.053563Z digest=sha256:99b178f670df492c9210f9df1dacc9bb5ab48b066fc305e64c64054b4f7cd2a7

Observation a6321269-a557-4ccf-98b8-5e6577c2a80b · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

GLAD: Generalizable Tuning for Vision-Language Models UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:22.126152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:22.126152Z digest=sha256:e838329df5e797f3103c0e1cace75a7ebe461abcdeda8565463acc92f8bb8a03

Observation 87838670-f2e1-4caa-8d19-a87cf28c806e · outbound

This paper cites Dropout: a simple way to prevent neural networks from overfitting.

GLAD: Generalizable Tuning for Vision-Language Models Dropout: a simple way to prevent neural networks from overfitting

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.158376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:22.207731Z digest=sha256:642d1462bd484f4f2d2f5875abcc761a4adf659f9942953c3794b324e05c0503

Observation 64b15ba4-b80f-4605-8475-87b9d93f4409 · outbound

This paper cites Rethinking the inception ar- chitecture for computer vision.

GLAD: Generalizable Tuning for Vision-Language Models Rethinking the inception ar- chitecture for computer vision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:22.304601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:22.304601Z digest=sha256:8c96b15149403715f7540aa2a1445aa865c521c2c3c1e0b7c381374735521c92

Observation f35469ff-89a5-4ed7-97a5-5237913a7af8 · outbound

This paper cites Saliencymix: A saliency guided data augmentation strategy for better regularization.

GLAD: Generalizable Tuning for Vision-Language Models Saliencymix: A saliency guided data augmentation strategy for better regularization

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.126994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:22.449012Z digest=sha256:dfb3682b1f8a5e80dee5121be4593174249b9d9da3448879d5dfd92281a6f8ab

Observation 19d8d1bc-bc71-4e05-ba88-830e49bcec3d · outbound

This paper cites Manifold mixup: learning better representations by in- terpolating hidden states.

GLAD: Generalizable Tuning for Vision-Language Models Manifold mixup: learning better representations by in- terpolating hidden states

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.113352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:22.547152Z digest=sha256:79bf2be058441572c396b3ec21b343b5f533f710237bd5ed2cc257d5ae3c3fb6

Observation c0af567c-bd48-4819-a7cd-1d224f61e30c · outbound

This paper cites Tuning multi-mode token- level prompt alignment across modalities.

GLAD: Generalizable Tuning for Vision-Language Models Tuning multi-mode token- level prompt alignment across modalities

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.098342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:22.642193Z digest=sha256:f9d321f97be28e84c49a6629e2b8da665e06751def2f4d601b2fccfe022fb307

Observation a9d23265-ea2c-46a4-b7c7-f71efadfa374 · outbound

This paper cites Learning robust global representations by penalizing local predictive power.

GLAD: Generalizable Tuning for Vision-Language Models Learning robust global representations by penalizing local predictive power

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.066014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:22.737119Z digest=sha256:25deb0a76231fc5c7209da1d04e5eec53fe0619cdafdf7a15bfe724770e35eb1

Observation c7bb7d1a-3b7b-46e3-8776-df01c172b0cf · outbound

This paper cites Sharpness-aware gradient matching for domain generaliza- tion.

GLAD: Generalizable Tuning for Vision-Language Models Sharpness-aware gradient matching for domain generaliza- tion

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:27.896307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:22.854498Z digest=sha256:150af59c96d272149bd735b853ad3915c907c9ab805e382f2c5d2257f4143e3f

Observation 93c1dd62-3c84-4a36-bd81-82d3afda935f · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

GLAD: Generalizable Tuning for Vision-Language Models Cogvlm: Visual expert for pretrained language models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:27.723422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:22.959808Z digest=sha256:7f42136147a6bd4227f28e84d07dcfdb356e5ac7e51e3d9c67fd864d188ecbc4

Observation d3ea47da-5c28-4003-ba0e-dcabcfd057ed · outbound

This paper cites Robust fine-tuning of zero-shot models.

GLAD: Generalizable Tuning for Vision-Language Models Robust fine-tuning of zero-shot models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:27.583409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:23.033154Z digest=sha256:f66056db95773c0b9875694b5be699d37519e0ffc2ad2ba019e62302f2d4b0c8

Observation 26048ba1-384a-4013-ab3e-dc289bd39204 · outbound

This paper cites Sun database: Large-scale scene recognition from abbey to zoo.

GLAD: Generalizable Tuning for Vision-Language Models Sun database: Large-scale scene recognition from abbey to zoo

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:27.403579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:23.079232Z digest=sha256:4dbb315b03aa96a60e95ee0d8ae6afe66a099ac3352cba60b7b3fa42213ca036

Observation 35131259-a89b-4284-bdad-790893de0a37 · outbound

This paper cites Tcp: Textual- based class-aware prompt tuning for visual-language model.

GLAD: Generalizable Tuning for Vision-Language Models Tcp: Textual- based class-aware prompt tuning for visual-language model

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:27.246670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:23.150721Z digest=sha256:1132287a774eaf23ee8e4f46d973230f47b223a1eb14462e4ebdc20bfa88859a

Observation 3806585f-bbcc-4e70-be8d-2fde26295261 · outbound

This paper cites Cutmix: Regu- larization strategy to train strong classifiers with localizable features.

GLAD: Generalizable Tuning for Vision-Language Models Cutmix: Regu- larization strategy to train strong classifiers with localizable features

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:27.078177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:23.220423Z digest=sha256:abd2968cf73ff0bf3552ce15e8454d2d83233d6cd97d9f81e31a1e7968e9e02e

Observation ebfd872d-fb0f-49e5-88fb-d918115b5910 · outbound

This paper cites Low-rank few-shot adaptation of vision-language models.

GLAD: Generalizable Tuning for Vision-Language Models Low-rank few-shot adaptation of vision-language models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:26.915990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:23.292963Z digest=sha256:e336a59cb15269f69c85d894f90ebc07f6a8765ad901f72ef48fd03b707d822b

Observation 626eb206-3728-4927-90aa-835e1c098562 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

GLAD: Generalizable Tuning for Vision-Language Models Lit: Zero-shot transfer with locked-image text tuning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:26.763594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:23.388218Z digest=sha256:851ea003d6c2105ecd887ebc22bab87fc53d50236c404ad08593295ab2ea2aa0

Observation ed262ff5-852e-41f0-9a30-dbb356a79355 · outbound

This paper cites Three mechanisms of weight decay regularization.

GLAD: Generalizable Tuning for Vision-Language Models Three mechanisms of weight decay regularization

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:26.620181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:23.425869Z digest=sha256:579f6f646490581707b506d342829e2adcd11fe59c6c0451336cfcf07b7e5975

Observation d28e2228-19a9-496c-8ffc-62e3c36aeeb4 · outbound

This paper cites mixup: Beyond Empirical Risk Minimization.

GLAD: Generalizable Tuning for Vision-Language Models mixup: Beyond Empirical Risk Minimization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:23.459308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:23.459308Z digest=sha256:59363f22d005f5da2f0d69c486fd19678c7f863684a89f46c86ce19c47f03627

Observation 4834b420-d026-4e38-86e4-76b17aadee68 · outbound

This paper cites Dept: Decoupled prompt tuning.

GLAD: Generalizable Tuning for Vision-Language Models Dept: Decoupled prompt tuning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:26.426058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:23.526271Z digest=sha256:1cf0a754ade631f34d60a6ed9fb5fa5e09510b2dfdbf91aaaeabcc7c562f0608

Observation a21fe1d1-d66b-4447-9f44-667a3af33382 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

GLAD: Generalizable Tuning for Vision-Language Models Adding conditional control to text-to-image diffusion models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:23.593178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:23.593178Z digest=sha256:028e9e55c3ff8e2169a426590740c734e31e99719a185bf19f6c124ef8e67b0b

Observation b7512f2a-6b89-4be1-b313-2e6847175b69 · outbound

This paper cites Llama-adapter: Efficient fine-tuning of language models with zero-init attention.

GLAD: Generalizable Tuning for Vision-Language Models Llama-adapter: Efficient fine-tuning of language models with zero-init attention

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:26.126021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:23.638487Z digest=sha256:a35a4e3dc1b6e52ddbabee60cb267fa756eb4ff56ae4f8e02536dda88c375bfd

Observation c04e7633-c62d-43a3-ba52-725bb69e0a6d · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

GLAD: Generalizable Tuning for Vision-Language Models SVIT: Scaling up Visual Instruction Tuning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:23.687949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:23.687949Z digest=sha256:e195b10576c02a54f9ff17ed439310ed6eb08183aa083d447f57199d839ffbea

Observation 2dadc935-ceed-472a-a3cd-822b7c5036d8 · outbound

This paper cites Regionclip: Region-based language-image pretraining.

GLAD: Generalizable Tuning for Vision-Language Models Regionclip: Region-based language-image pretraining

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:25.879634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:23.748806Z digest=sha256:f51828eee4ec0b6a0907cc69ab20a982030e3cbc10f6a4ba4bce3fad88181b31

Observation 7d993c5f-112d-4b83-b15e-97df888bc360 · outbound

This paper cites Conditional prompt learning for vision-language models.

GLAD: Generalizable Tuning for Vision-Language Models Conditional prompt learning for vision-language models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:25.618806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:23.856095Z digest=sha256:ce131eb480d3d1a751603c706d8698ce8eb16bef52859f8454c0bfa30088d391

Observation 2cbf78dd-d784-4289-a2ae-135dcd884f41 · outbound

This paper cites Learning to prompt for vision-language models.

GLAD: Generalizable Tuning for Vision-Language Models Learning to prompt for vision-language models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:25.237348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:37:23.927447Z digest=sha256:0fe8b31fa573e4849dc01d55e097407e3a1651233725577df7b3dd01dfdab346

Observation e2a1921c-ccf8-4a8d-a30d-298f18332a4e · outbound

This paper cites Prompt-aligned gradient for prompt tuning.

GLAD: Generalizable Tuning for Vision-Language Models Prompt-aligned gradient for prompt tuning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:23.993559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:23.993559Z digest=sha256:4443efb96ed64558c17768668a9aac482019bf4e316ddb635e55a92ec28d6400

Observation e9c866b0-7b5d-4199-be01-49ece54bf2c4 · outbound

This paper cites Surrogate Gap Minimization Improves Sharpness-Aware Training.

GLAD: Generalizable Tuning for Vision-Language Models Surrogate Gap Minimization Improves Sharpness-Aware Training

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:24.063644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:24.063644Z digest=sha256:516275942d97dd1032bcb1db9c4b7a2224ebf283c128e057e4e5f9d14fecc501

Pith citing papers

Observation bd78f140-02f5-4e16-b878-6d00bf653a39 · inbound

TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models cites this paper.

TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models GLAD: Generalizable Tuning for Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:53.437554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:53.437554Z digest=sha256:dc81cad69e357cfbc1cd489b34e629b588b3f7c49b1d8bc753a762c8da45acbf

Observation 30255c27-ef3b-4a30-b360-43510c56f66e · inbound

GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models cites this paper.

GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models GLAD: Generalizable Tuning for Vision-Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:40:26.018918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:39:15.820777Z digest=sha256:dc4cb05b55b86e2253c9f166095e36ebdbab102b4d22a829533169c1cb1ffd01

Observation 240109b8-82e3-4783-b64a-5230a8efe68a · inbound

GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models cites this paper.

GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models GLAD: Generalizable Tuning for Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T20:25:41.773345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:25:41.773345Z digest=sha256:61dae02835e970b551a5f707ba6e3f96c0c9d2ee2503e88fa77d092ea9f92022