Pith. sign in

Paper Citation Record · LEDGER

GLAD: Generalizable Tuning for Vision-Language Models

As of 13 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 3 inbound Pith citation observations for arXiv:2507.13089.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13089 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:37:24.063644Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:24:53.437554Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact3
  • verified fuzzy44
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 8c631083-f522-4e07-a307-b4355579d76a · outbound

This paper cites Food-101–mining discriminative components with random forests.

GLAD: Generalizable Tuning for Vision-Language Models Food-101–mining discriminative components with random forests

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:17.647963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:17.647963Z digest=sha256:59b34531a45ec906a8c3b65ed549c45dee568e9df5c5843a2877dfd4c6c2a649

Observation 18c46557-4c87-4e5d-9b2b-23ae3baf775b · outbound

This paper cites Domain prompt learning with quaternion networks.

GLAD: Generalizable Tuning for Vision-Language Models Domain prompt learning with quaternion networks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.642311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:17.697297Z digest=sha256:ab16486f82e70fd6b97c50a268dc8b1b57b77527e14137dd4ec5f357ea4e3847

Observation f5650e68-a17e-43f8-802b-123d6870fa79 · outbound

This paper cites Tokenmixup: Efficient attention-guided token-level data augmentation for transformers.

GLAD: Generalizable Tuning for Vision-Language Models Tokenmixup: Efficient attention-guided token-level data augmentation for transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.627929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:17.771183Z digest=sha256:7ff36319c595ab4adb61ab3b91b0fe7268df18c65e97e23676f13e0e0aec478b

Observation 7ef1785a-36f0-40fd-9ea4-f5bb33dd768d · outbound

This paper cites Describing textures in the wild.

GLAD: Generalizable Tuning for Vision-Language Models Describing textures in the wild

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:17.939549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:17.939549Z digest=sha256:5dd9acdb9d35d54113e8cd06773580f79d652e6b3dd46a8ae12ea6ab6b25023d

Observation 0ad76c32-0c90-486f-b8e4-3ea618b7fccf · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

GLAD: Generalizable Tuning for Vision-Language Models Imagenet: A large-scale hierarchical image database

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.603846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:18.028779Z digest=sha256:001fe2b45f7c7e2922a1e5705874ef2431db130ce3705c83b3dcbb1aa46f1c71

Observation 92e68a92-edaa-4485-8bfe-43b4805b65ee · outbound

This paper cites Learning to prompt for open-vocabulary ob- ject detection with vision-language model.

GLAD: Generalizable Tuning for Vision-Language Models Learning to prompt for open-vocabulary ob- ject detection with vision-language model

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.589929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:18.104529Z digest=sha256:e67e141c6137e09a372bcb8d1025743dae01db3fceb9a8705cf2614a948fe8e6

Observation 219ff661-2bad-498b-8dc5-b478deef58ff · outbound

This paper cites Learning gener- ative visual models from few training examples: An incre- mental bayesian approach tested on 101 object categories.

GLAD: Generalizable Tuning for Vision-Language Models Learning gener- ative visual models from few training examples: An incre- mental bayesian approach tested on 101 object categories

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.574755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:18.176667Z digest=sha256:bfd35b09998060c098a3552d1d0b70a14bc1bb180bd7b6fe72d41bfd65994d1b

Observation 1cdbdb46-701f-4237-be09-1ffb4e5a7ad3 · outbound

This paper cites Prompt- det: Towards open-vocabulary detection using uncurated im- ages.

GLAD: Generalizable Tuning for Vision-Language Models Prompt- det: Towards open-vocabulary detection using uncurated im- ages

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.555382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:18.272795Z digest=sha256:d58a8d57eab7742997e2610a505fc57d2a92efc46a3f6e0451459145cf3ca177

Observation 3123dc9a-92b4-4b59-a7cf-d4ce526c1dec · outbound

This paper cites Sharpness-Aware Minimization for Efficiently Improving Generalization.

GLAD: Generalizable Tuning for Vision-Language Models Sharpness-Aware Minimization for Efficiently Improving Generalization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:18.401393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:18.401393Z digest=sha256:3fad97a158503119b7e18196855b6884d1c99e0fa9890beb3f735f8909af8c11

Observation 062cd90d-afa2-40bf-b4ed-a5ecf1e1fef6 · outbound

This paper cites CLIP-Adapter: Better Vision-Language Models with Feature Adapters.

GLAD: Generalizable Tuning for Vision-Language Models CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:18.530702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:18.530702Z digest=sha256:7d4764a1b615e74f2d300d2964ea5a971ed113bb1eec7f1d6cec4edd52f98efc

Observation c8cdf682-f7dd-47c4-8bd7-fa1559d24ea6 · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

GLAD: Generalizable Tuning for Vision-Language Models Explaining and Harnessing Adversarial Examples

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:18.616605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:18.616605Z digest=sha256:fc48a64d3fdb96ad3c57a336117e1c175e09634f6fcbc3b0123a8209f21a7765

Observation c557da71-b43d-4aa4-ad2c-953260e28b21 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

GLAD: Generalizable Tuning for Vision-Language Models Open-vocabulary object detection via vision and language knowledge distillation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.540229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:18.662303Z digest=sha256:df340ef395da2a15637c0e918704159046473c2af682f48deea53ad5e8ee9a80

Observation 6d2940be-bf72-4e9f-98a7-d6b6929d039f · outbound

This paper cites Eurosat: A novel dataset and deep learn- ing benchmark for land use and land cover classification.

GLAD: Generalizable Tuning for Vision-Language Models Eurosat: A novel dataset and deep learn- ing benchmark for land use and land cover classification

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.525518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:18.791568Z digest=sha256:7d670709833b3c564706706600e7bd1fa474aec2dcecfc4cce9233ff735fb155

Observation f44a092f-0253-48bb-a48b-226d4d514e5d · outbound

This paper cites The many faces of robust- ness: A critical analysis of out-of-distribution generalization.

GLAD: Generalizable Tuning for Vision-Language Models The many faces of robust- ness: A critical analysis of out-of-distribution generalization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.509559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:18.857488Z digest=sha256:857dcdcdb39b2d391c8b74e4c1142bc73f3cd1cc018e81c78b0a78e6eedb749f

Observation 89a001d4-2771-45ac-b562-633982e981f0 · outbound

This paper cites Natural adversarial examples.

GLAD: Generalizable Tuning for Vision-Language Models Natural adversarial examples

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.495020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:18.963976Z digest=sha256:686d00b7c07c75f6f0711352b0213f42ecc6603478e384785528ad845116fbc9

Observation 6b185be9-37b5-46cb-a022-c7019faad37d · outbound

This paper cites Lora: Low-rank adaptation of large language models.

GLAD: Generalizable Tuning for Vision-Language Models Lora: Low-rank adaptation of large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:19.090863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:19.090863Z digest=sha256:d8bd4cac6cfac491ddb4be88b13a4df4d2bdc39658bbb0725590395082057344

Observation 00d50d34-2784-4f23-aa36-adebb894f1a2 · outbound

This paper cites Learning a Better Initialization for Soft Prompts via Meta-Learning.

GLAD: Generalizable Tuning for Vision-Language Models Learning a Better Initialization for Soft Prompts via Meta-Learning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:37:24.853475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:19.243859Z digest=sha256:0f408e23983afa1c22ff5f535c6279c915440b96e650c98658d81a3523586971

Observation 7f0920ef-a73e-4379-8f36-d74481668e8a · outbound

This paper cites Patching open-vocabulary models by interpolating weights.

GLAD: Generalizable Tuning for Vision-Language Models Patching open-vocabulary models by interpolating weights

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:19.347269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:19.347269Z digest=sha256:6e54c50ffc28e7e22b5e15b260da3b98dc096a988585da263b76baa83d77ea8e

Observation 5ff13c00-186d-45e9-97e9-fff6c9f2e3fe · outbound

This paper cites Vi- sual prompt tuning.

GLAD: Generalizable Tuning for Vision-Language Models Vi- sual prompt tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:19.448372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:19.448372Z digest=sha256:4d3aa115575bef570e4ffc4e879b170062e02b190ac0e92c2fe59a3cfd40e8e5

Observation adf384f3-9ef9-412d-9a1c-73dffa930e01 · outbound

This paper cites Maple: Multi-modal prompt learning.

GLAD: Generalizable Tuning for Vision-Language Models Maple: Multi-modal prompt learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.456352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:19.508974Z digest=sha256:91be1dae91e3873242e54fdf47ce4e0a11ef2780c0c0f7692f83e666a6dba1e7

Observation d58f73ca-6225-4079-8f74-8bc9337af873 · outbound

This paper cites Self-regulating prompts: Foundational model adaptation without forgetting.

GLAD: Generalizable Tuning for Vision-Language Models Self-regulating prompts: Foundational model adaptation without forgetting

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.440487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:19.619850Z digest=sha256:0b686913bc1c56808b9202e92a3f7f9fa59d0edb7a2235d1fb57f10d86c2cc73

Observation ec6763f4-2486-43f2-8d8a-cdec6d4795fd · outbound

This paper cites Co-mixup: Saliency guided joint mixup with super- modular diversity.

GLAD: Generalizable Tuning for Vision-Language Models Co-mixup: Saliency guided joint mixup with super- modular diversity

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.425528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:19.807655Z digest=sha256:880e97352dcfdbe83e566a71b1855fdedf96bb530eb5159cf5033b49035ddfd2

Observation 5b63c5e1-dc1a-4444-a1e4-dd0c43ae81e3 · outbound

This paper cites 3d object representations for fine-grained categorization.

GLAD: Generalizable Tuning for Vision-Language Models 3d object representations for fine-grained categorization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:19.982942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:19.982942Z digest=sha256:6a8f174bd9f4cba11d2fd4f958665de786497d3d35b32f63ebb84d033afaffe2

Observation b57eb404-5fac-42f4-ada8-4ea5bff4d2d0 · outbound

This paper cites Read-only prompt op- timization for vision-language few-shot learning.

GLAD: Generalizable Tuning for Vision-Language Models Read-only prompt op- timization for vision-language few-shot learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.396667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:20.056213Z digest=sha256:c6b96ec3ca8638064f7928a49bab1a2e2be33cc163dcdea875e934a4a9c868ab

Observation 8db438d4-ce38-4fc4-a3f7-f481a3744da0 · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

GLAD: Generalizable Tuning for Vision-Language Models The power of scale for parameter-efficient prompt tuning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.380993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:20.165333Z digest=sha256:13ed884267af84d571bd0f84d30ca95e762d7e1a9554584367c752d6be3efd9e

Observation 6825f01f-c45e-4115-805c-0a6b5cea8238 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

GLAD: Generalizable Tuning for Vision-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:20.278857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:20.278857Z digest=sha256:5d310e74eaf764f6e1af4faf51bdfb6de9ee3fc6067812adf83d762d32d2371b

Observation 8d1fffa2-0021-4fa0-a42f-194812283d3a · outbound

This paper cites Promptkd: Unsupervised prompt distillation for vision-language models.

GLAD: Generalizable Tuning for Vision-Language Models Promptkd: Unsupervised prompt distillation for vision-language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:20.463882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:20.463882Z digest=sha256:65a73e6d944ad2fa7347da13e97f6ea174fc6ae93d1801844d8b46e82a18abea

Observation 1db82ce7-f022-48a6-abd9-dca872383d73 · outbound

This paper cites Visual instruction tuning.

GLAD: Generalizable Tuning for Vision-Language Models Visual instruction tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:20.549960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:20.549960Z digest=sha256:7cd775c771b00e5a9696705c54baeba5a7eb6bd9bc883c174e1e9a7c4be624a1

Observation a3d9f104-71f8-4166-86f3-088edfee8a7f · outbound

This paper cites Pre-train, prompt, and predict: A systematic survey of prompting methods in nat- ural language processing.

GLAD: Generalizable Tuning for Vision-Language Models Pre-train, prompt, and predict: A systematic survey of prompting methods in nat- ural language processing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.333469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:20.621651Z digest=sha256:253b738e99ef6455555e2ed4f14592c1d31a71f8a24b772fa337b289049da0a6

Observation 9afda434-7a57-409d-b40c-f35ca7b1eb54 · outbound

This paper cites Decoupled weight decay regularization.

GLAD: Generalizable Tuning for Vision-Language Models Decoupled weight decay regularization

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.318282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:20.716825Z digest=sha256:a4221010b29f2a3db4c4e4d179f9bac76f1ccfbd62f70cd5b016cc7bfb600485

Observation ba63cdc1-477a-44e0-b77a-aef4a4a45e5f · outbound

This paper cites Prompt distribution learning.

GLAD: Generalizable Tuning for Vision-Language Models Prompt distribution learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.302878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:20.801162Z digest=sha256:fb3ff5e1c2ce52ec6299b7a0ae24d578a2f02fa4be8d086f57deb32a43e64e83

Observation 24a22a1d-c2eb-4b73-ad51-a3ba60b337d0 · outbound

This paper cites Image segmentation using text and image prompts.

GLAD: Generalizable Tuning for Vision-Language Models Image segmentation using text and image prompts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:20.912376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:20.912376Z digest=sha256:99ab01c99bccf141af256aea8e47e4c76e1517f7e5e86682606fb73e5a310f84

Observation a25130ff-2c5c-44c8-8ab4-ba7bad27944f · outbound

This paper cites Fine-Grained Visual Classification of Aircraft.

GLAD: Generalizable Tuning for Vision-Language Models Fine-Grained Visual Classification of Aircraft

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:21.038250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:21.038250Z digest=sha256:1c69ed817db91652289eff9a7d80530057df75a4eda42e0280dcb27445f1d61a

Observation 5772af9d-23eb-449e-94e4-c3813760953a · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

GLAD: Generalizable Tuning for Vision-Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:21.189934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:21.189934Z digest=sha256:e592a5d1535ce57f2e7090ad30bf495bec6f041bd544a874c4bdc741367ee6d6

Observation 715a6811-a62d-4ac5-b41b-be757be0b828 · outbound

This paper cites Lookbehind-SAM: k steps back, 1 step forward.

GLAD: Generalizable Tuning for Vision-Language Models Lookbehind-SAM: k steps back, 1 step forward

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:37:24.559955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:21.263196Z digest=sha256:bc5bd2866a2aa2338fd79c3f18c733a2e48fc57badff2b8c25590643a0f37d05

Observation cea1fbff-c8c4-4478-bdc4-839e46f9c63a · outbound

This paper cites When does label smoothing help? Advances in neural in- formation processing systems, 32, 2019.

GLAD: Generalizable Tuning for Vision-Language Models When does label smoothing help? Advances in neural in- formation processing systems, 32, 2019

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.275633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:21.345862Z digest=sha256:9470d42d47225a2c57c4b534e07019684b7176691c57d197a8f9eef02a62d30f

Observation cbe497ab-911b-43df-b53c-d286e638ee63 · outbound

This paper cites Automated flower classification over a large number of classes.

GLAD: Generalizable Tuning for Vision-Language Models Automated flower classification over a large number of classes

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.255053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:21.409595Z digest=sha256:c5bb1189f6a4cb0582cbbf98ab4c46add181d4cedcb8e591cb47a7b4059b2f34

Observation 50e15440-8eba-4475-8bd7-813569dbe26d · outbound

This paper cites Metropolis-hastings data augmentation for graph neu- ral networks.

GLAD: Generalizable Tuning for Vision-Language Models Metropolis-hastings data augmentation for graph neu- ral networks

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.239720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:21.560898Z digest=sha256:e70d1279b0787e6c900e71cb766ac556fdd4321c405577eb635861a4035adfb9

Observation 2dd8802d-63e3-42cc-8c26-ab2ba19371a3 · outbound

This paper cites Cats and dogs.

GLAD: Generalizable Tuning for Vision-Language Models Cats and dogs

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.224586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:21.659196Z digest=sha256:1055f2ce01556237928cd724d3fd46bf0b4a6cb9830ba4772333f683c2fd5ce3

Observation 4efa1e0e-858b-4470-91e6-cb0d0e004511 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

GLAD: Generalizable Tuning for Vision-Language Models Learn- ing transferable visual models from natural language super- vision

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.209884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:21.754468Z digest=sha256:3b41bdff52c473db12a0295ac8905cef12c1e13bc7261c17023703dcd626dc4b

Observation f66149b4-eb02-46e6-bea3-8b47a049627a · outbound

This paper cites Do imagenet classifiers generalize to im- agenet? In ICML, pages 5389–5400.

GLAD: Generalizable Tuning for Vision-Language Models Do imagenet classifiers generalize to im- agenet? In ICML, pages 5389–5400

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.193331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:21.848212Z digest=sha256:419b2bcbfb54893cff3beb1041eb9f65a53baf68325a608850e7422956e27d81

Observation 8ab4ee32-8f35-4cfc-b375-66f9e12e4cd6 · outbound

This paper cites Multimodal Instruction Tuning with Conditional Mixture of LoRA.

GLAD: Generalizable Tuning for Vision-Language Models Multimodal Instruction Tuning with Conditional Mixture of LoRA

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:37:24.375379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:21.969328Z digest=sha256:ac03317d5a1d098da0df2fc7c7d9d242bd771d316fb6f93b92c01d8ff35c8601

Observation 15bc974f-d69b-4c6c-bbac-a30160cdad02 · outbound

This paper cites Flava: A foundational language and vision alignment model.

GLAD: Generalizable Tuning for Vision-Language Models Flava: A foundational language and vision alignment model

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.175590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:22.053563Z digest=sha256:900c6a3f96eaba6a5bd17025cb59b615467f46a7b41e877f19def97c8d2978ff

Observation a6321269-a557-4ccf-98b8-5e6577c2a80b · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

GLAD: Generalizable Tuning for Vision-Language Models UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:22.126152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:22.126152Z digest=sha256:bd5f544bdb1268e7e13870c2899b52efa85660a2946b63bd6b7bca9723f2d039

Observation 87838670-f2e1-4caa-8d19-a87cf28c806e · outbound

This paper cites Dropout: a simple way to prevent neural networks from overfitting.

GLAD: Generalizable Tuning for Vision-Language Models Dropout: a simple way to prevent neural networks from overfitting

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.158376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:22.207731Z digest=sha256:4c5ec5799a0ecec8d1fb244b76ac313e702c60f0effa2edf8048319a579fd4b3

Observation 64b15ba4-b80f-4605-8475-87b9d93f4409 · outbound

This paper cites Rethinking the inception ar- chitecture for computer vision.

GLAD: Generalizable Tuning for Vision-Language Models Rethinking the inception ar- chitecture for computer vision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:22.304601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:22.304601Z digest=sha256:db56bf4e3746b43b0e58219cf00e3ee56225d31026cbffeab3b10f969a64b21f

Observation f35469ff-89a5-4ed7-97a5-5237913a7af8 · outbound

This paper cites Saliencymix: A saliency guided data augmentation strategy for better regularization.

GLAD: Generalizable Tuning for Vision-Language Models Saliencymix: A saliency guided data augmentation strategy for better regularization

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.126994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:22.449012Z digest=sha256:ffbcbd6a7034a00067b7886574825336452241e70fe2e30b4ef3836428758ce3

Observation 19d8d1bc-bc71-4e05-ba88-830e49bcec3d · outbound

This paper cites Manifold mixup: learning better representations by in- terpolating hidden states.

GLAD: Generalizable Tuning for Vision-Language Models Manifold mixup: learning better representations by in- terpolating hidden states

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.113352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:22.547152Z digest=sha256:82c48c5aa365a418ae9e4afdffa44950fb24d80b0881366c18ef9e4dd1511c63

Observation c0af567c-bd48-4819-a7cd-1d224f61e30c · outbound

This paper cites Tuning multi-mode token- level prompt alignment across modalities.

GLAD: Generalizable Tuning for Vision-Language Models Tuning multi-mode token- level prompt alignment across modalities

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.098342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:22.642193Z digest=sha256:da2fb05414e2774ad63e41a9cda6b180b6b71f07f003bfde4a9fc7cc7a526990

Observation a9d23265-ea2c-46a4-b7c7-f71efadfa374 · outbound

This paper cites Learning robust global representations by penalizing local predictive power.

GLAD: Generalizable Tuning for Vision-Language Models Learning robust global representations by penalizing local predictive power

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:28.066014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:22.737119Z digest=sha256:beb38e00bac3e33f72bfe21c61da3a7e754628824481051447dbf0b5828b4e36

Observation c7bb7d1a-3b7b-46e3-8776-df01c172b0cf · outbound

This paper cites Sharpness-aware gradient matching for domain generaliza- tion.

GLAD: Generalizable Tuning for Vision-Language Models Sharpness-aware gradient matching for domain generaliza- tion

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:27.896307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:22.854498Z digest=sha256:09a92ade10b62e24201de6d644c1cf560d2fbbc9d336054e056c53270ca1c9aa

Observation 93c1dd62-3c84-4a36-bd81-82d3afda935f · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

GLAD: Generalizable Tuning for Vision-Language Models Cogvlm: Visual expert for pretrained language models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:27.723422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:22.959808Z digest=sha256:1c724388baf3305203ae732d028f3f526b7ee05db1aa084875e0a6a71b4560fa

Observation d3ea47da-5c28-4003-ba0e-dcabcfd057ed · outbound

This paper cites Robust fine-tuning of zero-shot models.

GLAD: Generalizable Tuning for Vision-Language Models Robust fine-tuning of zero-shot models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:27.583409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:23.033154Z digest=sha256:d8268ebe7d8780446621de3984f328d0dfda16d21d2d8f393738ad979ef891aa

Observation 26048ba1-384a-4013-ab3e-dc289bd39204 · outbound

This paper cites Sun database: Large-scale scene recognition from abbey to zoo.

GLAD: Generalizable Tuning for Vision-Language Models Sun database: Large-scale scene recognition from abbey to zoo

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:27.403579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:23.079232Z digest=sha256:99ea128142847c6d099d43ae7703fe453bf2fecc149b70468eb36678b97cbf11

Observation 35131259-a89b-4284-bdad-790893de0a37 · outbound

This paper cites Tcp: Textual- based class-aware prompt tuning for visual-language model.

GLAD: Generalizable Tuning for Vision-Language Models Tcp: Textual- based class-aware prompt tuning for visual-language model

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:27.246670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:23.150721Z digest=sha256:a6adaebf17cae5dff1615fb69ddb1b32f14a1ec3553aa91e54146f7a1cf11cf8

Observation 3806585f-bbcc-4e70-be8d-2fde26295261 · outbound

This paper cites Cutmix: Regu- larization strategy to train strong classifiers with localizable features.

GLAD: Generalizable Tuning for Vision-Language Models Cutmix: Regu- larization strategy to train strong classifiers with localizable features

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:27.078177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:23.220423Z digest=sha256:b81c8fed207bbfd756ca71227ca8a0b449350961421d6635ad5b4673aca627cd

Observation ebfd872d-fb0f-49e5-88fb-d918115b5910 · outbound

This paper cites Low-rank few-shot adaptation of vision-language models.

GLAD: Generalizable Tuning for Vision-Language Models Low-rank few-shot adaptation of vision-language models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:26.915990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:23.292963Z digest=sha256:6f51b2f1d7e1e982d52a4f50df72c9bf0021e3feb3b3565e6e1049c302d27e00

Observation 626eb206-3728-4927-90aa-835e1c098562 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

GLAD: Generalizable Tuning for Vision-Language Models Lit: Zero-shot transfer with locked-image text tuning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:26.763594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:23.388218Z digest=sha256:23ad6562b10670584bbc1e8f10b2fad23f21db753b822f2177b9dda8230f51eb

Observation ed262ff5-852e-41f0-9a30-dbb356a79355 · outbound

This paper cites Three mechanisms of weight decay regularization.

GLAD: Generalizable Tuning for Vision-Language Models Three mechanisms of weight decay regularization

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:26.620181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:23.425869Z digest=sha256:fe42e4cace27363c18565700cfd81f3536ffff19c982c0b0251999698c3116f2

Observation d28e2228-19a9-496c-8ffc-62e3c36aeeb4 · outbound

This paper cites mixup: Beyond Empirical Risk Minimization.

GLAD: Generalizable Tuning for Vision-Language Models mixup: Beyond Empirical Risk Minimization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:23.459308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:23.459308Z digest=sha256:0da7dada6e27b46f3e83bb369248cde0403f1e8c368b02df86f4df01ab2bf476

Observation 4834b420-d026-4e38-86e4-76b17aadee68 · outbound

This paper cites Dept: Decoupled prompt tuning.

GLAD: Generalizable Tuning for Vision-Language Models Dept: Decoupled prompt tuning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:26.426058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:23.526271Z digest=sha256:bcaf7f1ee9806d31da88f01c9d5b516566b2df63734deb33ce9fa32d3c976852

Observation a21fe1d1-d66b-4447-9f44-667a3af33382 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

GLAD: Generalizable Tuning for Vision-Language Models Adding conditional control to text-to-image diffusion models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:23.593178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:23.593178Z digest=sha256:4e96340bc0b7c368f40c2fc412aafe8d6bf2cc9a70f436effe3139e829c08f55

Observation b7512f2a-6b89-4be1-b313-2e6847175b69 · outbound

This paper cites Llama-adapter: Efficient fine-tuning of language models with zero-init attention.

GLAD: Generalizable Tuning for Vision-Language Models Llama-adapter: Efficient fine-tuning of language models with zero-init attention

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:26.126021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:23.638487Z digest=sha256:bc9e20b797182b34ea9e24a8ba4fa01c77325c0c2d468541ef2391095f68d562

Observation c04e7633-c62d-43a3-ba52-725bb69e0a6d · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

GLAD: Generalizable Tuning for Vision-Language Models SVIT: Scaling up Visual Instruction Tuning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:23.687949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:23.687949Z digest=sha256:b0c6f8ddca189f72a9ca15dbe8afa945c7bc3a7d27906f68b201e5651fa54586

Observation 2dadc935-ceed-472a-a3cd-822b7c5036d8 · outbound

This paper cites Regionclip: Region-based language-image pretraining.

GLAD: Generalizable Tuning for Vision-Language Models Regionclip: Region-based language-image pretraining

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:25.879634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:23.748806Z digest=sha256:d663ff3451b3b944fbc6b8572a90ead33d6494d1b0fa33e5c98feeb54ec4e3d5

Observation 7d993c5f-112d-4b83-b15e-97df888bc360 · outbound

This paper cites Conditional prompt learning for vision-language models.

GLAD: Generalizable Tuning for Vision-Language Models Conditional prompt learning for vision-language models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:25.618806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:23.856095Z digest=sha256:fa9306ecae972e97f88ce7bf2f8fdf9351f27eae654e020329bff38ae6160267

Observation 2cbf78dd-d784-4289-a2ae-135dcd884f41 · outbound

This paper cites Learning to prompt for vision-language models.

GLAD: Generalizable Tuning for Vision-Language Models Learning to prompt for vision-language models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:37:25.237348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T16:37:23.927447Z digest=sha256:3f7fb0fc50c05a394e36834d4587b25f6a4d41b6a05a5b665d5504a18c9c4c88

Observation e2a1921c-ccf8-4a8d-a30d-298f18332a4e · outbound

This paper cites Prompt-aligned gradient for prompt tuning.

GLAD: Generalizable Tuning for Vision-Language Models Prompt-aligned gradient for prompt tuning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:23.993559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:23.993559Z digest=sha256:8b77e81d6ce5c6898eb1c18a88bc63af8bcf7c2a7a9fa7a5404d7d9493cbd723

Observation e9c866b0-7b5d-4199-be01-49ece54bf2c4 · outbound

This paper cites Surrogate Gap Minimization Improves Sharpness-Aware Training.

GLAD: Generalizable Tuning for Vision-Language Models Surrogate Gap Minimization Improves Sharpness-Aware Training

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:24.063644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:24.063644Z digest=sha256:09ab31c177ec0e7ca204d8ad5474f6b5ea0151700e0262dab2a4dcc5d28eb5fb

Pith citing papers

Observation bd78f140-02f5-4e16-b878-6d00bf653a39 · inbound

TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models cites this paper.

TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models GLAD: Generalizable Tuning for Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T21:24:53.437554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:24:53.437554Z digest=sha256:d5fa70f408e5134b79f19d960efdd9395d582069d22a798ed571cfa0150b9fc4

Observation 30255c27-ef3b-4a30-b360-43510c56f66e · inbound

GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models cites this paper.

GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models GLAD: Generalizable Tuning for Vision-Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:40:26.018918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T13:39:15.820777Z digest=sha256:2b06ff0bdcc03bf4f9a7adf28bb0ecee9d618abb9bb70df5dd93254d236b5281

Observation 240109b8-82e3-4783-b64a-5230a8efe68a · inbound

GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models cites this paper.

GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models GLAD: Generalizable Tuning for Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T20:25:41.773345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:25:41.773345Z digest=sha256:29d9b2954398ca36f2c69c2861a2651d63e0f6fcf5e8de8e5b8151c74e3e9328