Pith. sign in

Paper Citation Record · LEDGER

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering

As of 10 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2505.21755.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21755 v2

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:30:26.374735Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact5
  • verified fuzzy37
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99298249-d5ce-4d24-b7ed-7a2ec893d70d · outbound

This paper cites To- wards Causal VQA: Revealing and Reducing Spurious Cor- relations by Invariant and Covariant Semantic Editing.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering To- wards Causal VQA: Revealing and Reducing Spurious Cor- relations by Invariant and Covariant Semantic Editing

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:36.480682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:19.018950Z digest=sha256:af3b5478f6fc8bb3d931b7227f3072da7267b94223d5c02e787615cc7483a9cb

Observation 7c3886df-9f77-4bec-8bf0-cc141573074c · outbound

This paper cites Don't Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Don't Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:19.101688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:19.101688Z digest=sha256:f5ab466287014a74e3623ade572511052243b8701e42ed8fccb732408c424ae9

Observation 6ad4b3af-339d-4ff5-854f-99457823b23a · outbound

This paper cites Reassessing Evaluation Practices in Visual Question Answering: A Case Study on Out-of-Distribution Generalization.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Reassessing Evaluation Practices in Visual Question Answering: A Case Study on Out-of-Distribution Generalization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:19.222653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:19.222653Z digest=sha256:b5f4181e3bef3a1b9071cd2e4eae2151426babcafddf419c2bad79f6b477119f

Observation 4e779346-0c0d-4320-95e4-07b7788f6653 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering PaliGemma: A versatile 3B VLM for transfer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:19.360656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:19.360656Z digest=sha256:c2245848b42dab58dec98eabf824260259310d5778dc97bd5689782843b1d9aa

Observation 9ebf6be0-401f-4a07-97d5-2046c39be37b · outbound

This paper cites Cunningham.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Cunningham

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:36.279981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:19.440154Z digest=sha256:c7857fc52f5c6acdc6063029640df4689ff3808cc83a5da1d211357ba5207390

Observation 8ba14324-63cd-4c7b-a137-3ebabc3d09e7 · outbound

This paper cites VizWiz: nearly real-time answers to visual questions.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering VizWiz: nearly real-time answers to visual questions

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:36.028569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:19.536741Z digest=sha256:1fa21770b023b019991f28fc8ececfda1a86c7da0c48b219008322c69451f634

Observation e50859f1-9e25-4f6f-b1d8-9dbd59188da5 · outbound

This paper cites Behind the scene: Revealing the secrets of pre-trained vision-and-language models, 2020.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Behind the scene: Revealing the secrets of pre-trained vision-and-language models, 2020

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:35.887545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:19.664338Z digest=sha256:10f51c3a6c678ecfcec315778d53414173e668bd6c0a617438c756a38ca8d582

Observation 77e73c75-bb3d-49f4-b248-701bdd331e89 · outbound

This paper cites Benchmarking robustness of adaptation methods on pre-trained vision-language models, 2023.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Benchmarking robustness of adaptation methods on pre-trained vision-language models, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:35.718142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:19.809894Z digest=sha256:2c24d27d7d580fc4da94102dbd98f90454b268a56db8e1bb30d6d8331a73c1c6

Observation 44971840-6caa-4117-abe1-8330f81867b9 · outbound

This paper cites Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question Answering.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question Answering

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:30:28.679257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:19.956939Z digest=sha256:1a89a5dffdc8d7f3741f4d90d0fb8c32e499720d127c9a36ff90e3d233297c50

Observation c0f29e56-f5ff-4e9f-b049-7c03a6c03f70 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Imagenet: A large-scale hierarchical image database

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:35.581963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:20.044752Z digest=sha256:3699014b887b2686ee48ea5e3124903c2b2e33ef25b3bf4ca0f761b22ef67d51

Observation aea9155f-6a9a-4985-9e7e-b5502aed4e13 · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding, 2019.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Bert: Pre-training of deep bidirectional trans- formers for language understanding, 2019

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:20.156174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:20.156174Z digest=sha256:6c87fabf60ef35c6a4b4b49b8b3c5d49f4d93b0050ac3e7089a8f3329baad5a7

Observation 82cba0ee-3bdf-4bd5-9d4b-0a5e2146c87a · outbound

This paper cites an unresolved cited work.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:35.329020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:20.272974Z digest=sha256:dc6640f5acc4ce6fd291f6f2f62b396ae8bc0eddfea71d69e96f5d94c3947674

Observation 10ed4b95-718c-4c46-8997-a868a8d4c553 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale, 2021.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering An image is worth 16x16 words: Transformers for image recognition at scale, 2021

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:20.442632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:20.442632Z digest=sha256:b30b5869771e8f78f003982b3c8ab7de51b7b545fafd1d30f520178a5696edd7

Observation 3fd68671-db85-402f-9296-56412498f622 · outbound

This paper cites Ex- ploring the limits of out-of-distribution detection, 2021.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Ex- ploring the limits of out-of-distribution detection, 2021

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:35.140662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:20.550685Z digest=sha256:a1cac5d8abdc27899403820d88e2018f5a7e2bcb290862b28e45df190291c9c6

Observation 3bcc6fc5-54b4-4187-8022-96eedcfaa940 · outbound

This paper cites VQA-LOL: Visual Question Answering under the Lens of Logic.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering VQA-LOL: Visual Question Answering under the Lens of Logic

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:20.658990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:20.658990Z digest=sha256:d2789292f37ac0262468579090baea16f76a2a59e0faf52dd0ffa99960200d43

Observation c8732f46-27d1-4d29-8f49-7bf25efb7c6e · outbound

This paper cites paligemma-3b-pt-224.https : / / huggingface.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering paligemma-3b-pt-224.https : / / huggingface

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:34.953633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:20.880562Z digest=sha256:f7a8bc7b122f6697be0261cfa931b94438f6a2f3f7d8bce2385c3070d6fc9290

Observation be0a20f7-6c23-4b5e-9163-7a69bc48e666 · outbound

This paper cites Distance-Based Regularisation of Deep Networks for Fine-Tuning.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Distance-Based Regularisation of Deep Networks for Fine-Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:20.970889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:20.970889Z digest=sha256:4f4c09fa61efd965d5ea3f21624e8f5d0e41a6bc7bc6dd5dfb82bf54026e0718

Observation faf0a017-4266-436e-8b13-55f39c81540f · outbound

This paper cites Finetune like you pretrain: Im- proved finetuning of zero-shot vision models.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Finetune like you pretrain: Im- proved finetuning of zero-shot vision models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:34.763178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:21.092328Z digest=sha256:4ceaa362690ebcf4913f0045545faaa8e6f5375f117dc2aa1a4e97609d6a8fae

Observation f59dcad5-8993-44fe-b3f2-142718843a77 · outbound

This paper cites Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:21.172080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:21.172080Z digest=sha256:ce8b685614de950780be9dbc1e4b316606d4d8fae1bf2c86b54cd288d1be183b

Observation e87c518b-01a9-425f-9299-bd734b8ffbc4 · outbound

This paper cites Rasch, Bern- hard Scholkopf, and Alexander J.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Rasch, Bern- hard Scholkopf, and Alexander J

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:34.536315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:21.290993Z digest=sha256:55eb4a6bd67f6a736ca1879f514983ed36eed9e805d8826e1150dd790448109c

Observation c1e5c8ef-62a2-40d5-82f1-ce0e51eaea32 · outbound

This paper cites The many faces of robust- ness: A critical analysis of out-of-distribution generalization.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering The many faces of robust- ness: A critical analysis of out-of-distribution generalization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:34.369265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:21.394139Z digest=sha256:79d4008119807dc4e27fa16c994509b820191486ee43de77690ed2a29cde413f

Observation 2634c300-530b-41dd-a524-07c7847c98ad · outbound

This paper cites Natural adversarial examples.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Natural adversarial examples

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:34.128657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:21.479518Z digest=sha256:cd90d9dca31b3ec1857b835fbc4a8ffa34b680c0202d54c4b605b988035723c8

Observation c7664a7c-0b7f-4741-a1a9-3fcfcba3c88f · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:33.937969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:21.569817Z digest=sha256:d0ed072c187484f4be774d60db9232ac930360919815ede1f59e788d22e6649d

Observation 1a0fa88d-cea0-4893-9918-b7a371fa3a56 · outbound

This paper cites Llm-adapters: An adapter family for parameter- efficient fine-tuning of large language models, 2023.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Llm-adapters: An adapter family for parameter- efficient fine-tuning of large language models, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:33.789467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:21.734597Z digest=sha256:3dc4de4573dcddfc6590424dc33f3409cb87618561ea4eb3d561e180070c493a

Observation f819ba40-316b-4a04-af4b-b2d3997b083c · outbound

This paper cites Directional gradient pro- jection for robust fine-tuning of foundation models, 2025.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Directional gradient pro- jection for robust fine-tuning of foundation models, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:33.617956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:21.845558Z digest=sha256:18e53305b749ae76d2f4829e7069bc9efdd0ba1fb7d03a71ce744181851ae6ef

Observation 5b9d8667-9123-4559-9f65-d7330808bd6e · outbound

This paper cites Hudson and Christopher D.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Hudson and Christopher D

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:22.034752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:22.034752Z digest=sha256:b6a1aaca85ff160f8a77e2b3377f4c9ef6ba47d7ca8bfb817676a90650ae10b6

Observation 38927f64-8faf-4754-ba47-630122dc47f3 · outbound

This paper cites Roses are red, violets are blue.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Roses are red, violets are blue

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:33.411709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:22.147432Z digest=sha256:6ba26588712db3c011a9e5ef402c25580a6017dda718b3ba0c0c2fc02ab4c118

Observation 38c1c61e-f2ce-4191-878a-a44fc2f8c45c · outbound

This paper cites Fine-Tuning can Distort Pre- trained Features and Underperform Out-of-Distribution,.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Fine-Tuning can Distort Pre- trained Features and Underperform Out-of-Distribution,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:33.216026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:22.264829Z digest=sha256:6f579e0ee1bf65b3449c80fbc168a57647ebe642c0d1270884f31e4be4bb75ca

Observation 69669264-d650-4b15-a2e2-dae32b297438 · outbound

This paper cites an unresolved cited work.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:32.968130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:22.408385Z digest=sha256:cb9d8798358faa727d7acaa0804b5af6dc95d409f50a891adcd73958735d4fc4

Observation 1fd31f06-8625-42f2-9ae5-51f6f54e6a0d · outbound

This paper cites A Closer Look at the Robustness of Vision-and-Language Pre-trained Models,.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering A Closer Look at the Robustness of Vision-and-Language Pre-trained Models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:32.743187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:22.500847Z digest=sha256:66e155ce08c278e74068f8f4e3efd8fe186b8e75d5da8310e19ebd81ff4082b2

Observation fb99404b-2d19-48aa-870b-19bb764409bc · outbound

This paper cites Adversarial VQA: A New Benchmark for Evaluating the Robustness of VQA Models.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Adversarial VQA: A New Benchmark for Evaluating the Robustness of VQA Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:30:28.504421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:22.701978Z digest=sha256:fe3a3516667d9a8e2a129a14ad9d72acf214119d10da1559be35dbd9500fc3ca

Observation 898e22db-0104-4d4b-86ed-89407c199d5c · outbound

This paper cites Explicit Inductive Bias for Transfer Learning with Convolutional Networks.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Explicit Inductive Bias for Transfer Learning with Convolutional Networks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:22.787019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:22.787019Z digest=sha256:a600f1b64166b341631e0bd3cd615b7820aab63ce3e4a90cc7d2b57ef42ec720

Observation f660c6de-5279-4c97-8d8e-b59d828749fe · outbound

This paper cites A Closer Look at the Robustness of Vision-and-Language Pre-trained Models.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering A Closer Look at the Robustness of Vision-and-Language Pre-trained Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:22.632151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:22.632151Z digest=sha256:eb1d4178d6512c92237c7f62f1a8539ea5c67b6f658063cb1ec856c110f25062

Observation a1ffed09-94f3-46ab-9c52-9630f9bc08ce · outbound

This paper cites Robust Visual Ques- tion Answering: Datasets, Methods, and Future Challenges,.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Robust Visual Ques- tion Answering: Datasets, Methods, and Future Challenges,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:32.299018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:22.966162Z digest=sha256:c9d6034f2867a587005a037f759b77830a5be621297aafd0e4fc01dedc24b6f2

Observation 4269af3a-bb07-4034-97a6-0a7463af80b3 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.184037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.184037Z digest=sha256:1072619f9d2c8bcba0b9dd91d232a8aec0bc602e8c124d1c346d534a971d5c8e

Observation 90c00bb2-bd48-4863-b076-1d19ca4046e0 · outbound

This paper cites Visual instruction tuning, 2023.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Visual instruction tuning, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:32.543685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:22.844130Z digest=sha256:1dc5ec3c8349b366f5e3a06c7b975286f719ea4131c35d5186045d587b68bb42

Observation b7231ea7-fe95-46c9-9adb-f7da0ac9cfb9 · outbound

This paper cites Maximum mean discrep- ancy for generalization in the presence of distribution and missingness shift, 2022.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Maximum mean discrep- ancy for generalization in the presence of distribution and missingness shift, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:31.868985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:23.335849Z digest=sha256:b557bed0bf46077fb75cac9707c57328c7dcc49add698e86f855b5263d7e1687

Observation 0e73b822-b0da-4290-9ba1-a09fb3a64d27 · outbound

This paper cites Moment matching for multi-source domain adaptation.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Moment matching for multi-source domain adaptation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.427466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.427466Z digest=sha256:0264c2efdf64558b5c874cc47eb7b39c7106799fdab57cbc97c6a6147638bb13

Observation 88f7ff7c-038f-4bda-b3a1-b973613dad95 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Learning Transferable Visual Models From Natural Language Supervision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.504571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.504571Z digest=sha256:ec29c5324cf8d593aee2bf3260835cef3a7d253166a391ab627c16b66f0b7f7d

Observation 61fbbd8a-08f4-49c3-8cb9-3fdc7df2d140 · outbound

This paper cites Generalized out-of-distribution detection and be- yond in vision language model era: A survey, 2024.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Generalized out-of-distribution detection and be- yond in vision language model era: A survey, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:32.111032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:23.272303Z digest=sha256:f39cebf79ddcf26522cc050a11f3b756c7215e7af4f553659ed34283d3cc0b23

Observation d080134e-5c9b-43f2-8bcd-be03d72a3bab · outbound

This paper cites Cycle-Consistency for Robust Visual Question Answering,.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Cycle-Consistency for Robust Visual Question Answering,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:31.381355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:23.712779Z digest=sha256:97e9fc9523b24fa3006100bb89c47f87aaf101e0797ca1842073f2a41a831ab4

Observation d0823e45-7a42-42a4-b1a4-ca9660017ca5 · outbound

This paper cites Human-Adversarial Visual Question Answering.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Human-Adversarial Visual Question Answering

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.821325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.821325Z digest=sha256:27116c1ceb025cd62cf9e042ef66ac9fa533df35fea7ccaf604080e5b9b91e45

Observation fd9ff035-e16b-4adf-9f57-3b6b0094dd7f · outbound

This paper cites Benchmarking out-of- distribution detection in visual question answering, 2024.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Benchmarking out-of- distribution detection in visual question answering, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:31.214236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:23.931683Z digest=sha256:114deefaa03fbf14b1a9506c512c38293a2cc117ef224b5658579641230b6124

Observation 2388f9b5-1477-41d1-8439-7a3747386f0b · outbound

This paper cites Do imagenet classifiers generalize to im- agenet? InInternational Conference on Machine Learning, pages 5389–5400.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Do imagenet classifiers generalize to im- agenet? InInternational Conference on Machine Learning, pages 5389–5400

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:31.584883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:23.652233Z digest=sha256:cbd8321186fb26035c293a85084ea907e23778d95b9643295b3ee5f064e35247

Observation 8bbb6535-3567-4ef4-aa60-0102cf1f9303 · outbound

This paper cites Towards VQA Models That Can Read.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Towards VQA Models That Can Read

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.067003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.067003Z digest=sha256:db3d51fedc44ba150447a681c4807dc28e9e3f43f82bef2c51d0b0c3d8245071

Observation 0e42019f-cbda-4e04-ba92-79ac1ec40df9 · outbound

This paper cites Trainable Projected Gradient Method for Robust Fine-tuning.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Trainable Projected Gradient Method for Robust Fine-tuning

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:30:27.443352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:24.161750Z digest=sha256:88619c13bd4838bd0316479358fedd5a9d217f27d348f331aa1a1886d152ac16

Observation 33668d0f-2c8c-4cc1-ac66-6d6060229712 · outbound

This paper cites Fast Trainable Projection for Robust Fine-Tuning,.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Fast Trainable Projection for Robust Fine-Tuning,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.993713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:24.217760Z digest=sha256:1ebf9818ec7a37a240ca3306a51c90a6676f2754a1d5fa4b8fce384c50666816

Observation 47b5d286-8974-4be9-89b1-1f5097860ca0 · outbound

This paper cites Rethinking weight decay for robust fine-tuning of foundation models,.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Rethinking weight decay for robust fine-tuning of foundation models,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.903064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:24.338273Z digest=sha256:ee89491b13b7567ac500c361430f0512664d1bf9f7f8b2db023c3474f9f392be

Observation 811af034-14bf-4280-b2cd-fe90acc791f7 · outbound

This paper cites Language Prior Is Not the Only Shortcut: A Benchmark for Shortcut Learning in VQA.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Language Prior Is Not the Only Shortcut: A Benchmark for Shortcut Learning in VQA

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:30:27.964956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:23.986835Z digest=sha256:6362c8ee7e31e16db48d5998934c5400e33feec07552b988e498d7784aaf38c6

Observation 1904a24e-0b23-4a90-b896-316e078bdc05 · outbound

This paper cites Learning robust global representations by penalizing local predictive power.Advances in Neural Information Pro- cessing Systems, 32, 2019.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Learning robust global representations by penalizing local predictive power.Advances in Neural Information Pro- cessing Systems, 32, 2019

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.624114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.624114Z digest=sha256:d9ba06eee5863e0ee33bd5fedbc8865cd634db30f2e2ad29de9685cdb28b7467

Observation 384808c0-cd99-425e-bbff-20e5b2a95988 · outbound

This paper cites Ledsam, Patricia MacWilliams, Pushmeet Kohli, Alan Karthikesalingam, Si- mon Kohl, Taylan Cemgil, S.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Ledsam, Patricia MacWilliams, Pushmeet Kohli, Alan Karthikesalingam, Si- mon Kohl, Taylan Cemgil, S

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.671410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:24.721174Z digest=sha256:d253d5a5c9aa00259174192e45798d0c937620fcb1b843e6555aa519928122cb

Observation 4d9224da-b56a-4bd2-9780-c3c15738bbe4 · outbound

This paper cites Robust fine-tuning of zero-shot models.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Robust fine-tuning of zero-shot models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.804961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.804961Z digest=sha256:d2f930832be2964b3bd33ed47906e27fabe79e023cd70e9a749137eca645fcfb

Observation 4c6aae56-da85-417d-b12f-b6f063fb4c4e · outbound

This paper cites Domain-robust vqa with di- verse datasets and methods but no target labels, 2021.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Domain-robust vqa with di- verse datasets and methods but no target labels, 2021

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.489788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:24.935344Z digest=sha256:ff7a2855e164d04d45c991f232daa450bf010ffd538aac9b07e6453c065462b2

Observation 941f0fc7-0f4d-4992-9b64-cae1526f5f1b · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.323868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:25.034749Z digest=sha256:7497a00ecf77a8fd8a66fbbf499dbfb04e39afc9d5fb862d6d06c1aef55d984a

Observation d2431941-ff05-4f0e-b6d8-23218ac4ef36 · outbound

This paper cites VQA-GEN: A Visual Question Answering Benchmark for Domain Generalization.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering VQA-GEN: A Visual Question Answering Benchmark for Domain Generalization

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:30:26.964752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:24.515854Z digest=sha256:b87deea5a52f5b20a51494ca5248032d771f566667269626657a2377a0e6f336

Observation 80a62b49-695d-4d90-ba59-05d7e6b9cb98 · outbound

This paper cites We use the LA VIS [29] public repository to fine-tune all methods.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering We use the LA VIS [29] public repository to fine-tune all methods

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.195941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:25.244751Z digest=sha256:6100a8ed63200cc0afabce79de5fcbc3a5d30f5029ac4f876fca5302dbb40416

Observation 69fa2271-0d8d-48be-93be-d702805d9b76 · outbound

This paper cites an unresolved cited work.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:29.997695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:25.374892Z digest=sha256:28a68c45f08fe19f667c7480763f3bce9b54877086b22c92c377d7e70aaa7809

Observation fb31c073-9b8d-4691-bbd2-38ea04c2bb51 · outbound

This paper cites 7 shows the correlation between shift and performance for different embeddings under different fine-tuning meth- ods.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering 7 shows the correlation between shift and performance for different embeddings under different fine-tuning meth- ods

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:29.837989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:25.632830Z digest=sha256:99a15e6d45752f9259bbb528e2c64b7444fae0eaa4a29e4dba07389958f06b0b

Observation be9795bd-9761-486a-a85a-21852dd5a165 · outbound

This paper cites 5 shows the heatmap of the correlation between uni- modal and multi-modal shifts per dataset.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering 5 shows the heatmap of the correlation between uni- modal and multi-modal shifts per dataset

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:29.630166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:25.742953Z digest=sha256:b4f6017a568f3fa756e7de538547ee216ef4be652a897b9fcd567793c4f365a5

Observation b7f031ed-deeb-45c6-932c-10e582a2fabb · outbound

This paper cites 13 and 14 show the variation of MIv and MIq w.r.t.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering 13 and 14 show the variation of MIv and MIq w.r.t

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:29.448148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:25.865612Z digest=sha256:d23a359b6ddef8162e85698821dcd63bbae6318192b41b60339e2baab40bf618

Observation 84e2c426-a2b8-4bfa-bdec-637609601f3e · outbound

This paper cites 8, including LLaV A- 7B [33] with LoRA and PaliGemma-3B with full fine- tuning.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering 8, including LLaV A- 7B [33] with LoRA and PaliGemma-3B with full fine- tuning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:29.274521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:25.995036Z digest=sha256:7b27d37138dd18ab1e0819e38305ecd7605dd1401c42e454909fb41da3c96cc0

Observation 1692ec5f-f3e9-49f8-a3bb-aef9de954d75 · outbound

This paper cites The only exception, GQA-OOD [27] (based on GQA [26]), has only answer shifts.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering The only exception, GQA-OOD [27] (based on GQA [26]), has only answer shifts

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:29.064764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:26.134838Z digest=sha256:2a6de7de1efa5fec6ecd67aafb7e8e31f3d0319673ef3dd126be88e9b266d219

Observation 21b24e27-5652-4d7f-a8f0-b4f6b21c4e7e · outbound

This paper cites We further compare shifts using Maximum Mean Discrepancy (MMD) [12, 20, 37] with RBF kernel in Tab.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering We further compare shifts using Maximum Mean Discrepancy (MMD) [12, 20, 37] with RBF kernel in Tab

Reference 68

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T13:30:28.908369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:26.255127Z digest=sha256:b97f2669be0a20d5238869865f7348d7191ddc8fb972c5c05b2ad9c0f9709c37

Observation 58e707d3-ed14-4556-88d1-c455496b3fa2 · outbound

This paper cites This also serves as a veri- fication of the reliability in quantifying shifts via feature- based representations.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering This also serves as a veri- fication of the reliability in quantifying shifts via feature- based representations

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:28.802066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:30:26.374735Z digest=sha256:0161a4681a5c1c45f43375299312bf2e03a3ddec0d84529deefcf220a940a0b3

Observation 975ea182-8c39-49c2-a8b6-9a6b49c69ca0 · outbound

This paper cites Cycle-Consistency for Robust Visual Question Answering.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Cycle-Consistency for Robust Visual Question Answering

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.772858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.772858Z digest=sha256:cead9baa445a79266af0d54bf513bcd95b1491f10c038f934f47f15bc98edaa9

Observation 79e8ca76-f1a5-4db2-b092-1033396ce1e4 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering LoRA: Low-Rank Adaptation of Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:21.655216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:21.655216Z digest=sha256:26276c14d54ab8bb9fddeee0e003b8a8bb7590db1a84ff17a6384ee4a7bf0866

Observation e257422a-c06e-47bf-8cc4-76b650191457 · outbound

This paper cites Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:22.333864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:22.333864Z digest=sha256:318727bd82482190a8c92ffb2d84c83ab069211688e3832aa1cf970ae5740f82

Observation d014ed24-983a-45e9-8e21-1acf215453ba · outbound

This paper cites Fast Trainable Projection for Robust Fine-Tuning.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Fast Trainable Projection for Robust Fine-Tuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.269594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.269594Z digest=sha256:38c8073c056f63e9ff147ef82dc6d4ebdf1abee760b426e803d10e01361fb9ca

Observation 46c7e70c-9d2e-4cca-9013-9010f64a7775 · outbound

This paper cites Robust Visual Question Answering: Datasets, Methods, and Future Challenges.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Robust Visual Question Answering: Datasets, Methods, and Future Challenges

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.075341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.075341Z digest=sha256:f6d14647c2160226ea6a64359ddf33f37cdfcd051e886eecf104f586922df38c

Pith citing papers

No inbound Pith citation observations are available.