Pith. sign in

Paper Citation Record · LEDGER

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering

As of 9 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2505.21755.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21755 v2

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:30:26.374735Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact5
  • verified fuzzy37
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99298249-d5ce-4d24-b7ed-7a2ec893d70d · outbound

This paper cites To- wards Causal VQA: Revealing and Reducing Spurious Cor- relations by Invariant and Covariant Semantic Editing.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering To- wards Causal VQA: Revealing and Reducing Spurious Cor- relations by Invariant and Covariant Semantic Editing

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:36.480682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:19.018950Z digest=sha256:8e2c17167f3439bfaf5744002686bf4c9146698dcaee719508372b82f23ab2e1

Observation 7c3886df-9f77-4bec-8bf0-cc141573074c · outbound

This paper cites Don't Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Don't Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:19.101688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:19.101688Z digest=sha256:0e3267ac961952dbce782d5797d4f01606a0a1b785ca7c5d42a8facec34d8b4d

Observation 6ad4b3af-339d-4ff5-854f-99457823b23a · outbound

This paper cites Reassessing Evaluation Practices in Visual Question Answering: A Case Study on Out-of-Distribution Generalization.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Reassessing Evaluation Practices in Visual Question Answering: A Case Study on Out-of-Distribution Generalization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:19.222653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:19.222653Z digest=sha256:db642a8ffb85e50b9d0f004ec73f9dc8ed1698276661eef3fe3727867a11c8e0

Observation 4e779346-0c0d-4320-95e4-07b7788f6653 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering PaliGemma: A versatile 3B VLM for transfer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:19.360656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:19.360656Z digest=sha256:d4efa70d1df706d14b8ecf0e10dd4df7724be510144dbd440a8547d793d167e3

Observation 9ebf6be0-401f-4a07-97d5-2046c39be37b · outbound

This paper cites Cunningham.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Cunningham

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:36.279981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:19.440154Z digest=sha256:d3f552760849b13206e49876538299cfedfeb3b75290daf0ac04efe4ab436868

Observation 8ba14324-63cd-4c7b-a137-3ebabc3d09e7 · outbound

This paper cites VizWiz: nearly real-time answers to visual questions.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering VizWiz: nearly real-time answers to visual questions

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:36.028569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:19.536741Z digest=sha256:7e3643743b8545967ac0276b386cf203f809e1da778dd4b28cfacbb9800acfd9

Observation e50859f1-9e25-4f6f-b1d8-9dbd59188da5 · outbound

This paper cites Behind the scene: Revealing the secrets of pre-trained vision-and-language models, 2020.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Behind the scene: Revealing the secrets of pre-trained vision-and-language models, 2020

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:35.887545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:19.664338Z digest=sha256:6b3f1f17b9632a95a187d7a4334858718bf53076e698d6c540996bcb1f57c8f2

Observation 77e73c75-bb3d-49f4-b248-701bdd331e89 · outbound

This paper cites Benchmarking robustness of adaptation methods on pre-trained vision-language models, 2023.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Benchmarking robustness of adaptation methods on pre-trained vision-language models, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:35.718142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:19.809894Z digest=sha256:aa7c8939a8d458bbc6fab2f204caf32d3a354305047509e76dd667bf3931de6b

Observation 44971840-6caa-4117-abe1-8330f81867b9 · outbound

This paper cites Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question Answering.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question Answering

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:30:28.679257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:19.956939Z digest=sha256:94c59b7e826d2ef0c30c956a5e6a737d3f76161d7b0475ac76e6ec6d8246d8ed

Observation c0f29e56-f5ff-4e9f-b049-7c03a6c03f70 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Imagenet: A large-scale hierarchical image database

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:35.581963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:20.044752Z digest=sha256:8e8f713cbf7ae3c1c4c67534b460d36d368f94e85da299cd8557574fc56354a9

Observation aea9155f-6a9a-4985-9e7e-b5502aed4e13 · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding, 2019.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Bert: Pre-training of deep bidirectional trans- formers for language understanding, 2019

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:20.156174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:20.156174Z digest=sha256:f907ae565ad5ebdf1c1034345432a24c163f50bdc88cd577caaed11260a03510

Observation 82cba0ee-3bdf-4bd5-9d4b-0a5e2146c87a · outbound

This paper cites an unresolved cited work.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:35.329020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:20.272974Z digest=sha256:7f2b5ecc73808c40fccd86e8337e2a9ead3aeffb0ba5256f59934f6daa768a30

Observation 10ed4b95-718c-4c46-8997-a868a8d4c553 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale, 2021.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering An image is worth 16x16 words: Transformers for image recognition at scale, 2021

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:20.442632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:20.442632Z digest=sha256:f8da1ef8ffa85aac67fbc6bbc2ba564ac3d929c0d2b0ba9d0795278b04935ddb

Observation 3fd68671-db85-402f-9296-56412498f622 · outbound

This paper cites Ex- ploring the limits of out-of-distribution detection, 2021.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Ex- ploring the limits of out-of-distribution detection, 2021

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:35.140662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:20.550685Z digest=sha256:9fdde113c74ba4571febff7388650b298bbcec2e5bdb46c5b5d91f88b4c4ff84

Observation 3bcc6fc5-54b4-4187-8022-96eedcfaa940 · outbound

This paper cites VQA-LOL: Visual Question Answering under the Lens of Logic.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering VQA-LOL: Visual Question Answering under the Lens of Logic

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:20.658990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:20.658990Z digest=sha256:c17c48b110c04848b52e7bc25d097aed2a75c17a089501bab48f3168ebb85c8c

Observation c8732f46-27d1-4d29-8f49-7bf25efb7c6e · outbound

This paper cites paligemma-3b-pt-224.https : / / huggingface.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering paligemma-3b-pt-224.https : / / huggingface

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:34.953633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:20.880562Z digest=sha256:ee470985f851314097337856958aab64e80049c8096b74f2b1bd25264ead933d

Observation be0a20f7-6c23-4b5e-9163-7a69bc48e666 · outbound

This paper cites Distance-Based Regularisation of Deep Networks for Fine-Tuning.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Distance-Based Regularisation of Deep Networks for Fine-Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:20.970889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:20.970889Z digest=sha256:da617b0271a692a13f709a65f40ed167862f138ac56ee1017fa4c5a21bc7ff91

Observation faf0a017-4266-436e-8b13-55f39c81540f · outbound

This paper cites Finetune like you pretrain: Im- proved finetuning of zero-shot vision models.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Finetune like you pretrain: Im- proved finetuning of zero-shot vision models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:34.763178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:21.092328Z digest=sha256:f2b92fbf11de0f759653a66d8dfa56843a823f0b50ee5d75eb3eebad98468d52

Observation f59dcad5-8993-44fe-b3f2-142718843a77 · outbound

This paper cites Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:21.172080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:21.172080Z digest=sha256:8a9ef544f561b0e80efbb76f48af5c987122e3ee724ba45a87d60b0a6a965ea5

Observation e87c518b-01a9-425f-9299-bd734b8ffbc4 · outbound

This paper cites Rasch, Bern- hard Scholkopf, and Alexander J.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Rasch, Bern- hard Scholkopf, and Alexander J

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:34.536315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:21.290993Z digest=sha256:a21c57e74cc2f8a92936e34049a5b18840545c853437ad4c3a2a8db18af481f8

Observation c1e5c8ef-62a2-40d5-82f1-ce0e51eaea32 · outbound

This paper cites The many faces of robust- ness: A critical analysis of out-of-distribution generalization.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering The many faces of robust- ness: A critical analysis of out-of-distribution generalization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:34.369265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:21.394139Z digest=sha256:abdec5d4f13a345c32c372cfeb3b3e7da3e8d3d92364660dd13df19ca4468edc

Observation 2634c300-530b-41dd-a524-07c7847c98ad · outbound

This paper cites Natural adversarial examples.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Natural adversarial examples

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:34.128657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:21.479518Z digest=sha256:820fa94dc71b2a7d0645cc51efa3d76b34470b1648413672ae6494f9fba3865e

Observation c7664a7c-0b7f-4741-a1a9-3fcfcba3c88f · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:33.937969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:21.569817Z digest=sha256:fca34967bae8816dd57599d706a2723eaeacdc61a155a4d4ae6093c70b69b1e4

Observation 1a0fa88d-cea0-4893-9918-b7a371fa3a56 · outbound

This paper cites Llm-adapters: An adapter family for parameter- efficient fine-tuning of large language models, 2023.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Llm-adapters: An adapter family for parameter- efficient fine-tuning of large language models, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:33.789467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:21.734597Z digest=sha256:7188c4cdaed82f732392588b63a128c12fcbdfd447c296b05a9124a7917fc631

Observation f819ba40-316b-4a04-af4b-b2d3997b083c · outbound

This paper cites Directional gradient pro- jection for robust fine-tuning of foundation models, 2025.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Directional gradient pro- jection for robust fine-tuning of foundation models, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:33.617956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:21.845558Z digest=sha256:ff9386b3957bf793de2860fcbfaee437f6d46b51d34da33f6cb9038d9d40b638

Observation 5b9d8667-9123-4559-9f65-d7330808bd6e · outbound

This paper cites Hudson and Christopher D.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Hudson and Christopher D

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:22.034752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:22.034752Z digest=sha256:89d9919508daae823cfe4cf751db3a3f0b623b2bec083ca444a673dff6be9a4e

Observation 38927f64-8faf-4754-ba47-630122dc47f3 · outbound

This paper cites Roses are red, violets are blue.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Roses are red, violets are blue

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:33.411709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:22.147432Z digest=sha256:dfb809c54dce6bc339d7a7473fae930e222a37068a4111087293db93d8716097

Observation 38c1c61e-f2ce-4191-878a-a44fc2f8c45c · outbound

This paper cites Fine-Tuning can Distort Pre- trained Features and Underperform Out-of-Distribution,.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Fine-Tuning can Distort Pre- trained Features and Underperform Out-of-Distribution,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:33.216026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:22.264829Z digest=sha256:a1632be0cf4e05715db72df684c94e25ba258c23ee429cf52e9e7144b7796919

Observation 69669264-d650-4b15-a2e2-dae32b297438 · outbound

This paper cites an unresolved cited work.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:32.968130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:22.408385Z digest=sha256:38ae8376ac6ee0d171c7a1812c23d9f3753e153ff3dd5ace9ebd41d79cb4a718

Observation 1fd31f06-8625-42f2-9ae5-51f6f54e6a0d · outbound

This paper cites A Closer Look at the Robustness of Vision-and-Language Pre-trained Models,.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering A Closer Look at the Robustness of Vision-and-Language Pre-trained Models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:32.743187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:22.500847Z digest=sha256:82fc067e862ae83e067dd98c7c5eaf822bcad43600932211481e0558764e1303

Observation fb99404b-2d19-48aa-870b-19bb764409bc · outbound

This paper cites Adversarial VQA: A New Benchmark for Evaluating the Robustness of VQA Models.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Adversarial VQA: A New Benchmark for Evaluating the Robustness of VQA Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:30:28.504421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:22.701978Z digest=sha256:43f1efd92d8ff8773a8926833b8918fb618fae6aaeaf75324cefcd324325098f

Observation 898e22db-0104-4d4b-86ed-89407c199d5c · outbound

This paper cites Explicit Inductive Bias for Transfer Learning with Convolutional Networks.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Explicit Inductive Bias for Transfer Learning with Convolutional Networks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:22.787019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:22.787019Z digest=sha256:35e04dc903ea89b4d3183cf1c55cd445e60e6dcfa2f12fdf619c289ef4ac090b

Observation f660c6de-5279-4c97-8d8e-b59d828749fe · outbound

This paper cites A Closer Look at the Robustness of Vision-and-Language Pre-trained Models.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering A Closer Look at the Robustness of Vision-and-Language Pre-trained Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:22.632151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:22.632151Z digest=sha256:86035ae569a4d2492e1477e01bdf43eab5dacacaba9bce6b307242148994f8c6

Observation a1ffed09-94f3-46ab-9c52-9630f9bc08ce · outbound

This paper cites Robust Visual Ques- tion Answering: Datasets, Methods, and Future Challenges,.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Robust Visual Ques- tion Answering: Datasets, Methods, and Future Challenges,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:32.299018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:22.966162Z digest=sha256:7df490782301560040c33850ecb64159244bf1f0bff6470bcfe828daac1dfd26

Observation 4269af3a-bb07-4034-97a6-0a7463af80b3 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.184037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.184037Z digest=sha256:dee027cff3390992a1b2392f46ed418d58fad7732d581f4154f2ac088905ab43

Observation 90c00bb2-bd48-4863-b076-1d19ca4046e0 · outbound

This paper cites Visual instruction tuning, 2023.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Visual instruction tuning, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:32.543685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:22.844130Z digest=sha256:dde43d3411b5d8530180f472a21cdc1e092862acc13f0a8c7ca7510c0dd39746

Observation b7231ea7-fe95-46c9-9adb-f7da0ac9cfb9 · outbound

This paper cites Maximum mean discrep- ancy for generalization in the presence of distribution and missingness shift, 2022.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Maximum mean discrep- ancy for generalization in the presence of distribution and missingness shift, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:31.868985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:23.335849Z digest=sha256:12ec54e50dad7271fd34077ef837833558575337b11f8c7c0c8380421d3d98bb

Observation 0e73b822-b0da-4290-9ba1-a09fb3a64d27 · outbound

This paper cites Moment matching for multi-source domain adaptation.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Moment matching for multi-source domain adaptation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.427466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.427466Z digest=sha256:28c4961084a26c5da97e6437f35b37c7e63a3e2e941eaa3a11e80cd5ab642037

Observation 88f7ff7c-038f-4bda-b3a1-b973613dad95 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Learning Transferable Visual Models From Natural Language Supervision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.504571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.504571Z digest=sha256:b146afde9271fdf2b4510fef2b186bfa8a97e9a3afac1e3ffa1e8830072fcd1c

Observation 61fbbd8a-08f4-49c3-8cb9-3fdc7df2d140 · outbound

This paper cites Generalized out-of-distribution detection and be- yond in vision language model era: A survey, 2024.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Generalized out-of-distribution detection and be- yond in vision language model era: A survey, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:32.111032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:23.272303Z digest=sha256:8f634d33468a0688dff8ec2de7c2185a90def1ca748618ba35f840136669825a

Observation d080134e-5c9b-43f2-8bcd-be03d72a3bab · outbound

This paper cites Cycle-Consistency for Robust Visual Question Answering,.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Cycle-Consistency for Robust Visual Question Answering,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:31.381355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:23.712779Z digest=sha256:4489915f2883f71486ce304d4494bb99d2e2a4e382028213a2c436657b1221e4

Observation d0823e45-7a42-42a4-b1a4-ca9660017ca5 · outbound

This paper cites Human-Adversarial Visual Question Answering.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Human-Adversarial Visual Question Answering

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.821325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.821325Z digest=sha256:482be82f57416decebfcf7bbe9e780682a393e0774848f248fc2f4b843026539

Observation fd9ff035-e16b-4adf-9f57-3b6b0094dd7f · outbound

This paper cites Benchmarking out-of- distribution detection in visual question answering, 2024.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Benchmarking out-of- distribution detection in visual question answering, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:31.214236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:23.931683Z digest=sha256:dd17eb56caea4233a6e57ac4e38cc736a0e9f6f10df2d05c4f7fa0cb81b57986

Observation 2388f9b5-1477-41d1-8439-7a3747386f0b · outbound

This paper cites Do imagenet classifiers generalize to im- agenet? InInternational Conference on Machine Learning, pages 5389–5400.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Do imagenet classifiers generalize to im- agenet? InInternational Conference on Machine Learning, pages 5389–5400

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:31.584883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:23.652233Z digest=sha256:e7c8011e9297581b406f7316ee934e647f3c926f2b724105d92aaba6c67a940b

Observation 8bbb6535-3567-4ef4-aa60-0102cf1f9303 · outbound

This paper cites Towards VQA Models That Can Read.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Towards VQA Models That Can Read

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.067003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.067003Z digest=sha256:60a743f0f4cb0297d8b7f27a86e0a9bb4fae397e021f13ae5363f6bb8cb3abd0

Observation 0e42019f-cbda-4e04-ba92-79ac1ec40df9 · outbound

This paper cites Trainable Projected Gradient Method for Robust Fine-tuning.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Trainable Projected Gradient Method for Robust Fine-tuning

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:30:27.443352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:24.161750Z digest=sha256:f4fcd9b9f9e20e42bc915f3b0751af830769b7427cbcf3e290efb8ef60a7c3f3

Observation 33668d0f-2c8c-4cc1-ac66-6d6060229712 · outbound

This paper cites Fast Trainable Projection for Robust Fine-Tuning,.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Fast Trainable Projection for Robust Fine-Tuning,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.993713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:24.217760Z digest=sha256:02fba054c861841a021135bdc86f673b9d696e3e4047934e57cdd217c312eea2

Observation 47b5d286-8974-4be9-89b1-1f5097860ca0 · outbound

This paper cites Rethinking weight decay for robust fine-tuning of foundation models,.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Rethinking weight decay for robust fine-tuning of foundation models,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.903064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:24.338273Z digest=sha256:fc00b2f21b4343ca97f88cf350f62ed79df602ba77f6a5ee30f5c56d27d2f237

Observation 811af034-14bf-4280-b2cd-fe90acc791f7 · outbound

This paper cites Language Prior Is Not the Only Shortcut: A Benchmark for Shortcut Learning in VQA.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Language Prior Is Not the Only Shortcut: A Benchmark for Shortcut Learning in VQA

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:30:27.964956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:23.986835Z digest=sha256:d7652aa7c373117fe4975d3f559568ff104d2a152b1a2af5856a939a0016394e

Observation 1904a24e-0b23-4a90-b896-316e078bdc05 · outbound

This paper cites Learning robust global representations by penalizing local predictive power.Advances in Neural Information Pro- cessing Systems, 32, 2019.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Learning robust global representations by penalizing local predictive power.Advances in Neural Information Pro- cessing Systems, 32, 2019

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.624114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.624114Z digest=sha256:5eec76497b87022b7a2da1330118954f04d318af3f00d0b650b84f6ed37185e0

Observation 384808c0-cd99-425e-bbff-20e5b2a95988 · outbound

This paper cites Ledsam, Patricia MacWilliams, Pushmeet Kohli, Alan Karthikesalingam, Si- mon Kohl, Taylan Cemgil, S.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Ledsam, Patricia MacWilliams, Pushmeet Kohli, Alan Karthikesalingam, Si- mon Kohl, Taylan Cemgil, S

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.671410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:24.721174Z digest=sha256:9e06024b93bdd3a3d80b0e6b565ac3c7de1c5901c244f7c0ccf285ec1d8387f5

Observation 4d9224da-b56a-4bd2-9780-c3c15738bbe4 · outbound

This paper cites Robust fine-tuning of zero-shot models.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Robust fine-tuning of zero-shot models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.804961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.804961Z digest=sha256:bc94890bc09746c4fcf2373c9da419ebd3796b26308bb291a7abb5fc09b06913

Observation 4c6aae56-da85-417d-b12f-b6f063fb4c4e · outbound

This paper cites Domain-robust vqa with di- verse datasets and methods but no target labels, 2021.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Domain-robust vqa with di- verse datasets and methods but no target labels, 2021

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.489788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:24.935344Z digest=sha256:212feb5e7bb6e21befe505b7c17159fd15b4b24a7bdf6c4cf254f92a753796ef

Observation 941f0fc7-0f4d-4992-9b64-cae1526f5f1b · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.323868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:25.034749Z digest=sha256:a896c3a5498cd1a422db079e5cb04874df233e6639eb1a262f209bd2c5d7df6f

Observation d2431941-ff05-4f0e-b6d8-23218ac4ef36 · outbound

This paper cites VQA-GEN: A Visual Question Answering Benchmark for Domain Generalization.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering VQA-GEN: A Visual Question Answering Benchmark for Domain Generalization

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:30:26.964752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:24.515854Z digest=sha256:0b7bb44124159dca9bcb44853c8a4616014ea2d81dc565c21874d2c18e28fcd9

Observation 80a62b49-695d-4d90-ba59-05d7e6b9cb98 · outbound

This paper cites We use the LA VIS [29] public repository to fine-tune all methods.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering We use the LA VIS [29] public repository to fine-tune all methods

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.195941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:25.244751Z digest=sha256:8ef826d15957c9d935c1fc0c701ae770954a263d2e6ae07e018913a398db7d9d

Observation 69fa2271-0d8d-48be-93be-d702805d9b76 · outbound

This paper cites an unresolved cited work.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:29.997695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:25.374892Z digest=sha256:5ee0433f8dae78e0424009e7391c5e8771c7607a47fbcb1d2d0ee71364819253

Observation fb31c073-9b8d-4691-bbd2-38ea04c2bb51 · outbound

This paper cites 7 shows the correlation between shift and performance for different embeddings under different fine-tuning meth- ods.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering 7 shows the correlation between shift and performance for different embeddings under different fine-tuning meth- ods

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:29.837989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:25.632830Z digest=sha256:ece54aabbb992ebc43610b399ed6a17d806ee0b1105b12fe3e63c12611df800e

Observation be9795bd-9761-486a-a85a-21852dd5a165 · outbound

This paper cites 5 shows the heatmap of the correlation between uni- modal and multi-modal shifts per dataset.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering 5 shows the heatmap of the correlation between uni- modal and multi-modal shifts per dataset

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:29.630166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:25.742953Z digest=sha256:9a38f28721f3c48cda0e15949577c63bfc86667ed037338f10a358e9600ef4d5

Observation b7f031ed-deeb-45c6-932c-10e582a2fabb · outbound

This paper cites 13 and 14 show the variation of MIv and MIq w.r.t.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering 13 and 14 show the variation of MIv and MIq w.r.t

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:29.448148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:25.865612Z digest=sha256:d60084466b79c418ab267695a2352b16b6e3dea3a6ef5e465af79995f8647e58

Observation 84e2c426-a2b8-4bfa-bdec-637609601f3e · outbound

This paper cites 8, including LLaV A- 7B [33] with LoRA and PaliGemma-3B with full fine- tuning.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering 8, including LLaV A- 7B [33] with LoRA and PaliGemma-3B with full fine- tuning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:29.274521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:25.995036Z digest=sha256:bbbea658120dbb249bf65ddea11eac142bc25e329973f893a78338ec5c555ef7

Observation 1692ec5f-f3e9-49f8-a3bb-aef9de954d75 · outbound

This paper cites The only exception, GQA-OOD [27] (based on GQA [26]), has only answer shifts.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering The only exception, GQA-OOD [27] (based on GQA [26]), has only answer shifts

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:29.064764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:26.134838Z digest=sha256:f4b1e58f53fe4902423ba49230be21695416a34791e970f17f8188ed89a2ab12

Observation 21b24e27-5652-4d7f-a8f0-b4f6b21c4e7e · outbound

This paper cites We further compare shifts using Maximum Mean Discrepancy (MMD) [12, 20, 37] with RBF kernel in Tab.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering We further compare shifts using Maximum Mean Discrepancy (MMD) [12, 20, 37] with RBF kernel in Tab

Reference 68

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T13:30:28.908369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:26.255127Z digest=sha256:660488910ce81ad8b751a1eb9ee2fd062c1351f029131858afbb868eb5cb4f99

Observation 58e707d3-ed14-4556-88d1-c455496b3fa2 · outbound

This paper cites This also serves as a veri- fication of the reliability in quantifying shifts via feature- based representations.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering This also serves as a veri- fication of the reliability in quantifying shifts via feature- based representations

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:28.802066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:30:26.374735Z digest=sha256:2f00d687f3eeec27bf01149301714189009b64d24837571fbb350843a69b2fdb

Observation 975ea182-8c39-49c2-a8b6-9a6b49c69ca0 · outbound

This paper cites Cycle-Consistency for Robust Visual Question Answering.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Cycle-Consistency for Robust Visual Question Answering

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.772858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.772858Z digest=sha256:48abb31d816ec6c5ecf5ee64f6c946855531a618a1825d37810603aa333ee9f7

Observation 79e8ca76-f1a5-4db2-b092-1033396ce1e4 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering LoRA: Low-Rank Adaptation of Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:21.655216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:21.655216Z digest=sha256:c4a5185e4c3b685a4ca0236a56140c9629acd0d5fe4b5b273cb9ba1090df301d

Observation e257422a-c06e-47bf-8cc4-76b650191457 · outbound

This paper cites Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:22.333864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:22.333864Z digest=sha256:5147894e869027a21381de577ca5ff21cbc16b00cf1b6b722aa5b2b7d803a984

Observation d014ed24-983a-45e9-8e21-1acf215453ba · outbound

This paper cites Fast Trainable Projection for Robust Fine-Tuning.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Fast Trainable Projection for Robust Fine-Tuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.269594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.269594Z digest=sha256:9b977b500702c5607e25f48524562ac01505e701f4d16d0505607c22c3b961a5

Observation 46c7e70c-9d2e-4cca-9013-9010f64a7775 · outbound

This paper cites Robust Visual Question Answering: Datasets, Methods, and Future Challenges.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Robust Visual Question Answering: Datasets, Methods, and Future Challenges

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.075341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.075341Z digest=sha256:7d23f4c5d612111b9d8f25a8f1667680e8352d1ba7703194d840055b037d4918

Pith citing papers

No inbound Pith citation observations are available.