Pith. sign in

Paper Citation Record · LEDGER

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement

As of 9 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 0 inbound Pith citation observations for arXiv:2507.01643.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01643 v1

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:53:10.367278Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 105 outbound references displayed

  • verified exact3
  • verified fuzzy29
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87a779a4-5d37-4ccc-9100-ca4a1055b45e · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.174120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.174120Z digest=sha256:f1639948dc6f1d91f22a974fd9b5d24a3c7215a47b9a056a9ddcd3d350ca15a9

Observation f3f8b829-d3b8-4c9e-b49c-7a2c08ed4b7a · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.289384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.289384Z digest=sha256:ce08e1c793c821cb2a06e68e5c85863077b117c690cebc46b4db714e9e05d3fb

Observation 53466165-6d22-4a61-b980-3bf3b80cc3d9 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.354949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.354949Z digest=sha256:b7b2188914689b56e7aec4c41f4b39785c1981ba77fa740bc32eba6e54ef34d4

Observation 50929551-2aa8-4e2f-ad9f-43148dab93d1 · outbound

This paper cites Qwen2.5-VL Technical Report.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.473093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.473093Z digest=sha256:37377f97f6ce73387183a90ce5f7291b6ecfa5145094dd1c87b793135aecc757

Observation 4edcd2fc-7665-47a6-9aa1-72c48046aeb5 · outbound

This paper cites OCR-IDL: OCR Annotations for Industry Document Library Dataset.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement OCR-IDL: OCR Annotations for Industry Document Library Dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.530268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.530268Z digest=sha256:61147977b238512d471d92cb09074e6b0f8af8173d0360bfcf29851ffe2461d6

Observation c7f9a6f9-9427-4f31-81fb-2221bf4d07f5 · outbound

This paper cites Coyo-700m: Image-text pair dataset.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Coyo-700m: Image-text pair dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.606511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.606511Z digest=sha256:b4ea202f77c8d31d61be04b48938011f6af7d577f872ccc5d28ecaeed5f9ca77

Observation 4efb48bd-318b-4f97-9e72-279d1f7ed98a · outbound

This paper cites Reversible Column Networks.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Reversible Column Networks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.746148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.746148Z digest=sha256:c576f7d7094cfff1cf4ce9fcf1c842316d6706f806b1d9937966e9a7f8deda0a

Observation df37130e-1b88-4449-b312-380fe9e3d592 · outbound

This paper cites InternLM2 Technical Report.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement InternLM2 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.890756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.890756Z digest=sha256:2cef014950f4bcee34312e7a610c21f57568039d3db0a174daa0fb1b43fafc7a

Observation 76b273cb-d5bd-4289-b241-8eed1c6b3670 · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text encoding.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement An augmented benchmark dataset for geometric question answering through dual parallel text encoding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.992903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.992903Z digest=sha256:d0672279e84c0b3d8c6b8c4862c6d47abdd766719118ac00f30f670e5c639ce2

Observation 6603d20d-9e99-4651-bdeb-78240a99980d · outbound

This paper cites End-to-end object detection with transformers.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement End-to-end object detection with transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.082433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.082433Z digest=sha256:4ade0a97f1600aaed7b6df2320e9b3922bd920c53d4881b221ef7e72c8b0b98b

Observation 846e9db1-cacd-4b1a-ac71-bd61cb2742e9 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Sharegpt4v: Improving large multi-modal models with better captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.202211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.202211Z digest=sha256:ddcf6aadd4b4502cd52bff5a4f42422fe306e02df58749654f7851304113cf65

Observation 36a94428-a52a-4253-9157-6208eaf65a54 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.365902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.365902Z digest=sha256:fc50d09ddb0f78cc3040d80a871add107d0158246c43dba79e4d35d5f8282c04

Observation 190292f5-0b9c-450d-b19e-a5b00cc01660 · outbound

This paper cites Vision Transformer Adapter for Dense Predictions.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Vision Transformer Adapter for Dense Predictions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.512069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.512069Z digest=sha256:590756cd539661ac5f196eff241597b4c7b7aaf02caf6acef069d5df836b0eb6

Observation 3e4e7dab-6dad-4dce-a34a-f5ccd4100a79 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.696027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.696027Z digest=sha256:30ec89b47dd4af339c77bb85a29c252fd8e6a8c210a40a91a6e10b8cb9775c55

Observation ebb95258-c348-46e4-bc7a-d19ec1053c73 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.828141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.828141Z digest=sha256:f18122a9589819c730ae0d06a15880e1b0848addf60f035e3ed4c819a8b517f7

Observation f3351db5-7b83-4ae9-a3f6-4568eec2f45a · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.984634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.984634Z digest=sha256:74df06186aeebce232be4c0873c898230c3fef2069d703707679234c4f002be3

Observation 2da8fe5c-44cd-4d40-8531-8ef120e6c04b · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Opencompass: A universal evaluation platform for foundation models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.083800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.083800Z digest=sha256:664a5ce18b3fdbb6c03f8891f6cf0216c7c1b508da465d2f49ff9260880ff9f9

Observation 61c2ff98-d663-461d-8d30-7431fc90918e · outbound

This paper cites Grok-1.5 vision preview: Connecting the digital and physicalworlds with our first multimodal model, 2024.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Grok-1.5 vision preview: Connecting the digital and physicalworlds with our first multimodal model, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.213459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.213459Z digest=sha256:0431d198146f669524c94a0028c3f807f46d210f818ca635136530ffb613f589

Observation ed114bc2-240a-41d7-81a9-833564627524 · outbound

This paper cites Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.382480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.382480Z digest=sha256:91e321fcc58ab1c92707750b501ed1146786e25944eaaca16d18c72b6e090e06

Observation c982de7d-2989-4d96-a4aa-8479303dc6d1 · outbound

This paper cites Deformable convolutional networks.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Deformable convolutional networks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.460206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.460206Z digest=sha256:9fb3291204e53b1d8e68d4258cb3563169da5a8028419aadce6bf8140414cc3f

Observation 9942ddee-4829-4229-80b5-987fe6c51805 · outbound

This paper cites Scaling vision transformers to 22 billion parameters.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scaling vision transformers to 22 billion parameters

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.556501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.556501Z digest=sha256:052e1f076fff96a3199311f69ac11043586f63fe75081c8fe2e535f9dd972339

Observation 159447a5-105f-49d3-ae2b-f9316f254360 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.653350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.653350Z digest=sha256:6ba7c75210be77ef60df537eb7ebc868669c549180d0870e3d4b0d838f8695d4

Observation 04837be4-82b1-44a1-990f-3694e521845c · outbound

This paper cites Imagenet: A large- scale hierarchical image database.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Imagenet: A large- scale hierarchical image database

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.781183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.781183Z digest=sha256:3a8fbeb5a6a0ec808c35278b7d7135631a816f412e1415f171676fe8c997ed7e

Observation bd0b9c54-068a-4ccd-8543-fb75e41684bd · outbound

This paper cites Scalable Vision Language Model Training via High Quality Data Curation.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scalable Vision Language Model Training via High Quality Data Curation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.977919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.977919Z digest=sha256:68570728cb4aacd062eb891ece4a8eccd1008e10990ddf2bd9af2abc424a1d7c

Observation b9a8a1c8-dff9-438b-a88b-5f5ccc91fee6 · outbound

This paper cites Benchmarking and Improving Detail Image Caption.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Benchmarking and Improving Detail Image Caption

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.136534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.136534Z digest=sha256:cb449acf47c43f6ac629277f0ec71b2ebcc3d18573ee3ccc16ae633ea6e489bd

Observation dc0f21b2-c60c-4755-8734-e486f0e811b7 · outbound

This paper cites Adalrs: Loss- guided adaptive learning rate search for efficient foundation model pretraining.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Adalrs: Loss- guided adaptive learning rate search for efficient foundation model pretraining

Reference 26

Resolution
verified exact
raw_fallback, observed 2026-08-06T20:53:11.795734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:03.306529Z digest=sha256:0357422bf892558de8d5ade1ce85e7ebedeecccb142b0218691087e181321317

Observation de786191-9cd0-4b9f-ae4a-b2b735355d81 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.510387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.510387Z digest=sha256:6a4a58c997ca8890ef26d26be7155a5fc3f3bb82768a47f567d06e600cd03f37

Observation de626c13-022e-4fee-8d32-926d2276b58d · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.669288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.669288Z digest=sha256:e5bd882f96b18a5627ea0937a7c4d4576a9c25f7e2390d18be6bb02db71a57b8

Observation 4d8ec47f-687d-4fea-ab49-a7e1a0a11765 · outbound

This paper cites Scalable Pre-training of Large Autoregressive Image Models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scalable Pre-training of Large Autoregressive Image Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.825159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.825159Z digest=sha256:f6de4b359881a0180fbafe570832f43ee82cac0c04778cd2c0b94ccc84a93471

Observation ece4340a-d08d-4993-894a-dab7d2cd9aaa · outbound

This paper cites Scaling Language-Free Visual Representation Learning.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scaling Language-Free Visual Representation Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.943429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.943429Z digest=sha256:56ea2d8976bcdad753c138609f1cc8f52443a0235c07bdcf267d9b5978f0c19e

Observation c8b7ce0f-21d6-4ae1-8aca-36773b1fb843 · outbound

This paper cites Data Filtering Networks.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Data Filtering Networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.057607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.057607Z digest=sha256:6379d0d8375d2ac28f067bb3eb352d6226fc6e2e6c852db0eb5079164192a596

Observation da4f14f6-bca2-4478-8c1e-597d7811d6e1 · outbound

This paper cites Slowfast networks for video recognition.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Slowfast networks for video recognition

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.154213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.154213Z digest=sha256:85e218d06eb35c1e67274e63e0fe63bbd6845fd0e7bd6258bf2d40905ab81b8a

Observation 420672e1-1066-4026-9620-448ba0bb8c5f · outbound

This paper cites Multimodal Autoregressive Pre-training of Large Vision Encoders.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.264739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.264739Z digest=sha256:cf52cf56e0bd6e69f69f9bdaee25a0c3963c76e15750d3d621785321c8a2d82c

Observation 591ef97f-139f-478e-a5a2-38dfef806161 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.349270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.349270Z digest=sha256:920ef9ae951990abd8104a28cf17f5c6369a03ec4094671f832c9fa7cf3b23a7

Observation 974bae22-f802-40f6-a970-4f83a53ba6ec · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.453395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.453395Z digest=sha256:0b1a477169069b36b56c860499af6d2ef76d0b99942528dcbed8e9266a389b02

Observation a1e2ecc9-2e50-4db3-95b5-b9ad56713b35 · outbound

This paper cites Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.556702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.556702Z digest=sha256:402fcefb1394f7b0ac3165cb44166965208e6617938c50e511130050f3defe5a

Observation 6db352bb-2fd3-4d10-a4e3-c8c098c71fa6 · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.659113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.659113Z digest=sha256:1bf0f23816d47a41f94df1b58aa3a0150511402242d6ee348c1b80da79bbd439

Observation d9cc9776-df00-41ae-a9bf-2412b3bcc69f · outbound

This paper cites Mask r-cnn.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mask r-cnn

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.765807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.765807Z digest=sha256:353e2052ad61f46f704878b4ca036785242353db7d18249afbbb6fe608b921c8

Observation c9578615-533c-4f5f-9c72-b3e17aa48214 · outbound

This paper cites Deep residual learning for image recognition.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Deep residual learning for image recognition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.889813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.889813Z digest=sha256:f19f93426645c28ca33c6b43780764d9c80d2b8a76274cc571eee26286e3999a

Observation 7749181b-17cc-4cf8-844c-29f86273e8c5 · outbound

This paper cites The many faces of robustness: A critical analysis of out-of-distribution generalization,.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement The many faces of robustness: A critical analysis of out-of-distribution generalization,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.975078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.975078Z digest=sha256:62811c43ab356f2f96638bb8292ab9745173ddde955f88067bf54364c971018b

Observation fa1a1a87-8ac0-46cc-8df7-fc82cd529023 · outbound

This paper cites Natural adversarial examples, 2021.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Natural adversarial examples, 2021

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.049397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.049397Z digest=sha256:740427c7101402118266d1b0910ac871c061b5b1986aea3f0dc73579fc3b5d43

Observation f0de58a4-8897-4c39-adb4-2a867911ca71 · outbound

This paper cites Squeeze-and-excitation networks.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Squeeze-and-excitation networks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.144799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.144799Z digest=sha256:ff22ce06241caf15400df6e4d734c35ad774a7311fd658e336f2cbafc0d0c4d1

Observation f611ea19-3a0c-48f2-88f7-795668443c75 · outbound

This paper cites Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.239887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.239887Z digest=sha256:380b4f9d610ae69b25f8a1310d184579036e2d4e889372bca165871c64d6daf3

Observation 51a87678-09b0-4b3d-83b4-a3a6103152c5 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scaling up visual and vision-language representation learning with noisy text supervision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.355096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.355096Z digest=sha256:a22c13dd86b032e5e07d84e51d2b8f2edf441882347cf789197a77c7603660a9

Observation e5eab2a7-4141-450a-ad36-36ab8876a996 · outbound

This paper cites Deep visual-semantic alignments for generating image de- scriptions.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Deep visual-semantic alignments for generating image de- scriptions

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.440840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.440840Z digest=sha256:09a603871f4d0b9921403e823af7ef1ed9b7694458e64411c9f4bde93deeb636

Observation 74f37a5d-2c5c-4385-b53d-f0865f2b4502 · outbound

This paper cites A diagram is worth a dozen images.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement A diagram is worth a dozen images

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:16.976874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:05.499872Z digest=sha256:c7cb771d9f5bc3a4166ff80ecf20b21e5e027de3dac472dbf4306c68e56cd173

Observation 1b36cab1-6f83-4bf8-8687-e77075edc616 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Lisa: Reasoning segmentation via large language model

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:16.815444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:05.571684Z digest=sha256:9d72e613dd634c3793dc1ef64fc70506f1aad2abe60ebe49b3b132c07637ed88

Observation 827c4573-8d18-4ff2-bb34-487ed97466f8 · outbound

This paper cites Building and better under- standing vision-language models: insights and future directions.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Building and better under- standing vision-language models: insights and future directions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:16.676412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:05.672563Z digest=sha256:542ebc4867771c76dc69f491890589c650b9ac16e74091e5bce5031dda50081e

Observation 0044f049-24e4-4e83-863d-0ba8b5f52189 · outbound

This paper cites What matters when building vision-language models? Advances in Neural Information Processing Systems, 37:87874– 87907, 2024.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement What matters when building vision-language models? Advances in Neural Information Processing Systems, 37:87874– 87907, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:16.515187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:05.761968Z digest=sha256:1dd9c47691d68b45834f11737f51fe2abe24463e4ff081069b5aa58c62a68e40

Observation 8e3d7bf0-8038-4b4e-9817-32ed004715e8 · outbound

This paper cites The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.867758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.867758Z digest=sha256:2fa4a1a3f9a312a8445adfb80a585cc4b5a6a302e5069c82965373e9f60d2f1d

Observation fd89d39e-9699-41a9-9007-b31e3cb67f5c · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement LLaVA-OneVision: Easy Visual Task Transfer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.966643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.966643Z digest=sha256:ee874260d4cdd5a9d6788cc19ec129be449da1d57a2c889ce662ce3578bf1d6d

Observation 8edbc22b-7bf8-4438-9569-05985e23b1d5 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:06.041336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:06.041336Z digest=sha256:30b3d2530156e94dcbf2f40a9bbefff6065bf2bfe70bc576df8b23995e55e7fc

Observation 59f4a2d1-fba9-48e7-99af-b69dbd5588ce · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:16.332980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:06.099402Z digest=sha256:04e936b540f2767b0b39f294fb0cc85fa528353a734bc39a0eafd34ab2c9415a

Observation ebfdf35a-1998-4dcc-9cef-7e50202a6687 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:06.204273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:06.204273Z digest=sha256:08f6749a6d333aeb0c61d04d9799ffb7fe0278db78e4f31a447739d0eb1b2b0a

Observation 333b4891-09aa-4b17-bf68-fae1118550be · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Evaluating Object Hallucination in Large Vision-Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:06.310118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:06.310118Z digest=sha256:d2ed1667831b1c710027a3f1b1c6370310678d2305fdcaffae9d025a711128d6

Observation 0283545d-c372-412b-b1ce-0e4581157c35 · outbound

This paper cites Improved baselines with visual instruction tuning.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Improved baselines with visual instruction tuning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:16.185468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:06.413431Z digest=sha256:0f74e4ead17332998e01bc4b19efbf46e4f5a521ac8769d7ce90bdec021f4906

Observation 4a2fe0a4-b236-4c80-8df4-5eaf3dd23463 · outbound

This paper cites Visual instruction tuning.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Visual instruction tuning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:15.961535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:06.494248Z digest=sha256:b5576978634223d61ec404da3a90be48d284013d70010a11b8afbbcf19f3fa05

Observation a5a01265-787c-40d3-adc5-597f0b495cdd · outbound

This paper cites Muon is Scalable for LLM Training.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Muon is Scalable for LLM Training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:06.548354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:06.548354Z digest=sha256:30d71b3dc3f851f62a3065e3a4177ef3d7e9130dbc4fad3e9aa4d3f4bcf30083

Observation ae89d1a8-3bec-4090-ba17-80b3218be92e · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:15.741340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:06.649180Z digest=sha256:0f4f2a0268c7aa401098062d4576ac580dc340a48a91679d4fc863b85c4d866e

Observation 345b18e4-314a-4e36-8430-5d72da46ed77 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:15.552804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:06.750645Z digest=sha256:17fa2c4544a76161ece5e4b5365c0de5245b05ee3440242141f6025703647ace

Observation 2df5e4c3-e566-4e69-b955-7148d92b9383 · outbound

This paper cites Decoupled Weight Decay Regularization.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Decoupled Weight Decay Regularization

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:06.868876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:06.868876Z digest=sha256:22a2e5be116d06dd0f5a3d71507568b130e005f508af8e3bd33af7f116cd210c

Observation e37a86bd-b7da-4470-9ad4-db79d3c5e662 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:06.917722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:06.917722Z digest=sha256:9a55633e360da2d54644557f16f05a5acd00a0aef56d2199e141f0817bf789c7

Observation a29d00c8-d994-499b-a20e-78726d285db5 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:07.020742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:07.020742Z digest=sha256:55f799c68f60c356b5b5d689929293f14057d5a5ee0eb8535a80a536e0815aca

Observation bbbcbdee-0cf1-4a78-ac00-b63f7b284021 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:15.393314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:07.120075Z digest=sha256:2dbc1380cd8193ddd5c22fa59eaab8e2c273a1b5b33e1d3710c2905943df59cc

Observation ed625676-9ef6-457f-8452-fdcc4955ba73 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:07.221154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:07.221154Z digest=sha256:11fbf77375a4449a0efc720cf24cf0c8814a9bf4f44c133b113cac8ceb7d0adf

Observation aea5201b-43bc-4e53-925d-97fcc6f8ebe8 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:07.312239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:07.312239Z digest=sha256:210bfb1c50041935b8459eebbfda38b918c95e25bf9b2097edae98dd70b2158c

Observation 831dbd81-8405-4b6e-8b64-aefbf16eef2b · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:15.233277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:07.419963Z digest=sha256:d196c5fd9ec352ba8e65527bab226eae7a84c4ec7d35cfc2b81c357ddbf027f2

Observation 63e2f7dd-ad86-4244-a702-b01fb1db3d3a · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:07.511190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:07.511190Z digest=sha256:56240e1e3841e72f46398705d7e1100f652c7b5d486a4ad521ca6db5c04d76b4

Observation fde217ee-d772-4282-b2c7-64e7d15c94d0 · outbound

This paper cites Infographicvqa.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Infographicvqa

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:15.058601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:07.611803Z digest=sha256:764deedeb39b98c7aa3295c38b3d814ef7ec7cf70addf292cca942a2772e90f8

Observation 01ba0488-e905-41f9-87f4-b2895e1b06cd · outbound

This paper cites Docvqa: A dataset for vqa on document images.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Docvqa: A dataset for vqa on document images

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:14.878279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:07.684217Z digest=sha256:741b289734ca6a7f9182d9bb07aabc1306f4266e618fbdb2144a52fdc6109681

Observation 1afdc95a-9888-4be3-a8da-0751921568e7 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Ocr-vqa: Visual question answering by reading text in images

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:14.694685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:07.797962Z digest=sha256:e3d2e8abcba4322ea3b8743561b36f1a9393b4cb047fa86d937c0fe472718fb5

Observation d7e7995f-9a97-4699-b4fa-2833ed44db52 · outbound

This paper cites Introducing chatgpt, 2022.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Introducing chatgpt, 2022

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:14.509068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:07.848469Z digest=sha256:ce4fed1b45e629efb75319e176807b923cd6c555f7f6bb1283bf9cfbf7defe24

Observation 65d44a07-4e93-4451-bfc2-fb2aad7ba817 · outbound

This paper cites Learning transferable visual models from natural language supervision.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Learning transferable visual models from natural language supervision

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:14.329003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:07.959684Z digest=sha256:6597110ec9a5114b26fc00af1682cff8839384d043099984473d9dcd5111915e

Observation b61bca55-9d38-466c-a543-7801a7c5a1d9 · outbound

This paper cites Zero: Memory opti- mizations toward training trillion parameter models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Zero: Memory opti- mizations toward training trillion parameter models

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:14.195330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:08.067283Z digest=sha256:d48b1102d739e93489d7de608d823786bdb78e35811617baccbfab18faf9e932

Observation cbcbd198-0e40-451d-bb14-ea29fa46ad0a · outbound

This paper cites Do imagenet classifiers generalize to imagenet?, 2019.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Do imagenet classifiers generalize to imagenet?, 2019

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:14.040246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:08.176323Z digest=sha256:97442ed2f8f36e4cdc0b817fda28feec98f06b571e3639baf55a64b3de798ec8

Observation 67b0689a-37ba-48e2-97be-68b05328c0e7 · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:08.277308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:08.277308Z digest=sha256:0412a0eb87b4afed5338e9bed861ea62e4472daee068ac19dcf62168f133a2cf

Observation 9f4b5191-474d-4218-a490-47ba5b5a4763 · outbound

This paper cites The Curse of Recursion: Training on Generated Data Makes Models Forget.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement The Curse of Recursion: Training on Generated Data Makes Models Forget

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:08.389462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:08.389462Z digest=sha256:2038d64e58ef1ea2c2fb63e539c2fbdc4505d816c575d9c2e06b529e05afa662

Observation 129c05bc-c99d-480f-9f80-c2b7fbf1fc79 · outbound

This paper cites Hollywood in homes: Crowdsourcing data collection for activity understanding.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Hollywood in homes: Crowdsourcing data collection for activity understanding

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:13.840234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:08.498441Z digest=sha256:4547e9ab18322ed74aa70e033ec8cf385e9fe32b0a78e2b60494664e85289916

Observation d499bfdf-0ff9-40e4-b799-59f07d5dab77 · outbound

This paper cites Towards vqa models that can read.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Towards vqa models that can read

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:13.673354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:08.610364Z digest=sha256:2326f5a8ae07f7fdbe502f46a21b4b4698d3e40c5ed07ae559b9b14ecbcf378f

Observation 446b6f94-7e19-4f2a-bf2f-8b0c05c62892 · outbound

This paper cites Kimi-VL Technical Report.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Kimi-VL Technical Report

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:08.679854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:08.679854Z digest=sha256:08ecb980d9c2c56081d85805e188e22df20e8e555a9d969eb43abb10e79c55a7

Observation e03a2961-db65-42c8-b3da-05e24ef9ad89 · outbound

This paper cites Qwen3, April 2025.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen3, April 2025

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:13.554335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:08.740655Z digest=sha256:cda173f6c478c26342efd5adbd09a1edefbc327a933888e515ac7db4c1f04e13

Observation fcb05f98-913a-4806-b5f8-fbd495febd22 · outbound

This paper cites Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:08.830124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:08.830124Z digest=sha256:b2f3f0df3fe091d6929b031872cecc9c61b121f38ca04c80e9f4d03071be5004

Observation 366e44a7-412d-4023-a12b-e1e04d53176d · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:13.392724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:08.912731Z digest=sha256:ff2efcb883a201347b8a989435e1b6485ff6552a1244d61cfd9e32b2fc6be037

Observation 01b2c095-0adc-41bd-b71b-925321588093 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.011049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.011049Z digest=sha256:fc6c3343213c65548c62f6f04efa7a1744baedb3a539d42cbfc438fad6c89a0b

Observation 84feef7a-bfbd-46ec-adaa-a9571bf3c6ae · outbound

This paper cites VGR: Visual Grounded Reasoning.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement VGR: Visual Grounded Reasoning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.114205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.114205Z digest=sha256:5264280d5dda4d0ed2f869e425baa830f60edf8a9a2688a5bbe47f37d58a2947

Observation 68de02d4-03e1-4ab7-b0da-48d7c9ffbfc4 · outbound

This paper cites World to Code: Multi-modal Data Generation via Self-Instructed Compositional Captioning and Filtering.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement World to Code: Multi-modal Data Generation via Self-Instructed Compositional Captioning and Filtering

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.193598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.193598Z digest=sha256:98599e3c6296138e8c27bab5195459ec51ea6c5d0248ce47cdb78150162d668f

Observation 99cb87b7-ae69-4d36-a51b-54a0499edfd3 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.284570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.284570Z digest=sha256:a0198d046a327869aabd77943343ee77a0b1ce5e5cb2118ed5b63307694a7bd8

Observation 34bcd4d3-5a97-49c0-87ce-d15e2db08641 · outbound

This paper cites Pvt v2: Improved baselines with pyramid vision transformer.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Pvt v2: Improved baselines with pyramid vision transformer

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:13.252292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:09.384776Z digest=sha256:e386c6f2e842db08619d374fd15efcf90455c76ed93c1b3c9168a140658d18ba

Observation 275553a3-2d36-4427-becc-762e4e7daf42 · outbound

This paper cites Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.464295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.464295Z digest=sha256:019b28e0343dc793b1f97e51f9e8daae86aa1848ebe4f6998ad84c859d42a7e0

Observation df4f8766-c159-4818-9a53-645cf33e705d · outbound

This paper cites DocStruct: A Multimodal Method to Extract Hierarchy Structure in Document for General Form Understanding.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement DocStruct: A Multimodal Method to Extract Hierarchy Structure in Document for General Form Understanding

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:53:11.128260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:09.562208Z digest=sha256:081b02afa8db1aa90c20dd686f6989cfd9eb7e3e9b87a6eb30b6cee996cf10f9

Observation 42533c59-304e-48d5-b8b2-58c198e861a9 · outbound

This paper cites Learning, Reasoning, Refinement: A Framework for Kahneman's Dual-System Intelligence in GUI Agents.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Learning, Reasoning, Refinement: A Framework for Kahneman's Dual-System Intelligence in GUI Agents

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:53:11.031292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:09.643301Z digest=sha256:d672c68e0b37937721ac92c5de89ceb0e087e698546f7b8b2afe9a1815a3927f

Observation 21a80b97-285b-427d-8e2a-9042d2fba491 · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoencoders.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Convnext v2: Co-designing and scaling convnets with masked autoencoders

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:13.058543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:09.735785Z digest=sha256:f157cb6b5cfacc518c3a86998440f221c0a0f5f20b4bf83d19fdd550ff2a18eb

Observation 665e9818-9536-480e-a7e3-16e73fe9a998 · outbound

This paper cites Seeing the image: Prioritizing visual correlation by contrastive alignment.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Seeing the image: Prioritizing visual correlation by contrastive alignment

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:12.875815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:09.815315Z digest=sha256:7b5318f3b82e19cb575a12c56a245f7ec3396fde65e94f3b35083bfedb86365c

Observation 6e00b62d-a9a6-4216-aa41-1eb5feb23ef1 · outbound

This paper cites Aggregated residual transformations for deep neural networks.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Aggregated residual transformations for deep neural networks

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:12.734487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:09.911243Z digest=sha256:94e1be114ac6b97350dd075277b9afb624c6903244216aee695c370f8fcd8a10

Observation 3da9d2ac-8853-4031-9f32-69a7b568ddd4 · outbound

This paper cites Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.990394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.990394Z digest=sha256:830f3b04e9689004f2fc6cf7f239a68c178d9d75e47d4df2b774ff7f9f8f5ded

Observation b775d07d-d57c-4e11-ac09-ae766d755e34 · outbound

This paper cites Qwen2.5 Technical Report.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen2.5 Technical Report

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:10.097531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:10.097531Z digest=sha256:9260dea5c3ed71407b9ffa1a37f0ebcfbe32af4af50a022b8725cfe2bac172ce

Observation 7de236ae-333d-48c8-ab3d-74acdbdccaac · outbound

This paper cites Pediatricsgpt: Large language models as chinese medical assistants for pediatric applications.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Pediatricsgpt: Large language models as chinese medical assistants for pediatric applications

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:12.585192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:10.152994Z digest=sha256:17afcf8c90335723f43efb4a6c08bab85a8d4e0e998b3e7d3a4e26c7d12ee874

Observation 431a2d12-4355-440a-81a5-9cd058174da8 · outbound

This paper cites Improving factuality in large language models via decoding-time hallucinatory and truthful comparators.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Improving factuality in large language models via decoding-time hallucinatory and truthful comparators

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:12.432758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:10.220443Z digest=sha256:d4df69e5aabb62dc2c2f7fe616160db66e5a97071a9ec88a1644eb6a86e46c11

Observation 841289a6-f86c-493d-9105-fad874f1c0bb · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:10.285746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:10.285746Z digest=sha256:a1c405b2f65e78d91ce657e831a17acee72ad88be7f048d8df724ecdc10405a1

Observation 559193a3-2037-4663-b025-4d310ebbb974 · outbound

This paper cites Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:12.313580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:53:10.367278Z digest=sha256:2133069d35fc9f41b2e535a7a09e4c08874d097c7411384446548ff52cd9c83c

Pith citing papers

No inbound Pith citation observations are available.