Pith. sign in

Paper Citation Record · LEDGER

FILA: Fine-Grained Vision Language Models

As of 17 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2412.08378.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08378 v3

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:56:46.279061Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T20:43:16.364125Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T20:47:45.997187Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f813ec16-4a44-46b2-96c0-1376501fd228 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

FILA: Fine-Grained Vision Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.154642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.154642Z digest=sha256:d010fa7319f52d026aa0a2c732e8ffba983e41812d8f63541c7096e61b616775

Observation 4909a288-bb1f-4aee-b11f-df4b973d623e · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

FILA: Fine-Grained Vision Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.160077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.160077Z digest=sha256:05dfcc2d2f064a0f9db72182978eb69e3133c2614008e052c9e0fac81c69a13b

Observation 888df4e2-6bce-4c69-945a-8a2489a81a2d · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

FILA: Fine-Grained Vision Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.165232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.165232Z digest=sha256:4f913b2d89413b1f33b72b36d2e7353077ef65a0d45c4caecfbaf115c7dd8eb2

Observation ea4bc805-2138-49c8-b0d4-f514d986f7c1 · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

FILA: Fine-Grained Vision Language Models mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.175229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.175229Z digest=sha256:5dea1f004f9beee596a0e298ec11486640a870c9f069a364c82ef4b7dcd03bbb

Observation b3d3e998-e05f-4e26-97dc-c4565a33a5c6 · outbound

This paper cites Segment Anything.

FILA: Fine-Grained Vision Language Models Segment Anything

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.186496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.186496Z digest=sha256:ea8640b964059f6460ea365638aecc5ddfec66a070fa463251cb0e48657eef65

Observation 428fa63d-aa58-43e8-87c2-8b460dfb63c0 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

FILA: Fine-Grained Vision Language Models Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.192029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.192029Z digest=sha256:5bbfec05eb94206d81678881b0e908a3d09de3bf94d071b1f65f5ad7ca0b2b57

Observation 1175c106-796f-429e-8a42-686da30ca90b · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

FILA: Fine-Grained Vision Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.197387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.197387Z digest=sha256:96c86a96dc9f6f34fde71ab14aee8237ae9819757cf68c9114dfb2b951ff48ef

Observation 6ba66138-4fc2-4694-8843-2fba4401a95b · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

FILA: Fine-Grained Vision Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.202345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.202345Z digest=sha256:1d1e3d5750a4fa2e7955a36a44596b5dc7171cf37115f730084a53e583e61047

Observation 9bc0d697-dd1d-4f1d-87fe-7262da9f3897 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

FILA: Fine-Grained Vision Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.207284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.207284Z digest=sha256:23638d031757a5a60631f6d9db1a4fccae2b598a08b39b376c09cd7490aaf65a

Observation 8d3c4315-3ca8-4715-82c9-2179207d2bc1 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

FILA: Fine-Grained Vision Language Models Learning Transferable Visual Models From Natural Language Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.213549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.213549Z digest=sha256:7ca72a0df6399940e65dcd6dfa8a33f620ea3d16804921d513190859cd0434c9

Observation a0970d59-34c6-4972-a6b9-5388cb254a10 · outbound

This paper cites an unresolved cited work.

FILA: Fine-Grained Vision Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:56:46.770043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T17:56:46.218289Z digest=sha256:1c7bb55b523bd5d91d2aea455bd8547a68ec4b37d302e0ae9e25ec70475d241b

Observation f2e99b39-f197-4592-836e-3a76de13f774 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

FILA: Fine-Grained Vision Language Models LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.222949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.222949Z digest=sha256:bc052b331c76e54a7d477ae293a7b8c9c6ddaffa93909663ec06398647d0f619

Observation 44f6bc17-8221-4c0d-8520-aafe88bde415 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

FILA: Fine-Grained Vision Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.233399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.233399Z digest=sha256:53e0c44ec69091bfec1baa71af7fe1a7c059f7cba56bc98485d9ecf138c8f8db

Observation ae29ba6e-f9d1-4903-a0f4-529a03f08af7 · outbound

This paper cites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models.

FILA: Fine-Grained Vision Language Models Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.238546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.238546Z digest=sha256:5e78e44349d59364ce6a3c6722238bbad5b34201f26378fc44003ea4d32b6b9b

Observation 55c95798-9e2b-41f5-b04c-686960fd8bd6 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

FILA: Fine-Grained Vision Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.243482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.243482Z digest=sha256:c342fe35962535841adc8e63fe5fc353c8af82610f1e105916467357f2da83a5

Observation 8a2f6583-0adc-40a9-ad94-e54e32ebd7f6 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

FILA: Fine-Grained Vision Language Models Sigmoid Loss for Language Image Pre-Training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.248659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.248659Z digest=sha256:c8b71ee61f1a564f34fcbaabdde705de6ee0546bac56e4fe335803d7114a782b

Observation 7a54fad1-509b-4f1b-a290-bca8dd6e1302 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

FILA: Fine-Grained Vision Language Models OPT: Open Pre-trained Transformer Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.253816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.253816Z digest=sha256:1bdb2cb060a995896dcd715ab121962c27fa6666c2d3d65e1ea7360b1bb0ddba

Observation 13a0b67e-8927-45fa-b62a-1bb7f7b13556 · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

FILA: Fine-Grained Vision Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.258654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.258654Z digest=sha256:d4a0d239e56a3f7e1434b671884fbf7fbf7c1b03a214c1e868e1e9c7150bf96f

Observation d7a57c20-ba52-4629-bb46-b6dcbdfef324 · outbound

This paper cites LIMA: Less Is More for Alignment.

FILA: Fine-Grained Vision Language Models LIMA: Less Is More for Alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.263567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.263567Z digest=sha256:b5433b9527d253e57a7c19dc8269f37519a16e262e59594f78f424022c65b719

Observation 97de2a77-2931-475b-ba2a-127cf6c5935f · outbound

This paper cites We compared our model with Minigemini-HD and LLaV A-NeXT.

FILA: Fine-Grained Vision Language Models We compared our model with Minigemini-HD and LLaV A-NeXT

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:56:46.753400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T17:56:46.268813Z digest=sha256:13ea4026492de2b63e33d5a6f6842fb017c54636fdb32b94f95a084ab893b61a

Observation 3cc98f9f-d5eb-4096-94be-a52d521ab81a · outbound

This paper cites For the language model, we utilize LLaMA3-8B-Instruct (Touvron et al., 2023).

FILA: Fine-Grained Vision Language Models For the language model, we utilize LLaMA3-8B-Instruct (Touvron et al., 2023)

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:56:46.736611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T17:56:46.274343Z digest=sha256:f8391b59426cf685546d88fc0445463978c8f58badedc1da2b04b71418c0c289

Observation e10d6f62-724e-42c3-a318-5986a7073dfd · outbound

This paper cites 1 Published as a conference paper at ICLR 2025 C A LIGNMENT STRATEGY Conv StageInput Dimensions (D, H, W)Output Dimensions (D, H, W)ViT LayerViT Dimensions (D, H, W) 1 (192, 192,.

FILA: Fine-Grained Vision Language Models 1 Published as a conference paper at ICLR 2025 C A LIGNMENT STRATEGY Conv StageInput Dimensions (D, H, W)Output Dimensions (D, H, W)ViT LayerViT Dimensions (D, H, W) 1 (192, 192,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:56:46.720232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T17:56:46.279061Z digest=sha256:1651647ec1572982a78acc24277ee831034a46288b92a0353edfb7d5d0ec6167

Observation 90357c10-ccc6-4179-9542-bb05f495d320 · outbound

This paper cites DVQA: Understanding Data Visualizations via Question Answering.

FILA: Fine-Grained Vision Language Models DVQA: Understanding Data Visualizations via Question Answering

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.180959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.180959Z digest=sha256:4eb2ca8a4111190b6b1517a55de7ad2c6a88c6c9fe1bb41b965a8e04343a8251

Observation 4c3cc7a5-af74-4f29-a919-0133eb80ee98 · outbound

This paper cites Towards VQA Models That Can Read.

FILA: Fine-Grained Vision Language Models Towards VQA Models That Can Read

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.227473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.227473Z digest=sha256:071eb4b89340aab3fbaead15c4874fe3f39ba581286a1ad479726c6a236d3d7d

Observation 154c7030-8c73-4199-a715-a1fa32a30540 · outbound

This paper cites Language Models are Few-Shot Learners.

FILA: Fine-Grained Vision Language Models Language Models are Few-Shot Learners

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.136537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.136537Z digest=sha256:36eacf46bf3f198a7fd2a85f4a4d997ca8d24e35fcaa63cb9352b9eabe087ce0

Observation cc89895b-463a-4be5-adbb-c83152f070e5 · outbound

This paper cites Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts.

FILA: Fine-Grained Vision Language Models Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.142751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.142751Z digest=sha256:d7ca037a34327003948e12fcee3362d8cc9e204b4efea2130db5217f3dcad833

Observation 4094e0ec-f248-4279-add7-b8db8605f171 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

FILA: Fine-Grained Vision Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.130070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.130070Z digest=sha256:a181e2e4cfedd9a33130c3cc78fb357641501dc7ecc25913688e1930dfec627f

Observation c2f87344-0c1a-4a6f-8958-2bf4f66c849d · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

FILA: Fine-Grained Vision Language Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.148931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.148931Z digest=sha256:c09b7dd2a43f34f0f07da57b10295c49c290a3e30cdcc66e21f14fd870d787f3

Observation 81823777-0b65-4849-b7b2-a4ce37034255 · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

FILA: Fine-Grained Vision Language Models HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.170245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.170245Z digest=sha256:2b094e1adcca32a0e133a385520fee4d8cff4e59d509dc99186526a0881b8f15

Pith citing papers

Observation c76ac201-24c0-48b3-93cf-51a6d244060f · inbound

TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents cites this paper.

TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents FILA: Fine-Grained Vision Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:47:45.998928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T20:43:16.364125Z digest=sha256:1ef95588ff100e8918346f093d68ba835a160a1e497be48319a39527791df3ce