Pith. sign in

Paper Citation Record · LEDGER

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion

As of 14 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 2 inbound Pith citation observations for arXiv:2412.04424.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04424 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:27:52.652519Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:19:23.116337Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T19:52:01.828238Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b956506a-982a-431f-8170-27aaaa06b663 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.358895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.358895Z digest=sha256:4c5335207cc5ba9bab0e5ec717362bba97b62f1efe1e86ffcd862988d9896959

Observation 73d66cdb-0fd9-4603-88fb-82d5e93f9db0 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.365941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.365941Z digest=sha256:da36f014c84c284197421a90c341499d150905f62090ac414f71d57ca41949af

Observation e3dbb1f8-5959-44b1-bc98-6ab0b1857bba · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.373463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.373463Z digest=sha256:9e6c4eb2ef023e69c1b61eedd34f804e8b6f2ec1a7363dd07ff7295d425f6d11

Observation e545521f-1398-4fcb-acaa-3c8b045d2899 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.595830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:27:52.384569Z digest=sha256:8a33fa002f3e8bd781b23e9f6bc84ae4526c241ca3cf39253ed2c8cea8661f36

Observation c2634cb1-deab-4a63-af2c-963f94e4c157 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions,.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Sharegpt4v: Improving large multi-modal models with better captions,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.390586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.390586Z digest=sha256:0090590012012abee6f3168b20fc926762e1c8e6a269dd8e42843077206273bf

Observation 4bbbe195-f89a-480e-8f8b-41904477a200 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.396272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.396272Z digest=sha256:0f8cb319cf3e573bce990a57a405678e773703baa11db09481d3cba26c129959

Observation 705a2024-bc8b-4132-88a3-a692ecf4a82a · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.402068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.402068Z digest=sha256:d4d5b7f673f6286e63c794429beda46b37b531c62cd22733501de4ed10a71330

Observation 9414ee1a-b5f6-4f34-8654-3c2d883cd8f5 · outbound

This paper cites RedCaps: web-curated image-text data created by the people, for the people.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion RedCaps: web-curated image-text data created by the people, for the people

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.409590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.409590Z digest=sha256:d446acbb0dba0a2537008bd75e982fbdc523a10ee644dcf027c13f7a3e5a8061

Observation ecdf898e-0441-4504-ada4-9e1fbcfa386e · outbound

This paper cites Davit: Dual attention vision transform- ers.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Davit: Dual attention vision transform- ers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.416001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.416001Z digest=sha256:1a88f3a5fe04d6c982a71bcd94a4dac53e31744cfb04f03f1138958cbef0134d

Observation 65923e19-8764-488a-8f45-b407ced021d8 · outbound

This paper cites MouSi: Poly-Visual-Expert Vision-Language Models.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion MouSi: Poly-Visual-Expert Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.421540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.421540Z digest=sha256:5f7e1c3fefe1a8d68b0ef67284015fe9f7447fd83a700a783b2583203e84b41a

Observation d85dddf7-2316-4c01-aef1-2e9e7880a900 · outbound

This paper cites Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.550135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:27:52.427855Z digest=sha256:62fb5242b1b9f0f068c2b12964df0161c6dd88a55edd6f579fafa89da2a783f1

Observation 9949e921-a402-44e7-8569-72e9ea923183 · outbound

This paper cites 1 OCRBench ChartQA DocVQA InfoVQA Average Florence-VL 7B 41.4 24.3 44.5 29.4 34.9 OCR 40.9 22.9 44.4 29.0 34.2 (a) Ablation study on OCR features on OCR & Chart benchmark.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion 1 OCRBench ChartQA DocVQA InfoVQA Average Florence-VL 7B 41.4 24.3 44.5 29.4 34.9 OCR 40.9 22.9 44.4 29.0 34.2 (a) Ablation study on OCR features on OCR & Chart benchmark

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.530537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:27:52.433721Z digest=sha256:50acdf6953c51cb98c024516e94c5b6d3a5d17cefe9ee3c23c8c7649984e97eb

Observation 2a861900-2d06-4a04-aac0-5dff573460da · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.440844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.440844Z digest=sha256:b9409e485bc71ac8b97f8807a2035b1a05469767a82162f8824081febefec6be

Observation 1275089c-6f36-4479-b281-f93268b4aadc · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Vizwiz grand challenge: Answering visual questions from blind people

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.449811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.449811Z digest=sha256:6440eb17d0ad96a363b9ca6fbc0638de9d968f3ddab09ed82dcc101f99eb8b7e

Observation 97833a46-8d04-47f4-9ce4-7e47936c38dd · outbound

This paper cites Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.456702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.456702Z digest=sha256:793d97f6bb0efdd901adcbb59d1e0d96c79b0df6f6409e8fba7792e127732c7c

Observation 9057387d-297f-4da6-b097-d93f084b186b · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and composi- tional question answering.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Gqa: A new dataset for real-world visual reasoning and composi- tional question answering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.488827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:27:52.463327Z digest=sha256:a651025c2621c79d31fbc7c5bc882ba80d867023ccb1702271e44a2f4623aed3

Observation be1c5739-1cc9-4fc8-a3d4-d63acc3e9d10 · outbound

This paper cites https://huggingface.co/datasets/huggingfacem4/docmatix.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion https://huggingface.co/datasets/huggingfacem4/docmatix

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.471426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:27:52.469074Z digest=sha256:3b12af4d0439beb0db67242b32711e75b607edcc426e006ee9403792118f80b2

Observation 25e88791-2850-4d96-abf5-cb34e4ed394a · outbound

This paper cites BRAVE: Broadening the visual encoding of vision-language models.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion BRAVE: Broadening the visual encoding of vision-language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.474048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.474048Z digest=sha256:c78bbbe2a49110155f8668c8696a84c19a62d39a38defd1c09cb103a307f0fe8

Observation c81180e8-05f8-4fdf-b406-7af522edb8a9 · outbound

This paper cites A diagram is worth a dozen images.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion A diagram is worth a dozen images

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.480290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.480290Z digest=sha256:f96c8239d70d9d675137c4f7d229f4241f18aa9dfc1a42b3779efa4811f98402

Observation 69d6e750-e170-4d16-b40f-58e33aa5b120 · outbound

This paper cites Segment anything.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Segment anything

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.443083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:27:52.485430Z digest=sha256:849b3ac3314f92db8e7390139539962d039da2af64924add923a4679225c7698

Observation c18cdb79-ee2e-4d68-80ad-9c318057e82e · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.490300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.490300Z digest=sha256:df55040d05763ce3e1b43264f66d0c005ac7d6be678b56ecc754dc81aeb961c9

Observation 1c185393-b516-4b8f-a95a-a7f5e050ab82 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Evaluating Object Hallucination in Large Vision-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.495049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.495049Z digest=sha256:3b2b430506a2e0e05fd583edae26a13c094fb3e6216262f6b59b1158cdd228b3

Observation b7ada76e-a126-41bf-8e2d-2544309c0542 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.500877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.500877Z digest=sha256:7b0d40a12a98572262b44f0761207cb194fc5fd66ccb3db43ff0bd0e8465e6ee

Observation 585e9a9a-c40d-424c-af9c-6471ea6713dc · outbound

This paper cites Vila: On pre-training for visual language models.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Vila: On pre-training for visual language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.425426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:27:52.507812Z digest=sha256:5b114273fc7f20f312e7e8eaf557bb747bbc3ac325dceb498353a71b90e52fe6

Observation f82b0b6d-6db2-4632-b1e7-c07ee3860d84 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.407582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:27:52.513166Z digest=sha256:e160abb788237b8797edd9a3222a13a61bef23da742cc58736d72c6affc7f65c

Observation 0405be2a-9cb1-41b3-9e90-e9d964aad5f3 · outbound

This paper cites Visual instruction tuning.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Visual instruction tuning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.388823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:27:52.517698Z digest=sha256:b5e9359ce42f343232d6a0d560ff6bf473d5c2a1fdddc2842b985521d6ba41a2

Observation df400d4d-2b78-483d-a1cd-3d232cc529a9 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion MMBench: Is Your Multi-modal Model an All-around Player?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.522098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.522098Z digest=sha256:9ff64ca9c9ad6525aff1998673fee90f71c93f8b9cbab7e45ea45710958c63d7

Observation 66be9e77-1902-4cae-8bfa-ad0a2649ecd9 · outbound

This paper cites On the hidden mystery of ocr in large multimodal models, 2024.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion On the hidden mystery of ocr in large multimodal models, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.370325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:27:52.528019Z digest=sha256:8c59ea562453dfa3cd8c5627386927dd7395aeca2a2b1b61669db68ce18aa502

Observation b4cc4588-c9cb-406e-a003-5ad3c972f9f1 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.534535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.534535Z digest=sha256:9d0ece05611bf78a7250f4d3643e2a01d19fda73c1ae8440472d05dbcf9e1bfe

Observation 7882e73a-5805-4a99-a5e8-d8f1de9d942d · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.542222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.542222Z digest=sha256:ebc846e2dfc58ad7547b8790b132d9f2f4d42609f47d03514df8db2155e288d5

Observation 03014ae0-8ff4-4404-b100-0b1158232468 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.549213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.549213Z digest=sha256:3650f1fdbbb62f440c6946471e492b1ebb6affd061c91515e9940330d708f7ab

Observation 6b97a60d-041a-4f64-84d2-d8ab627bcdff · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Docvqa: A dataset for vqa on document images

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.340149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:27:52.555743Z digest=sha256:977ef5fb11350f01fcffe7ab58996329b70ec07455e074dd5abf66a06516d442

Observation 47f88643-2ab5-451e-a260-1f0eae8f096d · outbound

This paper cites Infographicvqa.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Infographicvqa

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.322795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:27:52.562395Z digest=sha256:7cf067642aea54248fa98a0d860cd6ff18cb3d43da095b8de7ea2d18b886f0f7

Observation a398bc26-f685-401e-bcd0-73a7cbafa363 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion DINOv2: Learning Robust Visual Features without Supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.567906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.567906Z digest=sha256:bae04c3151d4c78226f2563622e37490cda51348f39779758f2dcea4da1ee5a2

Observation d4943f97-7c2e-4cbd-808e-d00c1d060b6f · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Learning transferable visual models from natural language supervi- sion

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.576254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.576254Z digest=sha256:b9b15a3e03f852354aa091ab0f589fcdc82a966722add3f5481c2765e972e60e

Observation b88d5a6c-e5f9-46c9-ace0-dc8f350875f1 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion High-resolution image synthesis with latent diffusion models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.294408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:27:52.581690Z digest=sha256:cc5c61e488cdde3c4f2c63c3d913a03503a304eb46f9e3c0856ccdee0031001c

Observation 71e6ccaf-f75e-4a55-97a5-ec119f8c7d18 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion High-resolution image synthesis with latent diffusion models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.586544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.586544Z digest=sha256:be80baf0a14ca9c0290272ac06edf412a9b33eb8cb659f3d95e5477d9c69367a

Observation 09010bec-878a-4b04-b1d3-177e846291a6 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.592158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.592158Z digest=sha256:643966293d50d006510854e21e1223ac190e9402d3de9369d2651db5fcdfcf49

Observation 5d7b7ede-4e70-42f2-9375-2c3f2253565c · outbound

This paper cites Towards vqa models that can read.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Towards vqa models that can read

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.597176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.597176Z digest=sha256:55064fd8f66d287f9b61b86bbaa3bcdb330d61bbb3271f3292d24235b3a40c46

Observation 8a3a990d-15cd-4897-814f-c08f30a6125b · outbound

This paper cites From pixels to prose: A large dataset of dense image cap- tions, 2024.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion From pixels to prose: A large dataset of dense image cap- tions, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.245545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:27:52.601645Z digest=sha256:2cfcf7165df6f84b7709b83c670e5daed42004827302a48f9686575e3f715c19

Observation ae169a02-3cd8-42c9-8da2-c956e3896f11 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.606243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.606243Z digest=sha256:5c26e897739e3fc1aeba59c4e5d832a09517b7f1d750b6419d215afc9098b988

Observation 791cc9f6-8b4e-42c2-a0d9-998d402dfeb2 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.223712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:27:52.611668Z digest=sha256:f3960bb858052a9405ab3ec44eda9698c8178c2b89f6af934686d4c23bc9eeb7

Observation 539ee10d-03e0-4204-8786-9fd0a58bf643 · outbound

This paper cites Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language Models.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.617470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.617470Z digest=sha256:4df8e4d5e3a3cea154805bcd18bf2f5d5ef05350f53ab8980e0980ebf26c776c

Observation 50447d1f-36a5-4257-859b-53a1eb9bf355 · outbound

This paper cites Grok 1.5v: The next generation of ai.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Grok 1.5v: The next generation of ai

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.206047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:27:52.623951Z digest=sha256:7a67f1898331d19a1281025e9eb1485cd0737732cfaba6e96d9d56a2566bb21a

Observation 57ac3f18-f882-475b-b822-65dc159353f6 · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.189596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:27:52.630399Z digest=sha256:452b66edf4f75885b6d320e59116d5cd89c2129c7309e8480bc820ec21298925

Observation 297f77cb-0bbb-4f14-a6e5-a1b75c730c1d · outbound

This paper cites Vision-flan: Scaling human-labeled tasks in visual instruc- tion tuning, 2024.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Vision-flan: Scaling human-labeled tasks in visual instruc- tion tuning, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.170495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:27:52.635275Z digest=sha256:acb6e91a83f881baf61a5f1da97be0cd24d0ef103ac800d70acb93f2506b5fa1

Observation 92bde076-13c4-43fb-b283-aa4ef0b83d90 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.641065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.641065Z digest=sha256:6fafa80ee220b332d1b9c4552ae4c5bf190b7aad5c34376e8c977f303c9b8d29

Observation f0b8e686-718e-4e5b-9b77-3f0029a142b7 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:27:53.150271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:27:52.646737Z digest=sha256:50c139a3fe653e01e4f9ef864dbd744eccd820b17a9e81e83591680679d0c4d8

Observation 981cc79f-9a2d-4088-ab3c-0344007fd948 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.652519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.652519Z digest=sha256:9f1b2d455766ac5df567b776fd2060b811c9f6a616abcf6445cea6988c138cc9

Pith citing papers

Observation 64004760-0a21-4965-9e0b-4fc9dcc32781 · inbound

FastVLM: Efficient Vision Encoding for Vision Language Models cites this paper.

FastVLM: Efficient Vision Encoding for Vision Language Models Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:19:23.116337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:19:23.116337Z digest=sha256:0de43ef580a6cf12f10a556d662df041cb7cf199c18e687aa686ee38b46f19c1

Observation 218253da-2784-4767-9d0b-c3607ba54cfd · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:52:01.829787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:cf968dc8f2781d67b23b428a2ee9ebe2514f6101a4ef7659a7a78de26aaeeef9