Pith. sign in

Paper Citation Record · LEDGER

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression

As of 15 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 1 inbound Pith citation observation for arXiv:2412.04317.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04317 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:37:08.457373Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T15:40:05.730181Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:16:15.637500Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b8a476c1-9b94-4994-8048-469c0fea17c3 · outbound

This paper cites GPT-4 Technical Report.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.227561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.227561Z digest=sha256:b0ea09e8503655f23b26481cb6a1c161d16d95fc1155ad1e1fe1e2ecad988466

Observation 630a6798-b520-4f35-bebb-085cdcd1846f · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Bottom-up and top-down attention for image captioning and visual question answering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.231500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.231500Z digest=sha256:d018104e81923b2ba4eef6374cf1e61006ef1db9ef5c5f95e7da1266eb04f885

Observation 7153172c-ea37-4214-8c96-4f76675d4715 · outbound

This paper cites Qwen Technical Report.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.234380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.234380Z digest=sha256:908a6ac75063850bc89c0f1044d28380958447287cb42cef9e603615eb8b985a

Observation 1bbea82d-0bf0-4f1b-98f7-05fb89259d53 · outbound

This paper cites Gemma: Introducing new state-of-the-art open models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Gemma: Introducing new state-of-the-art open models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.237958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.237958Z digest=sha256:e160bcb41719f5d352c214d3b26780545c05f8b73d9c34856e084b5b42053b8a

Observation 48487edc-c0c6-4e0c-9400-a6d7a23f3fca · outbound

This paper cites Language models are few-shot learners.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.240691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.240691Z digest=sha256:eb2450c5be806630f6b7d9928b0ea470c5107f5f7b548d625a09408e601bcb3c

Observation 997d8738-1bd1-4fd5-9852-7e9055b7c5f3 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Honeybee: Locality-enhanced projector for multimodal llm

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.243430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.243430Z digest=sha256:5982e04579b974839039b8f6ca756d47b387931fef71ec9df2fed414b41a6a08

Observation c5ce3a98-1b80-4e69-8f83-431aa4261138 · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.246133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.246133Z digest=sha256:09c59e77d31a89d9638280451e54f34bb4f68fe00c8a25c941d8b78fe5b494fa

Observation a2f7a125-3085-4c86-8b39-27d816b7fb04 · outbound

This paper cites Lawrence Zit- nick.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Lawrence Zit- nick

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.249780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.249780Z digest=sha256:8f4f93c447f8dfaa2bccae1e78fafccfef4604859b235c318fede740946ac1ea

Observation 25bd50ef-cdc9-4fcf-95db-4c4bdd09e788 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.252944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.252944Z digest=sha256:1e455c6e231772a6ddf69618e20c84ee9150c89e183bae25cafb77c0a736dfea

Observation 4da9f6f9-2f64-4a7e-bc1e-72e72a807aee · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open- source suites, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression How far are we to gpt-4v? closing the gap to commercial multimodal models with open- source suites, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.256084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.256084Z digest=sha256:4acb56eb735f778e66df9630f26e6bbb9f5c298c15654dff2317545eada9b3ea

Observation 5dffd154-db90-4c29-a3ad-c7e60474eb54 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.258655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.258655Z digest=sha256:e746964f195a4fd602dac6958dabfc5a87158c006d27fe61147c23bd5f295916

Observation 2cfabdcd-b567-4d75-9956-05574f714485 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.261489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.261489Z digest=sha256:a3c8cc461aada2c2ca4e00063c2d00e7f9a94ab4b85d4380b80a6f8248fd1bcb

Observation 9544fdc1-b620-420f-9a3f-199c7d42de2d · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.264453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.264453Z digest=sha256:22c22dd6c5e45498eb261a67175aba03245a1fb548be01fcdc919c650cc3373d

Observation 402e63c1-ce06-472d-b65b-62a8131ba07d · outbound

This paper cites Mme: A compre- 9 hensive evaluation benchmark for multimodal large language models, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mme: A compre- 9 hensive evaluation benchmark for multimodal large language models, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:09.007536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.267203Z digest=sha256:634c789883e3107ac939fed925422abaccb42298eeef94b7f61b935da2540599

Observation 2e3c46bf-2041-457c-a0a3-71607bc6ddfc · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:09.000272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.269857Z digest=sha256:ba57645e2bb6b610e6cad0ccc960570886d70169e84b729f81246ad166389e56

Observation 527398e0-0bf9-48a6-a377-b88a7d565e53 · outbound

This paper cites Minicpm: Un- veiling the potential of small language models with scalable training strategies, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Minicpm: Un- veiling the potential of small language models with scalable training strategies, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.992917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.272901Z digest=sha256:84f000b4908eb6c1d1a8764bc34376e57b7f68f6f8b07da08effa586b8a3da1b

Observation eab121c9-2d36-4c35-a170-f3587b65e6eb · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.275372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.275372Z digest=sha256:fb29a8388781fe63883f0e31c726a275bf9159319e57282f5c9361c8ad98db18

Observation 6f0c000c-631d-492c-97a2-2c3ead4dec59 · outbound

This paper cites Token merging for training- free semantic binding in text-to-image synthesis, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Token merging for training- free semantic binding in text-to-image synthesis, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.985333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.278335Z digest=sha256:1b5704a2e1591e60cd969e92088c8b6def0223e98158fbda26941150d245351b

Observation dc1890a4-bb9d-4e1a-bc26-0cbe64f298a1 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.977404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.280814Z digest=sha256:c9443f70ba9022bf6449a201f76bc91151edf197da3d8e22a2307f752a6fbfe8

Observation e149fc0b-b43f-464a-b106-f780aadeb0cb · outbound

This paper cites Phi-2: The surprising power of small language models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Phi-2: The surprising power of small language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.969681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.283336Z digest=sha256:5c097ed448e811aa759da05bc0cdd27333e29a027fa8308bfc3b81f3c6953b8f

Observation 5200a819-4edd-424e-bdc8-2efbd06283bb · outbound

This paper cites In defense of grid features for visual question answering.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression In defense of grid features for visual question answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.962265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.285778Z digest=sha256:6e3e0e432bd83565574e293dced7c10948b2c7fea4d128b101a36e0b77917b3e

Observation 2f69eca5-7024-437d-b957-f3d20f9fb4be · outbound

This paper cites Contrast and classify: Training robust vqa models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Contrast and classify: Training robust vqa models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.955006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.288247Z digest=sha256:fa5ebeb99b668fe14b63ded86611490ccea1bab9592994eb639a55a5426fc032

Observation 295cde2f-b1d4-433a-9d22-a4aa1aad5b48 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.947365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.290817Z digest=sha256:f045af4a672b3a657b5061317799fc603539b1433e27b5c0b143186a4a7a4ae3

Observation 14175937-1d25-4def-9bdb-51590fdf9c75 · outbound

This paper cites A diagram is worth a dozen images.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression A diagram is worth a dozen images

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.293615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.293615Z digest=sha256:5b008176bf53202ce307f4aab8153a3cc429c717fe71c04b93ef8a2fecefcf7e

Observation 19186a00-6f9d-426d-a3ce-03d6c2c877e0 · outbound

This paper cites Seed-bench: Bench- marking multimodal large language models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Seed-bench: Bench- marking multimodal large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.936399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.296242Z digest=sha256:b868fdec7f3f64c06ca497552ec44aa9a7e851b1d12c5148cec399c60e1b0517

Observation 6fb39a02-0e10-4c36-befe-608ae9bd53d9 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.299006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.299006Z digest=sha256:68a2b0ae3ed927e29c6c083642e15ce248d491afeae0df19b916f7b58a1bf2e3

Observation f294ee1a-9caf-4144-95ce-b22068d88a12 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.301741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.301741Z digest=sha256:baa49eaefb95cabb1c4fe06fd2ac7abaf4bf8807baaf9c5a4346b2ffb4ee345a

Observation d3339ae5-77eb-42b9-97b5-e22a2112aad6 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.924912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.304194Z digest=sha256:fcbf54fe7b21c9d322e4be55364ead6a7421e18f27f92b5b47c29e8b8907a05a

Observation 64483dff-5527-4c2b-ba48-8dfb8a2a7b10 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.306619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.306619Z digest=sha256:a8093a441cd496d193152dae7056d9d124af633babf2115e941a7eeaecea675b

Observation d0e2b9f2-4dee-4358-9153-54d800348272 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Evaluating object hallucination in large vision-language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.917574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.309221Z digest=sha256:c21f35d893d6cbca242b9e89114a6156c12acfee76b22cbb9e5fbe703a0fd9ca

Observation be65e829-2862-4674-a628-60aeec5efeb0 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.311714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.311714Z digest=sha256:5894bfcf12b464d8be9ebc1c39f6e43015594b25ae003e27a7dad18f1df0acc5

Observation 6ec3cd88-490d-46c8-b9b7-102a66d57b8f · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.909721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.314473Z digest=sha256:73eb9ed0ecbae3b73afcad2ef89392bbd071c9c173ca394b1be73cbfbf2e0707

Observation 1a3aeb31-8b46-455b-aa92-cc18366b8b84 · outbound

This paper cites Improved baselines with visual instruction tuning.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Improved baselines with visual instruction tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.902388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.317072Z digest=sha256:a46d86f9d7bfda3c51d2ae4c896f483ef00fe6cc6242e726d03f409a441d58a4

Observation e5c29df2-6b30-4368-b898-6e8415d2ca4f · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.894198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.319771Z digest=sha256:e55fedf0da432d5e86e0a0fc05ed89ee44d3f01aea7fc39189b11d5afe4b5487

Observation b77bb695-f88d-4844-aa9e-2d08401f8d9f · outbound

This paper cites Visual instruction tuning.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Visual instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.885784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.322118Z digest=sha256:f3db223a63a32bee02ad0f62c197c1682b2c4fe843a9a1d88f98a5821cc39d60

Observation 8e6abf79-d0d5-4d55-a27f-6f3591ac9396 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mmbench: Is your multi-modal model an all-around player?, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.877386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.324392Z digest=sha256:a48a56a4f17bcf900a0fc12eceb11a0d634f1f28530f2c20238e930c616eeb00

Observation 4effb140-fbb8-4d2c-943c-6b91af32192d · outbound

This paper cites Decoupled Weight Decay Regularization.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Decoupled Weight Decay Regularization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.326784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.326784Z digest=sha256:c2513501cd26ec6fe54d6000354d1313dece9b999341a56965c54ccbb1acabba

Observation 26d39eb5-e572-43be-ad1b-f553186cd52e · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.329705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.329705Z digest=sha256:684f9e121077970c2e3b26fd99ff5aea32a8019fa5c05800bbf82bd36e03e423

Observation ab43a5b8-737a-4264-a66a-835219ebd03b · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.332592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.332592Z digest=sha256:4a8980e6369fe36730375716dd5a10f50121678be1749881374830cedf907ba5

Observation 72c5abab-e79f-4606-9c23-25d0eb12da07 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.335366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.335366Z digest=sha256:98f6f328a54ce15e293cf49c527bce5d0dd0b851303ad05c66f941e76e0da326

Observation 7327ee58-e4d1-4ef8-ac14-950aa62dff77 · outbound

This paper cites To- wards lightweight transformer via group-wise transforma- tion for vision-and-language tasks.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression To- wards lightweight transformer via group-wise transforma- tion for vision-and-language tasks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.865032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.338155Z digest=sha256:f2cbe0ea446b5db8981e8208fe0aee5b5171da292e9ae699d496d9512288aff1

Observation f496d380-de84-48cd-9559-242a5d46327a · outbound

This paper cites Cheap and quick: Efficient vision- language instruction tuning for large language models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Cheap and quick: Efficient vision- language instruction tuning for large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.857341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.341165Z digest=sha256:5b7b6348699c79d6dbf2f7ec121781375b87a73ef0c625d2aa23483cb850fed2

Observation f99125b1-42d1-4b03-b293-14644c5b97a5 · outbound

This paper cites Moil: Momentum imita- tion learning for efficient vision-language adaptation.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Moil: Momentum imita- tion learning for efficient vision-language adaptation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.343738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.343738Z digest=sha256:06d3c907c37736b3934eb470cd4b74b558acdc5dc59c432eb5bc2382f3c5444f

Observation c9fbdc75-f9fc-4890-b09e-5eb04d28568a · outbound

This paper cites Towards language-guided visual recog- nition via dynamic convolutions.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Towards language-guided visual recog- nition via dynamic convolutions

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.845990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.346149Z digest=sha256:0e3cddc79a8b081ec1d1fca176a6d7a6416ba8c81417a136e9582079c856ea86

Observation a0e638c8-4f5a-4ed1-ba4a-15faa67b0baf · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.348453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.348453Z digest=sha256:d9fedc7127bbd9662be609acadd1ebafa8c61b65ab80f1b743ecf6f3b8f8de79

Observation 8e27a5ca-7e74-46f2-ae78-61ad377c1b7f · outbound

This paper cites Chartqa: A benchmark for question answer- ing about charts with visual and logical reasoning, 2022.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Chartqa: A benchmark for question answer- ing about charts with visual and logical reasoning, 2022

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.838006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.351277Z digest=sha256:df8aa41e07432cea1429d7177c7290b82540ef4c809fdbd80a7681a8ed8f5120

Observation b94b76ec-e7bd-4786-92b9-4028b6f44007 · outbound

This paper cites an unresolved cited work.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:37:08.829927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.353618Z digest=sha256:344441e4ad75accc677342002b91c6c5e085e7b8bd7c087dfbef1dcd7d1e41fa

Observation 0cc2e6e6-aab0-40db-b992-db6c012216f2 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression DINOv2: Learning Robust Visual Features without Supervision

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.356012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.356012Z digest=sha256:ec48f8239ed6f02e2f45be0ed064b562f25d28d9b6ef2ea3feeb062941bb77d8

Observation 208567de-d3f6-45da-9e74-a99002bd76a9 · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Learning transferable visual models from natural language supervision, 2021

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.358592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.358592Z digest=sha256:ebd53b9983e411097b55a10d890dd5deb08e4f7912413a74a8fd7d329487944b

Observation f54f0555-8992-441c-8355-7a9e7a0297b2 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Learning transferable visual models from natural language supervi- sion

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.361676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.361676Z digest=sha256:6e3b8c73a9d9cf713412ecb5ad0cb5459cc686db90ac453cda81007d65e48d5f

Observation f09e7d82-b271-4736-aea6-4599f8fcc571 · outbound

This paper cites Imp: Highly Capable Large Multimodal Models for Mobile Devices.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Imp: Highly Capable Large Multimodal Models for Mobile Devices

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.364232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.364232Z digest=sha256:6e481e261509459d4d1d6fdd3c4b5539e0736f226e835d90a9dc954b85944ac6

Observation 873ba2fe-e190-4711-860f-6bf6827ed107 · outbound

This paper cites When do we not need larger vision models?, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression When do we not need larger vision models?, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.812922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.367027Z digest=sha256:20c6d3fdffced6c384b1c763866f125dbbbc074e1105891c2ba6b7cd07b09e12

Observation 16054fb2-0a0c-4f1b-8662-43a7cd5815a4 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.369438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.369438Z digest=sha256:1c700a837df2f5fd4243bad76678b5b2c9ad98653d473e3c29c467a710552c1a

Observation 60bbf9ad-f9a3-4482-b029-b3826a42ba27 · outbound

This paper cites Towards vqa models that can read.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Towards vqa models that can read

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.805615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.372029Z digest=sha256:f2f917c1eb2f4e39366af51feae7fc0affafb9b440c6b6cc2a6036eba0c6e60d

Observation 32e1e7cc-c75c-48c7-bf92-4d50e1954a75 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Gemma: Open Models Based on Gemini Research and Technology

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.374548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.374548Z digest=sha256:41ec3453b6e4017ff059ac48f6e27fc4dbf5065703cea34cbba4faf6dca95765

Observation 92d80c96-c396-4a7f-8abc-a8b4742a2691 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.377247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.377247Z digest=sha256:6ac9ac39f980552b859a373b2b1bb160b789f9bd9bae6c9072d9a18dc6e3c286

Observation ef2dddbc-3ab3-413d-bbb2-889b9be9d0df · outbound

This paper cites Well-Read Students Learn Better: On the Importance of Pre-training Compact Models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.379783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.379783Z digest=sha256:2c1f22c1d3bff957bc0c114dbe0027be909877d732a1ac7f23d255dec3438aa2

Observation a4b708de-7375-4531-97d7-9e73da4f9c3e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.382822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.382822Z digest=sha256:dfa32485505ce8d32e0d7e5a3df5e09d283303e6af24c2e650458748c49e08ec

Observation f271e797-ccf5-484f-8716-5a434963d2a9 · outbound

This paper cites Show, Attend and Tell: Neural Image Caption Generation with Visual Attention.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Show, Attend and Tell: Neural Image Caption Generation with Visual Attention

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.385543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.385543Z digest=sha256:1e3349ea7bdd40fa57d2ad60b45014908b4609c1b2ace18d80c6b564a137a790

Observation df808884-647b-4e16-bd72-646e0931afae · outbound

This paper cites Stacked attention networks for image question answering.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Stacked attention networks for image question answering

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.798043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.388241Z digest=sha256:4f1db3beeb2cdc2c1a1b893a0c543c687337af8ee2a850ffbec6dea5f6117d09

Observation 2bf59c01-1ee7-41e3-ad7f-a02e7a9590f1 · outbound

This paper cites mplug-owl: Modularization empowers large language mod- els with multimodality, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression mplug-owl: Modularization empowers large language mod- els with multimodality, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.790234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.390580Z digest=sha256:e9f90a951f73f3671093eb3e2e7fc26ac9db185c38cb98b83456eab11261a1f9

Observation 55a8b314-8652-4f49-8c5e-a1fc978e722e · outbound

This paper cites Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.393098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.393098Z digest=sha256:56295bacacb8dd9b7c4dba730b40363cb78902e485c3b7e55db5ab66bed789ad

Observation 92f936f9-cd1b-4c9a-9078-69b019d9c7ef · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities, 2023.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mm-vet: Evaluating large multimodal models for integrated capabilities, 2023

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.782501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.395868Z digest=sha256:12ccb5925f7826e359c58ec2d76c3983b6bb5be995fb9f07324daac2642a4712

Observation a8ab1367-712e-4e3e-ab80-7da9fa0e8e85 · outbound

This paper cites Deep modular co-attention networks for visual question an- swering.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Deep modular co-attention networks for visual question an- swering

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.775215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.398429Z digest=sha256:ec8826d5f7d4d10b4f26ad0425e32026b97539c5e9468d8779b92f73910b426e

Observation f72e0beb-c44f-48c2-96de-368aa8556536 · outbound

This paper cites Tinygpt-v: Efficient multimodal large lan- guage model via small backbones, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Tinygpt-v: Efficient multimodal large lan- guage model via small backbones, 2024

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.767930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.401027Z digest=sha256:7d0ef4bc9046251413fecfef932e95841bb7510f03866fcc8e2f8a32064aca68

Observation caa8085a-2a3c-4a56-a9cc-ff0e030ee692 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.760040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.403823Z digest=sha256:5d81c5a8dea85b386acab909567a262395e03eccdfc92e32ac829225e9b76be7

Observation b2f69a39-5f6d-4fe1-94ea-6661d0005158 · outbound

This paper cites Sigmoid loss for language image pre-training,.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Sigmoid loss for language image pre-training,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.407076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.407076Z digest=sha256:2109b2953d1959883a7a9d28ddf4ce02c8eafb3af7293b8c56aa4b3934f592fa

Observation acf37657-d8f1-4fec-a07d-c07e8a932943 · outbound

This paper cites MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.409547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.409547Z digest=sha256:e7cf12bffb2160ed78a0ca22847fda4e870d687cbad2556e5eb5ffc48ac3fbe4

Observation b0f53947-2e19-437d-93ad-4e44cb06403e · outbound

This paper cites Lmms- eval: Reality check on the evaluation of large multimodal models, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Lmms- eval: Reality check on the evaluation of large multimodal models, 2024

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.412907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.412907Z digest=sha256:1e9fb9534ab4911624a60aa8516d29868fcc68714de1eac07e5e171eb521004e

Observation 5c07de0f-ffba-412b-b950-8ba9bbd313f2 · outbound

This paper cites Vinvl: Revisiting visual representations in vision-language models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Vinvl: Revisiting visual representations in vision-language models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.415267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.415267Z digest=sha256:3b91aa1edf9e4fd17733d90886ac8d168b5ff6220229e5c6e7b807a13e043ac7

Observation 4536e021-a980-4ee2-8a99-b48b30bccffe · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression OPT: Open Pre-trained Transformer Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.417720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.417720Z digest=sha256:2813839016aac9fae2c8880995111c5c9046538cb39f35a24183fc6e7fb157bb

Observation 4dd6a1d4-33c7-4ec0-94cd-52a422597aae · outbound

This paper cites Free vqa models from knowledge iner- tia by pairwise inconformity learning.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Free vqa models from knowledge iner- tia by pairwise inconformity learning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.740215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.420493Z digest=sha256:96c6e53b05fbb92072dec4efe1730369c6a35df67a00753bd795f145fd2c3b5f

Observation 39a79b3a-3f8d-4519-9f6c-dfb5d8b3ebd6 · outbound

This paper cites Trar: Routing the attention spans in transformer for visual question answering.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Trar: Routing the attention spans in transformer for visual question answering

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.732876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.422900Z digest=sha256:8e53fa1cbc9ed18ce9f5af0001befdea0a09ee40c43f77dda739a6951b77e1cf

Observation 863d176b-39ce-4965-b82c-bc35082c741d · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.725697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.425431Z digest=sha256:35ab84c8933d7b0fdb1e714f32e862386e6bc4886953f01698a18e071f6372ba

Observation 6d31a72e-f5cf-467c-bda0-e859e0ea21ac · outbound

This paper cites Mipha: A Comprehensive Overhaul of Multimodal Assistant with Small Language Models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mipha: A Comprehensive Overhaul of Multimodal Assistant with Small Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.427822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.427822Z digest=sha256:21e8815352dfe9166e7799a382fd6f157bc63ef04a2e1bfbdc19f1b785021531

Observation 7a1cb42e-0145-42bc-93d2-d71b52a7757a · outbound

This paper cites an unresolved cited work.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:37:08.718218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.430767Z digest=sha256:611043a5a3b7028b2eecff05af667fd246bed83e553e86ef1456565b07e3c249

Observation 044164c0-8c90-40f7-8768-cc4f0c7ebb24 · outbound

This paper cites an unresolved cited work.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:37:08.711080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.433171Z digest=sha256:c58aff4a2525a4f104d49e8ccb78b4ecb5b99d514cdc3ca54d05e07513c62f6b

Observation 9803d311-a481-4c98-9916-a256a7ef60a0 · outbound

This paper cites an unresolved cited work.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:37:08.702909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.435461Z digest=sha256:9be98e6f6a351e4b5b4ee4be794cbb56f443bf9777f190ba5151bf24d1f1f262

Observation b8c945af-388a-4435-8980-7a542ac4e4a4 · outbound

This paper cites an unresolved cited work.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:37:08.695244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.437867Z digest=sha256:dfc28599b0e7e1a0377c777bf315d3c535293be7b008dc1f214f5e4e5ba550ec

Observation e7484725-fc2b-4fdf-8647-f81f95fc8b3c · outbound

This paper cites an unresolved cited work.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:37:08.686740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.440180Z digest=sha256:0046a0a83d84e53232f714aa34806e2bc341550e35cfd62ec2338592e2df2170

Observation 5d5ce4cc-e1d8-4f57-97bb-d4ec7709e76a · outbound

This paper cites Therefore, the value of the square in the figure is 2.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Therefore, the value of the square in the figure is 2

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.679023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.442544Z digest=sha256:b0023e5ac57380cce19ee724406a1c5fe11ac723503521bc58f549e49de823d9

Observation 823a836a-6e37-43b4-bf38-f2db8912ba61 · outbound

This paper cites control center.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression control center

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.670988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.445034Z digest=sha256:c1753143ba43da4a5cd21e3c4281dd05882034fd65cece37ec97ae71077bdc87

Observation 3ab4ddba-1331-4c01-9e1b-b3940d8881e0 · outbound

This paper cites an unresolved cited work.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:37:08.655902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.450057Z digest=sha256:1b3c9aca5d9d60a51b4bc22126d65fdd070b0e4d038e31b74d68b824adbef9f0

Observation ad67f5a0-d0b2-4b9b-95da-b78400b582a0 · outbound

This paper cites an unresolved cited work.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:37:08.648609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.452450Z digest=sha256:7ba00ef59782dcec356b2f6b4a7adb5c34879e14e218ff83b186eef1ccbe0ca2

Observation 275209cb-2bf7-4178-aac4-6d48a670f5a5 · outbound

This paper cites an unresolved cited work.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:37:08.641489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.454943Z digest=sha256:f21293bfc212322f079ebb47ea78298aa6a5653f309dcd25667a937687167c03

Observation 9c492150-ab38-4202-8025-e9e1cf4e7473 · outbound

This paper cites RIGHT",.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression RIGHT",

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.633594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.457373Z digest=sha256:e70a6589a52c5e42ee153233be6fce2aa1c9af6d5444aea5a9db12d1c9ab5355

Observation 6c5fd70e-35c4-484c-aaea-fffc2a260bd9 · outbound

This paper cites Decreased.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Decreased

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.663257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:37:08.447590Z digest=sha256:6d080b32e78b133864b6b0d3e8028da92f70a49100d7228c2524cfeb27c743fb

Pith citing papers

Observation a4d841cc-d0f7-44db-9149-201bb63d1284 · inbound

Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning cites this paper.

Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:16:15.639326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T15:40:05.730181Z digest=sha256:4fec1e6cecf10ab647252fb2627b39213ef513d7ad0a262054279408f307336d