Pith. sign in

Paper Citation Record · LEDGER

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning

As of 11 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:2508.01540.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.01540 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:36:21.235533Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved65
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3509e6fe-17cc-4f38-a96f-c7bc8453f957 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:14.734761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:14.734761Z digest=sha256:ae52a5970d5095bd849bea1460463db41d3fd3c3c05e24aa14f56f4a6724d5ad

Observation 4107bf47-e436-4f5c-8fd6-bd4191b4c2f1 · outbound

This paper cites write newline.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:14.860484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:14.860484Z digest=sha256:4dfc402fd5eae2e8b42274f75b7c4ba63ed68b691b28e0cafc670bb4ba3bac67

Observation 6f2ad7b1-59c7-4c4d-a721-e5c1dc70fc6b · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:15.006639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:15.006639Z digest=sha256:653cc7955f2a3597b6f50e9d3a9d4837a748f492f7bd7ae12b9a1289df5a0267

Observation 8354bbe4-7984-416d-96f8-b20b65d1236e · outbound

This paper cites an unresolved cited work.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:36:23.241059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T05:36:15.249216Z digest=sha256:c8f7d6208b82e404d8539de612cccabe8ece50f617e5fa5946af85e71991ffaa

Observation 997a9521-839a-4566-9bda-ddc669f88f25 · outbound

This paper cites PaLM 2 Technical Report.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning PaLM 2 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:15.392288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:15.392288Z digest=sha256:3ac04551f6d167e3615b98282cf04b72eeff00205ab75ba7a4e852d666668785

Observation 54d73ddc-ad49-4d92-a802-17d5dcb5f9c6 · outbound

This paper cites Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:15.587037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:15.587037Z digest=sha256:ab87bd348cf08692b8b732a25ea24506ac52176b8b468f7347779023af90720f

Observation d220843a-790f-4b94-93b2-9ab424255fec · outbound

This paper cites Computational Bottlenecks of Training Small-scale Large Language Models.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Computational Bottlenecks of Training Small-scale Large Language Models

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T05:36:22.608519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T05:36:15.746964Z digest=sha256:96f84d4c3d13dfc8f742d5b7999c7fb5663e71afcdfe890b18e634d4f76b8a71

Observation fdf5abf0-e0bd-4f9a-bdae-f098965e3f2c · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:15.887455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:15.887455Z digest=sha256:c4596cfa16b788e37b2e4a93f726605db7ec456a0ffa7f00242fcf88daa3a0f9

Observation b742f291-5502-489f-b166-d41b08f40875 · outbound

This paper cites Qwen2.5-VL Technical Report.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Qwen2.5-VL Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:16.008492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:16.008492Z digest=sha256:16bacee291f53bd954d467bc9330c9bb715dffad9c406209fdafa1e94dbf6f0a

Observation a759d4f8-f3aa-4161-8f46-b64515a7f480 · outbound

This paper cites Language Models are Few-Shot Learners.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Language Models are Few-Shot Learners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:16.131300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:16.131300Z digest=sha256:a1768173558598517c4c1523a6f10bc3a42baddfa3956e32831f0766c95d4b89

Observation 2f5c417c-18a9-4dc0-9b18-847fb8334728 · outbound

This paper cites Matryoshka Multimodal Models.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Matryoshka Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:16.252253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:16.252253Z digest=sha256:4af161316efbe209e319018f22ce350d1fceba0f2abec3b96354e02b739e0d32

Observation ee21ce58-989d-4df0-9fc8-e76010ad85e2 · outbound

This paper cites an unresolved cited work.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:36:23.203133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T05:36:16.396835Z digest=sha256:2f3b97398de53fb5e95c8c8015983749d206973814f978a6f0762dcac38acdf1

Observation bf91c920-1571-41c4-88ac-1e15fe993083 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:16.550704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:16.550704Z digest=sha256:550ed69d3d0654d359d86694badd0d04c2db6dd54e229d78e904c1000f18dcd1

Observation 5094cf2c-ae6e-4a02-b93b-d71c7eb784b2 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:16.705692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:16.705692Z digest=sha256:a1d870ef17bcc654af7476c7d3686c33ef2a7d6911ba39b803e1eadbb92bb250

Observation 9f2a78f9-20fd-4b6f-a7c4-36f3ba711a7f · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:16.823297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:16.823297Z digest=sha256:7080e8c7e12653cae4bf388ebc26fc6c1805e02c93c98a80f47a79e8a99c4c65

Observation fc3ad882-5ad8-4422-9638-ed4c17654ccc · outbound

This paper cites an unresolved cited work.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:16.916546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:16.916546Z digest=sha256:86796e5f693d5ef853be1ed019e44c270735bfb2b7193c575bd0afe77ad9a178

Observation beb49cf1-6177-438f-96b7-487fb8f11626 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:17.046484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:17.046484Z digest=sha256:255c393092a2608de652b210d8b20315484ed0abd7ae3895cb4b007b4b3ca6ea

Observation d4b12cfc-f810-4ba5-b98f-1a3ff685a3d2 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:17.214634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:17.214634Z digest=sha256:8b296edb8e3b54d6581bff7ff20c4f2be9f4caf218f1a2134dcb3125974c855e

Observation ff91f547-a92f-4e11-9933-3f2843efc94a · outbound

This paper cites PaddleOCR 3.0 Technical Report.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning PaddleOCR 3.0 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:17.378776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:17.378776Z digest=sha256:19023475d8114b6df517588e9e03f16f3b652929abe387f79633f9b9fe99327b

Observation cec6b9c8-65b2-49c3-a151-68e7aaf1cc5b · outbound

This paper cites an unresolved cited work.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:36:23.149836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T05:36:17.575878Z digest=sha256:4e87392f7dcc794f0d9cb27b3358340d290c81d8082792df0be9e74eeaf67fe6

Observation f842255c-0dea-46c9-b7ec-e60f46cf450c · outbound

This paper cites an unresolved cited work.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:36:23.121631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T05:36:17.725893Z digest=sha256:0849b2a72fe112c89eceb03006499c9fe8073980a24cea9955ae0f6a34a70ac0

Observation 99691d76-b85b-4bb9-8efa-7e0ccdc350e8 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:17.899494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:17.899494Z digest=sha256:02c213aa169fbc8063ab924ba147da86e9beea208389574e143ef1ef59db4450

Observation 931622a0-1b76-40a3-87c8-948eb6e1f4f6 · outbound

This paper cites Data Filtering Networks.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Data Filtering Networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:17.999665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:17.999665Z digest=sha256:367bcd197294a09ee04d051732502469b4990ae6d4cab8595d51c7b87175af0b

Observation 04379f03-058a-4e4f-a9fe-3456385065b4 · outbound

This paper cites Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:18.071945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:18.071945Z digest=sha256:733bee99ef007fe4e0c08bc8eafd76681fc0035745acf37695e71dca5f3cb84f

Observation f02266ad-6a60-4fe1-b239-5f8f1b6c1394 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:18.156872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:18.156872Z digest=sha256:e4d893b7d0052e5c870c2f5c3206b750c6842a120a80feb106e3fd97e13af40a

Observation 168dbd3f-8959-4949-b2c4-8fe6331c5555 · outbound

This paper cites H.; Kamath, A.; Peng, N.; and Chang, K.-W.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning H.; Kamath, A.; Peng, N.; and Chang, K.-W

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:36:23.096804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T05:36:18.237291Z digest=sha256:0aa4a7b2cbd4d080bd585ebb6a4a0e4bf9c3fae8d370c58f933f7eb14f862106

Observation e56ac321-f401-4885-848f-3fc84bc9692f · outbound

This paper cites Interactive Speculative Planning: Enhance Agent Efficiency through Co-design of System and User Interface.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Interactive Speculative Planning: Enhance Agent Efficiency through Co-design of System and User Interface

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:18.309017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:18.309017Z digest=sha256:b2b2868cb21fffa120c938cbf8f40786fbab72393a2768a89d6c5c6b42b0f9d0

Observation fe8d2891-1f27-40ee-8a9b-e9f96cbd28d8 · outbound

This paper cites Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:18.365077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:18.365077Z digest=sha256:d1aadf12c33ec729e28f103ad9d986850a29655b16ddeee2a7c7e4ebe8df02f1

Observation d827b38b-0df4-4468-88e2-93219cf436a5 · outbound

This paper cites an unresolved cited work.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:36:23.069223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T05:36:18.432523Z digest=sha256:15cb8904266bb0c4818512a7a2e1ac07142bae2a8ebfb223ebc4ff2ec8277cf2

Observation 8e3065cc-a1d9-4c35-b585-c887a0bebd03 · outbound

This paper cites an unresolved cited work.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:36:23.021894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T05:36:18.482100Z digest=sha256:702f9c7d663bf84e2084adeb060c83a5ce92006523a23c394cf7272cd429724a

Observation 1f110178-b429-413f-a5fc-6688fa611a17 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning LLaVA-OneVision: Easy Visual Task Transfer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:18.580953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:18.580953Z digest=sha256:1158e9439b58e52e41c6126a7468931e862d30dee61aee3959e680e4b70d05be

Observation 7e9c12d6-0d9a-4c28-8714-fb48617210fe · outbound

This paper cites an unresolved cited work.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:18.666785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:18.666785Z digest=sha256:8d6a5e987d79eb2073d9bf1c7cbcbc20abbd7b76105f6f827ac3dfd7fb074069

Observation 4d74209a-9cf5-4e68-97fb-7456a785e78c · outbound

This paper cites Transformer-Lite: High-efficiency Deployment of Large Language Models on Mobile Phone GPUs.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Transformer-Lite: High-efficiency Deployment of Large Language Models on Mobile Phone GPUs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:18.779333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:18.779333Z digest=sha256:ce2018ee643f81e031c835db94b288c6dff2b7e1805604b65d7d36e2f00d6976

Observation 5abc18a0-57c1-41e4-9b75-a1c7165fae67 · outbound

This paper cites VILA: On Pre-training for Visual Language Models.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning VILA: On Pre-training for Visual Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:18.864386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:18.864386Z digest=sha256:8f5891d6a33bd36d3aae807d474f696689cb0fa8e0b8c9a899ea6fc3d32665bb

Observation af964230-1ca9-4adf-8b52-3362aa47d2c6 · outbound

This paper cites an unresolved cited work.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:18.949544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:18.949544Z digest=sha256:d4c4673e6abf6b389407d909e0c18a29d095c75aed87e997ae6522af706b20a7

Observation b8ca1f5d-f3fc-429d-a9ff-8e258dfbf45e · outbound

This paper cites an unresolved cited work.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:36:22.955843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T05:36:19.055602Z digest=sha256:0212296c858682644ba61f1af720e174f9abf6a41afb55ac8f374d3c27300e91

Observation 6420b49e-5ddd-4161-a852-8fe376e436b7 · outbound

This paper cites an unresolved cited work.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:36:22.926566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T05:36:19.221480Z digest=sha256:f12c5c2462c38bddba9fbcde69b1d8b742f3a66bdf7f80b4e30bf7c2dc153fbf

Observation cdf3d3c9-7768-4d90-a61e-bf7e1904a373 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:19.341740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:19.341740Z digest=sha256:fbae4763ecee6c0443849fdef2440801806ac06b858bab84168df9b87068708c

Observation 16851500-9e0d-433d-a0a2-3a6a0604864e · outbound

This paper cites an unresolved cited work.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:19.420247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:19.420247Z digest=sha256:4d22c5473f587dd9dca63d292c2c8aef06dedb0bce43a45097386afb93a66b4c

Observation dcae26e5-9d74-4f1b-b725-9d1175c24896 · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning NVILA: Efficient Frontier Visual Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:19.512240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:19.512240Z digest=sha256:03556b0ee5c9453364c70bbba942e7c6431282979578b9ba7bd34ecc3da3d868

Observation f482c604-dbed-4c4b-a88a-e2bfae8f60fa · outbound

This paper cites Decoupled Weight Decay Regularization.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Decoupled Weight Decay Regularization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:19.605928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:19.605928Z digest=sha256:9d7b91e557f1c55f8f50df3969d276bedcbd6253439921c7801b7dd3fc7d6bc8

Observation 76420909-a476-4ec6-8430-31d05aa5cc23 · outbound

This paper cites BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:19.716602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:19.716602Z digest=sha256:2a7268771686085098882cb45b918c0c2c3ed457c0e042816c4bf8bca5c17a8a

Observation cf96d5d1-f44a-4e86-b065-e6a114834a41 · outbound

This paper cites Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:19.867128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:19.867128Z digest=sha256:25db3f45491a735955fb9820c27d2f45c2fc5fa4cc3d06d4dcbc342f17aab1ea

Observation 5a8c0152-c4cd-4521-b86a-ea08825139dc · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning SmolVLM: Redefining small and efficient multimodal models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:20.025302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:20.025302Z digest=sha256:a0994440b88f182541aa3989a2fac0d4583799d42ed58b8e25ebf1fd024eb6a1

Observation 340c3a01-ac8a-4fda-9dd5-eb288decbba2 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:20.191131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:20.191131Z digest=sha256:0241c49e146acec61c9084d6abab7da3c90018ffb41d7771bcb7c1c13a327657

Observation 2e9096e5-a688-4227-b0e2-f1970f4eeb9d · outbound

This paper cites H.; Cao, Q.; Horton, M.; Jin, Y.; Sun, C.; Mirzadeh, S.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning H.; Cao, Q.; Horton, M.; Jin, Y.; Sun, C.; Mirzadeh, S

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:36:22.877808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T05:36:20.282045Z digest=sha256:30b2e109e000b8be0635673934f2920d8c86d8aad976aea5eed566a8ad346bc9

Observation 03bac653-3936-4554-bc9b-68a7324d64ef · outbound

This paper cites GPT-4 Technical Report.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning GPT-4 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:20.402653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:20.402653Z digest=sha256:edc711ecdc189f686a25232474f95b5d5c2e5f61df653ec95eea790299b286f5

Observation 6e791c27-2820-460d-9d6f-b51d36decc48 · outbound

This paper cites Mobile Edge Intelligence for Large Language Models: A Contemporary Survey.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Mobile Edge Intelligence for Large Language Models: A Contemporary Survey

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:20.568916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:20.568916Z digest=sha256:44bf15b617e00e2b2278570f1e293bf031f93b5523cd3efe272caf3f3ff50ca4

Observation fb4932a0-a3c4-4001-8fc4-5405e72c1adf · outbound

This paper cites W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:20.695127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:20.695127Z digest=sha256:e89a40eda0f08f06c53d4a909be485d5e95926a0cf2cee8925d33d73e7e614b4

Observation 9f396a20-9baf-49e2-9c7e-c2c8924d86c4 · outbound

This paper cites J.; and Yan, Y.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning J.; and Yan, Y

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:20.806605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:20.806605Z digest=sha256:bcb391f80d5821a052387fc5ce5bc41931592a1ffe9f1de63120c26d5a0dcb2a

Observation bcf31caa-61fe-4cf8-8152-0b803eaec2a8 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:20.951429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:20.951429Z digest=sha256:72a068cb0d28a539d48a6fff309c1b750f0b0a1224938efcf07b8c1f5f2ac46a

Observation 204367c9-8c6e-4356-8b10-1497e7e306f2 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:21.049168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:21.049168Z digest=sha256:af9de6da27a9010954ea172584e94241c8f4d2d7e5ee1d586e2cba22e8c92f9d

Observation 09744846-b49f-4163-ac75-237c81675511 · outbound

This paper cites C.; Yang, J.; Yang, S.; Iyer, A.; Pan, X.; Wang, A.; Fergus, R.; LeCun, Y.; and Xie, S.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning C.; Yang, J.; Yang, S.; Iyer, A.; Pan, X.; Wang, A.; Fergus, R.; LeCun, Y.; and Xie, S

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:36:22.817175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T05:36:21.078411Z digest=sha256:2179941002dd525a1c7fddbc8878b50dee09f7b27b3fff39d25d0f7c7db86f6d

Observation 035bffb9-8d16-4026-bd1b-4c5e03287061 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:21.086476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:21.086476Z digest=sha256:74c088e96e1f07c5bfdcdadfb6881eb6ef8fd00e63f948c05f3afb526954d438

Observation af17a881-ca72-412c-8b27-d32eca853e68 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning LLaMA: Open and Efficient Foundation Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:21.095817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:21.095817Z digest=sha256:8ed34388197e0a2c299d3970a7ed856515ae028acc41e96ebf9e13d8e68f7fa2

Observation 7affba21-a93b-40e3-a295-dde3a9805345 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:21.108741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:21.108741Z digest=sha256:5e53b564c2f9435ec573a47b3996958217aa725376c01b37f29cfa11f6ef1aba

Observation bbc1fc66-c4da-4a1a-95e0-cf6b897993fc · outbound

This paper cites H.; Wu, Y.; Le, Q.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning H.; Wu, Y.; Le, Q

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:36:22.789437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T05:36:21.117062Z digest=sha256:e4a0e429b3b414bf24cd0c75cdfed0cae6a394150ac222294f9250a8d8f11aea

Observation c1bb5954-51c3-4ae1-a02b-6b51003b0a03 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:21.122527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:21.122527Z digest=sha256:626bd820a4304b0bc2b871c8f3bb0826538078422447f6b6db9ef2cefb19461e

Observation 9701cd37-bc72-4fbe-aab4-75211c41a85a · outbound

This paper cites an unresolved cited work.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:36:22.765889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T05:36:21.127917Z digest=sha256:fa128681f4b6105fe685a6ed6e3c14d302f6f88061a5cfceed535084c8597965

Observation 9fdfb4da-b154-47ef-9913-30c37e565476 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:21.132522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:21.132522Z digest=sha256:7388706c08240757ee5c63ecafd2395dc430d0de291f816e2bbb96495a3eca97

Observation 1bec38ff-b4af-4ab2-98ed-3816e185f836 · outbound

This paper cites T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:21.140462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:21.140462Z digest=sha256:cb8af302faa1120db182eae484eeb8d99f05d1bd3b002cb7242a1dc7f5988854

Observation 726c944d-d1f8-44ed-8119-4f8b9ca0bc80 · outbound

This paper cites Emergent Abilities of Large Language Models.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Emergent Abilities of Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:21.148753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:21.148753Z digest=sha256:c3bb98258169314a79adb0dd5d2c90922dce1da80a827955dcdae5c493bb9919

Observation dbb96f30-34c4-438f-addc-dc9902738ff9 · outbound

This paper cites PowerInfer-2: Fast Large Language Model Inference on a Smartphone.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:21.155515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:21.155515Z digest=sha256:9eaa87d688896e77d5d3b583725fffc0f2854193c3c7836e48a5747225ad5b6a

Observation fbdfc907-9b57-400a-91e1-9337da8c6fee · outbound

This paper cites Qwen3 Technical Report.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Qwen3 Technical Report

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:21.163843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:21.163843Z digest=sha256:a21c32c387983a5befdd9da1ef634ea1f9b3603ba1c60eac91b53dcff4b9e318

Observation 6e13913e-690f-415a-808b-5dd2ad160b16 · outbound

This paper cites Qwen2.5 Technical Report.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Qwen2.5 Technical Report

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:21.183605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:21.183605Z digest=sha256:85dff1bd60cee65c705863d1da6ac12a3fcfb21a2e8f4ee830b4b0ddadd92ce9

Observation 1b1d00b8-d9ea-4111-95b0-e46810108f96 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:21.199932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:21.199932Z digest=sha256:ecc4ca0df3fb3ef368b002da57bb63504d60a0efea65f9e2856ce1520671467c

Observation 47ecbe81-0a8c-43db-871f-ab1f316f792c · outbound

This paper cites an unresolved cited work.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:36:22.740875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T05:36:21.212688Z digest=sha256:f81903cfb8f619fe7da14da08da3e8045c4065ae4970cdb571d4b694027b5206

Observation aa0571e1-fee6-4f6e-abeb-7fea0f01e71d · outbound

This paper cites MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:21.221692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:21.221692Z digest=sha256:cb9cf8ecec9fe0488b029979576eba4004f459a85e780d324b225c5cf134dd22

Observation 3880fc57-0d21-4329-ad7d-e5250514dd13 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:21.228925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:21.228925Z digest=sha256:6c09b56fc6a095aca935a5c19d8c382feeec333cc2ec2980ce90f95b9650544b

Observation 4c98efeb-8b81-40eb-80e0-7bc001648977 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:21.235533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:21.235533Z digest=sha256:0321f84014d20bcb052396244adef1b029ef72181fdcd1565e95aa95e9edb972

Pith citing papers

No inbound Pith citation observations are available.