Pith. sign in

Paper Citation Record · LEDGER

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models

As of 20 August 2026, this Paper Citation Record lists 100 of 137 outbound references and 2 inbound Pith citation observations for arXiv:2507.12566.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12566 v1

Coverage vector

measured 100 of 137 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:50:05.568237Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T05:12:37.339084Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T05:17:18.750827Z

Reference resolution

100 of 137 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved98
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 235b395d-36c4-4a87-86a8-8c3e66eb72d8 · outbound

This paper cites GPT-4 Technical Report.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:55.260899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:55.260899Z digest=sha256:ec28d6596e5214ac44e4e5e49722cf98cf3148956ac00860800792df396eca7d

Observation d9774078-184d-4460-83d9-23b4956496b9 · outbound

This paper cites Qwen Technical Report.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:55.431651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:55.431651Z digest=sha256:a3e6c02b423719b104d43fe6a29a3354a39ab03d9281b30f40a7d9b616936d46

Observation b0b86d26-1227-476b-8a16-ea8ab9088e16 · outbound

This paper cites InternLM2 Technical Report.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models InternLM2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:55.597995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:55.597995Z digest=sha256:4bdd4ea018d341161975c017690c2f7635965318d60b1fb9228e79f4d1fb66a5

Observation c183d979-6b75-4a40-bd9a-0dafa73d6666 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Learning transferable visual models from natural language supervision,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:55.726783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:55.726783Z digest=sha256:7ea614e42cf04899a1aed24c57b7e1257cf32f0b6262152a69f121c77de75398

Observation e0901f71-40b2-4641-908e-c290254308de · outbound

This paper cites Visual instruction tuning,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Visual instruction tuning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:55.834849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:55.834849Z digest=sha256:1aeae274ff3547c1d78157b50e7a46a092a737faa598f3096ec56852c35a05c7

Observation 3ae42cdb-16e2-4436-966e-725c985a3d53 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:55.905181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:55.905181Z digest=sha256:c3456cd547c3f54b2c6d8a6b290e16d8d1cb5651fe203310723d4a596707b022

Observation ac57a490-84dc-4135-9514-b86f724a33e8 · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.005313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.005313Z digest=sha256:e49c9208aeeceb7ecb782fe3646fc227657a9b91ecdf864352b018f4262a492d

Observation 9c7033a1-d62d-4e9a-9189-9f5e541158b9 · outbound

This paper cites Introducing our multimodal models,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Introducing our multimodal models,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.084702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.084702Z digest=sha256:95f012623cd0a43ae5b42621a460613dbf3ad5467f8b14946f1370b1ab7b8a6e

Observation 31e4b70e-3842-4761-9ffd-58b30ea63fba · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Unveiling Encoder-Free Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.218152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.218152Z digest=sha256:daf4fce1c828890ed30f7a67ab66241288c870f5314f367d9c1a97c69bf3cad7

Observation 1c525bd1-5ec4-4b4f-9005-06ab4929f066 · outbound

This paper cites SOLO: A Single Transformer for Scalable Vision-Language Modeling.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:50:10.351624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:49:56.373596Z digest=sha256:7c01c536272acd182a64793b1ac9afc72f8f58f585b78cdeddbb411f140ed5a8

Observation 23b62d2c-270b-4f8c-8d6d-96db9f90abcd · outbound

This paper cites Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.460558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.460558Z digest=sha256:472467ee012dd85d5d1a7ce817397fbb383007a4404139d03f66f06e259d8f52

Observation d4dfeccd-4672-44a9-8db0-229275532d35 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.572596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.572596Z digest=sha256:d2515def22484aa391b0bb1e56881fad0d52f873146a358046efcd9d5e3fc03e

Observation df089f74-f30b-44f3-a10f-2de6df635cdf · outbound

This paper cites Investigating the Catastrophic Forgetting in Multimodal Large Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Investigating the Catastrophic Forgetting in Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.665670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.665670Z digest=sha256:fe1c7c08eaa7c160a05e636734ca4f45f99fed61d6c234d8d61d7822b74171e2

Observation 2cffd4b2-6780-4103-aa7f-85eca601f8e6 · outbound

This paper cites Delta Tuning: A Comprehensive Study of Parameter Efficient Methods for Pre-trained Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Delta Tuning: A Comprehensive Study of Parameter Efficient Methods for Pre-trained Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.785510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.785510Z digest=sha256:4cc4d6218506f1ce6275c0cc1e03bc9b0c5362fb16fb4b2c497b26a443d13bc9

Observation 73aee5ed-f683-4926-a808-79f3ecb4aa9b · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.896912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.896912Z digest=sha256:18f39ac53d1091ad8cde1803b50dee73cd6a39f3e87f19357c6dd19e1839875f

Observation 93868ee2-dba0-452c-8fd9-2d19ae9d9976 · outbound

This paper cites Lima: Less is more for alignment,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Lima: Less is more for alignment,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.999030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.999030Z digest=sha256:ad68a045fe463663583e8f4ceecefe8fc386d01b724b01cb6daf9fc963198eac

Observation 27ad4242-fc85-46d9-a397-83c596813eed · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Emu3: Next-Token Prediction is All You Need

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.069428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.069428Z digest=sha256:27e86f5d8518c0f54bcafaca809dc57bb0c0e00ce2ec88866cfab8d362b69a97

Observation 901c8ced-76a7-47a9-93e0-abdc860ade0e · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.153253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.153253Z digest=sha256:6fc221f282203e08c4f27a8136b197485a68afc91f082a876c01e316851f1c5b

Observation 4c338e29-89b5-40fe-980e-4969d4907720 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.248846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.248846Z digest=sha256:72a00acb8c100e844c6d478525cd3eadb947fa0316f2f0d4cde0955fac3cb754

Observation 220026a1-5a60-44c4-8099-462afadca072 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Improved Baselines with Visual Instruction Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.353167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.353167Z digest=sha256:8f4cfd2d9bc577bb8fce44fe40fa5e67e46fa46d27e3259f5802dec567212141

Observation 0f625dc8-cdfc-42c1-9448-9668a4a65cfd · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.463091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.463091Z digest=sha256:f56142aaed67b8ceb801164e128ea4347724e2e1bf2a3ae392de243e03f46439

Observation cc75399e-c724-46e9-8338-dbd706c97401 · outbound

This paper cites Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:50:10.051759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:49:57.683627Z digest=sha256:774b802408ad287a2e7817030c2517dbf821315f2dfe3e86852260b84f264037

Observation 718b4b77-42d6-4f16-baa8-b702afb4b3c8 · outbound

This paper cites SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.751825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.751825Z digest=sha256:47cf980c3c180cf5f819326d8d1525f2b7c697b08ada8ef72ae1f715f879687b

Observation f5a625a0-61af-40c6-b69e-ed10033673fa · outbound

This paper cites BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.862972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.862972Z digest=sha256:2b02f928bcedfa0adf8b7194900260d67d63817dbefbe3b57997ae8fef111e7e

Observation b6ab71c8-d911-40a6-8d89-4ef5fba23175 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Instructblip: Towards general-purpose vision-language models with instruction tuning,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.985266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.985266Z digest=sha256:cb1379f0fb7a08a14944b3ea8183f92bd933b2f51083da46a483ae286935a3e8

Observation dab453f1-892f-41b9-9259-186e59988c55 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.074139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.074139Z digest=sha256:e44549b4554c1397bf42ed2f94c110def2d5b199d0e7d37564b806eed8498259

Observation f68fdac2-7f06-4dfa-a0b8-6f5c56f104dd · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.168347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.168347Z digest=sha256:86c0f7ae035f6a161fa007b196b8a3dfb2c35eecd00b8b7aa7ce62f09a3c0661

Observation df391df8-6293-45c1-b131-2b8d35f1e3a7 · outbound

This paper cites Qwen2.5-VL Technical Report.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Qwen2.5-VL Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.311277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.311277Z digest=sha256:4a2374754460259bd8139f5be420b9da50c33e67a01cf2327a53daefc614c545

Observation 202f1580-fc12-439a-8af0-3029d2195b70 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.390094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.390094Z digest=sha256:cfa143783019bfc20cc0ce3a81fb390a500bc2d64c86331170b74188b987afb6

Observation 871361c5-bde1-4422-8f1b-ebb842a6a0b1 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.597304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.597304Z digest=sha256:063e97bd91f4b2f0f8329ce75027ebaca63b7fd1deb2bf97aeb28471127fa59e

Observation 8eba6607-61ef-454b-a467-b2be0c2d9381 · outbound

This paper cites Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% performance,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% performance,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.697634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.697634Z digest=sha256:c21eaf15e8e2010f10fdbca286c1123d19e05dc31006ab2a782d7bb1f7a629f1

Observation 63a1e2a6-3313-4851-9e28-bdd64ec2e4a3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.800042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.800042Z digest=sha256:bf2204f484c32d1b3b3e3eb7927c18358470292f9bd617e8404cc80db120b4c8

Observation 730f0235-ff65-47e0-b3e4-822e238e42da · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.881346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.881346Z digest=sha256:e158cb444898f9adbeb2434084f118d0981d2b25d7f4d967779fff0bdc20d61a

Observation 0244e7a9-46d2-49e3-86a4-682e461f0b46 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:58.977983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:58.977983Z digest=sha256:e709e9524c2c99c85a404600327c588f3d196082322168867cfca75e63ef3be1

Observation 7e7d78cb-6aef-4d4a-8d55-1fea0f4b1312 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.055806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.055806Z digest=sha256:f05da98b256802da712650863722708d25472763eeae60fb04581d94d33f1d77

Observation bed0b7ad-391e-4c43-bca4-ac6c35dda569 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.135515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.135515Z digest=sha256:929cad8060a0f691460b79beac8442ff260f4f2ed06a8a9c09ab07cdee96518a

Observation 2dde65db-d26c-4439-b13c-a083fc6b3b8d · outbound

This paper cites Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.224871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.224871Z digest=sha256:9d6aaa313754eb02ca414ebed4a4b5d1d13949e302c17ad0028f1fd335eac7a3

Observation 1fe92c26-84cd-40aa-97f4-5e7386379e26 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.336128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.336128Z digest=sha256:64cf052358c3a718f4599c12b40a7f3e4f3a2f8925959e9db6e965750052102d

Observation 3e3f5af5-d0c2-4d93-89e5-8a1fe41310a5 · outbound

This paper cites Scaling Vision-Language Models with Sparse Mixture of Experts.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.424967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.424967Z digest=sha256:251731f7a5d8d69ddc43d85a3d195de75d3a71e5b85964ae93043a0690a1d377

Observation f5284dae-5a57-4d0b-9d5d-eb743de78905 · outbound

This paper cites Twenty years of mixture of experts,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Twenty years of mixture of experts,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.548683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.548683Z digest=sha256:e340701281e737cd7d2a1b96c3a3c8e15bba5a170918a5ba42f39dc8aeb61e5d

Observation 0a5d6262-037d-431e-83b5-cee3b08144d0 · outbound

This paper cites MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.656307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.656307Z digest=sha256:af44e2818aa2770767ca38c9f337cbe6e9fdc116f65bbdee278544cb93a0f9bc

Observation fc70db5e-76e7-4e1b-af06-22a874c99c41 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.760100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.760100Z digest=sha256:57815126daa8cb770dde83914964374b536d6e13f0d36cb7391f152919a72a1c

Observation 5512c647-879c-48fd-b611-5cd6e6d1ced4 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.827141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.827141Z digest=sha256:a6aa66d95b2ccf883ba775de841230b84fe3e8215cd609312fc22908b4bcd311

Observation 855a26ff-2ca3-4c63-a529-07a4d654281f · outbound

This paper cites Attention is all you need,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Attention is all you need,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.936912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.936912Z digest=sha256:a21198fdd4fe88098183f23291ca2e27854e50934ed88118117274df5b2e6105

Observation eb7f29c9-3b3a-430d-98aa-32245d0eae88 · outbound

This paper cites Root mean square layer normalization,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Root mean square layer normalization,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.080685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.080685Z digest=sha256:c1d6a13a0cd8b54c57fe8f485fd9966cd2cde0b921c4ecea0752a054f4761b64

Observation 3cec4baf-2c06-4943-b6c5-51583007169e · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Laion-5b: An open large-scale dataset for training next generation image-text models,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.195147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.195147Z digest=sha256:dd1615d3b0b1ee4cb233fb05307e52bb100367ba194a51b78311d3ce0c0d0cd0

Observation 7cdd1300-3cee-46af-9e71-46005c45e6b0 · outbound

This paper cites Coyo-700m: Image-text pair dataset,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Coyo-700m: Image-text pair dataset,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.302668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.302668Z digest=sha256:cd935e8ddb701e13c14d4a1f94f83146fc5cdbebd0ce2b5ea12b340c77c6a7ae

Observation 55bf65fe-a0b4-424d-a3f3-aa0da6d66bbf · outbound

This paper cites Segment Anything.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Segment Anything

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.388194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.388194Z digest=sha256:d5d5e4ab45b01f18d532e15a52d538b7761fad13ced72bbc1599a1ffc9aae566

Observation b8d07cf3-fba4-4f8d-9360-55cb2ec9d4a8 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.517198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.517198Z digest=sha256:4caeb36a4b0bfbc2ad1ac261ab4bb757a1fd1508d39b7945daa300e2aa052c5e

Observation ffe36186-4645-44ad-8511-f3eb6f815ab9 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.619189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.619189Z digest=sha256:5a94d94482149b0e3eeb2320a88169353e5b1ea6bc4f0b4c4200352ad4b5eb30

Observation 712be2aa-2719-4a1b-b345-2678d7220474 · outbound

This paper cites Textcaps: A dataset for image captioning with reading comprehension,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Textcaps: A dataset for image captioning with reading comprehension,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.734000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.734000Z digest=sha256:9cde248df8196ba0904e281445ff9fa257d5834b95dd7c0b29be98338176ed11

Observation 86908a5d-546f-45a5-9e28-ddc473bd2a39 · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Objects365: A large-scale, high-quality dataset for object detection,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.805492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.805492Z digest=sha256:678d22232089a92604165dfdaa2e2dd4974f20e949ebffd26311dca8b477b89e

Observation 776abb4f-f5de-465b-9e53-745ae1e26aa0 · outbound

This paper cites The all-seeing project: Towards panoptic visual recognition and understanding of the open world,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models The all-seeing project: Towards panoptic visual recognition and understanding of the open world,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.874286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.874286Z digest=sha256:0afadd96c03a1e896f5fa107550820bd19319f8ed524f02c432152f8c4e4647e

Observation ba30b0f4-3e61-476c-b798-e2e44d305b27 · outbound

This paper cites Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:00.947847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:00.947847Z digest=sha256:9d29dfaebe52d866e479e6d36b44d8f050b62d4e1c87f70e61c74cf9f8a1e70c

Observation a0b97b9e-f3a1-473d-bb14-7c57e43644a3 · outbound

This paper cites Laion coco: 600m synthetic captions from laion2b-en.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Laion coco: 600m synthetic captions from laion2b-en

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.031594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.031594Z digest=sha256:6b23b11fea8a0191a4f19fb841bce35a521595ede92acca2594770318a381b90

Observation 4ea5e301-bf72-41f3-ab76-9789c6925236 · outbound

This paper cites MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.162563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.162563Z digest=sha256:7b38432f724c49a2e547249a4bf670c4e67b46b9bedbb3b51df9c820a6b1a2e4

Observation d2e3ed08-d8f3-43f3-a7ca-f6d5c4d79193 · outbound

This paper cites Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.240559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.240559Z digest=sha256:9b8c7a12dfddae06d260311c03fa84c6470f008207a57d2228d12c408c7f5f66

Observation 4912a34c-18da-4fa0-97a8-d82f53569c69 · outbound

This paper cites Scene text visual question answering,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Scene text visual question answering,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.315416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.315416Z digest=sha256:71d4294768898de299bfff24a427156bc700f5afa518da96cc8b04b99647fac0

Observation c7493250-004c-4ba9-b345-df73855b3bcb · outbound

This paper cites Icdar2017 competition on reading chinese text in the wild (rctw-17),.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Icdar2017 competition on reading chinese text in the wild (rctw-17),

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.390445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.390445Z digest=sha256:734f1d556294a66933d8414fb42af72ed1a75b60b74217dd4924640372c4836c

Observation 66f3aa88-0aee-4f2e-bd2b-d4dffbf4ddec · outbound

This paper cites Icdar 2019 robust reading challenge on reading chinese text on signboard,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Icdar 2019 robust reading challenge on reading chinese text on signboard,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.495683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.495683Z digest=sha256:4e2d628c49924c3ecdd71a1761449544589987e8883157044867c9b95539a732

Observation b330ee7f-0072-4b50-895e-b9ee0691b6ea · outbound

This paper cites Icdar2019 robust reading challenge on arbitrary- shaped text-rrc-art,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Icdar2019 robust reading challenge on arbitrary- shaped text-rrc-art,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.578838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.578838Z digest=sha256:aa8888b7dcb0effb991d9890f0482a745d3fbdaed9c30c6cd22c874b4fdd1f9a

Observation 3756c26d-5fd7-4b25-929a-db601c9dbe4f · outbound

This paper cites Ocr-free document understanding transformer,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Ocr-free document understanding transformer,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.712181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.712181Z digest=sha256:746bd3881fa9552c3ff9400c6c03f6cf111eccb6fd3a541d4a27fb500a5ff0d4

Observation 0ec19cce-6c0a-4f41-be1e-fd52bd161f61 · outbound

This paper cites COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.817911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.817911Z digest=sha256:e041d0384ae928fa7ac1758870997a4646a0e4c9112c3ef0711189cadd5b9355

Observation 7c8d6861-01c7-4da6-9a3c-1890fc2068bc · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Chartqa: A benchmark for question answering about charts with visual and logical reasoning,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.924520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.924520Z digest=sha256:f13d2ef89dfc5bc707a8700c3fa501cfae53fc7c30e9a798152eb0e606ae8238

Observation 61e6366c-974c-407f-8fcb-4b8263ade4fe · outbound

This paper cites A large chinese text dataset in the wild,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models A large chinese text dataset in the wild,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.002148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.002148Z digest=sha256:6282d35ac0e213deda1fd6d1393944057f15154e4c1dc25738a8955a5866e678

Observation d043d554-ecdf-429c-88da-cf42cc6088d7 · outbound

This paper cites Simple and effective multi-paragraph reading comprehension,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Simple and effective multi-paragraph reading comprehension,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.099560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.099560Z digest=sha256:a8b6cdd7b3a0f8ee8d89be295165cd8d892fd892779da1ef7232da02bffe4684

Observation 0b0c7140-c968-45e5-a601-f6060c883f63 · outbound

This paper cites Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.203648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.203648Z digest=sha256:a96d9c63c3833e219747972222facc968045aa76b7a71524bfed6fe406cf9882

Observation 07cde5a0-af22-42d9-a8cd-9c43ab5f7c76 · outbound

This paper cites Plotqa: Reasoning over scientific plots,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Plotqa: Reasoning over scientific plots,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.320729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.320729Z digest=sha256:476a610db6281a2ce4f0db7cf6e42280f62b8fff7cc39d798ad0a25c243519fd

Observation c83be487-e7ae-427b-bf7a-9f520ec9f738 · outbound

This paper cites Infographicvqa,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Infographicvqa,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.423626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.423626Z digest=sha256:d1720e7d5e26ee360ede70ae5c2cc9d8b1263ddd595337a9056595d9a0975997

Observation 66675f87-dc47-41fa-b6ab-4ab661215fed · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Making the V in VQA matter: Elevating the role of image understanding in visual question answering,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.536884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.536884Z digest=sha256:e395f6b76dfdf5363f0102948cc1e509e8458ff6a910d741244276233db5e797

Observation 546425af-4490-472f-8619-4c156080f53f · outbound

This paper cites GQA: A new dataset for real-world visual reasoning and compositional question answering,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models GQA: A new dataset for real-world visual reasoning and compositional question answering,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.641238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.641238Z digest=sha256:91dfa46cc376d7bbe70ca4e8373186e175c871bfe6821f3406122e41e876ae99

Observation f900c643-6726-47c9-aa9f-549eb9f80e8f · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Ok-vqa: A visual question answering benchmark requiring external knowledge,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.746998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.746998Z digest=sha256:3569b3ae0ea520fc0ecbb6e77b8b0ce164519d493237fc71786c3d0efd41a226

Observation 2bc39bd7-4222-4de2-9972-5cd6c418f854 · outbound

This paper cites Visual spatial reasoning,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Visual spatial reasoning,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.852857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.852857Z digest=sha256:77e190792af9bc2f9341aaa3f652736a1f8b87a3be01b0be87c416a5112f8772

Observation d9264504-9076-436e-8b1e-bc1716e767a3 · outbound

This paper cites Visual dialog,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Visual dialog,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:02.963424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:02.963424Z digest=sha256:c509075780bf263c58751875406788e92fec2b1b58283487f85001b147c9710d

Observation eeba70c5-126f-4a87-bc5f-c70595bee6fd · outbound

This paper cites A diagram is worth a dozen images,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models A diagram is worth a dozen images,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.027579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.027579Z digest=sha256:cc7b34602c0e078974eeecdee9ba80c8fd68e684649e526646748cc96d044977

Observation c19cef38-137e-4204-beaa-d73fa8976722 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.153797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.153797Z digest=sha256:4fefab776321a987a6cbee11799dd069fbf96cc87d4c821c6cd3461acbce8743

Observation 4b270660-b3fa-4df9-b4fd-4cbb7a7e6ffd · outbound

This paper cites Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.236776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.236776Z digest=sha256:3bfdff971e65ef5cdcf75ba89d7cf2c9946c41dec82f74eb3e627f26a28d2b1f

Observation 28539a3e-357e-43f1-a2e4-1a229e7f10a5 · outbound

This paper cites Dvqa: Understanding data visualizations via question answering,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Dvqa: Understanding data visualizations via question answering,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.384359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.384359Z digest=sha256:9978f90a6387e39a1936d89960ae04198a41f940ead6b5b6a251b2a0ae10971b

Observation 3f8cb091-aeeb-40d2-97cc-477d8c8b0a9f · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.529613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.529613Z digest=sha256:7c4d8b0ad76f3d5e7d989dd281076646306f1dc0a6ec76a229ab8deb57d876a4

Observation 7382f2c5-4d5a-431e-bb17-b95211031b07 · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text encoding,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models An augmented benchmark dataset for geometric question answering through dual parallel text encoding,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.654147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.654147Z digest=sha256:9cd585acf4cec7e6f8b44c8d2744262cd32fde3f939c24e8f6a89a1c79175e1b

Observation f8b413f7-7d5a-45e3-8655-2a7e0c828785 · outbound

This paper cites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.783517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.783517Z digest=sha256:73471252d4a150c6c76bec408fcef5425dbad211baaa1778a031c50260698223

Observation 33c867b7-c587-4172-a55b-d3a85925e403 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.926291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.926291Z digest=sha256:f51e795b433b909e7d95cea777e99159b34d5dd9b74247039c0bfbe93feea6e5

Observation 65f50570-1ed8-4434-a038-0a613d7da31b · outbound

This paper cites CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.072385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.072385Z digest=sha256:d8528dd92e0242b4b0a7ddb1e4c7c52ab9fbecf2a670cc846d8afc04a4ac7474

Observation 26aebb84-91f1-4696-a69d-57638124a7e8 · outbound

This paper cites Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.211430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.211430Z digest=sha256:004ce831f7b2a8597876468c235e3bfcc4f7a2cfc2c80c7b4fc6bdd6b029a848

Observation 3eaa7f89-402a-4e8d-bff3-c1d1320a9647 · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.330336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.330336Z digest=sha256:bdd9613ce82004f4572d68ab93e4820bc5974bb4f867b5251053d60482565348

Observation 9408ddf6-13f9-46d8-9640-f0bca390a041 · outbound

This paper cites Kvqa: Knowledge- aware visual question answering,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Kvqa: Knowledge- aware visual question answering,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.443260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.443260Z digest=sha256:145d7568abc5aa114d45dfa7000c5624335f80c820e61b9a7f445e5079ef1d10

Observation e2063723-1016-463b-a2ca-3d5ac2691921 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models A-okvqa: A benchmark for visual question answering using world knowledge,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.560616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.560616Z digest=sha256:5334ff71d6f42ab86c09fe7b58ce21e13db5b92c999158f71feddbc3d40314f3

Observation 197b7174-653d-4cb5-9496-32b7ee46a3a6 · outbound

This paper cites Viquae, a dataset for knowledge- based visual question answering about named entities,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Viquae, a dataset for knowledge- based visual question answering about named entities,

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.660448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.660448Z digest=sha256:2d7c28b12c70eb12df2c2ee6dee41b18dd3318cfb2733a71c5cdcc409a2410f2

Observation 3d17bb68-d896-410f-921f-b62d7d2c94aa · outbound

This paper cites WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.722409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.722409Z digest=sha256:9bee9c4ac64834dfb123ae2631bfe194cd0185cbb8a3f8201071a9d52ae81b5b

Observation 7cb3cda1-178f-4ead-a61f-9f8600364ab3 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Ocr-vqa: Visual question answering by reading text in images,

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.785021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.785021Z digest=sha256:d1a2c4327c84d07251476b615e0700d53b1d1adccd5fb1cdf350c8713d077d6c

Observation b50417c2-8ed8-42df-b13a-983ad4fd2128 · outbound

This paper cites Towards VQA models that can read,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Towards VQA models that can read,

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.861787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.861787Z digest=sha256:e06ae09deaa539b5e274ee517b773a4bdc01f3307aa5560174659d29e580b8de

Observation 581d54b1-e307-452a-bfe9-6335283eaa70 · outbound

This paper cites Modeling context in referring expressions,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Modeling context in referring expressions,

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.943013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.943013Z digest=sha256:04bd3ab0a256b2ebf8110765fb5e8eec4e166d0e9b2fd8d0f8ad8e0f06553d8a

Observation 8f850fd4-aafe-4cb7-9acd-c1616efcf848 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Generation and comprehension of unambiguous object descriptions,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.027353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.027353Z digest=sha256:e8a5d0d02d4094781189ee63a07b004594ce8153f80a0fb2bc6061eb245770d7

Observation ed4402f5-07e4-43c4-bc10-7c2124596bd3 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Visual genome: Connecting language and vision using crowdsourced dense image annotations,

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.105073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.105073Z digest=sha256:bcb6b3d60f3af99172cc3b3f5231e77d2bfb99aea611cd83dd442cfcab878119

Observation b970ecbc-9311-4a1b-bd3c-f83feadd74b4 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.174737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.174737Z digest=sha256:0b7dde05937f086a73679b8957534842af3e197c4265f34cf969a2226ef6e1f3

Observation 244696a9-e355-4869-9b6a-18fa1cf73075 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.268162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.268162Z digest=sha256:462f8530a26fd523f2b2665ef13055078d81c6b05491c2b285f761d7dc19d250

Observation 65664252-de60-4c8d-b58a-736bf1bdcfb6 · outbound

This paper cites Gpt-4v dataset,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Gpt-4v dataset,

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.335533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.335533Z digest=sha256:eb12805ad09738dbb91a333b66663fc199b07e5741ec8494aafb3ad140c7d7fc

Observation 6d38fdb5-0f7b-4210-b379-69b149ae4558 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.406294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.406294Z digest=sha256:1bc6f483964d782a842a2e791aa65076d1120d8e7472d8a2fb70de6779512308

Observation d86c7702-3db9-44f1-ab53-65377a19e076 · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models SVIT: Scaling up Visual Instruction Tuning

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.494156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.494156Z digest=sha256:140f6a495559abd7d4b7161e1d05f653903ce5b4d06761687edf77637d63aadf

Observation 24e6065c-6e91-4ccd-ade8-cada6d4633bf · outbound

This paper cites Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.568237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.568237Z digest=sha256:a2578872e0be9dd235ab73c69b99129ab7bfd4cdac27d99d429373553f6c0839

Pith citing papers

Observation fe67a2ae-8c0f-4042-abb3-5246f5c6fee4 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.202761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:f121e192d62d6760eeb7255e1bed6f26df9331a0dca6428e82c3933df89393dc

Observation 85ce1315-55db-4859-9f18-b28e63813223 · inbound

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture cites this paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.752838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:cece8d8e17e95cd71be7a719c46ba5c3d635831c76b59db29780b664e82faf24