Pith. sign in

Paper Citation Record · LEDGER

MBQ: Modality-Balanced Quantization for Large Vision-Language Models

As of 21 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 4 inbound Pith citation observations for arXiv:2412.19509.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.19509 v2

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:21:52.377583Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:12:56.131001Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:18:37.366607Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2ccd005e-c2e7-4511-bd86-d3ef88c200f8 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.105585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.105585Z digest=sha256:46790daf8ebaf0eb10d073679fa70b1d278faa72737b836369bf104b154d15c4

Observation e8f2e767-f845-4a23-854d-5852311f48d1 · outbound

This paper cites Vqa: Visual question answering.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Vqa: Visual question answering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.110703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.110703Z digest=sha256:52f17827acccdc57691725f29d7e847be72a5da4805fbf618d7e11b50c4f7f48

Observation af04e1b6-f081-448d-9815-fc87e0416989 · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.115070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.115070Z digest=sha256:da6df75196a4e39795e897fbe764fecc5329586c4b9cb3fd0865249fcc16b01f

Observation 78b32ce7-ef0b-4309-9ad8-d0918eff8080 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.120684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.120684Z digest=sha256:a60d04323fdb688d93494f612b1fea152769f4db5b80eb6c6b64349d23418ff3

Observation 8a62efcc-4947-44b1-9d88-508cdefd994f · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models PaliGemma: A versatile 3B VLM for transfer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.125017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.125017Z digest=sha256:a447b827bb7bab19f75df744ed7bd87b0041f022d96e39d324d24753d8a3c89f

Observation 18e799cc-3e96-4d80-bad7-3be4a4b751a9 · outbound

This paper cites Madtp: Multi- modal alignment-guided dynamic token pruning for accel- erating vision-language transformer.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Madtp: Multi- modal alignment-guided dynamic token pruning for accel- erating vision-language transformer

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.478870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.129207Z digest=sha256:c57b2bfdd98b5e5c304064afd281ee516dcb80b8fd07b3d35ff24869bb797f71

Observation 2f70b025-08aa-45a6-86e5-f71421afb842 · outbound

This paper cites Quip: 2-bit quantization of large language models with guarantees.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Quip: 2-bit quantization of large language models with guarantees

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.464849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.133190Z digest=sha256:e21037eca2bd00616e180c87267eadcf066115bd208f3bfb1e51440a3f02f2f2

Observation 4649ac03-06bf-4d4e-a620-172b4beb1571 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.137468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.137468Z digest=sha256:c95098535a3bb0fc1b77dd12603b5edad0b01b796621a1d6c8cfe8638ff75d84

Observation 6ff507a7-7fdc-46ce-9903-dc28985ad3b3 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.450320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.141610Z digest=sha256:f485a23091c51de6a1035ab39759c9fc2d76e6f85e685d06d4a305af297ecce2

Observation ab70f971-27a1-43e3-847c-74cbf10be103 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.145313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.145313Z digest=sha256:af731a85daf3c3caccc28201fd928543656c988f70e6bb2dd37dee358a0106e4

Observation 3fa41131-0df4-46ab-997e-05de94b2030e · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.435429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.149879Z digest=sha256:9965486b81d1b777331d7a06aecb4224588c827868f16b256590de37c888d920

Observation 87bd6ebe-a169-4823-a15a-41161d57b92b · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.154461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.154461Z digest=sha256:9d4b6f385deb47afe0583cbe5d2402914601b6a6b4aa3646081e4280c945f7f5

Observation 0527c9fd-9b03-416e-ba9c-4a9ca20110e1 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.159156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.159156Z digest=sha256:4209b69669b39e73cf58959f4506d62abb5a025b13472e05b54394946fffc97d

Observation a81c7215-b073-4b70-a25e-9b81be08aff8 · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.163502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.163502Z digest=sha256:6d5b32ee88031e1e3622c68e9323faba8aee8e905f19380e39e3ac5fafabbff3

Observation 0c08c9cf-1851-4048-a4b3-a95242fd1e98 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.168181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.168181Z digest=sha256:9aa0f8a5cac95d549c3f4799b1614e49f6af49a9d8098ecb97401cfc905861f1

Observation 1f9b97e6-56b7-4c7a-9ed8-26436026e99c · outbound

This paper cites Moa: Mixture of sparse attention for automatic large language model compression.arXiv preprint arXiv:2406.14909, 2024.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Moa: Mixture of sparse attention for automatic large language model compression.arXiv preprint arXiv:2406.14909, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.172820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.172820Z digest=sha256:e8d4a642ab775df6c7e90927dea083cc358ce42b985c0fd718bed571a9824f0f

Observation c2f72d89-ac3b-4c8f-8360-2ca3d5b06f43 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.177269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.177269Z digest=sha256:5c0605946246b34f11ae6ea09f6470660a2ab12ae554e403611af69b8f938053

Observation fc66d3c8-698b-4521-a5a1-ea89041de2cb · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Vizwiz grand challenge: Answering visual questions from blind people

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.182684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.182684Z digest=sha256:5c88225b7354dc05702912f48b79142dc2ed2183e4b6c3c499b59257e34e31ba

Observation 80130e5f-f41b-416d-aeaa-6b63ddd9dc19 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.187123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.187123Z digest=sha256:5b39ba293d6ca1fe5a7fac1e723be6f2da06a044def3c637badd540e2671ef8f

Observation 298ad498-c48f-42f5-9415-68f218437624 · outbound

This paper cites Llava-next: Stronger llms supercharge multimodal capa- bilities in the wild, 2024.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Llava-next: Stronger llms supercharge multimodal capa- bilities in the wild, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.411165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.191778Z digest=sha256:53a4797a1cbcdfa9c55e33c35da82e6905684d7ff8f890449527d30a03db3e0c

Observation 00ee3d80-6cce-4d23-a9b5-4a1f09edcb9d · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.195975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.195975Z digest=sha256:d9645914b22bb091079791e7694c448af43590a324c2e741678e9d2f38c9b061

Observation a71ab9fc-5a5e-4ef9-862b-b3fea50e9344 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.200549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.200549Z digest=sha256:bb7ed361c4b426e84f7bf2b7e7cd741556b36b49a14e0104f18a9b891d8b9c21

Observation 5492805a-4c72-4c41-aa82-6a7a0658dcf1 · outbound

This paper cites Llm-mq: Mixed-precision quantiza- tion for efficient llm deployment.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Llm-mq: Mixed-precision quantiza- tion for efficient llm deployment

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.386113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.204809Z digest=sha256:fbf6392ad6f98bc4fcb19f9f86288f4b4dfdd95001747f45be840471c996fba9

Observation 7e8fcf70-d007-408d-8ef1-4d849fa730ed · outbound

This paper cites Evaluating quantized large language models.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Evaluating quantized large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.370020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.209140Z digest=sha256:2126436c22fbf56aa30f954c5b6bd81765eccc377685a8902e2dbfb20f2e3c34

Observation 0e393eaf-cb02-4872-bcfd-40bf8cff4b9d · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.213328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.213328Z digest=sha256:282145d6b202e2412a9193938edf1522e83d0c933a713687512f5f075f074af9

Observation 457cc0aa-c337-4925-99da-fc2f404f3bf1 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.217708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.217708Z digest=sha256:9af35365ecab739c2f2ad39c94a087a96aa6d3e81155576b620018361cb2faea

Observation eb250ba1-e5e7-4766-9cd1-548bb4100176 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Awq: Activation-aware weight quantization for on-device llm compression and acceleration

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.354930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.221996Z digest=sha256:4b51de0d4e2c053a128a5bd7912632e4731b423e6d1eee172f1a0f045ac22b71

Observation e5e05113-4ff1-48e2-8401-565607cab953 · outbound

This paper cites Vila: On pre-training for vi- sual language models.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Vila: On pre-training for vi- sual language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.226055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.226055Z digest=sha256:909a67dd4a85127118e47f223eff0eabd402c98d5ef4036d3b96299febdfe93a

Observation 1248ba68-d2f8-48eb-9cc1-147a40f1b0af · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.229717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.229717Z digest=sha256:675f79feeea5a6a276c7de088b3b2194107a6bf932d957fe07380e33a420eaa7

Observation e31ec82e-f815-4370-bf47-b7610183a5ff · outbound

This paper cites Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.233673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.233673Z digest=sha256:c297d3c6a1bc04e26c35c22250c620865c76514f79dd705f2fb19e3cd765199a

Observation c6f2fe14-8d04-457a-83c6-f411447120fc · outbound

This paper cites Visual instruction tuning.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Visual instruction tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.328362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.237690Z digest=sha256:ac980a091b5a8048c027b29bf61dec44ef9faeaeda6ef0ab064c7348d1ee9eef

Observation f6438b1d-8b42-40fa-bbe1-d64fc72be943 · outbound

This paper cites Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.313941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.241932Z digest=sha256:1c6f0d6a4d7e083c798e1e1d7b2f90afbe2bc2c29a8e009ebaaadfc5a678dc08

Observation b7552357-7818-412b-ae44-394082910cfa · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models SpinQuant: LLM quantization with learned rotations

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.246328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.246328Z digest=sha256:7168ebe28dc2d883b51a60d5c2135e37400512e0f2023a56538827c3f7f02613

Observation d6e2d7be-32a6-4958-a20e-3651839a077c · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.299425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.251680Z digest=sha256:cd739076009440663ec7404219dc4bed883411974b9511913c40538080c47dfd

Observation fc5e0330-9cfa-46ee-9090-0804226388ec · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.255831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.255831Z digest=sha256:74dad18cf3bdfcb23feedb575049957a7c52fb42234ce005ca74f508a34a94f0

Observation 999bfc33-81c3-4292-91f9-62db3437da67 · outbound

This paper cites Skeleton-of-thought: Prompting llms for efficient parallel generation.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Skeleton-of-thought: Prompting llms for efficient parallel generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.285355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.260834Z digest=sha256:9bc96060748976484bf5617bdeab6c727bd42817f94371bedf29cdd542accabc

Observation 87dfa7f9-5307-44c3-b52d-374d49bd534e · outbound

This paper cites VL-Mamba: Exploring State Space Models for Multimodal Learning.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models VL-Mamba: Exploring State Space Models for Multimodal Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.265439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.265439Z digest=sha256:fd323aa64a8ecd330b7cfc9a92b52ce6f7231873e14dddfa5017275a40b2603a

Observation f595520b-1b07-41a3-a7cf-a5541f737394 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.269865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.269865Z digest=sha256:f469f6b62b8a2e3b7a7a0bed6337ae02e2d9efccd6654ca3721e19769c80dea2

Observation d7a20f75-d9d6-46fa-9a54-149ee760f6ab · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.274118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.274118Z digest=sha256:c4f57107c663633e43f0660c72ab2278a40f021ae88b5d46a6c1d76b15d5a285

Observation ec444c70-24b4-43c2-a44e-a7bb796c0678 · outbound

This paper cites Towards vqa models that can read.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Towards vqa models that can read

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.270017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.278599Z digest=sha256:352ad238dcf997a176d51f98e1f921dd525f4d64d02900d24edfdaf3324a4e4c

Observation 84f6cc41-98bc-4567-8db8-f5aa81e0f34e · outbound

This paper cites FlatQuant: Flatness Matters for LLM Quantization.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models FlatQuant: Flatness Matters for LLM Quantization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.282753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.282753Z digest=sha256:a90a69a8a9ac1b42314cfc472995759466df26966b91f6e55213bdc50914557a

Observation 63b7a57d-0ecd-4844-8583-3ab40b3369d3 · outbound

This paper cites Attention is all you need.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Attention is all you need

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.287293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.287293Z digest=sha256:925c93303cb8d0ece53b2a627de3c4b38b2463d1f6ce66e35462fd18b1d1a09b

Observation 0999704a-26e5-418d-9632-184d2537b9a9 · outbound

This paper cites Smoothquant: Accurate and effi- cient post-training quantization for large language models.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Smoothquant: Accurate and effi- cient post-training quantization for large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.245362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.291453Z digest=sha256:4d4e1b2d8dc0faaeab405f46521ca436f4227cde502cee83e5288938c2b0ea21

Observation bb5e6cbd-ce04-42d2-8e04-52df6428f9a5 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Efficient Streaming Language Models with Attention Sinks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.295904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.295904Z digest=sha256:2b327a45226697a98c539108a287ddbaf290c047addaff7da6629ef49f9107ac

Observation ca19e749-1942-4d8e-87fb-4d3f901b9cce · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.300528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.300528Z digest=sha256:2e4ea583df1c2002fd6c9f5d413419bfcd1b5aa687ada6f1b94e6f57b69c957d

Observation e13bac50-265a-49da-8a76-8133c812febf · outbound

This paper cites ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.305041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.305041Z digest=sha256:19d1d5669ed3ed934e7b8b3a1cc5504645e256dd9ab645d1d25e04172193458e

Observation f1f7cbd6-4511-47d1-a7eb-990d43074858 · outbound

This paper cites Mmmu: A massive multi-discipline mul- timodal understanding and reasoning benchmark for expert agi.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Mmmu: A massive multi-discipline mul- timodal understanding and reasoning benchmark for expert agi

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.230644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.309090Z digest=sha256:cbaa1c3378eb614ab99517dcec6fb7e7847dddc99a850e10e265f7cba1514083

Observation 92f982ef-741a-470e-9b61-2fe9c039edb3 · outbound

This paper cites Sigmoid loss for language image pre-training,.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Sigmoid loss for language image pre-training,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.312700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.312700Z digest=sha256:0b23d881a0648cf5b415dc916e2cb9bab95218f680442233d8398c51c7e09bb6

Observation b95bd41c-7765-48b4-953f-85419d44d3b8 · outbound

This paper cites Lmms- eval: Reality check on the evaluation of large multimodal models, 2024.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Lmms- eval: Reality check on the evaluation of large multimodal models, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.205977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.317025Z digest=sha256:476a2a8a1cb0283c99de4142cc59f71bfeef2d28e72d2b83f29bd3d0d90e3b3f

Observation 3af66ba1-ac2d-419c-b818-cf3e37abc2ac · outbound

This paper cites Debi- asing multimodal large language models.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Debi- asing multimodal large language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.192263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.320617Z digest=sha256:c6c95b4a45474812c4d0589a6ed3d6aace4890d1ff4e37539ebf97158e448a91

Observation 13e58d4b-495e-40be-ab0f-72f490fca7d7 · outbound

This paper cites H2o: Heavy-hitter ora- cle for efficient generative inference of large language mod- els.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models H2o: Heavy-hitter ora- cle for efficient generative inference of large language mod- els

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.178261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.324291Z digest=sha256:073cc37b652a5b54b33961ff66ac794a39504acb7d0b8de8b3d4fa93bd659514

Observation 7b56f215-0911-45c3-8102-17d9289bdd3a · outbound

This paper cites Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.328600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.328600Z digest=sha256:bd33bc71049a6ed39e9cc816ede8488be6257925bb99e9b26686dbcd7d4b2afd

Observation aa11f565-5bc6-4306-853e-cbbd9ca7db2e · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate llm serving.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Atom: Low-bit quantization for efficient and accurate llm serving

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.163852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.332913Z digest=sha256:6f3bf8c532e649919b20763a433aadb81f68c56394cefe6f386ed1c871f223ec

Observation 968423d1-a085-480d-974d-62db0ea1534b · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models A Survey on Efficient Inference for Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.336947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.336947Z digest=sha256:9a99dc6de30c15b36ea996c2372b74b7b1437e39910a89245881d218386efe3a

Observation a67f440b-5e8c-4cb4-89e8-23d5b4647d28 · outbound

This paper cites Llava-phi: Efficient multi-modal as- sistant with small language model, 2024.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Llava-phi: Efficient multi-modal as- sistant with small language model, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.149037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.341302Z digest=sha256:2ca17130476f1e60f9a16f30fa7f4ed1589b814e16b584041b515738198bf7d7

Observation 1b8a54a9-3e88-4268-9074-7d3abfb44a92 · outbound

This paper cites The Inference Process of VLMs The inference process of VLMs is shown in Fig.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models The Inference Process of VLMs The inference process of VLMs is shown in Fig

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.134311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.345353Z digest=sha256:0c674d7e360233027f8c53b9dd2eb227dc0d4afe554099da3988c30dad0d9288

Observation d82b6987-bed0-405f-b22f-4d2aff834e60 · outbound

This paper cites LLM Quantization Post-Training Quantization (PTQ) techniques are widely used in LLMs to accelerate the inference process.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models LLM Quantization Post-Training Quantization (PTQ) techniques are widely used in LLMs to accelerate the inference process

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.119372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.349657Z digest=sha256:195e7fdc8200fef5877f64b0692c73213de9ba3f24fcedcb7cd94a6459155f8c

Observation 2b75c01d-38d6-41ae-84f3-e19e94dd071d · outbound

This paper cites W4A16 and W8A8 Results on Large VLMs As shown in Tab.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models W4A16 and W8A8 Results on Large VLMs As shown in Tab

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.102289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.354164Z digest=sha256:4082131c07c694227612a066c39ab15a45863ae2113f1269ddf4a7d9e6251b58

Observation 87b1608f-7d4c-4052-be24-b5c0131895ec · outbound

This paper cites an unresolved cited work.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:21:53.084796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.359493Z digest=sha256:1e2a68a58dfe1a6cdcb3d0e63aad6531f50ed7b4255c3c439d35b18b27faf470

Observation 50a5ff3d-43dd-42e9-bfb1-2d8e8d236efd · outbound

This paper cites an unresolved cited work.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:21:53.068672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.364020Z digest=sha256:c2edda99e18cc50132c72f7b07b505b1da32abbeb3e81285de81c093d81cb845

Observation 30f51fb5-687d-40c2-b914-3f0bbb37c015 · outbound

This paper cites an unresolved cited work.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:21:53.054182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.368471Z digest=sha256:e3abaaaefe8f7ee41e65764fbf3afeaac063915d046e295c5b9e747d4b0c681b

Observation 94d6c4e6-3ad1-4783-a0cb-4fb8dcde253c · outbound

This paper cites an unresolved cited work.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:21:53.039055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.373114Z digest=sha256:1b4692c078972ec67831ba92d0006148ea0f2466f24902def26fa00dabefd4ba

Observation 21771865-8bd4-4057-b2af-4dd4a109a19f · outbound

This paper cites No Output.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models No Output

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:53.023986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:52.377583Z digest=sha256:90e7088a0f307c3f3bc962a9c8c5d2ceab13829bd825e0ce55e27d5b7912d635

Pith citing papers

Observation ed410a1c-caa0-4802-9fcd-db2d6b9c6bf6 · inbound

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization cites this paper.

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization MBQ: Modality-Balanced Quantization for Large Vision-Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T19:12:56.131001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:12:56.131001Z digest=sha256:b3b78c65379409757d71acc0630406e2ee5cbb984fe81261c32277494e5eb87d

Observation a21d352c-9f33-497b-9c6f-b426317f3ed6 · inbound

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation cites this paper.

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation MBQ: Modality-Balanced Quantization for Large Vision-Language Models

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:02:34.572979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T18:55:51.474956Z digest=sha256:3ca0655ed9d19aaef8285188300e3e9a257a28ae220760c12548bd412486e6d5

Observation 96b75b99-8981-44aa-b58e-d14b31774a73 · inbound

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models cites this paper.

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models MBQ: Modality-Balanced Quantization for Large Vision-Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.712644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T19:09:29.347284Z digest=sha256:2230cacd00f014fe8fbed1161b046b636dda8752761465b7a3f211fb06a6a64c

Observation 95409af6-cc6c-4d04-8135-01f72d2376a3 · inbound

SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models cites this paper.

SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models MBQ: Modality-Balanced Quantization for Large Vision-Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:18:37.368114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T16:14:03.717787Z digest=sha256:c38724254fa7c4eb71e56319311a3d4c65328ae9ce70c42bf2a8765b4bb5d299