Pith. sign in

Paper Citation Record · LEDGER

Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2408.02034.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.02034 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:21:54.423993Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T04:50:21.359486Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d4560ee1-b806-42fd-be9a-873ccef88b5d · inbound

BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices cites this paper.

BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T19:33:00.706742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:33:00.706742Z digest=sha256:0a5a72476c77e86b98dfc35ee628494d247e7143b5a50b810154a1b9d3a1cce3

Observation f4b447fe-d351-4dde-a838-546c1f5c83ef · inbound

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay cites this paper.

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:31:44.394774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:31:44.394774Z digest=sha256:c6c31af5f007e4989f71bc63ece422fc65017ef1e279ac20cd2205f235fece40

Observation 7828f2ba-fc4f-4316-ac5e-3fb15c75e4c4 · inbound

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs cites this paper.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.707003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.707003Z digest=sha256:570e2b11ba8e08845dc4a1422f9e4fe433b9b24049f4c2f7d90f12711573cf9f

Observation e5b3535b-2a3b-4b74-bbe5-529146704571 · inbound

Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization cites this paper.

Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:54.423993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:54.423993Z digest=sha256:2e7ee54bfec1953dff18143564bde767b3f25c4c355d60985506875cacf5a607

Observation 015b35cf-7bf1-4dfc-947c-46164b3e9309 · inbound

WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild? cites this paper.

WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild? Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:41.670215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:41.670215Z digest=sha256:3d27a287a3a1d24ff4c27bae452cef03c25c05f8947bab3bc50edc583c3b2de1

Observation 7a46b53b-59b1-4e10-a856-f00b73567ab5 · inbound

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs cites this paper.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.334971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.334971Z digest=sha256:2fbc4d9bc603365defee49333ee83ea20de6f4547f8a1ee40c199678b3b4c38e

Observation f611ea19-3a0c-48f2-88f7-795668443c75 · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.239887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.239887Z digest=sha256:93dcb3c43a461d8ccd4f10a99486f90a52f2476386ec76732fd514a9a166a812

Observation 2a234bd9-2c76-4500-b588-a30ff55ea878 · inbound

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation cites this paper.

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:49.933618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:49.933618Z digest=sha256:864098942842914eb7fd83e3a4e4618030822ec61a26027a28ec74a21df2da54

Observation 2b79dd09-ef27-4f58-ae5b-f6dfefd523b9 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:01.142507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:01.142507Z digest=sha256:007e87e900c7eaf8f67ffc9e80801d581f6855211f20a64e8a1b96c5ea8c7154

Observation fe8d2891-1f27-40ee-8a9b-e9f96cbd28d8 · inbound

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning cites this paper.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:18.365077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:18.365077Z digest=sha256:0f26d08eadb31cb5e13f702d5f1ad113ba67824f2a2eba63dbed3aee20352b1f

Observation 8639ac15-6145-4cde-b036-f7ba2bc56403 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.686268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:a5c79fd19ed9fdd587ce975b98cf49805cfbc932b7a035e55f9b57f1ababf6eb

Observation 85000c5e-19d0-4438-ade5-10153eaed7b5 · inbound

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models cites this paper.

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:00:51.526642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:46:26.869644Z digest=sha256:2461c6e7f4a1389b29edf64d0d33faeadc25dc08b371ff36cc1c09574231b2a3

Observation 4c8d35b5-1b68-42e7-bc82-7f5d0d3929c5 · inbound

GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing cites this paper.

GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:51:00.475208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T17:56:12.097628Z digest=sha256:5e13070c7482f8eb25b19049a6d3de8508faf5ac86af1bea45430ad82c63d309

Observation 7eb3020c-6987-47bc-a5ca-18fa8b34fcc0 · inbound

CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception cites this paper.

CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:50:21.363120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-25T04:46:52.071071Z digest=sha256:a435c1169c038843017fe7acffd4b03e9fe47c56a1a7a039a938ad769781ef33