Pith. sign in

Paper Citation Record · LEDGER

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning

As of 9 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2607.24424.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.24424 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T15:16:36.659461Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 265850aa-de40-408b-a3a6-99916840bc9e · outbound

This paper cites Building Lightweight Semantic Segmentation Models for Aerial Images Using Dual Relation Distillation.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Building Lightweight Semantic Segmentation Models for Aerial Images Using Dual Relation Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.452784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.452784Z digest=sha256:d6f391bfc07d95fd4f05ac667a425582196d056d20479d36a4909639bf53f86b

Observation a40eecd5-d6ad-41c4-a083-8102051bed19 · outbound

This paper cites Yingen Liu, Fan Wu, Ruihui Li, Zhuo Tang, and Kenli Li.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Yingen Liu, Fan Wu, Ruihui Li, Zhuo Tang, and Kenli Li

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.666344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.666344Z digest=sha256:785aab42ba701cd42db084b4a876b5a003a54938ad16473444d69e719e346203

Observation 8d1b3508-6f4c-4e2b-8fb4-820f22c94017 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning MMBench: Is Your Multi-modal Model an All-around Player?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.713332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.713332Z digest=sha256:842e5e857d1719c279057dfc230a4718b183ec67cb12f765e5443e9c1e806645

Observation 5c263949-b6bc-42f6-ba5f-88e8d9d2d287 · outbound

This paper cites LLM-CoT enhanced graph neural recommendation with harmonized group policy optimization.arXiv preprint arXiv:2505.12396,.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning LLM-CoT enhanced graph neural recommendation with harmonized group policy optimization.arXiv preprint arXiv:2505.12396,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.776370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.776370Z digest=sha256:517a81923959d1a919795c5f83b61f9b3eb8f114b4a2095f2054866407f315fc

Observation 076e2494-6d70-4159-b50e-7c46c16574ee · outbound

This paper cites Synthetic Lung X-ray Generation through Cross-Attention and Affinity Transformation.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Synthetic Lung X-ray Generation through Cross-Attention and Affinity Transformation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.827292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.827292Z digest=sha256:95024478dc5b6fcea6aba7a9c96d6ad23a12a27a8392cb4df12eb3e5be60d320

Observation 2f5f2ab3-e589-4d00-a86b-9f6df844c650 · outbound

This paper cites Boosting General Trimap-free Matting in the Real-World Image.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Boosting General Trimap-free Matting in the Real-World Image

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.885573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.885573Z digest=sha256:08bddc1521f393d743728f984a6a8ad6686ae9684d6ca0574c8b9a7f6d92ea12

Observation 2860f3ab-5cea-41f0-840f-f2d73964f7e7 · outbound

This paper cites Edge-guided and Class-balanced Active Learning for Semantic Segmentation of Aerial Images.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Edge-guided and Class-balanced Active Learning for Semantic Segmentation of Aerial Images

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.961854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.961854Z digest=sha256:de75d95789ae11b7758ebecb442b4417fd621f2cc1f612503fe925559de38ac9

Observation 5b72cc80-e18a-441d-bb84-bd50cac27d7f · outbound

This paper cites LLaV A-PruMerge: Adaptive token reduction for efficient large multimodal models.arXiv preprint arXiv:2403.15388,.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning LLaV A-PruMerge: Adaptive token reduction for efficient large multimodal models.arXiv preprint arXiv:2403.15388,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.036325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.036325Z digest=sha256:65e54d28549491b7e571cbd6cbf18d3ea3253f7b7de5f65b2fd67c34b6a518ea

Observation 615aaa49-0563-4914-a4ef-9059c1386717 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.166786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.166786Z digest=sha256:8735fe231f3d0f763c05d1d11ccd7393e515da3958f41cfc498e1afd3074e3c2

Observation aff33c9b-f53e-4232-b2ac-e8a09bce3b99 · outbound

This paper cites GDGS: 3D Gaussian Splatting Via Geometry-Guided Initialization And Dynamic Density Control.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning GDGS: 3D Gaussian Splatting Via Geometry-Guided Initialization And Dynamic Density Control

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.231359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.231359Z digest=sha256:4880cb25840323cae9ef7e02af17e4ecd90b758b539728496d10a7631f3c81b4

Observation 256a2bc7-e757-49fd-98e6-8408da8096ee · outbound

This paper cites RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.284072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.284072Z digest=sha256:e1f616171496c5058600bbd29c73f5daa9ea8d34fea84b998491f450b3bacdd2

Observation c955fe02-0de6-4385-83ac-b075f505bd63 · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.341288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.341288Z digest=sha256:b0ad42980290c546a8b4ed48c3b22a7192d6203d8b32a56cdde23ca9b9175ae6

Observation 7df803d8-1062-441e-9490-50f8a674284e · outbound

This paper cites KV-Efficient VLA: A method to speed up vision language models with RNN-gated chunked KV cache.arXiv preprint arXiv:2509.21354,.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning KV-Efficient VLA: A method to speed up vision language models with RNN-gated chunked KV cache.arXiv preprint arXiv:2509.21354,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.403095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.403095Z digest=sha256:d63994a13e67e85c8a2347248c11f7d5854e71f058838819789e61fa11ccdb57

Observation aa5ec07b-db4a-4d1c-b66a-904a0e21d061 · outbound

This paper cites A Global-Local Cross-Attention Network for Ultra-high Resolution Remote Sensing Image Semantic Segmentation.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning A Global-Local Cross-Attention Network for Ultra-high Resolution Remote Sensing Image Semantic Segmentation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.496400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.496400Z digest=sha256:603c57cac7910b5b96af446e31f1aec56afe0e5c07a0b002a23634df032f2143

Observation b13f99e4-8ea2-4f11-961c-0038e2117e76 · outbound

This paper cites Asymmetric Mamba–CNN collaborative architecture for large-size remote sensing image semantic segmentation.IEEE Transactions on Geoscience and Remote Sensing, 63:2002419, 2025a.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Asymmetric Mamba–CNN collaborative architecture for large-size remote sensing image semantic segmentation.IEEE Transactions on Geoscience and Remote Sensing, 63:2002419, 2025a

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.589098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.589098Z digest=sha256:358bd3224515766439dc2ccedb361d5c6f3958f38c71cb0035d397966f2f47bf

Observation 8485d1c8-2336-4c7e-ade7-65d6837eeaea · outbound

This paper cites DynRsl-VLM: Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning DynRsl-VLM: Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.659461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.659461Z digest=sha256:f6bda70cd55c2b80ff739faa1250fa5a9ae74d282508e4c4292384cdf109f16e

Observation a33bd187-fc65-4803-86b4-289a43d4f062 · outbound

This paper cites NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.591649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.591649Z digest=sha256:497db4331fc498b6f8bbf77970471a3ac6be2bf134311b6139a681dbac9edf68

Observation 1504ed36-23b4-4c05-9f33-67006ed22c3b · outbound

This paper cites GMM-Based Comprehensive Feature Extraction and Relative Distance Preservation For Few-Shot Cross-Modal Retrieval.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning GMM-Based Comprehensive Feature Extraction and Relative Distance Preservation For Few-Shot Cross-Modal Retrieval

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.098226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.098226Z digest=sha256:15fd532bd91d6bd9bcb892f988ec59dc5c557fe0b9084cd6f8cb4fedac87c1b5

Observation 0c31cf5c-e39d-475f-a4b4-9f95d8461d36 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.257881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.257881Z digest=sha256:a7819bf682f52edc0adbfa0afe4b67ed5ae01f00290b77421208703ba0a2bf07

Observation f86ea83f-0f57-4ac1-88c1-8fa8535132cb · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.069570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.069570Z digest=sha256:7b5d4669fbb21f785d2b77201c43c58b885d224b837cc0d21741689b40525379

Observation 0fca96e3-fc26-49c9-a068-05d487e41777 · outbound

This paper cites The Binary Quantized Neural Network for Dense Prediction via Specially Designed Upsampling and Attention.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning The Binary Quantized Neural Network for Dense Prediction via Specially Designed Upsampling and Attention

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.157744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.157744Z digest=sha256:bba637512733e74a7ac4d7695023643f17572bab087cde2fec4fb1627ceeba6f

Observation 37815ebb-7506-4f76-af7f-0b0ad3507636 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning LLaVA-OneVision: Easy Visual Task Transfer

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.320995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.320995Z digest=sha256:c7fa34be2a2e6f36632d839accbf560558719f2103884f0308c6401ff7f04ff6

Observation 8f953d7f-c32b-4170-910c-9dc353963321 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.532712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.532712Z digest=sha256:eaf2cb432aa6aecab3dfb686dee2a8ac31b6b2179b5f4986a3d7bcaab33e383d

Pith citing papers

No inbound Pith citation observations are available.