Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 90 inbound Pith citation observations for arXiv:2405.20797.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:40:26.552528Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T19:50:10.268440Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 82093b3b-5904-4d0a-9070-69f69e35f8bf · inbound
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1fefffe4-841f-4476-ad1b-aab72a973087 · inbound
MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a30158a-bbd4-454f-a035-56074b10d43a · inbound
VAGUE: Visual Contexts Clarify Ambiguous Expressions Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02504af0-6a21-40c6-8cd6-918760059e41 · inbound
Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2599b4e7-255d-4427-bf93-cc868e3785a3 · inbound
Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92301f46-ef64-48b7-b9d5-80f8c10b6a9f · inbound
AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10691041-26b8-4e48-b5e5-e17e4edce9ba · inbound
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4edee02f-3b63-410e-bc20-a9cd18be64e9 · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 169
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a8fbda6f-fa91-4fa9-b3d8-097936025f9a · inbound
Chimera: Improving Generalist Model with Domain-Specific Experts Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e226fb37-2f60-4d15-a93c-75819721536c · inbound
POINTS1.5: Building a Vision-Language Model towards Real World Applications Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33171c3b-2af3-44ee-89cc-a1671fa3703a · inbound
COEF-VQ: Cost-Efficient Video Quality Understanding through a Cascaded Multimodal LLM Framework Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0878f72-9289-411e-b773-26e5a6ef4bc4 · inbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 87e98038-a2f1-4336-a8f9-7a6806ca83a3 · inbound
LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2df41e0f-89b4-4108-a781-7705033c0986 · inbound
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dfb87a1a-2332-4769-8d31-8c1dc71dc8c7 · inbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faf8283a-7e18-46ca-a27c-1f8a2aeb734b · inbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac100909-eebd-449f-a3ac-de9c8ccd516b · inbound
Compositional Generative Model of Unbounded 4D Cities Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 120
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fc3821e-0ff2-44bb-84c7-380649c8191e · inbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4186bfa-ba20-4805-b984-86377a61e97d · inbound
Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c57eab0a-b3ea-448c-9325-a90e0a335d38 · inbound
MobileA3gent: Training Mobile GUI Agents Using Decentralized Self-Sourced Data from Diverse Users Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4566433-df2b-4547-be0c-63e97c43124e · inbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4cef6bf6-12ac-4b35-b56b-21fb16ccee8a · inbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f564e75-13ed-4af2-a92e-1d7985f7c840 · inbound
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e38f2089-40de-436e-890c-ef9043706c2c · inbound
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0d3637f-053d-4f92-8518-2f05c4747acb · inbound
RePOPE: Impact of Annotation Errors on the POPE Benchmark Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bed95706-12aa-4353-b922-d6ced562f7f0 · inbound
WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild? Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d565fb3-4a9a-4d30-9dc0-c791df6aad28 · inbound
NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76420bee-e81f-491f-86c1-0dd85a6789b6 · inbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 66051f3d-87e3-4bc9-be8c-dd4563c51b30 · inbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49dbe8c7-b28d-4ddf-b0db-d885e34f02f6 · inbound
mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 464bd6fb-d219-47f4-bb85-680b63386024 · inbound
MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f99d2ac-6ad3-429a-b55a-21d3195c5188 · inbound
Affordance Benchmark for MLLMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3ea50c4-9b66-4532-954b-befd96fe0118 · inbound
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e591ffba-6bbc-4c1d-97ef-c3745c0836c4 · inbound
Multimodal Tabular Reasoning with Privileged Structured Information Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e4548a8-5541-4013-8f38-477bf4129879 · inbound
SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e33b33c-1f92-4ff7-b52d-f59fb216fe01 · inbound
Towards an Explainable Comparison and Alignment of Feature Embeddings Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7ea8119-c001-4313-b8b4-2dd6293465b7 · inbound
Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f4d212c-0e72-4016-a0dd-0c62c367eb1e · inbound
WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a3e6cea-898e-40e9-91fe-c3d3a86793f0 · inbound
LPO: Towards Accurate GUI Agent Interaction via Location Preference Optimization Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cc074f68-bcd9-4daa-901a-fa301cb8dd23 · inbound
Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2eb594c4-84b6-4122-868d-e00f4d52204c · inbound
Text-Aware Image Restoration with Diffusion Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20fd1c83-8f59-442d-a989-204ca7a10f17 · inbound
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9fd3091-8793-423c-a21f-8becf134416e · inbound
Taming the Untamed: Graph-Based Knowledge Retrieval and Reasoning for MLLMs to Conquer the Unknown Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9e15661-e791-481c-8847-73d935275cfd · inbound
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9249863c-eb5c-4438-82bb-b1938613a098 · inbound
Med-Art: Diffusion Transformer for 2D Medical Text-to-Image Generation Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aad56e1-aa05-4038-b813-07d979938afa · inbound
Ovis-U1 Technical Report Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed625676-9ef6-457f-8452-fdcc4955ba73 · inbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4713053-4307-437d-9fb0-c1d5bf82f1c1 · inbound
Describe Anything Model for Visual Question Answering on Text-rich Images Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37af84cf-e2e8-4547-a004-8aab72ea8307 · inbound
In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73c5e2f0-5c99-4a30-8ff4-cb2338632a10 · inbound
MMGraphRAG: Bridging Vision and Language with Interpretable Multimodal Knowledge Graphs Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71701df6-149d-4416-8faa-27b3e2ccb689 · inbound
Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72f64ce5-8e37-4f3a-b84e-f3d780d1ae00 · inbound
Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8543d23-cf53-4291-9d69-f9b6bb28ff82 · inbound
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation adb73585-05b2-45da-8192-0ca009ca97cf · inbound
KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4cd68da-f164-4e7f-8d36-51d711937cea · inbound
Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53aec416-a58f-4094-8e9d-41da5ae12f38 · inbound
Measuring Epistemic Humility in Multimodal Large Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa648e32-13c7-4ed4-be3f-195e6326d6d0 · inbound
Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55144bd6-3362-419a-9d27-1e85742fc5c8 · inbound
Diagnosing Corruption-Induced Reliability Failures in Vision-Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d87c5108-92ad-47a2-8078-87dce0dddb10 · inbound
SCLARO: A Dataset for Grounded Scenario-Level Scene Understanding and ScenarioCLIP for Benchmarking Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d081467a-fba5-47ed-8fed-6128430195bb · inbound
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afc7c333-08f8-428a-9dd9-e77e1b97476e · inbound
Grounding Everything in Tokens for Multimodal Large Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f2ec3220-6993-45e7-afe4-3c0d2f6efae1 · inbound
OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 405b0bd7-b878-4b6b-80e6-77fb86d9fd85 · inbound
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8055677a-4504-4f9f-aeb7-9700d90d47ed · inbound
VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 383e7ce0-1622-4682-b2ce-6b40bb973922 · inbound
VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a0ef8d2-a537-406a-a6e0-7c2c54522585 · inbound
ST-BiBench: Benchmarking Multi-Stream Multimodal Coordination in Bimanual Embodied Tasks for MLLMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d68a784c-ddc3-4120-ad0c-919370ae5623 · inbound
TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c24cc6b2-a876-4c11-a939-82b8a62d6942 · inbound
AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c5a2c74d-25d1-474d-85cf-87a7b22718e4 · inbound
Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a6b4084a-6e5c-4be7-b4d8-b60f0d7ccc78 · inbound
DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1dde09a9-7fce-4a46-b732-8dc915e1c2b5 · inbound
SenseBench: A Benchmark for Remote Sensing Low-Level Visual Perception and Description in Large Vision-Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f549e748-2bac-4aa8-b48a-289107c1c9d5 · inbound
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0018550c-f79e-4a6e-8bba-94f75c0d394f · inbound
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d9f9bc5f-8391-455c-8134-3286b2c01be5 · inbound
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 614f2864-8cd5-4007-b080-24761b762d57 · inbound
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d201d336-c9fa-445d-9f65-12ab9b1d7268 · inbound
Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation faa0fa58-f5bf-4400-a2d3-40f4c955cfc0 · inbound
PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 758b02c2-9bbd-4c5b-96f8-58ff6ff3930e · inbound
Improving Multimodal Reasoning via Worst Dimension Optimization Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 55bd621e-e782-4ba8-816a-259ff541926e · inbound
Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 04329fb6-cc2c-49bc-bb8d-c18d94ef6f20 · inbound
Vision Language Model Helps Private Information De-Identification in Vision Data Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bda50666-f002-45bb-93eb-488050348969 · inbound
MMGist: A Comprehensive Multimodal Benchmark for 2027 Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1d106ec9-fef4-475b-81b4-adb180e853be · inbound
SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 49bfc17c-8e25-4814-b5a0-0d1bdc7b8123 · inbound
Learning to Deny: Action Denial in Multimodal Large Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0a7e38af-4132-4b7f-8c47-ce4f2a0758e6 · inbound
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 169
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cfa9f04-8c0d-42f4-9872-845e2e18b834 · inbound
CritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7c78af9-1734-4328-8a70-714b749a41ad · inbound
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34d2f83b-6417-421b-976b-1456d15465a3 · inbound
ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1836c33e-29ea-4d80-8e5c-30e8cb31dd43 · inbound
VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99db7e44-8e9e-4f4c-94de-25f2a8a3cad7 · inbound
LookBack: Where and How to Score LVLM Responses via Visual Reference Usage Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a3d6da-d642-43e6-9e99-291332c9eaba · inbound
NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.