Pith. sign in

Paper Citation Record · LEDGER

Ovis: Structural Embedding Alignment for Multimodal Large Language Model

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 90 inbound Pith citation observations for arXiv:2405.20797.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.20797 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 90 of 90 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:40:26.552528Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:50:10.268440Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 82093b3b-5904-4d0a-9070-69f69e35f8bf · inbound

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step cites this paper.

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:35:25.922253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T11:35:25.894465Z digest=sha256:e4494bf390a139e618bc7676ca2ff08bc22cde3bbd56c953fd44b1970ae6aa37

Observation 1fefffe4-841f-4476-ad1b-aab72a973087 · inbound

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective cites this paper.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.605366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.605366Z digest=sha256:33be63ee5bf393d8fae8a3ff9b5fa253b9e5efaa4658082304ab87cdb16beaac

Observation 4a30158a-bbd4-454f-a035-56074b10d43a · inbound

VAGUE: Visual Contexts Clarify Ambiguous Expressions cites this paper.

VAGUE: Visual Contexts Clarify Ambiguous Expressions Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.453481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.453481Z digest=sha256:0b1df8e6d9a8852d5c04622565a37cc8a3d769363d40fb4e5c589cf27ced48c4

Observation 02504af0-6a21-40c6-8cd6-918760059e41 · inbound

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models cites this paper.

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:06.869007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:06.869007Z digest=sha256:5853d13f24dc690743221c2f762cbd33cee53c9b13a80044cec09ce1443ff3c4

Observation 2599b4e7-255d-4427-bf93-cc868e3785a3 · inbound

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens cites this paper.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.301248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.301248Z digest=sha256:8e516ea62110e6e8161d26dbf2b7ba91da3f41b5297c300957405d904c1be0b4

Observation 92301f46-ef64-48b7-b9d5-80f8c10b6a9f · inbound

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks cites this paper.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:27.913396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:27.913396Z digest=sha256:ab275710188683383f7877677116a100abe928da2922b169c8341c644be66036

Observation 10691041-26b8-4e48-b5e5-e17e4edce9ba · inbound

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features cites this paper.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.750576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.750576Z digest=sha256:8dff9c4e70c00a535e016621f8946d9e32f08924c55444e42f66db4b9a7d2a3e

Observation 4edee02f-3b63-410e-bc20-a9cd18be64e9 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 169

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.792749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:f415e2bf13cf939eccbf6f2507d07f64ce034b6684c5a1dadb79ba8cadd76d1e

Observation a8fbda6f-fa91-4fa9-b3d8-097936025f9a · inbound

Chimera: Improving Generalist Model with Domain-Specific Experts cites this paper.

Chimera: Improving Generalist Model with Domain-Specific Experts Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:47.617136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:47.617136Z digest=sha256:c3b8190ee7651ef9041b53b69323c4a908be7b3ebf5506064aa4f15d2469971d

Observation e226fb37-2f60-4d15-a93c-75819721536c · inbound

POINTS1.5: Building a Vision-Language Model towards Real World Applications cites this paper.

POINTS1.5: Building a Vision-Language Model towards Real World Applications Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T17:53:35.992520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:53:35.992520Z digest=sha256:558d090baaed194f9209986ca384633526e65921d025680a1b8aa39b1787c766

Observation 33171c3b-2af3-44ee-89cc-a1671fa3703a · inbound

COEF-VQ: Cost-Efficient Video Quality Understanding through a Cascaded Multimodal LLM Framework cites this paper.

COEF-VQ: Cost-Efficient Video Quality Understanding through a Cascaded Multimodal LLM Framework Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T18:10:44.715962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:10:44.715962Z digest=sha256:6b5ca10252c51e16a52a8c5e0c70047bcb9e1314f3c1b20c411820edaf74ab88

Observation b0878f72-9289-411e-b773-26e5a6ef4bc4 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.826379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:fca84d1ba411d19b3e6d6f1cbd78d8604c94331c8488a674db0d835be0ec7ec0

Observation 87e98038-a2f1-4336-a8f9-7a6806ca83a3 · inbound

LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models cites this paper.

LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:10.984594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:45:10.984594Z digest=sha256:3f0af1aa14daf16d11c82fc661ca456948873625c08ad889b539c35aa15533a7

Observation 2df41e0f-89b4-4108-a781-7705033c0986 · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.727652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:2f0d6a7259ef19e4dc3fe18659ac0b1aea163d02d033f87f95cd806ab1f96c5e

Observation dfb87a1a-2332-4769-8d31-8c1dc71dc8c7 · inbound

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs cites this paper.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.176720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.176720Z digest=sha256:a974c084d6507aa6e1c92b918639a33ca641154a1f23cd600a8b7a65a3260612

Observation faf8283a-7e18-46ca-a27c-1f8a2aeb734b · inbound

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design cites this paper.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.086305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.086305Z digest=sha256:eb7a0a7fca496db26152b40253d7a631ccb3373b520d2e515fced95a665cb725

Observation ac100909-eebd-449f-a3ac-de9c8ccd516b · inbound

Compositional Generative Model of Unbounded 4D Cities cites this paper.

Compositional Generative Model of Unbounded 4D Cities Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 120

Resolution
unresolved
no resolver link, observed 2026-08-10T20:17:26.551329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:17:26.551329Z digest=sha256:a581d4e9937bab43ea7447fd3e98d1f2d01a2209cbd1ec86e44f217652958fbd

Observation 5fc3821e-0ff2-44bb-84c7-380649c8191e · inbound

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models cites this paper.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:49.974575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:18:49.974575Z digest=sha256:578644984f4afc4b90b999491aadf47dafc03a312d5cf3a3570cb5a2c4dce384

Observation b4186bfa-ba20-4805-b984-86377a61e97d · inbound

Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMs cites this paper.

Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T21:10:05.920921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:10:05.920921Z digest=sha256:5e6138e92279be28a0c6141c876811ea7729222110a913b86700374f112626da

Observation c57eab0a-b3ea-448c-9325-a90e0a335d38 · inbound

MobileA3gent: Training Mobile GUI Agents Using Decentralized Self-Sourced Data from Diverse Users cites this paper.

MobileA3gent: Training Mobile GUI Agents Using Decentralized Self-Sourced Data from Diverse Users Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T10:29:38.525159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:29:38.525159Z digest=sha256:8ee0f925b65aa3c0127e09ed1896ab5e68bf4bb7a5228a2798814ac9897097d2

Observation f4566433-df2b-4547-be0c-63e97c43124e · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.046852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:889725aae617f3caf71745d29bc00c826b453c305e468fd7bf7d6b762cf304d7

Observation 4cef6bf6-12ac-4b35-b56b-21fb16ccee8a · inbound

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching cites this paper.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.552528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.552528Z digest=sha256:b61d85fb08e576f71f4cf2a1ca167fab68e8c44c6e923055c0bfe2776334c834

Observation 5f564e75-13ed-4af2-a92e-1d7985f7c840 · inbound

VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models cites this paper.

VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:47.223722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:33:47.223722Z digest=sha256:c4b8ced258e741d42811f84d29f7bdd41f24bfff2aaf2922612848666ea5c984

Observation e38f2089-40de-436e-890c-ef9043706c2c · inbound

Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs cites this paper.

Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:32:17.438888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:32:17.438888Z digest=sha256:50384c1d1696eb765ab218bebdbf24034a909ddbf387fb67e5c4a9c01484a290

Observation b0d3637f-053d-4f92-8518-2f05c4747acb · inbound

RePOPE: Impact of Annotation Errors on the POPE Benchmark cites this paper.

RePOPE: Impact of Annotation Errors on the POPE Benchmark Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:22:34.522788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:22:34.522788Z digest=sha256:8f6ecb425247caaef522e802e2074671de598dfe33712203f83cdf19009a7717

Observation bed95706-12aa-4353-b922-d6ced562f7f0 · inbound

WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild? cites this paper.

WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild? Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:41.877246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:41.877246Z digest=sha256:9454db88a5bb50e4f1b6fca0fd42d0d745b66d90aeaa762aaa123a5d489cdb90

Observation 8d565fb3-4a9a-4d30-9dc0-c791df6aad28 · inbound

NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment cites this paper.

NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:47.031949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:47.031949Z digest=sha256:5e1e880499b1657c9389c45665753625da883a8ba40b27d8273686243e73aa49

Observation 76420bee-e81f-491f-86c1-0dd85a6789b6 · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.788591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:dcf03c48dc4468997030d5e33b3bfe7b11ccf1937a0507a3e1e2a074301c737a

Observation 66051f3d-87e3-4bc9-be8c-dd4563c51b30 · inbound

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs cites this paper.

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:22.917151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:22.917151Z digest=sha256:6037d5480c08ae0bc4aa47555f0e7dbe0ad9b5079254e2eda157e1280ef9eb5e

Observation 49dbe8c7-b28d-4ddf-b0db-d885e34f02f6 · inbound

mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation cites this paper.

mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:25.374172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:42:25.374172Z digest=sha256:c50f864a2865014bcf1b797ef05cf770bf5f809c9ea33fc02d9898e537da0138

Observation 464bd6fb-d219-47f4-bb85-680b63386024 · inbound

MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book cites this paper.

MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:04.167429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:04.167429Z digest=sha256:4bd57f396dea76ff756c66d19eabd88087447e6b85acc7cdc91e5fe978208827

Observation 3f99d2ac-6ad3-429a-b55a-21d3195c5188 · inbound

Affordance Benchmark for MLLMs cites this paper.

Affordance Benchmark for MLLMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:54.201377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:59:54.201377Z digest=sha256:4bfe1ccdfac75c76fcd7c0f70ae76c71c509aa0f7b1c010b8a66e133fd1e0db4

Observation f3ea50c4-9b66-4532-954b-befd96fe0118 · inbound

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking cites this paper.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.202412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.202412Z digest=sha256:eae8c2d79659c2e3a1c42f7940163ccf3991f281b9e51acec41c6bf662569238

Observation e591ffba-6bbc-4c1d-97ef-c3745c0836c4 · inbound

Multimodal Tabular Reasoning with Privileged Structured Information cites this paper.

Multimodal Tabular Reasoning with Privileged Structured Information Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:00.675957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:54:00.675957Z digest=sha256:e7eacb879f4f5a63b2cb963986675f342a808ba281d0c7ee602a9d8de49ead73

Observation 9e4548a8-5541-4013-8f38-477bf4129879 · inbound

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs cites this paper.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.565388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.565388Z digest=sha256:eee288ed252c4541ddd836ffd4d12bf968f1a8493ba679cb95a088bec05610c1

Observation 6e33b33c-1f92-4ff7-b52d-f59fb216fe01 · inbound

Towards an Explainable Comparison and Alignment of Feature Embeddings cites this paper.

Towards an Explainable Comparison and Alignment of Feature Embeddings Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:39.512287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:39.512287Z digest=sha256:cb41b6726782ef1af05d16ed0740bd59f0c99387074dff441ced4db39b036b31

Observation e7ea8119-c001-4313-b8b4-2dd6293465b7 · inbound

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding cites this paper.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.920387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.920387Z digest=sha256:d0896625652966147b211154afe38ed25746f350bc7806a3ac7ddd739f9c0d39

Observation 1f4d212c-0e72-4016-a0dd-0c62c367eb1e · inbound

WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code cites this paper.

WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:50.136340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:50.136340Z digest=sha256:eacd2917bdfaebe0b2237c1c9d8e1bacd95d3cf719d29c57fc3c3b7a80f6fdef

Observation 7a3e6cea-898e-40e9-91fe-c3d3a86793f0 · inbound

LPO: Towards Accurate GUI Agent Interaction via Location Preference Optimization cites this paper.

LPO: Towards Accurate GUI Agent Interaction via Location Preference Optimization Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:47:13.940530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T09:45:55.674155Z digest=sha256:a7e85bbee65c937af9d08ee69555631ae7c1c15b035eb74db0bf5a71244aea7c

Observation cc074f68-bcd9-4daa-901a-fa301cb8dd23 · inbound

Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding cites this paper.

Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:43:02.189331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T09:42:14.171289Z digest=sha256:bf74cd00e17188583bfa1895db533b32671ac17a3add39a32bafd0aec46a612f

Observation 2eb594c4-84b6-4122-868d-e00f4d52204c · inbound

Text-Aware Image Restoration with Diffusion Models cites this paper.

Text-Aware Image Restoration with Diffusion Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:41:07.869503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:41:07.869503Z digest=sha256:66884bf641f114dde1f2fc49216704108b81d94ea4b8b783416b82cbd53c3066

Observation 20fd1c83-8f59-442d-a989-204ca7a10f17 · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:24.586499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:24.586499Z digest=sha256:2c46c57e9ccfb55e5a7b8beb760e7c40735a208604b75823b7d2095508a4d953

Observation b9fd3091-8793-423c-a21f-8becf134416e · inbound

Taming the Untamed: Graph-Based Knowledge Retrieval and Reasoning for MLLMs to Conquer the Unknown cites this paper.

Taming the Untamed: Graph-Based Knowledge Retrieval and Reasoning for MLLMs to Conquer the Unknown Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:13.859150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:13.859150Z digest=sha256:2294c9f275bee767a7918c30d5e2d799fb91915475a36ff978901313f2a4ec92

Observation b9e15661-e791-481c-8847-73d935275cfd · inbound

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation cites this paper.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.723984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.723984Z digest=sha256:82adfb4aa638ef95a38499f31b5210f787945f9c99d1833640b7429890597324

Observation 9249863c-eb5c-4438-82bb-b1938613a098 · inbound

Med-Art: Diffusion Transformer for 2D Medical Text-to-Image Generation cites this paper.

Med-Art: Diffusion Transformer for 2D Medical Text-to-Image Generation Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:53:13.548511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:53:13.548511Z digest=sha256:866a698b325e0f547d5a8ddd5284a2e40f55811e527cbc2786de6b2dbe76ef68

Observation 2aad56e1-aa05-4038-b813-07d979938afa · inbound

Ovis-U1 Technical Report cites this paper.

Ovis-U1 Technical Report Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:20.362999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:20.362999Z digest=sha256:d8a6ff5e797365c2486cef370d6f116375c97c84b0e9773645e00fafae66b128

Observation ed625676-9ef6-457f-8452-fdcc4955ba73 · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:07.221154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:07.221154Z digest=sha256:ce0b0f30ec9a5c3c8c2147266cb5931b76022f1308d4ceb3c207c5929ffe889b

Observation f4713053-4307-437d-9fb0-c1d5bf82f1c1 · inbound

Describe Anything Model for Visual Question Answering on Text-rich Images cites this paper.

Describe Anything Model for Visual Question Answering on Text-rich Images Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:25.555322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:25.555322Z digest=sha256:61adc9ffbff1f090b0cc90ed4a28afd185cab9d8f3967c9b5c191716e0c5d374

Observation 37af84cf-e2e8-4547-a004-8aab72ea8307 · inbound

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems cites this paper.

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T15:43:16.151049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:43:16.151049Z digest=sha256:ab39382c07ff9196e3ba631d923757f2569c6226e95525b108972b8ba0d24b5a

Observation 73c5e2f0-5c99-4a30-8ff4-cb2338632a10 · inbound

MMGraphRAG: Bridging Vision and Language with Interpretable Multimodal Knowledge Graphs cites this paper.

MMGraphRAG: Bridging Vision and Language with Interpretable Multimodal Knowledge Graphs Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T13:20:35.862491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:20:35.862491Z digest=sha256:e52e345b4bd2f47e74d5afbd1ef267711aa642b074025cce6bc04bd79633b9af

Observation 71701df6-149d-4416-8faa-27b3e2ccb689 · inbound

Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis cites this paper.

Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:30.234395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:30.234395Z digest=sha256:e9dd33e1df70a81cb974340cd322734dd418a13b0f311106c4bbfb2ba1e3fdee

Observation 72f64ce5-8e37-4f3a-b84e-f3d780d1ae00 · inbound

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation cites this paper.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.347677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.347677Z digest=sha256:0b2c11a3499331d728913723d080458d75131b29c08c6300463afd88812fc008

Observation a8543d23-cf53-4291-9d69-f9b6bb28ff82 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.194870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:c280590da819157e75adbc00f9e0918e0698e5b6fd0a933cbbb3739d099586c8

Observation adb73585-05b2-45da-8192-0ca009ca97cf · inbound

KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts cites this paper.

KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:23:53.395808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:23:53.395808Z digest=sha256:5c4626789b6560636857960b9ba1fd35bceae7c69666b84ba3611ceeaf9fb79b

Observation d4cd68da-f164-4e7f-8d36-51d711937cea · inbound

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis cites this paper.

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T19:27:40.738360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:27:40.738360Z digest=sha256:e717e7ff4e7b05a09e2f286972517d4d36d786ccf0ca19cf352d552301c145ad

Observation 53aec416-a58f-4094-8e9d-41da5ae12f38 · inbound

Measuring Epistemic Humility in Multimodal Large Language Models cites this paper.

Measuring Epistemic Humility in Multimodal Large Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:39.216124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:39.216124Z digest=sha256:7c3d72ab46da6ed2305c4e2e2f33737040c57f3f10ba8ebd4b096297646c5584

Observation fa648e32-13c7-4ed4-be3f-195e6326d6d0 · inbound

Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance cites this paper.

Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T22:36:09.874162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:36:09.874162Z digest=sha256:d9586062084ad7694b5e55390a750beeae5a375aabf06779c6886cf17f7d7fff

Observation 55144bd6-3362-419a-9d27-1e85742fc5c8 · inbound

Diagnosing Corruption-Induced Reliability Failures in Vision-Language Models cites this paper.

Diagnosing Corruption-Induced Reliability Failures in Vision-Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T20:41:50.695776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:41:50.695776Z digest=sha256:a04f310e582f34b9029fb127ae2b35b7f9b154d0e9c88b58bb70cd6be60f26ec

Observation d87c5108-92ad-47a2-8078-87dce0dddb10 · inbound

SCLARO: A Dataset for Grounded Scenario-Level Scene Understanding and ScenarioCLIP for Benchmarking cites this paper.

SCLARO: A Dataset for Grounded Scenario-Level Scene Understanding and ScenarioCLIP for Benchmarking Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T20:24:00.699557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:24:00.699557Z digest=sha256:75d5d67476763e1c618e0f201323895e7228154c939f27469aa024e41dba63a2

Observation d081467a-fba5-47ed-8fed-6128430195bb · inbound

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension cites this paper.

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T18:15:06.083870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:15:06.083870Z digest=sha256:fb2238c87bcfd0749f9a871f2eb982e3c4f7857d3c70d025db96b0dacf86b80e

Observation afc7c333-08f8-428a-9dd9-e77e1b97476e · inbound

Grounding Everything in Tokens for Multimodal Large Language Models cites this paper.

Grounding Everything in Tokens for Multimodal Large Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:31:21.863850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T23:31:05.422935Z digest=sha256:da5dd15679c51bd365b59dd2ab8a6f002ddb35d106cea949ddf9cda7772d3d02

Observation f2ec3220-6993-45e7-afe4-3c0d2f6efae1 · inbound

OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving cites this paper.

OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:38:37.851700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T22:34:00.895252Z digest=sha256:f815d64f618cb2d17dc6acfef8946218f9279642be3a5bcc40a49ca4954efc3a

Observation 405b0bd7-b878-4b6b-80e6-77fb86d9fd85 · inbound

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning cites this paper.

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:02:42.549876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T10:02:20.477517Z digest=sha256:0ad678fa616aaa12a6c10de6a5d1a933e9970664e9968e4cb07b6b5e876c0c36

Observation 8055677a-4504-4f9f-aeb7-9700d90d47ed · inbound

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? cites this paper.

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:40:12.519458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T13:36:28.844714Z digest=sha256:d601830973bbcbbe67ac28267fabe290be65208699e889de1e41334012f5a35f

Observation 383e7ce0-1622-4682-b2ce-6b40bb973922 · inbound

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? cites this paper.

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T04:31:21.095784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:31:21.095784Z digest=sha256:6214e9259bac3bbe7d0b8d8efa9de0cf5b89d181fac8cd4924f7f776bfd65c3d

Observation 2a0ef8d2-a537-406a-a6e0-7c2c54522585 · inbound

ST-BiBench: Benchmarking Multi-Stream Multimodal Coordination in Bimanual Embodied Tasks for MLLMs cites this paper.

ST-BiBench: Benchmarking Multi-Stream Multimodal Coordination in Bimanual Embodied Tasks for MLLMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.614819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T06:00:02.043029Z digest=sha256:21ce6e140f804a8f06a8016c6838c25bd4cc5d0b27cfc0b3bb4db6b40a206b2f

Observation d68a784c-ddc3-4120-ad0c-919370ae5623 · inbound

TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios cites this paper.

TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T17:05:14.221995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:05:14.221995Z digest=sha256:4029af6c3c63fb4245b5e45bd05f07900b025f4f12d55224443da7beb485d260

Observation c24cc6b2-a876-4c11-a939-82b8a62d6942 · inbound

AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis cites this paper.

AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:05:48.191772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T19:22:08.306946Z digest=sha256:8bf6f6e7eb95e0a602c9c19cfb56ab1da0e4548dd1117a9c6686fa06d61aa430

Observation c5a2c74d-25d1-474d-85cf-87a7b22718e4 · inbound

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization cites this paper.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.076778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:2dea59068b8171978fb8c913cc401e3c5d687b0b5f525956cb3d0fc72f362174

Observation a6b4084a-6e5c-4be7-b4d8-b60f0d7ccc78 · inbound

DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing cites this paper.

DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:36:14.874224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-07T16:44:23.162726Z digest=sha256:a07511c13f6bc26d9af7dd4ada2085be8410ee7e3f2989a104f76711103e517c

Observation 1dde09a9-7fce-4a46-b732-8dc915e1c2b5 · inbound

SenseBench: A Benchmark for Remote Sensing Low-Level Visual Perception and Description in Large Vision-Language Models cites this paper.

SenseBench: A Benchmark for Remote Sensing Low-Level Visual Perception and Description in Large Vision-Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:34.976766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T04:00:50.940530Z digest=sha256:44d2bcf9f38d3bb5fe3327bbcb0a80cc40cdbfec9d3e08ad6e9a77c79de58a01

Observation f549e748-2bac-4aa8-b48a-289107c1c9d5 · inbound

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation cites this paper.

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:37:03.536685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T01:35:09.568623Z digest=sha256:a56fb1d75d9a3788840d4241309596968fc822bfd2baa860185293f67ef168da

Observation 0018550c-f79e-4a6e-8bba-94f75c0d394f · inbound

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation cites this paper.

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:18:00.096827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T21:10:49.059686Z digest=sha256:2505aa0dc84bcbdfdd33161682bafe29cdd091f35eb2a1bd1d04c054d2bc77f2

Observation d9f9bc5f-8391-455c-8134-3286b2c01be5 · inbound

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation cites this paper.

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:02:40.601291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T16:58:55.817334Z digest=sha256:a05634f98741190cbe3e9f27f47270c41cb7bcdb78959661b52a28cf4bf0912d

Observation 614f2864-8cd5-4007-b080-24761b762d57 · inbound

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation cites this paper.

Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:45:46.453550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T22:41:13.164195Z digest=sha256:d9dcf0067699fd2de8d5a4a20341464a7b38c6edff3e71d0eba703d6fda93e22

Observation d201d336-c9fa-445d-9f65-12ab9b1d7268 · inbound

Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding cites this paper.

Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:28:17.150176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T12:24:47.111809Z digest=sha256:95f5b113f0da80730fb03c049290f25090b63476cb7e43f4234cce00bb6e7196

Observation faa0fa58-f5bf-4400-a2d3-40f4c955cfc0 · inbound

PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training cites this paper.

PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:28.874283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T10:39:14.444486Z digest=sha256:af5d0c721ba250e7ac33a9c6c9754cd735bcf367a77788aaa49a0d8d22125102

Observation 758b02c2-9bbd-4c5b-96f8-58ff6ff3930e · inbound

Improving Multimodal Reasoning via Worst Dimension Optimization cites this paper.

Improving Multimodal Reasoning via Worst Dimension Optimization Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.785734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:ef1632e2071699dbcaf8a9163929ee412ea3825471a226503f55ab992dafe86d

Observation 55bd621e-e782-4ba8-816a-259ff541926e · inbound

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems cites this paper.

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:47:22.877488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T20:11:02.445626Z digest=sha256:464068a50a959e46f36464a2f63197cf0f1921e667a8ce89f79762af388ca853

Observation 04329fb6-cc2c-49bc-bb8d-c18d94ef6f20 · inbound

Vision Language Model Helps Private Information De-Identification in Vision Data cites this paper.

Vision Language Model Helps Private Information De-Identification in Vision Data Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:27:31.571951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T16:29:33.294960Z digest=sha256:925d9c3945e2e75ac8063660e006f053d0ce8d37bc571fc7df1c6b63865f7785

Observation bda50666-f002-45bb-93eb-488050348969 · inbound

MMGist: A Comprehensive Multimodal Benchmark for 2027 cites this paper.

MMGist: A Comprehensive Multimodal Benchmark for 2027 Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:49:41.624746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T11:05:14.573386Z digest=sha256:9d9ad4bb0fd36834449ba0babb823e44709aeb7e806d5b86bea0a9986f78be78

Observation 1d106ec9-fef4-475b-81b4-adb180e853be · inbound

SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity cites this paper.

SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:50:10.269996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-25T21:02:44.441202Z digest=sha256:a243792c3b01f326e21edfe1e4aaacb78a7cec27d79db0a469cd78aa5418be55

Observation 49bfc17c-8e25-4814-b5a0-0d1bdc7b8123 · inbound

Learning to Deny: Action Denial in Multimodal Large Language Models cites this paper.

Learning to Deny: Action Denial in Multimodal Large Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:39.575996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-01T06:21:09.996386Z digest=sha256:9108aee9a624a47167d1d076ce6e4829e1f60027f6fce07e6cefc6e15d25b4d7

Observation 0a7e38af-4132-4b7f-8c47-ce4f2a0758e6 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 169

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:cc7b9f5f37195bde067ed4ba72f5e9493ff8cc5951cc491f2513915f644af125

Observation 3cfa9f04-8c0d-42f4-9872-845e2e18b834 · inbound

CritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving cites this paper.

CritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-11T21:09:01.431209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:09:01.431209Z digest=sha256:ed43692cdad7bb1a5f68ba38d710cc02e912d78faed4987588709340ebcfd26f

Observation e7c78af9-1734-4328-8a70-714b749a41ad · inbound

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget cites this paper.

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-02T06:14:00.278259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:14:00.278259Z digest=sha256:99f1097fcf84e69b63a13bc143fcca25e454dbfddc8cf401762a322a2d6ffe6a

Observation 34d2f83b-6417-421b-976b-1456d15465a3 · inbound

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs cites this paper.

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T04:16:08.332153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:16:08.332153Z digest=sha256:9c53a732082497a1074d46b1d5595ab852b66f9b528f61c83144a5d6f525c0d2

Observation 1836c33e-29ea-4d80-8e5c-30e8cb31dd43 · inbound

VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling cites this paper.

VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:15.914749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:15.914749Z digest=sha256:0054016b9408bf4f6e195f9215a3c078b2618f045b9663606f9804a79198dcfa

Observation 99db7e44-8e9e-4f4c-94de-25f2a8a3cad7 · inbound

LookBack: Where and How to Score LVLM Responses via Visual Reference Usage cites this paper.

LookBack: Where and How to Score LVLM Responses via Visual Reference Usage Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:29:57.236863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:29:57.236863Z digest=sha256:3f8c14152e0757368b6cf9be6c256d7aef8706b45e344b26ce5866ec578acaae

Observation 41a3d6da-d642-43e6-9e99-291332c9eaba · inbound

NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents cites this paper.

NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:04:29.967519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:04:29.967519Z digest=sha256:d410b3e5cef77fe43ad7ec46bbf0af5b1a1297e5b8562cda735f96ce34d0d248