Pith. sign in

Paper Citation Record · LEDGER

Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 50 inbound Pith citation observations for arXiv:2409.12961.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.12961 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 50 of 50 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:47:39.207697Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:09:50.847114Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e0925f8d-efec-4110-8251-c6b1daa0f532 · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.235839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:66fe73bd89b57d60f5c1cd1c841f4e7b6091999cf4ee8f89fb25d1d57d6f3758

Observation a0f969e2-d8f5-486e-bdf4-ec96e7c62893 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.747383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:3b00f30c7fed3bfd6c8584ffec2d37290c5c84147f1835ec6ac3beee94ec959f

Observation 25818a3a-9cf6-4172-babf-0805141c6282 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:45:28.369891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:e845d419a08e322739302f63139ebfa2a13ff313413b5863e027b264938e5bfe

Observation 96ae2de7-6e16-41b4-8353-9851ffc6d606 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:02:37.421236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:2d1beb89416c20a421f4f325e60abdc7cd02d2a108b360c9f66033c8de65461f

Observation 5c2840ee-d5b0-4714-8645-3ddf4a95a080 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:19:59.722872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:13203f738d72d4dcc34490ce3a6e4085785f012a2aaf07793933c188f12880d8

Observation 0d28cb45-166f-4728-9025-ba01d2f44225 · inbound

Ola: Pushing the Frontiers of Omni-Modal Language Model cites this paper.

Ola: Pushing the Frontiers of Omni-Modal Language Model Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.207697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.207697Z digest=sha256:f550a5d7233f9d00ff486d4f8aea224719d92fd77ffc76109d21de2e37624f50

Observation 78546f28-51b5-41e9-b238-f48203b0cc7f · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.381147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:0d0c64ee5834e521ce7ceef410c4b4c879e60f2b8582b61e0c97ad0267c6b924

Observation c77b11ea-bf67-4cfe-b28a-8fe3ed3e4a48 · inbound

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction cites this paper.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.995688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:a7ef985a119c9d2707b75f9f0020f07129ed6d2c9f814f1ecf1adf1c08dcfd83

Observation 7356ba04-1818-41c7-a294-66f14449f1a2 · inbound

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation cites this paper.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:02.748763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:02.748763Z digest=sha256:d6d7fb98b027f412d4957db9de47f3c8a30561dbc4f639b3b1e4aa4b2412a1cf

Observation 5d95c883-b5b7-4803-9ca9-0b08cdbbd3b0 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.933917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:569183b4d267c64e30c10ada864a2e9521d2a006d99c83b8c68e360868fbe609

Observation bdd27281-da2f-4e45-85c9-51cdef6f0107 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.255586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:cabd62766ca3af4896e08f2c7a930fd3f4aa7a569d679d4de02b049e54f37c10

Observation 76f94cf2-f665-450c-bbba-a82b44d337a1 · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.957257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.957257Z digest=sha256:ef36df010bd6a46ef5447717ea15311ac969bb14c15fcc6565ec5c1fa167bb02

Observation 899732fe-3962-44d9-9f9a-12075b47872f · inbound

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs cites this paper.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.556727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.556727Z digest=sha256:2ed1f5103a16007499e907753603d226be51fec7f9b20cddae421b84ab4baca4

Observation 521945c6-9181-4c18-bd95-6359df42064c · inbound

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos cites this paper.

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:36.281736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:36.281736Z digest=sha256:e758640be7486365d40df3c6679a814b78ba18be03ad832e84914f53533f9594

Observation 9ebd11f4-cb7d-4fc7-9870-6582a89258f8 · inbound

How Important are Videos for Training Video LLMs? cites this paper.

How Important are Videos for Training Video LLMs? Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.860377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.860377Z digest=sha256:d5e89fb899d5ef3a5d24ba07bcc017e929c0bea8f8b84669c10fa0bc5fca1dca

Observation f2c93c76-ca12-46eb-8d29-e800355ccd76 · inbound

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs cites this paper.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.692121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.692121Z digest=sha256:c7a3a7ddbc0fc7c0e1acabe0c54b93b81fc5b9a93ec41f88afbf5fed59ee87e9

Observation 98ea8d80-91eb-45a2-b744-d3f1e36301d0 · inbound

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning cites this paper.

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:32.189262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:32.189262Z digest=sha256:733ae3467d86031d9dea955f44acc1beee0a5078bf36a2cd26fdf970236355b5

Observation 04efda21-51c5-4fd5-9208-784a884467c3 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:29.684209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:29.684209Z digest=sha256:8d1fb4ee5c25337732ca663109cbe7c77e719fe8d299627ef87b41e9efe217da

Observation 9a7864cc-bbaa-41c0-bd7b-dca40e139fa4 · inbound

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams cites this paper.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:05.734732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:05.734732Z digest=sha256:97208d89108e8fa8c2060e2d34b367f09559e47c63895ccfa6b290002e23ad02

Observation 478d1767-13d5-4ab3-886f-a740cdfadc8e · inbound

Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis cites this paper.

Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:00.411216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:12:00.411216Z digest=sha256:f5885e1f397f48ceb193b70e7bea80c92f8aa259f17a0aa0176ec5f9626644eb

Observation f5616df8-bba1-4eb2-a19e-34372d632036 · inbound

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning cites this paper.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:07.016238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:f07ddd4f88400a2ca21e38bb8faaaf47f8245d05632df11ac0ff9254f029554f

Observation 28d07efb-3f67-463c-a1ed-a519d52b2c8c · inbound

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation cites this paper.

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:52.831200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:52.831200Z digest=sha256:3fb0acb55cff6bfed3187ef2272e3e80e8b377e30392bd8428f1312d1c0d818a

Observation ef22287e-6b45-409f-83b5-5da365626f32 · inbound

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding cites this paper.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:33.180292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:33.180292Z digest=sha256:ffff2f676f3a23935ba1b8080e7fe6e92296df80f859033f7108f5ada58eaaa4

Observation b95fbbae-b9c3-43f3-8cf9-7f2c29168dbd · inbound

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting cites this paper.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.696608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.696608Z digest=sha256:7f1509d7f43c404351750f0712c8db82f45c6667b57a2038536c753838005035

Observation 88aa53dc-00ca-49c3-b000-fd682ddfb9e3 · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:10.952463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:10.952463Z digest=sha256:a2210a89b35f6bedef00e2dc9faa97ade3b8869c907ce7f9e9626ef89a1bd13e

Observation 658ce853-647e-4c11-9558-ed2cef15393e · inbound

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark cites this paper.

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:07.966651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:07.966651Z digest=sha256:ec7c6c33483509749f954904aaa07edbb62722e7b9044cdba5f18ede58830627

Observation 782a660f-92ca-44c3-bae5-d9aa665a2db4 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.187761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:e9d13477f46c849ff88ec1ee9bc1f485676eb0179aeda4fbf020ae1197c3091e

Observation 8291684d-8999-4361-96c2-dda099eae3b5 · inbound

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models cites this paper.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.684255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.684255Z digest=sha256:d9e45a0fba92af4b5c6a81e7a6c83ae5abbe43add2a8ba132355b5bb57dde8ab

Observation 79b3cbe0-4f3b-4d46-b8a0-4f9896202924 · inbound

Vision-Language Memory for Spatial Reasoning cites this paper.

Vision-Language Memory for Spatial Reasoning Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:34.081170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:34.081170Z digest=sha256:163f995afcd48dd335b4c1df3533973df0233418e17dce7761ab6ec746bfb1e0

Observation 470f744a-a41c-44e2-9102-190dfa7b92dd · inbound

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding cites this paper.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.612461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:8fe59e00a4fb46988f61bdf447fa818d3c32a63e387db67602a811cb0ccb14a0

Observation 011af0e0-4d45-4789-8742-668caa3108ca · inbound

Streaming Video Instruction Tuning cites this paper.

Streaming Video Instruction Tuning Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:48:21.851540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T19:44:11.032898Z digest=sha256:25b96e643927d818cc83e0faecb911f04e3e9394d7bce0d14eddbc0dd084800f

Observation 30140cd9-f595-41b2-a384-a90470bebc5a · inbound

Social Caption: Evaluating Social Understanding in Multimodal Models cites this paper.

Social Caption: Evaluating Social Understanding in Multimodal Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T09:12:18.037151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:12:18.037151Z digest=sha256:610a3e6a131022f900e2d154c9903e549e593a1937e886f5e31bf10137de106c

Observation a45b1f45-aa83-4b36-b733-fce8c1a1d963 · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.467842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:576e546418820f9698146833e8e8659813e72efe169ee93b61842cfa02ca8373

Observation 147c87bf-0038-4bb9-9f12-5d276f99f031 · inbound

Stateful Token Reduction for Long-Video Hybrid VLMs cites this paper.

Stateful Token Reduction for Long-Video Hybrid VLMs Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:05.151986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:05.151986Z digest=sha256:24a3d133c2d958be044524c2f4ce121b4ed9198163c698070e8c3d12d992ea62

Observation d74697c3-69bb-4ccf-9ece-b3f48fee6c83 · inbound

3D-IDE: 3D Implicit Depth Emergent cites this paper.

3D-IDE: 3D Implicit Depth Emergent Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:38:11.440996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T22:34:04.833557Z digest=sha256:2e958f61d24d213bbe07fbf94c97c01c3213a7783703c091ee7462d645ab7193

Observation 6df034c9-fb8f-4325-b306-44d5a82d595b · inbound

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding cites this paper.

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.088181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T14:48:39.444933Z digest=sha256:d5ca87a32fa7eb7840756b28aac8d277c8af3f50f41e7acc71c2bc6694d2979c

Observation 6b37bb7d-5986-47b8-b93e-0126796083d6 · inbound

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding cites this paper.

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:05:56.390960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:57:42.822121Z digest=sha256:a3cb045c2a50aa1b64c57e9e2d1f284cc88e279c7e458088f5682d0c8f7bbacb

Observation b90458eb-4969-41dc-b8a5-c7ce70b27228 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:08.785397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T14:06:27.953376Z digest=sha256:dffbc4fd16d38ea954ed91e87aba4b7f0637ed867292d9946e0f89ac853c99fe

Observation 343d0087-eb2b-4d90-97ed-fb4a66216af0 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:50:51.275095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:49:41.654207Z digest=sha256:1eb1ead589d3043b304cf708afd3505dcf389bf5d2a87e84bf9e4758cf2f8eeb

Observation 06b676f5-4ff0-45b7-925e-51e09e45b1d5 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:28.052362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:35:59.553683Z digest=sha256:6a76f7e3b43a7c6ed2a1f86fe0f22c7024e7f8c0d765f9197dee84c0d96e148c

Observation c7e94571-41a5-4ef3-9a69-a6b7ab33be83 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:10:24.064572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T06:08:19.956833Z digest=sha256:d1687e60c4d5028ec61b51c334c0a099b14c18e51b6b909feab761768802ff32

Observation 45513e9d-92f6-4ba8-bc30-ada40750a6a2 · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.244738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:3e6ae21687e77f56d74d356ab1cffe939ee2b5fc656608d68f2d760f0ebf24c5

Observation 8e43fe16-f2a6-4541-ad1d-63cb1e5b1499 · inbound

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly cites this paper.

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:21:20.651921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T09:20:32.920925Z digest=sha256:f04e9ab51fd8b55792bfc175084fec82b20af5058b02fbe03522f40747e40dbe

Observation bf5a937f-eaf8-4524-924f-2ebead1d02e1 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 211

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.738324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:bfa7ce9668a2bddd9979b391e656152f356f80fe755107cb8e472009f2adb5de

Observation 9fc1f3ba-b6a6-45d6-87c6-c59c5d4c378d · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.596700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:cab01debad04412fa046413cfdfeb6056a75acf4948f5d2a93a75dc82a5445e5

Observation 687ae146-f056-42c1-bd57-d3b9df6a016b · inbound

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding cites this paper.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.842034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:6cd5daa9e5e849aa5ff079b4f6df5c339bb0387236b5d3209db7e79f4d1a1a8d

Observation 89a1a066-f3e4-47fc-9d6e-9f3370537c85 · inbound

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams cites this paper.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.170346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:182e4df0ae27d195efbb5738ea176a4c5a62bb384ce116aa77c15c9123406510

Observation d7608799-4708-4045-b623-9ca07ce51d91 · inbound

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution cites this paper.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:50.849452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T05:27:41.578243Z digest=sha256:40852aec64b1fc652a9455c8e8bc21da2ca452068c8fd3b2b1f987d5dbadde84

Observation 1fa8c5de-9e72-4954-ac08-d538123a1901 · inbound

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution cites this paper.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:39.749751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:98745bdecfdac4ad440dfeb5ce67b5e0b252ce5077620eb4b19da80b41d93c0e

Observation acfb088c-807c-40ce-a2ea-8892e3ff6b96 · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.233288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.233288Z digest=sha256:e0a626c2563a57b338159a27cd464dbce6ba140ee8b41346cf11f85722820294