Pith. sign in

Paper Citation Record · LEDGER

Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 68 inbound Pith citation observations for arXiv:2409.12961.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.12961 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 68 of 68 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:16:27.002786Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:09:50.847114Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b423cda2-9aac-4e76-9042-e228b4f43f00 · inbound

VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation cites this paper.

VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T16:42:55.844186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:42:55.844186Z digest=sha256:a70f805a99c03faf75d2c5f9f06a9386dd66e7553203a21a4fb86c668d41bdec

Observation 564b53b9-86db-4229-8470-fbfe2bb65200 · inbound

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models cites this paper.

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:06.854087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:06.854087Z digest=sha256:3de106dd31043026ad4748b6c32981c8d2d1d2e6416ce7962d96004491f72f52

Observation e9ab469b-1570-4b36-a5ef-5d045d1d3b1d · inbound

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding cites this paper.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.660952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.660952Z digest=sha256:8eef9fbd4d40cf55c205cc821563241cf7d85948f8c1d5328c796071bcf6c2a3

Observation 03d4faa6-4cb6-4059-8a3c-72a8cc06bd6d · inbound

SEAL: Semantic Attention Learning for Long Video Representation cites this paper.

SEAL: Semantic Attention Learning for Long Video Representation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T01:00:10.590689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T01:00:10.590689Z digest=sha256:e408e6b929df985cfea6fc0713e013808b1f7f85a24530efe3e35dc6377fd47c

Observation 438170e9-4897-429d-99a3-811a3693fd8b · inbound

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding cites this paper.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.607317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.607317Z digest=sha256:9906907f7eb9f134aab33b981d2767bc69d5e4d50dffd7d61f139e4f30a904df

Observation 383b0178-ecc3-447c-a64e-0fcef2d8ff7e · inbound

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay cites this paper.

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T21:31:44.472046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:31:44.472046Z digest=sha256:bf552906021aec2e16ff33f819a35d82abce1d17a5f1d0e7aa02843e74b4d28b

Observation e0925f8d-efec-4110-8251-c6b1daa0f532 · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.235839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:ab782a3eeeb00421584d5576fa65c6bbb3bfcd15bb2c1b2eb7d4232cc6b7437c

Observation 945ebb97-a10c-4d9b-97af-c503e35b8d95 · inbound

LinVT: Empower Your Image-level Large Language Model to Understand Videos cites this paper.

LinVT: Empower Your Image-level Large Language Model to Understand Videos Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.295108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.295108Z digest=sha256:a00df24107ef042f0a42a15896b6c0bbabc1b311c7d6b0692728309aa949844c

Observation a0f969e2-d8f5-486e-bdf4-ec96e7c62893 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.747383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:72525e8bf6a68d08c9c32d1bd18bc78b202d1a142c22d1684f0a6aaf99c13717

Observation fb5ea9b2-ab62-4346-bb97-54565482d7fc · inbound

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models cites this paper.

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:42.528279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:42.528279Z digest=sha256:82a69bd8b717141911eda1786dd152eee6e5aa2bb85f9a03d81eb0df0ef95bea

Observation 58188ea2-4386-4301-b92b-2a63edf7b967 · inbound

Apollo: An Exploration of Video Understanding in Large Multimodal Models cites this paper.

Apollo: An Exploration of Video Understanding in Large Multimodal Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T16:11:10.569245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:11:10.569245Z digest=sha256:163d91b56f5b3b9648c28ae4390aa594c1a7c5e5d04772dc478286c961791be0

Observation aa778b62-8f11-4fac-8743-75aea584390e · inbound

Do Language Models Understand Time? cites this paper.

Do Language Models Understand Time? Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:17.574353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:17.574353Z digest=sha256:97bc77b7930404e6f6ca440ce6c5b67346f1acc3ca41f7b41de733068a2c4009

Observation b9d2e19d-f8bf-4f2d-b833-61b343c53b28 · inbound

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer cites this paper.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.752481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.752481Z digest=sha256:7f01448a33d930b0951371fab9e41ba84d715391115b7adead6d36105582c6fd

Observation 25818a3a-9cf6-4172-babf-0805141c6282 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:45:28.369891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:341e28e132acee2c8cb711f6337e541546a899d3e86cf5c0e67dcb6017c20e02

Observation 50435518-a08e-4d16-bc62-a55863b19e30 · inbound

Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives cites this paper.

Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T21:44:55.870643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:44:55.870643Z digest=sha256:5295b672ee09b05ccb58d300ed4aaa1fa2892c1fd1714d173d5f3b850eef38d9

Observation 96ae2de7-6e16-41b4-8353-9851ffc6d606 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:02:37.421236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:5d33d2081a01f5ad2e326099c7f3eb49b58a47b6c449f8db1f55b14252447153

Observation 5c2840ee-d5b0-4714-8645-3ddf4a95a080 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:19:59.722872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:989b1333965035d607e1d60554e34d5baf1fb742d5cb4b1cf4a4765dc3112768

Observation 83e485c8-b2cd-4efe-bf0c-123b9c521a66 · inbound

Temporal Preference Optimization for Long-Form Video Understanding cites this paper.

Temporal Preference Optimization for Long-Form Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.203011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.203011Z digest=sha256:1dcdff281337002a906967bd85926798edec18cc2eaa61678713a90350f98773

Observation f5d14e7f-84f7-4876-b759-c49fcb9189a5 · inbound

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding cites this paper.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.266841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.266841Z digest=sha256:644477928c57cc6db8484f38fc175da2c5f5b6c0daefb25cf3dcceada41e6668

Observation 0d28cb45-166f-4728-9025-ba01d2f44225 · inbound

Ola: Pushing the Frontiers of Omni-Modal Language Model cites this paper.

Ola: Pushing the Frontiers of Omni-Modal Language Model Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.207697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.207697Z digest=sha256:d4846bd356f8323120d2fe8e110701b44f27345135eb31f85d19328b8438e63b

Observation 78546f28-51b5-41e9-b238-f48203b0cc7f · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.381147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:1ddb024274145e9ba8a4e74c05fc01b2d2586b21d32ca840c4bf7710f1f41ee5

Observation e984bf2f-6b36-497b-a987-99e4c1bbab05 · inbound

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale cites this paper.

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:27.002786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:27.002786Z digest=sha256:232906c9d2d379209259956e9ce21e5a734570d6565c051ee8f23bc74a7a707e

Observation a73585c4-cc32-4249-9408-473bea2aa5c0 · inbound

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension cites this paper.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.549842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.549842Z digest=sha256:30c5b04debec3ea5d7e3d6e700c6e1fde41d9617d688ad2c64e7cf7b471f1f95

Observation c9a6248b-b904-44db-aff0-78109d8466af · inbound

VEU-Bench: Towards Comprehensive Understanding of Video Editing cites this paper.

VEU-Bench: Towards Comprehensive Understanding of Video Editing Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:29.819644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:29.819644Z digest=sha256:44d532e3eb7f2649fa35c6b2c00b6f7668efe52d0cd7f3755ddd50d015142211

Observation 85d49c75-0db9-4a8b-a2c5-6836a03b46f8 · inbound

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model cites this paper.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.059117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.059117Z digest=sha256:f19ff4523d5ee5ed145bee2dce844ab31a0063de68aa46aaca9592939dff1457

Observation c77b11ea-bf67-4cfe-b28a-8fe3ed3e4a48 · inbound

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction cites this paper.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.995688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:99d9e3c355a224d2cea01f29b12447cdee9cd076b83be4f485353f1cad336f68

Observation 7356ba04-1818-41c7-a294-66f14449f1a2 · inbound

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation cites this paper.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:02.748763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:02.748763Z digest=sha256:7aa8d9ab07f6877edfa2341712f5eb2dd09238d03f1a9cd032704f65acdead3e

Observation 5d95c883-b5b7-4803-9ca9-0b08cdbbd3b0 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.933917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:843b68d1dca698e5056aae76a84c82b2fe2f2bc0b6b10b363ada286af51e9144

Observation bdd27281-da2f-4e45-85c9-51cdef6f0107 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.255586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:0b23290a121c7263c31edf38f25b5270f90b02e84260602c910105d84a53a88b

Observation 76f94cf2-f665-450c-bbba-a82b44d337a1 · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.957257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.957257Z digest=sha256:59b4a67c5d7bfbaa2509f6577361466cca0afeebd3c06bb7c80e9fc9956111d5

Observation 899732fe-3962-44d9-9f9a-12075b47872f · inbound

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs cites this paper.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.556727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.556727Z digest=sha256:5fbd64c7c426d315149cfddf9e50375490f4ef15cffa5e9404b0570cdf7680a7

Observation 521945c6-9181-4c18-bd95-6359df42064c · inbound

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos cites this paper.

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:36.281736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:36.281736Z digest=sha256:8fced646e3d4466c920592b8bee6489c671fa869ff98f52db2230ebd5e7f6156

Observation 9ebd11f4-cb7d-4fc7-9870-6582a89258f8 · inbound

How Important are Videos for Training Video LLMs? cites this paper.

How Important are Videos for Training Video LLMs? Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.860377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.860377Z digest=sha256:9956ba1c9639712c4b111d42c431846158329cee94e93407a4bd96cb5fd4ea44

Observation f2c93c76-ca12-46eb-8d29-e800355ccd76 · inbound

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs cites this paper.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.692121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.692121Z digest=sha256:90e546d061e0dae0a39db314cbd468bb2a110b82e2e17b93bfb77360c100d6fc

Observation 98ea8d80-91eb-45a2-b744-d3f1e36301d0 · inbound

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning cites this paper.

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:32.189262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:32.189262Z digest=sha256:cd8524b1f7676379c05df8b5b6d7ea8555813418f555a54ccd1d569b84238125

Observation 04efda21-51c5-4fd5-9208-784a884467c3 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:29.684209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:29.684209Z digest=sha256:045859312ead5f4e0804ad30c6103be85b6e5a0e29458ec012dc61226a504253

Observation 9a7864cc-bbaa-41c0-bd7b-dca40e139fa4 · inbound

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams cites this paper.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:05.734732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:05.734732Z digest=sha256:85a629bd4811d8cefb6d76159645fbcf4637d251adcabaa40efcd1898ed656a2

Observation 478d1767-13d5-4ab3-886f-a740cdfadc8e · inbound

Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis cites this paper.

Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:00.411216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:12:00.411216Z digest=sha256:0b2e11c2108c8dbf7f9db984673e3042410a068dae2c46de438bae21ada3b51e

Observation f5616df8-bba1-4eb2-a19e-34372d632036 · inbound

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning cites this paper.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:07.016238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:e08dc821132e9c94f6c4d6990f41f1432d6e84552d97593c764fc9c3253541cd

Observation 28d07efb-3f67-463c-a1ed-a519d52b2c8c · inbound

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation cites this paper.

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:52.831200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:52.831200Z digest=sha256:42bde8a4202827546b18806ea01260e80f80b959b1381ec455e36582bf61a4f0

Observation ef22287e-6b45-409f-83b5-5da365626f32 · inbound

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding cites this paper.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:33.180292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:33.180292Z digest=sha256:f6a227368444332c3574be7ae80f851133236a47035c9d7424b2bf0c2cacc31d

Observation b95fbbae-b9c3-43f3-8cf9-7f2c29168dbd · inbound

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting cites this paper.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.696608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.696608Z digest=sha256:786282e1475f4312e37c25cb0c1040b8d5a9738a4a3eacf8acf0c84b9698e665

Observation 88aa53dc-00ca-49c3-b000-fd682ddfb9e3 · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:10.952463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:10.952463Z digest=sha256:9bcf482129988adf7c5bcdcb666754c96c3bb1ba8d3713ca198219ea967f08d2

Observation 658ce853-647e-4c11-9558-ed2cef15393e · inbound

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark cites this paper.

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:07.966651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:07.966651Z digest=sha256:cf86cd38620d3cf1d7625f3f70bba83928025a723b259d51574a44ac8f79abfb

Observation 782a660f-92ca-44c3-bae5-d9aa665a2db4 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.187761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:2c3fe932009753483f84c5c1736d9e7bb0be9b0185ab18a926706742ca3cc98f

Observation 8291684d-8999-4361-96c2-dda099eae3b5 · inbound

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models cites this paper.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.684255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.684255Z digest=sha256:9e3c0db7de3b137d82074c3d83ffe511d68b57b00b29c365c8410987f0f241bc

Observation 79b3cbe0-4f3b-4d46-b8a0-4f9896202924 · inbound

Vision-Language Memory for Spatial Reasoning cites this paper.

Vision-Language Memory for Spatial Reasoning Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:34.081170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:34.081170Z digest=sha256:a8204f36c5bff4163b4d7059e8d1d5965f7a50584555b03c5ecbd7f7b996a542

Observation 470f744a-a41c-44e2-9102-190dfa7b92dd · inbound

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding cites this paper.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.612461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:1706d13879b5922f0464baabf55753eaf282f08685b6b82c257120e96e6d6ac7

Observation 011af0e0-4d45-4789-8742-668caa3108ca · inbound

Streaming Video Instruction Tuning cites this paper.

Streaming Video Instruction Tuning Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:48:21.851540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T19:44:11.032898Z digest=sha256:527ec73dfa9f23f46bc002883d0c624b747afd1c56bd33c5d0acf8e01a7696dd

Observation 30140cd9-f595-41b2-a384-a90470bebc5a · inbound

Social Caption: Evaluating Social Understanding in Multimodal Models cites this paper.

Social Caption: Evaluating Social Understanding in Multimodal Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T09:12:18.037151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:12:18.037151Z digest=sha256:6dd7d0ad9940d5d620d97a5fc55d00e1dc8b4aa810d9884b203046747a6db7d5

Observation a45b1f45-aa83-4b36-b733-fce8c1a1d963 · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.467842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:e3189e1733fd2178be61b0de091c9bf83368adee001e6d801f365ac58de25cba

Observation 147c87bf-0038-4bb9-9f12-5d276f99f031 · inbound

Stateful Token Reduction for Long-Video Hybrid VLMs cites this paper.

Stateful Token Reduction for Long-Video Hybrid VLMs Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:05.151986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:05.151986Z digest=sha256:76a96a13f5d2dff58e988a0926559a0e9d5096318f07550052c24fee8e404bcd

Observation d74697c3-69bb-4ccf-9ece-b3f48fee6c83 · inbound

3D-IDE: 3D Implicit Depth Emergent cites this paper.

3D-IDE: 3D Implicit Depth Emergent Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:38:11.440996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T22:34:04.833557Z digest=sha256:728df7660de737e3f29ddd74b5d02dc8f44fde538dd4f960e59307acc3caa698

Observation 6df034c9-fb8f-4325-b306-44d5a82d595b · inbound

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding cites this paper.

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.088181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T14:48:39.444933Z digest=sha256:2ffef9d7605ca57a0303edddb61ba68a256986cb2e14f38fdec257d2c1f2b7b5

Observation 6b37bb7d-5986-47b8-b93e-0126796083d6 · inbound

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding cites this paper.

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:05:56.390960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T01:57:42.822121Z digest=sha256:24b5933082fc631f60a52ac82b9e4b60dcd1684c225858f7dbf5c1af92565fc8

Observation b90458eb-4969-41dc-b8a5-c7ce70b27228 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:08.785397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T14:06:27.953376Z digest=sha256:2e195c4b844984da041c8e1b827fa4365ae6a87425424e621bad4c3b8b9c4a5a

Observation 343d0087-eb2b-4d90-97ed-fb4a66216af0 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:50:51.275095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T01:49:41.654207Z digest=sha256:6ab6c091bafb70ba83b94b20cff595c1f9121be0881205cf6f4b8cef214550c7

Observation 06b676f5-4ff0-45b7-925e-51e09e45b1d5 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:28.052362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T03:35:59.553683Z digest=sha256:4f563c820ceb168c1c8944bedf747f84d578ebd9844f1eb55d14a293f5f6c917

Observation c7e94571-41a5-4ef3-9a69-a6b7ab33be83 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:10:24.064572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T06:08:19.956833Z digest=sha256:afbeaedb39e82cf883c04b71c2fddaf696084954d2ad3a2e9e85a36e19e5bdc2

Observation 45513e9d-92f6-4ba8-bc30-ada40750a6a2 · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.244738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:425df7ae7064c4b8c93787e12636bc1011a2b18cf73cd98e27ba71421e958b65

Observation 8e43fe16-f2a6-4541-ad1d-63cb1e5b1499 · inbound

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly cites this paper.

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:21:20.651921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T09:20:32.920925Z digest=sha256:f28e32ef85511eed942a2724d9faa139f3b78fb08f8f3f9b73778bb7a4d21a9a

Observation bf5a937f-eaf8-4524-924f-2ebead1d02e1 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 211

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.738324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:b649e9347a8c36810d8ee3f72f219bfa1f25c73a6e555cbd6453d7ed460852d9

Observation 9fc1f3ba-b6a6-45d6-87c6-c59c5d4c378d · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.596700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:69b8a016d95f60271eccbb7b580762a98f9e5462104635e15e5215c5ca3cca9c

Observation 687ae146-f056-42c1-bd57-d3b9df6a016b · inbound

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding cites this paper.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.842034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:3fb6e091ca6c6dcdd3ee5dd15f916caff07bd003e63635acb5d89376bebd668a

Observation 89a1a066-f3e4-47fc-9d6e-9f3370537c85 · inbound

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams cites this paper.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.170346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:dd05343752260b94adc3afd0c70f961ea39170fb275cce7faa9c0de4b148c304

Observation d7608799-4708-4045-b623-9ca07ce51d91 · inbound

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution cites this paper.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:50.849452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T05:27:41.578243Z digest=sha256:dd169f06d94f924690143121d19b6ca57ac2023a33f9b78513eed23b58e56b92

Observation 1fa8c5de-9e72-4954-ac08-d538123a1901 · inbound

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution cites this paper.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:39.749751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:9d3d62db033f4a99137352bd468eee64e3bf9c5c9b5000aabc2ba131094ddeec

Observation acfb088c-807c-40ce-a2ea-8892e3ff6b96 · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.233288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.233288Z digest=sha256:19fd90f60eaf3c792bf4e39be2c34287b7ad186a338386b81db0d1a5e467b5bc