Pith. sign in

Paper Citation Record · LEDGER

Aria: An Open Multimodal Native Mixture-of-Experts Model

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 49 inbound Pith citation observations for arXiv:2410.05993.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.05993 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 49 of 49 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:08:55.016786Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:49:39.565893Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 63d04862-f1a2-4fe4-9032-23211b9dec67 · inbound

Large Language Model-Brained GUI Agents: A Survey cites this paper.

Large Language Model-Brained GUI Agents: A Survey Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 220

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:08:27.787101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T11:08:27.472508Z digest=sha256:ecbadf7600cd495cabfb4926509adf1818568d593de6b2ae8f05eb336cfdbd96

Observation 6ae4961c-ef96-458c-9340-fd90ed497a19 · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:09:23.665883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:fdc94c45ac3f0954d2c44c4dc22b2696451c64087db1d0056d0f4f5dae954610

Observation 7f7d654c-d3c0-4c4c-a437-1618750f0f72 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:45:28.351105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:94aeea3a0bd0f3e96d6a52ce7ada7075c8ffa3fdb5178eb1d3af92d862c195b4

Observation 760a6262-90f3-4eb9-8b22-62e863cfc59e · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:39:22.513505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:1fdb1ada8408a3f8081389b0a070531662fcf36310c357c862503bc23296729d

Observation c5f32c4e-ddd6-47aa-8090-8408a836f07f · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.254763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:624c63b484844c8d0713056ac3b05aa58dd880b9cbad614d318cb7c79af3f3c7

Observation 8659e668-aea5-46f6-af9e-87a2a11fa040 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.720298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.720298Z digest=sha256:a9abf8c6bd184a0b61f2ba8c16d3c10ed277742e86d00292760b19fecefd7a22

Observation 8821343c-b7a2-4f6e-811a-0b539bf251f3 · inbound

TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments cites this paper.

TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:15.480907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:46:15.480907Z digest=sha256:8227457f9fbeab532e17e0bbc4d74ab30c6020a387f834f097423804dccec3d1

Observation 7e7399e0-866e-4842-a160-4c0336e41f05 · inbound

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation cites this paper.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:02.870434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:02.870434Z digest=sha256:5bcbe6936662e62ae3f28c31f3aeff40d6e9fca1b341de4de52997dbb40117c9

Observation 21cc2dfd-dfd4-44a1-b1c5-5e5e52635d85 · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.178112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.178112Z digest=sha256:1732f7e89dfab3edc3781f437e69508ef6d40ed587f6414b42c732783d80d14c

Observation 264fb417-cc39-45a4-a218-e19a0cc9fa8d · inbound

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding cites this paper.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.321429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.321429Z digest=sha256:90a6b227c77d2cef32fbacd36e3d7acbf41e98b42ac2e22f1b982f36f52feccf

Observation 956df445-02d3-43c9-ba84-4aa92a14787f · inbound

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning cites this paper.

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:14.699207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:14.699207Z digest=sha256:84382f0dfc56e22a408fcc5010d673d61eb16c290053551c791d6fd513e8cb89

Observation 5aa974a3-aacc-4d02-a416-d2cbbb1fc859 · inbound

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos cites this paper.

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:36.269793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:36.269793Z digest=sha256:e8d1a75205bf3d978ead0a94873c8dc528b107bee93d53849ad72f7abb3a6e3b

Observation 61a95390-c468-4b95-992e-c6f6636ede24 · inbound

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding cites this paper.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.895831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.895831Z digest=sha256:a0a54cf8ecee09bbdf9d7f49684cd1ae725706eff90854bf037c4e62dfb58d01

Observation f1bce6f4-2bd2-43e1-b54c-0a8033a0e85c · inbound

SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities cites this paper.

SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:18.685096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:18.685096Z digest=sha256:8d4051296ca69cf476989934fb23ad61f5edd18546648f5c192436189e8de174

Observation b92b4578-1639-44c9-8c4e-488ab0c0a9d5 · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:47.142398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:47.142398Z digest=sha256:6e7b8980bb2ec122966d29ce6498ce13b8db34a5399fba9cd59de33f6c604af2

Observation bdeedada-c8a1-4732-8f62-8830251ce4c6 · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.888224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.888224Z digest=sha256:e116842eff79358bbcd91a8a75844f88a5fe1e86ee230a9f502721ab47afdfe1

Observation 03a9cf9c-51af-453e-893c-cbad769da65f · inbound

SeqPE: Transformer with Sequential Position Encoding cites this paper.

SeqPE: Transformer with Sequential Position Encoding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:03.582120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:03.582120Z digest=sha256:ec30ddd1f19a7126ce67f85f89b069e107228020ed7c38a15b00a896f2e42f87

Observation d45607f7-0eeb-4f4a-9271-0dce9d0007ae · inbound

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training cites this paper.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:03.076564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:03.076564Z digest=sha256:fd63aee3f6f7c3ac57fdb4bf741f4fe969d8494d1cc04bbdad827e58d7bbe9b7

Observation e9d3eefd-96c8-4e8b-ad0a-a00b44db0101 · inbound

MMSearch-R1: Incentivizing LMMs to Search cites this paper.

MMSearch-R1: Incentivizing LMMs to Search Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:27:04.381238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T15:27:04.228144Z digest=sha256:59cf7cb8ba5960784d766fa1bdcd9709dfa1c04809a8f039b34bc9297fb872e6

Observation 84e3669f-09fa-4e89-bf38-aecdd5b59f7a · inbound

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering cites this paper.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.651138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.651138Z digest=sha256:cca305a22575c12cb4c33652397c5333d7f989c2730c2512749540431eaba5cb

Observation eead32a8-c627-48ac-b545-2cbbaea125dc · inbound

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs cites this paper.

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:03.532425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:03.532425Z digest=sha256:aa45d0d70476fbb4601955662412d79f0724fa910912d1785cbc4b074a325a72

Observation 81842245-3584-42af-815f-d44d8e3626c9 · inbound

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization cites this paper.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.062315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.062315Z digest=sha256:1ac7b0b7a0c57fdee2083e1d290d4065f17f51d4a10c4004e1ba700105d36093

Observation 67cf1236-7fa1-4118-8cbe-6711b0775924 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.704165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.704165Z digest=sha256:a47200e66ee00037e10076fdeb913d40f3cd58e378b62bb898cca6c0febe8aa4

Observation 9e98d311-ebb1-4b89-86f5-713a3c31b6e6 · inbound

Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis cites this paper.

Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:00.379646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:12:00.379646Z digest=sha256:45578cb3c7212d04ec8f79e41ac1c47441d47d8f9aba57ed722bd3210feccfa9

Observation a3e25417-5864-4888-a920-35b88ceee824 · inbound

VIP: Visual Information Protection through Adversarial Attacks on Vision-Language Models cites this paper.

VIP: Visual Information Protection through Adversarial Attacks on Vision-Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:28.576320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:16:28.576320Z digest=sha256:4a40adb9ec0c1b35d33d58336a92dbaf0a8bf097d8af4d8e474c4698035f4d52

Observation 5512c647-879c-48fd-b611-5cd6e6d1ced4 · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.827141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.827141Z digest=sha256:451fe7ecaea41469fea0e94002298a43aefa76e03ee284f3efc6e088a3a01e8c

Observation 25639214-899c-434e-956e-1357c905573c · inbound

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding cites this paper.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:35.065340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:35.065340Z digest=sha256:89d55567da6cc6af3a8a3744763ad14984314b9190fa90163ba001f2b35e5027

Observation 27fc9e3e-772a-4fb5-8274-f53805fd91b9 · inbound

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth cites this paper.

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T05:15:42.093183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:15:42.093183Z digest=sha256:8ddb83f60e9998e69b91bb69eaf368f59aa97f1ebea8152e7cf2ecd38dd4a129

Observation 967cacbc-4c71-454c-96d2-8db4b59de9d3 · inbound

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models cites this paper.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:35.830990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:35.830990Z digest=sha256:ed5a9c3bee8ffc317fcfef999c88811b10586d333254c98369b344f388601995

Observation 525d0f11-90c8-43dc-9609-5537117c5b0f · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:10.590321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:10.590321Z digest=sha256:0cd1c497f6e6e87de183b5911151c07e070251db7b1b76df724511d95b39bcab

Observation 748abe4c-ecc3-4ce2-9b2c-bce63cc6da9b · inbound

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation cites this paper.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:46.702453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:46.702453Z digest=sha256:b2cb447465cbece956deab163cce92465847ba67b464c95bb01faa438aab2f94

Observation 4decabbf-88ca-4b45-9f26-4fc0e6e38fec · inbound

Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search cites this paper.

Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:17:55.677423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T01:17:55.500268Z digest=sha256:87b10dc3be44aa46e6d3553a1d754d0eff9d704ee2b0e0ce08e597a9a335dcb8

Observation a133ce72-f350-4e46-b0ab-30c9f0976478 · inbound

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models cites this paper.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.586319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.586319Z digest=sha256:570328d3a6b9e2444bc59bb1153eb17b3cc1f280e1505511d31084660de7e83c

Observation f84a2837-f958-4a6f-b14a-661a3cf0974f · inbound

InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning cites this paper.

InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T17:46:56.236585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:46:56.236585Z digest=sha256:754cccf1c68f0bbab2cdd5083e50b42637d987193df8b10aded53b8542f1d241

Observation 6a12fce7-745b-41bb-aa36-4dfe8bc923bf · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.431023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:3ea6503995e5b8cc8a2c430cdf3b46aa0eb4b57a6ca5f64bc80f80dbe67921e2

Observation 20fe5682-f600-44dc-a371-8c6b83588456 · inbound

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG cites this paper.

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:50.236502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:15:02.124035Z digest=sha256:cd4e4b6571a5314fee416d8f70f82ed1fc0dc0120ccaed283d8430a31b4efeb9

Observation cf2041fa-7470-46f5-a5cf-576a911d4d48 · inbound

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents cites this paper.

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:31:08.078362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T03:21:30.732925Z digest=sha256:96c0cf20010916f1ae149a9e911571fb708d564bc5479cbdebb14674edf6499c

Observation 207c8141-ec26-4d31-a20d-2de2f3daa16c · inbound

ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards cites this paper.

ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:41:10.087841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T01:12:17.469552Z digest=sha256:bb45fe736bd4b9a9bb493a48e6cc7454e6319937e1941b79fadee8f8d4287cbe

Observation 8c1a6100-ea69-4415-b286-51f3fb6d7d16 · inbound

OProver: A Unified Framework for Agentic Formal Theorem Proving cites this paper.

OProver: A Unified Framework for Agentic Formal Theorem Proving Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:48:23.469466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T14:43:46.517807Z digest=sha256:70904e867e00a500658e358a7685f0dd31288eb0fab63170f739f7f9e83c6cd7

Observation f31884c4-788a-4c9a-a17b-e330fee1ded4 · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.644704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:bd7a7e6277a7ab689a50e8383ab5fdcdcbf4b2d3667d543c21468dadd8d29b7e

Observation 75332c6c-2cce-4d14-97c2-1754d2c6854d · inbound

MobileMoE: Scaling On-Device Mixture of Experts cites this paper.

MobileMoE: Scaling On-Device Mixture of Experts Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:53:51.431424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T18:48:50.656971Z digest=sha256:0329685709959e153de0c1500319c72b775ad19c065f854527c25f9a5b8f425e

Observation 36825118-971b-4ca7-b7a7-0ee70fab0eb3 · inbound

Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models cites this paper.

Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:26:00.060381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T22:43:33.929871Z digest=sha256:2ad8405068f705f8f16ecf15bdcb00c0058b8b66119c30ff01e99026606700a3

Observation 78e9bca3-f6e7-40d1-9b2c-92288425112e · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 291

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:28:32.081900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:803d46bcf9e7a9992f70e16af38f77fb01ab1711329f09eb2f163c80b066ddc8

Observation c308a88b-ff1d-4416-a119-d94f25a1d378 · inbound

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models cites this paper.

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 150

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:30.343170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T17:37:11.371892Z digest=sha256:b360ce8e2e99870f740d075656f173d4ab6b082d10baead59afcbc5c8b9a2e49

Observation 11cbdc20-f013-4cad-adfc-486ee162293d · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:37.708094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:1d1826e12f3627f17ee3afadfd819026e6d6ea68af61e400282bf35bd8885f59

Observation 834e8346-8de0-4ee5-8cd5-66661c36e10b · inbound

WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware cites this paper.

WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:49:39.567377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T12:40:53.103233Z digest=sha256:52279c7b56db00f87dd5e946cfeea74dfd3f36758d2f532be4079ec60219735a

Observation af36094b-e3bd-4a0e-94ea-65fee3fc4322 · inbound

Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models cites this paper.

Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:36.918993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:51:36.918993Z digest=sha256:77f44fc4ccdea1268e8b682773609760d49ce905b92c695ab59334117ee144e5

Observation 73900df1-0d7c-4843-b1c9-72d742d9c704 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:27.943328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:27.943328Z digest=sha256:4408d90a8b0d8bf8646e52de6b15072b799ad07a26c5dca86e6745af0f6ac2e7

Observation 81818fbe-04cf-4da4-b587-3226a56f54e5 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:55.016786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:55.016786Z digest=sha256:1eda1d04bf1e04fbd70fcd29f545020dc555f9e4e03706b026add70a7c721691