Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:46:18.874616Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 2 inbound Pith citation observations for arXiv:2505.21079.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:46:18.874616Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:28:28.887051Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-17T22:15:22.050578Z
93 of 93 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation abe1037d-31f3-4496-9bcd-17cc3458f866 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Hierar- chical open-vocabulary 3d scene graphs for language-grounded robot navigation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e020615-500f-4fae-9edb-73a28b292d0c · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Sg-nav: Online 3d scene graph prompting for llm-based zero-shot object navigation
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66875d4e-66bb-4eda-9f02-e8ec46ad5fd4 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b088345-0db1-4df2-872f-155d2b587552 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Multi-modal data-efficient 3d scene understanding for autonomous driving
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fae8e8b4-7b35-46bf-9ab1-af8a33b52cd5 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Dme-driver: Integrating human decision logic and 3d scene perception in autonomous driving
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0106f186-2589-4936-9d5f-189914c03086 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Drivinggaus- sian: Composite gaussian splatting for surrounding dynamic autonomous driving scenes
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d46c641b-4510-45e9-98ae-738e371fd414 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Editable scene simulation for autonomous driving via collaborative llm-agents
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93596bd4-263d-4455-8adb-1f4d7365138f · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts How to enable llm with 3d capacity? a survey of spatial reasoning in llm
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53c7b625-754b-420c-a522-4ec004070c13 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Scenecraft: An llm agent for synthesizing 3d scenes as blender code
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8adc1b71-2718-4873-932d-b18a1222c50c · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55ce2427-d730-489e-881a-b95412075880 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Grounded 3D-LLM with Referent Tokens
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31d7a077-b571-46c9-ae5e-87c0b1ef5c3a · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Comp4D: LLM-Guided Compositional 4D Scene Generation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f3c1414-9e44-478f-9482-700086ccecaf · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80c3b379-7431-4a6f-a63a-360908578eb8 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f8a92f6-16a5-411d-929d-e212a398ab36 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Chat-scene: Bridging 3d scene and large language models with object identifiers
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59c48915-716a-45af-a6a2-7de31346977a · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31acac97-89dd-432e-855f-aef3ea07a7bd · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c01c3fa4-4a1b-4a66-967e-dce7171274df · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Scanqa: 3d question answering for spatial scene understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b4c3c6c-2eaf-420f-8276-b0542cabdb5d · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Scan2cap: Context-aware dense captioning in rgb-d scans
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c12676b-a1a9-41a5-b575-7c99e3bc2948 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts SQA3D: Situated Question Answering in 3D Scenes
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df72ca07-5be0-496b-8169-c3e3f170f8d9 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Pointllm: Empower- ing large language models to understand point clouds
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fb51303-b186-4311-81fd-ed0387cc064f · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc4dd6aa-075c-4139-82bc-ad69408057e7 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b8ae0e8-de17-4292-8b25-e5a4f5566296 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Objvariantensemble: Advancing point cloud llm evaluation in chal- lenging scenes with subtly distinguished objects
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 56a74583-e0c9-4f22-b7a9-585c3006e47b · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Gpt4point: A unified framework for point-language understanding and generation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6c9804c-9211-426a-9349-80c5d5df308d · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Unifying 3d vision-language understanding via promptable queries
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b2b6703b-7b06-483e-a62b-1b25f607e2a7 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Lidar-llm: Exploring the potential of large language models for 3d lidar understanding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2cbf34af-9941-4054-a2e0-7d34e679fe44 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d49dffac-20ce-4382-8ebd-b66ed401d2c1 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Liba: Language instructed multi-granularity bridge assistant for 3d visual grounding
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 36cd438f-6601-4aec-b961-09c236eeb544 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts 4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object Understanding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 71c8085a-3bc9-40f3-acc4-0a4af09b6440 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Space3D-Bench: Spatial 3D Question Answering Benchmark
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfa2a3c7-da90-4b0f-a8a3-343e810bad04 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b26fd254-490e-46e6-b794-35981a26791a · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58d23528-1edd-4483-a4f0-a17c2c013d05 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts 3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a9cbfaa-84a7-43f0-a20b-aa0e4d0e0786 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Sceneverse: Scaling 3d vision-language learning for grounded scene understanding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2fc252e-45b6-4916-941b-f3cca1a8531c · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Image as a foreign language: Beit pretraining for vision and vision-language tasks
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0abf9120-0f6b-4356-8a78-4edca9689486 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Uni3dl: A unified model for 3d vision- language understanding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c5ee0eaf-dda9-43a5-bfbd-d1e485fde89b · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Vision-language pre-training with object contrastive learning for 3d scene understanding
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cf7acec0-b0f4-4617-8774-0a8ffe20e460 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts When llms step into the 3d world: A survey and meta-analysis of 3d tasks via multi-modal large language models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 974b4501-a380-4b02-837b-c29974481829 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Mixture-of-experts with expert choice routing
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f4c6c5b7-6629-4ee2-a0fa-36051bc411c8 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts A Survey on Mixture of Experts in Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e6411fb-704c-43ea-b942-30c9f3e33ffb · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Revisiting MoE and Dense Speed-Accuracy Comparisons for LLM Training
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6dc91f9-957d-438f-8e46-7b78ec8f0fbc · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts ProMoE: Fast MoE-based LLM Serving using Proactive Caching
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f67659c6-7273-469e-8866-95eb01f56ba6 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6b45c4f-df13-47d0-8e59-5727b1ed5834 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b25f7791-3c29-4b06-a152-e42c32f6c2c6 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Scaling Vision-Language Models with Sparse Mixture of Experts
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44d5141c-8100-4e17-9fbf-3a6a0fe632e8 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a776cff8-8824-4d97-b82b-6dd8664a8bb9 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Ada-k routing: Boosting the efficiency of moe-based llms
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f745a5f-f77a-49d4-a003-4884aa126729 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd29c3b7-32ed-4da1-be2b-23990d2161ea · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Llama-moe: Building mixture-of-experts from llama with continual pre-training
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab168ca7-d01a-44f8-9e8c-3d866fc399aa · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Uni-moe: Scaling unified multimodal llms with mixture of experts
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ec264051-b15b-4672-ad89-4c4289841fd2 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts 3D-MoE: A Mixture-of-Experts Multi-modal LLM for 3D Vision and Pose Diffusion via Rectified Flow
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b39e2d56-872c-4d1e-953c-087506ebb4cf · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Minigpt-3d: Efficiently aligning 3d point clouds with large language models using 2d priors
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9d0e7d08-ea5f-4681-8d38-ac38b5b22f3c · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts DINOv2: Learning Robust Visual Features without Supervision
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9900a261-95af-4cf3-96b6-77ae5aa9c000 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Learning transferable visual models from natural language supervision
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a50b727f-c45f-49bd-a102-77120c739898 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dd7b16f-a899-4752-a93f-7586404e5a80 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Mask3d: Mask transformer for 3d semantic instance segmentation
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4fc414e-856d-4cc6-be66-fcf3f8a21dc4 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Scannet: Richly-annotated 3d reconstructions of indoor scenes
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e62bd513-2842-4d28-9f94-4c3d807f284d · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Scanrefer: 3d object localization in rgb-d scans using natural language
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f138a48c-270f-4626-abab-eb9a6e134db9 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Multi3drefer: Grounding text description to multiple 3d objects
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 41cd88d9-8585-4d71-b565-5a69a56e54f6 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8504db6e-dd20-4c52-a550-76c5b5c5afae · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40031675-6144-4b68-a221-1d878f6e210a · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Bleu: a method for automatic evaluation of machine translation
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee138cc5-f44e-4883-ba44-1dbf94c066b0 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38de929c-ac86-41a9-9fc0-e410279ffcab · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Rouge: A package for automatic evaluation of summaries
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d724260-e5ea-4175-8c9d-ca14fe7e6c5a · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Cider: Consensus-based image description evaluation
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fab14a4f-c41b-4382-9de3-42f4c6b02548 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Context-aware alignment and mutual masking for 3d-language pre-training
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a461d938-a693-48f4-b239-95aa8b05a8b0 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts 3d-vista: Pre-trained transformer for 3d vision and text alignment
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a11f7e75-7d53-4f32-8fc8-0fb0710c46c6 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d9e9fbe-b150-4716-a71a-1cf7ec52cda5 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 123279b9-48f2-448e-98d8-9f2e0c7aed8f · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12c600fc-3b43-4561-a1ad-6e931994fde7 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fde67467-ac28-41b9-8fff-4ab2b418bd52 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts 3d-llm: Injecting the 3d world into large language models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1d2de7e-81d4-459f-9cda-97874fd4f2d5 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b896d076-9b48-4ca8-b16b-c194431e4349 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 59aa549d-4755-4f82-9b60-f94b16bdc241 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts An Embodied Generalist Agent in 3D World
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20c4dede-de2a-4823-9071-1ec931fa9b06 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Principal components analysis (pca)
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5b455741-a976-4bd0-921d-51a0d97dbf53 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4f9e6749-68c8-4694-aa86-00fe14be757c · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts End-to-end 3d dense captioning with vote2cap-detr
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 407a0d21-38d9-435a-b446-64c737fc95eb · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts X-trans2cap: Cross-modal knowledge transfer using transformer for 3d dense captioning
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8cdb7196-3760-4303-a22a-0f95b1a3480d · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts MVT: Multi-view Vision Transformer for 3D Object Recognition
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7438619f-e002-4e53-97e7-4dff0e2a6e5b · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts 3dvg-transformer: Relation modeling for visual grounding on point clouds
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3f666bd-9bc5-431a-83a9-3e99b6aa9ec4 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Language conditioned spatial relation reasoning for 3d object grounding
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1a8eeff-6005-43f8-a02d-35249b7e51d2 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Less is more: Clipbert for video-and-language learning via sparse sampling
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d502179a-f70d-436b-beb5-3d0aed48b469 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Text-guided graph neural networks for referring 3d instance segmentation
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fb3fdcde-5000-4ef4-b0d5-4f987f3f9322 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts In- stancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 19ee3c74-9ecb-4d4d-a1cb-2f1b8e355a85 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts 3d-sps: Single-stage 3d visual grounding via referred point progressive selection
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9595cc4c-1e4c-4f43-b63d-3fcfb8f4460a · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts D3net: A speaker-listener architecture for semi-supervised dense captioning and visual grounding in rgb-d scans
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e307ec8a-2d47-4424-a4bc-f77ca43fa1c3 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Bottom up top down detection transformers for language grounding in images and point clouds
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7a183bd7-3cec-47e5-9641-643389b45c1f · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Learning Point-Language Hierarchical Alignment for 3D Visual Grounding
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75982f4f-e0f7-44c8-aaef-8312fa98d3bc · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts 3DRP-Net: 3D Relative Position-aware Network for 3D Visual Grounding
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fa273ce-e506-408c-8999-0442de193403 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Eda: Explicit text-decoupling and dense alignment for 3d visual grounding
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bdfbdf3c-1840-450a-abc8-6f0425535a05 · outbound
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts the bed, which is rectangular in shape, is located adjacent to the door
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5b933114-d814-4c29-94b1-37285e3b4ebb · inbound
DoReMi: Bridging 3D Domains via Topology-Aware Domain-Representation Mixture of Experts Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 15f2ecf4-f69a-407a-b325-57b9f5e01fa5 · inbound
SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.