Pith. sign in

Paper Citation Record · LEDGER

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

As of 15 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2608.02980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02980 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:27:29.029373Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact1
  • verified fuzzy43
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aeb6dbff-f449-4e49-b91a-9442bef094cc · outbound

This paper cites ReferIt3D: Neural Listen- ers for Fine-Grained 3D Object Identification in Real-World Scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding ReferIt3D: Neural Listen- ers for Fine-Grained 3D Object Identification in Real-World Scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.183649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.719317Z digest=sha256:45a68bbce7ff512bd5fef634e8e6e1a381283f2723eeb97e95a27c55fabc6216

Observation 1d6a1d8e-9086-4208-94a4-ffcdb6f39801 · outbound

This paper cites Llama 3: The llama-3 herd of models.https: //ai.meta.com/llama/, 2024.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Llama 3: The llama-3 herd of models.https: //ai.meta.com/llama/, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.165375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.724810Z digest=sha256:c782db86a113daf7888f1fa19f8a8181eda956e9a52c8ada530b673c0a519891

Observation 2705b699-df6a-4def-ba9e-90ca31a5b5a5 · outbound

This paper cites Locate 3d: Real-world ob- ject localization via self-supervised learning in 3d, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Locate 3d: Real-world ob- ject localization via self-supervised learning in 3d, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.148181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.729999Z digest=sha256:7578cbdadb7145e07a611828cb210393f64999b46b05880c467121fa14be1b6a

Observation bf16414f-d807-45ae-885d-41f2b0be43cc · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scanqa: 3d question answering for spatial scene understanding

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.129808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.734941Z digest=sha256:e6edda0ae661fe0b0aba1dcda4c029f0dbdabcf84e60fa4ff7a19b87421b4f2f

Observation a53fa677-6872-40d5-97c4-0fd172862f17 · outbound

This paper cites an unresolved cited work.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:27:30.112824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.740493Z digest=sha256:e36512c2f4f527ffd5b817fd58075306bbc6e0c499b720c1f7a85a6e30e9d15c

Observation 07d077cc-e702-40ff-b6d3-9e1d824d8e99 · outbound

This paper cites Token merging: Your ViT but faster.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Token merging: Your ViT but faster

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.095012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.745801Z digest=sha256:848520dcffe9f14b00b90c6e52fddc3461d28eb548117983362025bed179f297

Observation 0c2c2a72-6b94-49ed-823a-92146925c91d · outbound

This paper cites From thousands to billions: 3d visual language grounding via render-supervised distillation from 2d vlms, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding From thousands to billions: 3d visual language grounding via render-supervised distillation from 2d vlms, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.078294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.751420Z digest=sha256:8656628b12a91fb37ef13a38b368f1138c5698c2447a200a54d1aef66de53d9f

Observation 558c4219-0243-4c7c-a0e5-5490ec30a8e7 · outbound

This paper cites End- to-End Object Detection with Transformers.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding End- to-End Object Detection with Transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.059219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.756349Z digest=sha256:b3c52c7d79f96e76425eeda661ddcd9a5db22ecbdd6060fc817498c6efac7fc7

Observation ec62c190-df38-4cbf-8493-7fc304c59ed1 · outbound

This paper cites Matterport3D: Learning from RGB-D Data in Indoor Environments.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Matterport3D: Learning from RGB-D Data in Indoor Environments

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.761505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.761505Z digest=sha256:82f9121529871018e1aa7da5c198d36ad3a391b8b42a40c995b7d045f96739c5

Observation 9bb4ba43-8281-40f9-b5dc-512086bd13d1 · outbound

This paper cites ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.039282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.766679Z digest=sha256:23641f29c812630282c6b94f0f26385721c58418d4ff5dc6a10fa198742453e0

Observation ca779a4b-4832-461b-8f06-d871b5697ef7 · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.017462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.771739Z digest=sha256:886295245c0e76ce55348b01856830e25aa29668c26c85f11f7304e0c9b732e9

Observation 619f4d38-021e-4b34-a0e4-bcd3f7a9c974 · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Grounded 3D-LLM with Referent Tokens

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.776677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.776677Z digest=sha256:9a450c7eeb806c604c225ee483871b9caa2627bc03e813e134bc6258646e0ecf

Observation 7196c203-57d3-4550-bb01-9805a5c02bfb · outbound

This paper cites Schwing, Alexan- der Kirillov, and Rohit Girdhar.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Schwing, Alexan- der Kirillov, and Rohit Girdhar

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.996076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.782179Z digest=sha256:1a24295bc4604cde1de01c4b0c8af91c8606c92b5d5b85e84825ed4384e7f037

Observation ff2b045f-ce34-4731-8b73-0618518d85e3 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.979160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.787025Z digest=sha256:4134ac499feb516520c15713375a37c7d1da19ef3d118a0923d49e3f51a2919f

Observation e479c723-70a6-4685-a861-2e2752513adf · outbound

This paper cites Bundlefusion: Real-time globally consistent 3d reconstruction using on-the-fly surface re-integration.ACM Transactions on Graphics 2017 (TOG),.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Bundlefusion: Real-time globally consistent 3d reconstruction using on-the-fly surface re-integration.ACM Transactions on Graphics 2017 (TOG),

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.962201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.791806Z digest=sha256:8dd42672370b7f78d9317c5cc2b107d33fb3c134ca03c00ec22bc0aa3ab2d1ba

Observation 335a98b9-e5cf-4641-8d30-f44a2c5d3c9b · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.796534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.796534Z digest=sha256:8daffe0fde8f182499f759fab5c0277ad2b5e732866c03a2f9a1884cc0d8e964

Observation 9d43c921-b730-4e1a-a0e7-c6ff566238d8 · outbound

This paper cites 3d-llm: In- jecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3d-llm: In- jecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.802169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.802169Z digest=sha256:28831a8c1ca9f15c1b524dcefa72836b0bb1ffeaa5c2c8e754438996c71139f6

Observation 4bcbb5c7-fc2b-41e5-a0ee-43312f09a4b7 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.807289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.807289Z digest=sha256:1d9ba7a633f22b132e73105755fa0606e307e35c6f45a391abdedcbcdd10d0c8

Observation 0639deb9-e224-41af-832e-f8b83e7fa876 · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Chat-scene: Bridging 3d scene and large language models with object identifiers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.921345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.812043Z digest=sha256:b919bcc56230cf089c4c79d291d88a28f3ef58e05b63f48314b5534426953b6b

Observation 826b1c76-0114-4b09-a29b-1337a1ccb579 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding An Embodied Generalist Agent in 3D World

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.816998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.816998Z digest=sha256:7554ce982613e822f55878d59e77f50f4833d3fa17692d3ba5aa7c2ceb4f93e4

Observation da6b087d-9517-436b-b3b5-7d670b149057 · outbound

This paper cites Revisiting multimodal positional encoding in vision-language models, 2026.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Revisiting multimodal positional encoding in vision-language models, 2026

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.904945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.822679Z digest=sha256:2029dcf2dfa4234b4c057dedade1cd92b73d355895845045652bcb1e16a18b77

Observation dc93f922-558f-492f-95c5-56de9aac65c1 · outbound

This paper cites Reason3d: Searching and reasoning 3d segmentation via large language model.3DV, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Reason3d: Searching and reasoning 3d segmentation via large language model.3DV, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.887836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.827621Z digest=sha256:44b64e8b2d1889c1e3df70a60a28d351b32795bcb1c8d6dff0c6642013c2a971

Observation 6d18d69e-bbb6-403e-a26e-21a3aeeb3ec8 · outbound

This paper cites Bottom up top down detection transform- ers for language grounding in images and point clouds.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Bottom up top down detection transform- ers for language grounding in images and point clouds

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.869456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.832603Z digest=sha256:ecbe58dd2580fd4a55fbb4c5df4c5350dcf1e91b83899d1640147d8e224d1eeb

Observation 00114ae5-a656-48af-af38-86faf6e5ad96 · outbound

This paper cites Odin: A single model for 2d and 3d segmentation.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Odin: A single model for 2d and 3d segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.850869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.837664Z digest=sha256:b5b262148860f941129b6cc0207683338052284944c84aac33eb1e18cf9b12e6

Observation 1f15f975-3b82-4750-8445-9ef82ff9df72 · outbound

This paper cites Unifying 2d and 3d vision-language un- derstanding, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unifying 2d and 3d vision-language un- derstanding, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.832652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.842550Z digest=sha256:68b50550d6f9f1d5be8c1967322aee75338771cc2b6a9003ac8482cda248ff85

Observation 6594ca3a-8222-4c0b-9389-eafbf8bf2361 · outbound

This paper cites MDETR - Modulated Detection for End-to-End Multi-Modal Under- standing.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding MDETR - Modulated Detection for End-to-End Multi-Modal Under- standing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.814536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.847319Z digest=sha256:862246838065a0fcb511462e104d3b13cf7c89ee3775bbb8f2d7fd8d2eb53640

Observation 54add548-3d69-4b75-b0ed-14a7a4c44850 · outbound

This paper cites ReferItGame: Referring to objects in pho- tographs of natural scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding ReferItGame: Referring to objects in pho- tographs of natural scenes

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.797281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.852077Z digest=sha256:f40424f447e034a32d5aff441565bee518751eddfc5e9ab633151a69bb952e25

Observation 4df9d2ee-444b-4b06-8fcf-80a493fc90c8 · outbound

This paper cites Restr: Convolution-free referring image segmentation using transformers, 2022.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Restr: Convolution-free referring image segmentation using transformers, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.780428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.856930Z digest=sha256:23c93827e7e60f8bba349f086df8a1328856c302669aaeb566b767255a548b3a

Observation 7a97f72d-02bf-433a-9bf6-2406accc3ba8 · outbound

This paper cites Mask-attention-free transformer for 3d in- stance segmentation.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Mask-attention-free transformer for 3d in- stance segmentation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.862157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.862157Z digest=sha256:16e2ea1f228ea9635f6af2aa00aa8c8fdbde50e234ead8e16382077543ca9274

Observation 1fbc27d2-3930-4d67-b19e-c3dcb341710d · outbound

This paper cites Lisa: Reasoning segmenta- tion via large language model, 2024.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Lisa: Reasoning segmenta- tion via large language model, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.752413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.867247Z digest=sha256:9f53a75ce3d34322dae11748be5f8c517d29dab224ffc589dfae923f6109b646

Observation 2bc1063e-20fe-48e6-a533-28a0ba3899a5 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.872639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.872639Z digest=sha256:c88b666007c30b3fd8800630ce934336c4deb6353b1db677266fa6fe22bc46ca

Observation 990060e3-85a9-4452-b100-14f7a6892a78 · outbound

This paper cites Grounded language-image pre-training.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Grounded language-image pre-training

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.723467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.877707Z digest=sha256:fe74bf8e87d65ab807d18fc24648dabfaf21e034b1a5e811e116bb570e2b4470

Observation d6cdb1fe-4044-4bc4-84f5-41148dfb81d4 · outbound

This paper cites 3eed: Ground everything everywhere in 3d.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3eed: Ground everything everywhere in 3d

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.702798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.882564Z digest=sha256:d59d182458d9311beb1d07af64a9a7c37a8815cc07a0140c5eb28da999517287

Observation fff96ced-d235-4551-8a5b-736904880d32 · outbound

This paper cites Microsoft coco: Common objects in context.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Microsoft coco: Common objects in context

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.682245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.887595Z digest=sha256:65665ee94c546c2f098955947f6f2041ec00569b66b6627596dd45028722cc73

Observation 679bd081-9eab-4124-b646-e150e493fdae · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.662921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.893221Z digest=sha256:b80b1b646b009246694def14213ce420604f7e3057e2d04c686a809a39991bc0

Observation 86f5f433-2743-4c66-9336-3161c0ddaef9 · outbound

This paper cites View-on- graph: Zero-shot 3d visual grounding via vision-language reasoning on scene graphs.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding View-on- graph: Zero-shot 3d visual grounding via vision-language reasoning on scene graphs

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.646105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.898167Z digest=sha256:0e5af4398c74f38f94288b3789e8e2d15288987011be355a6c12eafeb3c099f4

Observation e7e3f48c-5888-47b3-99cf-02e104416982 · outbound

This paper cites 3d-sps: Single-stage 3d visual grounding via referred point progressive selection.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3d-sps: Single-stage 3d visual grounding via referred point progressive selection

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.627699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.903472Z digest=sha256:0801dad54d0b153c1025a309d31fe14b9b96ab8e33061e0afe20b686e747bc33

Observation eaa2cf8c-ab49-4f40-9e22-9dd632d5699f · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.908340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.908340Z digest=sha256:2fa27acec18f8211fa36ccc1ae3da82d5fe5fd1970cb2ec56474ac20ead047dd

Observation 9f5db08d-fb08-4d91-97e2-a62a87f4d3cf · outbound

This paper cites Goucher, Adam Perelman, Aditya Ramesh, and Aidan Clark et al.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Goucher, Adam Perelman, Aditya Ramesh, and Aidan Clark et al

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.610276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.913803Z digest=sha256:166258774af0724d145e6e631e496a4297026e2b55eee4eb53d27e3255a859fd

Observation 89d94f3b-0fd8-4117-84e4-7eb4187182f3 · outbound

This paper cites Languagerefer: Spatial-language model for 3d visual grounding.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Languagerefer: Spatial-language model for 3d visual grounding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.918804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.918804Z digest=sha256:a3980eea6868c95b537e94a9498f0ba5c5a57cb02518eee3b1d481b6d364b51b

Observation 804b1176-f87a-4488-82aa-f34154d5fdb3 · outbound

This paper cites Language- grounded indoor 3d semantic segmentation in the wild.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Language- grounded indoor 3d semantic segmentation in the wild

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.578334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.923771Z digest=sha256:56ff691c6838956292def0dc3caf3b16d1d9a9f406bf3cf2eb9b0a01354537e1

Observation e8181c31-468e-47a5-9aaa-621d58963501 · outbound

This paper cites Mask3d: Mask trans- former for 3d semantic instance segmentation.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Mask3d: Mask trans- former for 3d semantic instance segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.558961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.928579Z digest=sha256:b39dc6f7c7fef909e3843e78599d1a9b15eee0693c67ea297580b6a3a3c4bc59

Observation e85d7520-83ca-4c62-a919-c53c0a1894e0 · outbound

This paper cites Evaluating zero-shot gpt-4v performance on 3d vi- sual question answering benchmarks, 2024.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Evaluating zero-shot gpt-4v performance on 3d vi- sual question answering benchmarks, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.540015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.933509Z digest=sha256:08eaece9cdb18848d9cc1db7fb90691b60d122522aaa34d519ee3ada451d36ac

Observation fb31858b-2724-4037-a0f0-14fc4a0e0baa · outbound

This paper cites Hashimoto.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Hashimoto

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.521967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.938741Z digest=sha256:ab8c39ab4d69a6a82ec4d94fb454f8032ec1f085bb37fb943dfb4adc1a12fe45

Observation e7588881-b2a1-403d-981b-8ab81995c8c4 · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Gemini: A family of highly capable multimodal models, 2025

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.502906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.944625Z digest=sha256:c3671dca8d55aa71892a32d9e90a667719b0f183efd8346597bbd4047f35162c

Observation 00d5c209-b687-46ef-930b-6564ef3fe0c7 · outbound

This paper cites an unresolved cited work.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:27:29.486307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.949602Z digest=sha256:7dae5dcdbdbf0b0a377b6c9524b06cf5d4ac47052b9900a6e7576ceb26fc5600

Observation 40bfa89e-34f3-4e62-89db-af3f72fd4069 · outbound

This paper cites MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.954483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.954483Z digest=sha256:97b53a531fd75100c14383c8700f41cf0636ea313cfede55822d4494238ee7d9

Observation 9dfa92bb-95b8-4e8f-a1ee-e45f1e9792f5 · outbound

This paper cites Realworldqa.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Realworldqa

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.469572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.959712Z digest=sha256:739a63e3986379934ca1f103d21f62341348f2ed27fff3038b5ab26db8a45a98

Observation 6889bd48-7bd7-473d-ad12-834ff21fc8c2 · outbound

This paper cites Qwen2.5 Technical Report.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Qwen2.5 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.964720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.964720Z digest=sha256:ab540d7be14fe5e8fb44fb2d88080b790daa41535a42cf2934dd46dc335c9a3a

Observation 73561748-6326-40c7-8411-bbce61381ead · outbound

This paper cites Sat: 2d semantics assisted training for 3d visual grounding.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Sat: 2d semantics assisted training for 3d visual grounding

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.452483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.969841Z digest=sha256:3158c93e67d3c80973b5f2284c3459469861a0a64169a1f97898ffdd456ea750

Observation 08c79efa-76bc-4065-9f52-83fed13b6322 · outbound

This paper cites an unresolved cited work.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:27:29.435741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.974617Z digest=sha256:2840f7f468936ce4fc6b575bddf16352b9bf5cb436703692ffb3c88dc79830db

Observation 8e754b19-6361-4677-b4d2-9b35de80d58c · outbound

This paper cites Scannet++: A high-fidelity dataset of 3d in- door scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scannet++: A high-fidelity dataset of 3d in- door scenes

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.979475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.979475Z digest=sha256:9f4510234494dbbd0e754d2d11da39f9c70759650028f203686acb6340c20d5e

Observation 72dadb02-5187-48be-8e89-950c586ccd5d · outbound

This paper cites Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.408883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.984880Z digest=sha256:b49cf4c95212e90732a8cb4ae2d6250f041b777cf038593342df1756c7501bb3

Observation a4be2bb5-2751-4d74-a4b4-c5f4ac0c264a · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d ob- jects.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Multi3drefer: Grounding text description to multiple 3d ob- jects

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.391572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.989830Z digest=sha256:27098c47735b4d58e35b932e0af1b635ea7be0b41d58150f001d7543f9d5d1e7

Observation 7ba9a399-a68b-4a16-b07c-e8d9b19d776e · outbound

This paper cites Towards learning a generalist model for embod- ied navigation.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Towards learning a generalist model for embod- ied navigation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.371586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:28.994481Z digest=sha256:69ff5ae9f0d7a50cf24be0b99b2b3c7561837283e03152e98a717f5878cb0577

Observation 473c6c45-a62f-48ee-946f-aeb34c1cb832 · outbound

This paper cites Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625,.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.999121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.999121Z digest=sha256:e39790557f555e50db234d11930601291d7bea7b07c85c30c563ec67fd73767e

Observation 1d49da31-7603-4865-8b38-9e7ca9fa0cb9 · outbound

This paper cites Video-3d llm: Learning position-aware video representation for 3d scene understanding.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Video-3d llm: Learning position-aware video representation for 3d scene understanding

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.353145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:29.004503Z digest=sha256:c9919b6188e31ff0a6ba8e4b96ac7272b0a04ec7f6a2a146e09d4d2434b293e7

Observation bc1e7350-31e4-4b27-9f50-6a1eb06190e7 · outbound

This paper cites Llava-3d: A simple yet effective pathway to empowering lmms with 3d-awareness, 2024.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Llava-3d: A simple yet effective pathway to empowering lmms with 3d-awareness, 2024

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.336725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:29.009579Z digest=sha256:a38753d551fdc4132ac26426eb54667826f53164df89d7aa5905f34de8ac36fe

Observation aede57e3-b847-485b-8cae-c13d3ab279a2 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.320100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:29.014371Z digest=sha256:2e3d76bd63793658df9a00ce283445826ff3af417a260b541518f5156dd5a24f

Observation 87b2bc23-7840-435c-aa55-b96ad2a15573 · outbound

This paper cites Unifying 3D Vision-Language Understanding via Promptable Queries.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unifying 3D Vision-Language Understanding via Promptable Queries

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-08T04:27:29.079215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:29.019048Z digest=sha256:6110c3d2416a80975ea6598e7c61c155c3741edef42996638a07ed8220fc3dd8

Observation 5d76e898-d0c7-45e5-a8e8-862ac983013b · outbound

This paper cites Generalized decoding for pixel, image, and lan- guage.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Generalized decoding for pixel, image, and lan- guage

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.303684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:29.024405Z digest=sha256:0378dee649cc2ae7b421a8377f1b80a41675b8f09f5543e66a5f028f915b2908

Observation 78a29522-3b1a-456e-9cc1-91dc6ac60d28 · outbound

This paper cites an unresolved cited work.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work

Reference 62

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T04:27:29.286999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T04:27:29.029373Z digest=sha256:e8926838603c57dd9c586b2d8b09c30a704a94346689344e4d7599013250c5bf

Pith citing papers

No inbound Pith citation observations are available.