Pith. sign in

Paper Citation Record · LEDGER

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

As of 19 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2608.02980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02980 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:27:29.029373Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact1
  • verified fuzzy43
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aeb6dbff-f449-4e49-b91a-9442bef094cc · outbound

This paper cites ReferIt3D: Neural Listen- ers for Fine-Grained 3D Object Identification in Real-World Scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding ReferIt3D: Neural Listen- ers for Fine-Grained 3D Object Identification in Real-World Scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.183649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.719317Z digest=sha256:876e5a9580feef10010371855291a2f82987acaa00c65844cd145928de40390b

Observation 1d6a1d8e-9086-4208-94a4-ffcdb6f39801 · outbound

This paper cites Llama 3: The llama-3 herd of models.https: //ai.meta.com/llama/, 2024.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Llama 3: The llama-3 herd of models.https: //ai.meta.com/llama/, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.165375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.724810Z digest=sha256:e2ed931552f09876408a7474c14201a09feacab343a91d958acd60a38ecf5799

Observation 2705b699-df6a-4def-ba9e-90ca31a5b5a5 · outbound

This paper cites Locate 3d: Real-world ob- ject localization via self-supervised learning in 3d, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Locate 3d: Real-world ob- ject localization via self-supervised learning in 3d, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.148181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.729999Z digest=sha256:9570982c653b5f761944254f364b7af85a876956881de61469538eb4039587b7

Observation bf16414f-d807-45ae-885d-41f2b0be43cc · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scanqa: 3d question answering for spatial scene understanding

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.129808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.734941Z digest=sha256:ce25c00c3409846867d954781ccd0645b2c2c94406b24528be955dc61f43749d

Observation a53fa677-6872-40d5-97c4-0fd172862f17 · outbound

This paper cites an unresolved cited work.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:27:30.112824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.740493Z digest=sha256:2d0e3c7bb69f5ca4ee7cc732312258bbbb0faac1aeda33fb10b0549979e8c22a

Observation 07d077cc-e702-40ff-b6d3-9e1d824d8e99 · outbound

This paper cites Token merging: Your ViT but faster.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Token merging: Your ViT but faster

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.095012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.745801Z digest=sha256:4940513d8b90973e33d3ad7384facd76166ae102547902315d23d9db8688244e

Observation 0c2c2a72-6b94-49ed-823a-92146925c91d · outbound

This paper cites From thousands to billions: 3d visual language grounding via render-supervised distillation from 2d vlms, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding From thousands to billions: 3d visual language grounding via render-supervised distillation from 2d vlms, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.078294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.751420Z digest=sha256:c0d6a46c1a60c29fae07b42c72725eff2c08fb61d04565699f0e23bf2cc22fd6

Observation 558c4219-0243-4c7c-a0e5-5490ec30a8e7 · outbound

This paper cites End- to-End Object Detection with Transformers.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding End- to-End Object Detection with Transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.059219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.756349Z digest=sha256:ed169ce4f67ecd5498c3052fbeb7e965ae55d9dd29b0acee7f13723e0bb978c1

Observation ec62c190-df38-4cbf-8493-7fc304c59ed1 · outbound

This paper cites Matterport3D: Learning from RGB-D Data in Indoor Environments.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Matterport3D: Learning from RGB-D Data in Indoor Environments

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.761505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.761505Z digest=sha256:b43498621a12440a7d300f773b199c39492ce6ee62564962d06cd685fd87f796

Observation 9bb4ba43-8281-40f9-b5dc-512086bd13d1 · outbound

This paper cites ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.039282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.766679Z digest=sha256:261c0fc02bc3116b8186506e20b74b3cbdf18b24cdcda0adbdc1aa0dca7e7916

Observation ca779a4b-4832-461b-8f06-d871b5697ef7 · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:30.017462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.771739Z digest=sha256:b90aeba0d227c636c444721f0b053aa7db40af942cb8e00946ff9f3e2e4332ec

Observation 619f4d38-021e-4b34-a0e4-bcd3f7a9c974 · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Grounded 3D-LLM with Referent Tokens

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.776677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.776677Z digest=sha256:0219249836c986ba90bb012be19feca2bf1c644c4969d1b61ba49f51271c59f1

Observation 7196c203-57d3-4550-bb01-9805a5c02bfb · outbound

This paper cites Schwing, Alexan- der Kirillov, and Rohit Girdhar.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Schwing, Alexan- der Kirillov, and Rohit Girdhar

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.996076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.782179Z digest=sha256:e5f9d0114398889fb2a983ec2dd14bbe4c9e5638306f036639f1678e4589107a

Observation ff2b045f-ce34-4731-8b73-0618518d85e3 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.979160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.787025Z digest=sha256:c82ebf7c94f8b3ecab6ff22a3cd47e4013acc9fb37a379fff9304fdb48e4986e

Observation e479c723-70a6-4685-a861-2e2752513adf · outbound

This paper cites Bundlefusion: Real-time globally consistent 3d reconstruction using on-the-fly surface re-integration.ACM Transactions on Graphics 2017 (TOG),.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Bundlefusion: Real-time globally consistent 3d reconstruction using on-the-fly surface re-integration.ACM Transactions on Graphics 2017 (TOG),

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.962201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.791806Z digest=sha256:24e228b248fc64f3e282433ec126a9ca416bafc2bb9c7cf6e58a1daf146386fb

Observation 335a98b9-e5cf-4641-8d30-f44a2c5d3c9b · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.796534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.796534Z digest=sha256:69020298c32a23f30fa958c52233c9e054d45ff69d90693e012e0a9c6b0f8ca5

Observation 9d43c921-b730-4e1a-a0e7-c6ff566238d8 · outbound

This paper cites 3d-llm: In- jecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3d-llm: In- jecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.802169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.802169Z digest=sha256:444cd25def9e8c9f700b15b719836d33ab7b5c181515defa91c5e55335595f49

Observation 4bcbb5c7-fc2b-41e5-a0ee-43312f09a4b7 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.807289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.807289Z digest=sha256:2ca0b2bc42962d788568c3b9290ed22f38e50905f0b0cd1f7c4c0f222156d8a9

Observation 0639deb9-e224-41af-832e-f8b83e7fa876 · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Chat-scene: Bridging 3d scene and large language models with object identifiers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.921345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.812043Z digest=sha256:45f925f57a72bf8d829f8bb13cb1d85c9c019c143533e2273dde2dd80083742b

Observation 826b1c76-0114-4b09-a29b-1337a1ccb579 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding An Embodied Generalist Agent in 3D World

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.816998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.816998Z digest=sha256:f743a37231c3e8eebe963024543e3d21464bd12edd0ef91b0af78cd888f6ea0e

Observation da6b087d-9517-436b-b3b5-7d670b149057 · outbound

This paper cites Revisiting multimodal positional encoding in vision-language models, 2026.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Revisiting multimodal positional encoding in vision-language models, 2026

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.904945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.822679Z digest=sha256:9d6fce562da76b5cee361bff55577d3627f71f45431f805de92405485e2e2979

Observation dc93f922-558f-492f-95c5-56de9aac65c1 · outbound

This paper cites Reason3d: Searching and reasoning 3d segmentation via large language model.3DV, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Reason3d: Searching and reasoning 3d segmentation via large language model.3DV, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.887836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.827621Z digest=sha256:95ce69bfdd217c9122de7097a193e538882542c71d0c13f36f5f815c4659046d

Observation 6d18d69e-bbb6-403e-a26e-21a3aeeb3ec8 · outbound

This paper cites Bottom up top down detection transform- ers for language grounding in images and point clouds.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Bottom up top down detection transform- ers for language grounding in images and point clouds

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.869456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.832603Z digest=sha256:3c1736faaabc534fab735bf414e24485a61e2af317522e989be330daa7aff338

Observation 00114ae5-a656-48af-af38-86faf6e5ad96 · outbound

This paper cites Odin: A single model for 2d and 3d segmentation.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Odin: A single model for 2d and 3d segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.850869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.837664Z digest=sha256:164ddfaa54eed6c6bde628d534dc1c04b5f360d84548ba451f3f44791065f5f6

Observation 1f15f975-3b82-4750-8445-9ef82ff9df72 · outbound

This paper cites Unifying 2d and 3d vision-language un- derstanding, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unifying 2d and 3d vision-language un- derstanding, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.832652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.842550Z digest=sha256:69335ee5a984a5a90a1b86dccbb4ec1472abf85598d8fb15d9a195dddf8d9c7f

Observation 6594ca3a-8222-4c0b-9389-eafbf8bf2361 · outbound

This paper cites MDETR - Modulated Detection for End-to-End Multi-Modal Under- standing.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding MDETR - Modulated Detection for End-to-End Multi-Modal Under- standing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.814536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.847319Z digest=sha256:98182daab673332e486d89d451900f99e7200e227010dc5258e87c8305a10c9b

Observation 54add548-3d69-4b75-b0ed-14a7a4c44850 · outbound

This paper cites ReferItGame: Referring to objects in pho- tographs of natural scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding ReferItGame: Referring to objects in pho- tographs of natural scenes

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.797281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.852077Z digest=sha256:711a4e838c151af35113b4eabdb42a1431f88bb7730a2df293db7ed3b0a19183

Observation 4df9d2ee-444b-4b06-8fcf-80a493fc90c8 · outbound

This paper cites Restr: Convolution-free referring image segmentation using transformers, 2022.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Restr: Convolution-free referring image segmentation using transformers, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.780428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.856930Z digest=sha256:4c82320dafbd7990ff092d5a4eebb2286bd0bcf6b58e428f1a9d4a815e010aba

Observation 7a97f72d-02bf-433a-9bf6-2406accc3ba8 · outbound

This paper cites Mask-attention-free transformer for 3d in- stance segmentation.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Mask-attention-free transformer for 3d in- stance segmentation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.862157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.862157Z digest=sha256:4a3bb88e34b8dee1d2d96702be691ed7d63a4ef222d5f279efb6f7457ca837aa

Observation 1fbc27d2-3930-4d67-b19e-c3dcb341710d · outbound

This paper cites Lisa: Reasoning segmenta- tion via large language model, 2024.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Lisa: Reasoning segmenta- tion via large language model, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.752413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.867247Z digest=sha256:ce1c9178a18f878cf1ba47a98d4b91647871b5c29a3a0af99a988864eb64fa88

Observation 2bc1063e-20fe-48e6-a533-28a0ba3899a5 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.872639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.872639Z digest=sha256:774e69dac448bcb786eec7d3e9a8eedb0a5bdaefe36375b35dfefb0534d55f98

Observation 990060e3-85a9-4452-b100-14f7a6892a78 · outbound

This paper cites Grounded language-image pre-training.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Grounded language-image pre-training

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.723467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.877707Z digest=sha256:ba1d90c4885ffb6fe3d14d1f728f22c266142defe283b47ef25a67ce2b3141d2

Observation d6cdb1fe-4044-4bc4-84f5-41148dfb81d4 · outbound

This paper cites 3eed: Ground everything everywhere in 3d.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3eed: Ground everything everywhere in 3d

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.702798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.882564Z digest=sha256:1dcf0acbd0bd522d9fc7105c86e79891758bfde9dce6b001f9dcd73201554e4c

Observation fff96ced-d235-4551-8a5b-736904880d32 · outbound

This paper cites Microsoft coco: Common objects in context.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Microsoft coco: Common objects in context

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.682245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.887595Z digest=sha256:652f04468aa947bac8cd02f91ea190eb75b3be6279e8a745ac18dc0fa30733a1

Observation 679bd081-9eab-4124-b646-e150e493fdae · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.662921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.893221Z digest=sha256:b1819bd825ff0d5483260a01eaf2010123bd00438d8a91ddb6b234fbd9fabf5b

Observation 86f5f433-2743-4c66-9336-3161c0ddaef9 · outbound

This paper cites View-on- graph: Zero-shot 3d visual grounding via vision-language reasoning on scene graphs.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding View-on- graph: Zero-shot 3d visual grounding via vision-language reasoning on scene graphs

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.646105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.898167Z digest=sha256:4500fceaa30735dc70c0a805b514d9d9ee2008c54fbc1113bcc62153261f021b

Observation e7e3f48c-5888-47b3-99cf-02e104416982 · outbound

This paper cites 3d-sps: Single-stage 3d visual grounding via referred point progressive selection.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3d-sps: Single-stage 3d visual grounding via referred point progressive selection

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.627699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.903472Z digest=sha256:5046febe1377b8cab94afef1575f1787d47dd5f4be302228b134e12c1b38a749

Observation eaa2cf8c-ab49-4f40-9e22-9dd632d5699f · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.908340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.908340Z digest=sha256:440d2931804509a6a388011047f924ec140aa078f66d21ee615802ef60605d3f

Observation 9f5db08d-fb08-4d91-97e2-a62a87f4d3cf · outbound

This paper cites Goucher, Adam Perelman, Aditya Ramesh, and Aidan Clark et al.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Goucher, Adam Perelman, Aditya Ramesh, and Aidan Clark et al

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.610276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.913803Z digest=sha256:a9cabe4660cae557764a733c81c2aea7d3346a0a0fb0fee16266d821113f88a3

Observation 89d94f3b-0fd8-4117-84e4-7eb4187182f3 · outbound

This paper cites Languagerefer: Spatial-language model for 3d visual grounding.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Languagerefer: Spatial-language model for 3d visual grounding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.918804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.918804Z digest=sha256:f0e9f1514b82a3197817e936565c094c3dbf6afde3ab34c0671a42ea0ddf2eb3

Observation 804b1176-f87a-4488-82aa-f34154d5fdb3 · outbound

This paper cites Language- grounded indoor 3d semantic segmentation in the wild.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Language- grounded indoor 3d semantic segmentation in the wild

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.578334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.923771Z digest=sha256:6535c8e4646df3083f7657d95f2f8b493d16c73bff90b9ac7bd2194ee1dcce82

Observation e8181c31-468e-47a5-9aaa-621d58963501 · outbound

This paper cites Mask3d: Mask trans- former for 3d semantic instance segmentation.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Mask3d: Mask trans- former for 3d semantic instance segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.558961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.928579Z digest=sha256:0bffc356ff488c6169d24998eb1076dddc3d73f4f7f7cba2d8ba8651079da8d6

Observation e85d7520-83ca-4c62-a919-c53c0a1894e0 · outbound

This paper cites Evaluating zero-shot gpt-4v performance on 3d vi- sual question answering benchmarks, 2024.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Evaluating zero-shot gpt-4v performance on 3d vi- sual question answering benchmarks, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.540015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.933509Z digest=sha256:91c27b688f63fc923095fa5d5356a1dbb523dbd9c26e964af7b32e51a193b4d1

Observation fb31858b-2724-4037-a0f0-14fc4a0e0baa · outbound

This paper cites Hashimoto.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Hashimoto

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.521967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.938741Z digest=sha256:b2445ddae1e0e80082c193f179cc7d015f142c4c6dc2e1f75bfdf59097849720

Observation e7588881-b2a1-403d-981b-8ab81995c8c4 · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2025.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Gemini: A family of highly capable multimodal models, 2025

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.502906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.944625Z digest=sha256:b0316697d7b096670d883510b55f196c4c18024b62b66797a6bf95e5e7759ee0

Observation 00d5c209-b687-46ef-930b-6564ef3fe0c7 · outbound

This paper cites an unresolved cited work.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:27:29.486307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.949602Z digest=sha256:a40335e3b71a87a870225dc161eae5d1c856e12dfb62e527769c6b73befc8e8f

Observation 40bfa89e-34f3-4e62-89db-af3f72fd4069 · outbound

This paper cites MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.954483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.954483Z digest=sha256:fdb6a1f0e651d0fea871dbef085309c36a1ed8dd54c96b3e2115d4eaaac9d71f

Observation 9dfa92bb-95b8-4e8f-a1ee-e45f1e9792f5 · outbound

This paper cites Realworldqa.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Realworldqa

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.469572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.959712Z digest=sha256:ce0fcd98ab77177ddd00ae957eedea62f97a1c14d5374c6def326c1d59c71b44

Observation 6889bd48-7bd7-473d-ad12-834ff21fc8c2 · outbound

This paper cites Qwen2.5 Technical Report.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Qwen2.5 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.964720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.964720Z digest=sha256:8c4063944b65ebf7d65b34c33f1186d20efc98d74aaead225f5b45cf4dcd34c7

Observation 73561748-6326-40c7-8411-bbce61381ead · outbound

This paper cites Sat: 2d semantics assisted training for 3d visual grounding.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Sat: 2d semantics assisted training for 3d visual grounding

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.452483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.969841Z digest=sha256:d8afcf4a49d5caa22e1c3618a1addf5314a2f2f1a16d44a7735c3fa2feea2b0f

Observation 08c79efa-76bc-4065-9f52-83fed13b6322 · outbound

This paper cites an unresolved cited work.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:27:29.435741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.974617Z digest=sha256:9924240de98176c726c86216db3f4bef03fbca7736f0588772c4f7a12b2050bb

Observation 8e754b19-6361-4677-b4d2-9b35de80d58c · outbound

This paper cites Scannet++: A high-fidelity dataset of 3d in- door scenes.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scannet++: A high-fidelity dataset of 3d in- door scenes

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.979475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.979475Z digest=sha256:8691339a0d3f9dc46b1760c9d42fefb043d97f63f8a25f4a53548a7c264abeb9

Observation 72dadb02-5187-48be-8e89-950c586ccd5d · outbound

This paper cites Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.408883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.984880Z digest=sha256:c078b2e273a2199e9fd704cf34625ae93fcdea3ccefb88d0344c2c6dd6c3a826

Observation a4be2bb5-2751-4d74-a4b4-c5f4ac0c264a · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d ob- jects.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Multi3drefer: Grounding text description to multiple 3d ob- jects

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.391572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.989830Z digest=sha256:6860db7376a20b3b44787ee7d08878a6942f2f6c75fe37281b3ceee6d0f1e392

Observation 7ba9a399-a68b-4a16-b07c-e8d9b19d776e · outbound

This paper cites Towards learning a generalist model for embod- ied navigation.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Towards learning a generalist model for embod- ied navigation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.371586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:28.994481Z digest=sha256:5a3a74a0385b60b8dd296511fd5318c7adf555611803ff653a2eff4189f72360

Observation 473c6c45-a62f-48ee-946f-aeb34c1cb832 · outbound

This paper cites Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625,.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.999121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.999121Z digest=sha256:333f2c28e2da5a2dc73f75406022a0b06936207febccc260af5d2171924f4c95

Observation 1d49da31-7603-4865-8b38-9e7ca9fa0cb9 · outbound

This paper cites Video-3d llm: Learning position-aware video representation for 3d scene understanding.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Video-3d llm: Learning position-aware video representation for 3d scene understanding

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.353145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:29.004503Z digest=sha256:5aa2dd7af045362a0e6ee1a66b2b992c15590aa9b574fe07014af98da0cd0e61

Observation bc1e7350-31e4-4b27-9f50-6a1eb06190e7 · outbound

This paper cites Llava-3d: A simple yet effective pathway to empowering lmms with 3d-awareness, 2024.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Llava-3d: A simple yet effective pathway to empowering lmms with 3d-awareness, 2024

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.336725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:29.009579Z digest=sha256:b3706869b5a79ebf369f768639fb8c86a923d5e69827d8e49687f066dd3bb1c3

Observation aede57e3-b847-485b-8cae-c13d3ab279a2 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.320100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:29.014371Z digest=sha256:c3cf2bf6e648a263d7d8778eab09f71e61104b7200d91a061ac47c98755dafc6

Observation 87b2bc23-7840-435c-aa55-b96ad2a15573 · outbound

This paper cites Unifying 3D Vision-Language Understanding via Promptable Queries.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unifying 3D Vision-Language Understanding via Promptable Queries

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-08T04:27:29.079215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:29.019048Z digest=sha256:63aee232ba677938782dc4f246b5336eaeb50cd56ee970a3a3d7d7500271b636

Observation 5d76e898-d0c7-45e5-a8e8-862ac983013b · outbound

This paper cites Generalized decoding for pixel, image, and lan- guage.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Generalized decoding for pixel, image, and lan- guage

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:27:29.303684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:29.024405Z digest=sha256:6bf2477f8d771398c39cff18ff65bfe460f5a4bfdefe74393206b6a713cdc30a

Observation 78a29522-3b1a-456e-9cc1-91dc6ac60d28 · outbound

This paper cites an unresolved cited work.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work

Reference 62

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T04:27:29.286999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:27:29.029373Z digest=sha256:3e6529947621aa4c6941867db1cf9920484e4b27805314a3c0722a76da5a394f

Pith citing papers

No inbound Pith citation observations are available.