Pith. sign in

Paper Citation Record · LEDGER

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering

As of 9 August 2026, this Paper Citation Record lists 100 of 132 outbound references and 3 inbound Pith citation observations for arXiv:2502.00342.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00342 v2

Coverage vector

measured 100 of 132 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:24:18.126915Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:46:14.341180Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T20:30:11.608435Z

Reference resolution

100 of 132 outbound references displayed

  • verified exact2
  • verified fuzzy22
  • unresolved76
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 70d23fbd-d842-4de4-a65b-46e7eb36c2a1 · outbound

This paper cites an unresolved cited work.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.183064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.183064Z digest=sha256:b0dbea39bfae80d67679ab9c299de6773e6c4f3b84df12d9dbdd513c0772a9e9

Observation 54b6e8f3-17c4-40a4-a222-f3ffc0d03c45 · outbound

This paper cites an unresolved cited work.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.194824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.194824Z digest=sha256:9201213e9f32442876acfb9ab46cb5dd5bc04843a3ac89f7d19e4f322382af7e

Observation 942209e0-0db5-4783-b096-2471d72e78bd · outbound

This paper cites an unresolved cited work.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.200835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.200835Z digest=sha256:d0f67cab0fb04474f448344bf7f1f39324b5d0681b1b571bf11c708506b98a7f

Observation 62c83fb3-28f7-49ff-9c7f-7258bbb05153 · outbound

This paper cites Vqa: Visual question answering, in: Proceedings of the IEEE international conference on computer vision, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Vqa: Visual question answering, in: Proceedings of the IEEE international conference on computer vision, pp

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.206308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.206308Z digest=sha256:16aea3788a334ea37824bb98f0bf0da633d49c6be41379d35eedb96d5fd82fc4

Observation 95ae9a82-bc37-4ade-b610-17a73c773f54 · outbound

This paper cites Circle: Capture in rich contextual environments, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Circle: Capture in rich contextual environments, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.211674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.211674Z digest=sha256:e8a1a2b43746849a7ab167bd4dfb63a8ad000fdf4103e5b4dfa93d04b962531e

Observation 5d9e4135-bd41-472d-9de3-37310d84dd14 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding, in: proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Scanqa: 3d question answering for spatial scene understanding, in: proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pp

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.217404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.217404Z digest=sha256:483a9f3ddfd9652607dd4627fcbddafc9b8b9bfa4bff54dc8a58429a462b8aec

Observation d0828396-d8b8-43c9-9f28-b485ea0f4e6c · outbound

This paper cites Surgical-vqla++: Adversarial contrastive learning for calibrated robust visual question-localized answering in robotic surgery.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Surgical-vqla++: Adversarial contrastive learning for calibrated robust visual question-localized answering in robotic surgery

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.222917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.222917Z digest=sha256:a096a60aa732437187fcfd793ac3e4eb9ec8702ac785ba8d68c490c7ce65919c

Observation 58c7dd4a-8cac-415f-ab02-8d02dc927460 · outbound

This paper cites an unresolved cited work.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.228472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.228472Z digest=sha256:932ea83baaf067c2f2e157e17d027b1694905fbd1493322db29a88c5154e0fe1

Observation 1872f29e-ae7a-44aa-8de5-9f8d2e4590c4 · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.251847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.251847Z digest=sha256:ae84240767b79069aa318f435da8f62385324cabe9edffb4d1a8eac8fe5b151a

Observation 36245d22-d728-4b3a-8da3-bfe70f1a3c7f · outbound

This paper cites Models for multiparty engagement in open-worlddialog,in:ProceedingsoftheSIGDIAL2009conference, the 10th annual meeting of the special interest group on discourse and dialogue, p.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Models for multiparty engagement in open-worlddialog,in:ProceedingsoftheSIGDIAL2009conference, the 10th annual meeting of the special interest group on discourse and dialogue, p

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.280301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.280301Z digest=sha256:39afc3aeeba8614a240a84ad0ab13053167a32c7bec0741b0e19654beb3b6d4b

Observation d1fcc83f-a521-49f7-832c-13153b8991ad · outbound

This paper cites Language Models are Few-Shot Learners.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Language Models are Few-Shot Learners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.328533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.328533Z digest=sha256:ad045846e00cf92e1fd729ab3e5b5f4bb0393d56af4de32a06fe922238c48a1b

Observation 55af83d5-ae2f-4339-bf0a-4ef3e1216bd1 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language, in: European conference on computer vision, Springer.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Scanrefer: 3d object localization in rgb-d scans using natural language, in: European conference on computer vision, Springer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.351180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.351180Z digest=sha256:ac70225dd5a722deee435c4b73784b91bf5a9a12b5d2d7e515793d4389ff6e59

Observation a2df58f7-2892-4d84-940c-b79308b9d3ff · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3dunderstandingreasoningandplanning,in:Proceedingsofthe IEEE/CVFConferenceonComputerVisionandPatternRecognition, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Ll3da: Visual interactive instruction tuning for omni-3dunderstandingreasoningandplanning,in:Proceedingsofthe IEEE/CVFConferenceonComputerVisionandPatternRecognition, pp

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.356602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.356602Z digest=sha256:73af6fcad3ea50cf3cb01acf7ab3dc616b313aa4efbc9e9c32ad71bf65a00a7c

Observation 460349e1-c418-4f34-b4b0-100fea5d91dc · outbound

This paper cites Lan- guage conditioned spatial relation reasoning for 3d object grounding.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Lan- guage conditioned spatial relation reasoning for 3d object grounding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.361165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.361165Z digest=sha256:08abea1d01d03b2dcc89452f9b32bea8dcc27847f80221e85dd4dfc0ca49fabe

Observation b5baedf3-7d99-4864-997e-cb71562b3929 · outbound

This paper cites End-to- end 3d dense captioning with vote2cap-detr, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering End-to- end 3d dense captioning with vote2cap-detr, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.366419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.366419Z digest=sha256:ac8b935386a7703688ee9c971dc2a69d0b7909d3f6842ba690332b71404bab3a

Observation 399a2083-abf1-4944-9d15-e54f4930ca75 · outbound

This paper cites Scan2cap: Context-awaredensecaptioninginrgb-dscans,in:Proceedingsofthe IEEE/CVF conference on computer vision and pattern recognition, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Scan2cap: Context-awaredensecaptioninginrgb-dscans,in:Proceedingsofthe IEEE/CVF conference on computer vision and pattern recognition, pp

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.371311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.371311Z digest=sha256:23033434fde4deae73ce71c0012c997e8d3b6fe674c3cbc1f1a73fd9d8b01d8f

Observation 8d1917cd-0b08-455d-96f4-afe8c01f069f · outbound

This paper cites Vicuna: An open-sourcechatbotimpressinggpt-4with90%*chatgptquality.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Vicuna: An open-sourcechatbotimpressinggpt-4with90%*chatgptquality

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.376771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.376771Z digest=sha256:9b2e32bd8ea512ab6dc64b433bdbd804aeebed87faf6298232d2764fcd6acfde

Observation 9376e65e-9f53-47f3-b08c-96dbb702b48a · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Scannet: Richly-annotated 3d reconstructions of indoor scenes, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.381656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.381656Z digest=sha256:6cec92c0627030b381b6b73373bdae52450b0a8ba73d4db73c0045f83870d29a

Observation 08f626b6-ed15-4e81-8005-8f7587b644ec · outbound

This paper cites an unresolved cited work.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.387211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.387211Z digest=sha256:673034993e98023dd5ee74b61217c03399e0a4a19457bb2bfd287c09bba77553

Observation 03f42869-94da-4364-ac63-8d3ea5716903 · outbound

This paper cites an unresolved cited work.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.397068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.397068Z digest=sha256:9ed2671c4204583a3195ccd4ce13a7f1734f700e9f8a6356428aace59870f8b3

Observation 3594b23b-b802-4e45-826e-28eac4984134 · outbound

This paper cites Multi-CLIP: Contrastive Vision-Language Pre-training for Question Answering tasks in 3D Scenes.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Multi-CLIP: Contrastive Vision-Language Pre-training for Question Answering tasks in 3D Scenes

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.408439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.408439Z digest=sha256:9668575a58692f9e3b57e905d11d02d742f8c1a97372cca28dd206b9110ea307

Observation 4f38fffc-246b-4034-9930-77a253c8c086 · outbound

This paper cites an unresolved cited work.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.414063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.414063Z digest=sha256:f854dbce63cc832f563d693c9a380b97f0576ac3ded2de04d8fb1ebebe8c8681

Observation 5c27add2-0984-4c01-836d-dc3e635fcf39 · outbound

This paper cites 13142–13153.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering 13142–13153

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.402751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.402751Z digest=sha256:46d342e93d068d160f7f1a750573449fec9bee9021bec0faf0ea03bd865b335c

Observation 29c94ed9-1f01-4783-acf8-2ce0d30e9fba · outbound

This paper cites 1billion: A large-scale benchmark for general object grasping.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering 1billion: A large-scale benchmark for general object grasping

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.424306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.424306Z digest=sha256:d589a8b6a151ef346f2988e0e1a2061caa2c2891a5f27065605ad7a387ff9747

Observation 19ef5905-3ec6-477c-a1eb-36f824b37983 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.429259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.429259Z digest=sha256:ecd35d1257a1fd1eb723c28f954c6843ab969b88f5e9870e412586945a5f37ab

Observation f0bff6a9-dfff-438d-a643-3a3fde851008 · outbound

This paper cites 3dvqa: Visual question answering for 3d environments, in: 2022 19th Conference on Robots and Vision (CRV), pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering 3dvqa: Visual question answering for 3d environments, in: 2022 19th Conference on Robots and Vision (CRV), pp

Reference 26

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T19:24:19.418916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:17.419316Z digest=sha256:a8957eb213a8c97d6ad30e216e4dc9c0b0bcbfb14930f0c124ca6fc4564fe8f0

Observation f967d846-56d6-4966-987f-8f151ee4f9ea · outbound

This paper cites Long short-term memory.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Long short-term memory

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.439965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.439965Z digest=sha256:0257e905ee0b0cbd1cc6e5a47779051147ae546656b191bfff5f61558d4407b6

Observation a049acd6-2159-4f38-abfa-8f88fe71d9a6 · outbound

This paper cites Deep learning for 3d point clouds: A survey.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Deep learning for 3d point clouds: A survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.444583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.444583Z digest=sha256:45a18a248670a804793499e4d7bbe9ebc3b44c8c910c62a955f4625b9433c947

Observation 05ae10cf-944b-4b54-9c80-f102398baa2b · outbound

This paper cites Iqa: Visual question answering in interactive environments, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Iqa: Visual question answering in interactive environments, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.435118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.435118Z digest=sha256:904b6c607bb344b1b42d22ac37574e8d5a5a3a1f67bd17a91a10a7b77e530c41

Observation b0911f5d-1bdb-4156-939e-1d6a3ac76104 · outbound

This paper cites Stochastic scene-aware motion prediction, in: ProceedingsoftheIEEE/CVFInternationalConferenceonComputer Vision, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Stochastic scene-aware motion prediction, in: ProceedingsoftheIEEE/CVFInternationalConferenceonComputer Vision, pp

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.454260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.454260Z digest=sha256:8ed38191270b98d2434ca84c64ceb858f6b7386e330087fea76518296fdae5f9

Observation 810e9e81-9094-42d8-a5aa-f81905b1ebb5 · outbound

This paper cites Resolving 3dhumanposeambiguitieswith3dsceneconstraints,in:Proceedings of the IEEE/CVF international conference on computer vision, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Resolving 3dhumanposeambiguitieswith3dsceneconstraints,in:Proceedings of the IEEE/CVF international conference on computer vision, pp

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.459249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.459249Z digest=sha256:aadfc80dda23caffeabdd76af0dffac77df5e7840d8a7f794ad5080dc995a144

Observation e12b5f64-5bca-4850-9abe-a2c708a73eb6 · outbound

This paper cites A review of algorithms for filtering the 3d point cloud.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering A review of algorithms for filtering the 3d point cloud

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.449717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.449717Z digest=sha256:2e176e02751833f519982e05250c7362d95d26d5536b218d5bfa132cfb46a442

Observation 8bb0bd13-9232-4514-b069-22ccf4c855a3 · outbound

This paper cites Mask r-cnn, in: ProceedingsoftheIEEEinternationalconferenceoncomputervision, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Mask r-cnn, in: ProceedingsoftheIEEEinternationalconferenceoncomputervision, pp

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.492801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.492801Z digest=sha256:3e46110fe36569a0610bf11317ffea78e6e478cc51025937cbd9fd49b86db041

Observation 95692027-649a-465d-9170-a4fc77d6099f · outbound

This paper cites Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.519144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.519144Z digest=sha256:c1f54fd09ff2bd7391b20e1c7bfc20757255abc70f892159d2eeee95f3066587

Observation 738f1f3f-8a2a-4c51-9b9a-5eaf850f2bfe · outbound

This paper cites Transrefer3d: Entity-and-relation aware transformer for fine-grained 3d visual grounding, in: Proceedings of the 29th ACM International Conference on Multimedia, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Transrefer3d: Entity-and-relation aware transformer for fine-grained 3d visual grounding, in: Proceedings of the 29th ACM International Conference on Multimedia, pp

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.464174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.464174Z digest=sha256:6c269715225b65a6c5ede8f670b18990c5c0521acedaff8595a0a32f7ffd021e

Observation a50ea0db-3c77-4091-9499-c4eb0b3df067 · outbound

This paper cites Long short-term memory.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Long short-term memory

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.572933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.572933Z digest=sha256:e9f835b98393015e7f4ef0393ab0dcb15d8621bcf094082e9e7d36a039cb8a2b

Observation 0c2ca99c-c36e-4b7c-8667-92527f6f645c · outbound

This paper cites 3d concept learning and reasoning from multi-view images, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering 3d concept learning and reasoning from multi-view images, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.597230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.597230Z digest=sha256:eebda060fdec080755910bbcf273c498b8f28454ad6c5fab7f8541d87aebba17

Observation 05180ece-d3ed-4883-9ef5-49394ec39af5 · outbound

This paper cites Deep learning based 3d segmentation in computer vision: A survey.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Deep learning based 3d segmentation in computer vision: A survey

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.544576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.544576Z digest=sha256:554ed983aca15614905b61f84ba7b665cd0749208608fcae5fa8bdda7b055b70

Observation 83b04d9b-96e5-4569-b047-132e7fed3dfd · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers, in: The Thirty-eighth Annual Conference on Neural Information Processing Systems.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Chat-scene: Bridging 3d scene and large language models with object identifiers, in: The Thirty-eighth Annual Conference on Neural Information Processing Systems

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.606269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.606269Z digest=sha256:bfb2746d579f1b111dcc76830033bd5310ccd23ab477a3e35f9abdd07d3c47ee

Observation 6dcb2c80-9a8c-402b-9ca3-a23c8a3547a2 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering An Embodied Generalist Agent in 3D World

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.611071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.611071Z digest=sha256:5e65f371c614588f28354f87e4938e5624bfa26e34ddf88b02bf73ea4eca9eff

Observation dbe06fa9-2b03-4478-9163-e0c1406bf99c · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering 3d-llm: Injecting the 3d world into large language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.601645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.601645Z digest=sha256:622b2805504d688d67e76e1ff35decc62299c83d009e1a1ebae2cfa7870d05e1

Observation 328ca192-d5b5-4264-89dc-11055a38dca7 · outbound

This paper cites an unresolved cited work.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.621384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.621384Z digest=sha256:99af49c762e4a7babe89e2b2f9d0022a31287d3335580f81e525686993e18602

Observation c08760af-071b-45af-8053-183cf4dcf0bd · outbound

This paper cites Clevr:Adiagnosticdataset for compositional language and elementary visual reasoning, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Clevr:Adiagnosticdataset for compositional language and elementary visual reasoning, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.631066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.631066Z digest=sha256:6bde25c794b8826fc0c5584053d9f31611382f3b17e6586de82107af90c92e79

Observation fe8e1fc7-8bf5-4926-9440-7b9fcfa46a4e · outbound

This paper cites From image to language: A critical analysis of visual question answering (vqa) approaches, challenges, and opportunities.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering From image to language: A critical analysis of visual question answering (vqa) approaches, challenges, and opportunities

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.616432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.616432Z digest=sha256:054fd08f827d94e41f559d45be881dd4bfb9c3e095b24017f27a83c27ffbe6c4

Observation 30a559e7-2bf1-478c-853b-a462c84f49e3 · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.640967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.640967Z digest=sha256:0e50ca86dfb03c05ea080d8011cd01aac93cc1eb4a576916a506e62ca7188222

Observation 8f9cbe6d-9475-418c-b98b-a3e166acf154 · outbound

This paper cites SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.647894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.647894Z digest=sha256:41dd31b7bf43664a2b8128376a49c381cbc6259ace5f39c31256bf05a2c496ce

Observation faae1030-4bc9-4f4a-abee-82d6a14190f5 · outbound

This paper cites an unresolved cited work.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.704008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.704008Z digest=sha256:2f69c3c407eab28c4d10f4212790216f65b8a9a43a19dbb5562ff6548c2e8bc5

Observation 90ea8b41-51f2-4a23-927d-c3a02a80d9a9 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding, in: Proceedings of naacL-HLT, Minneapolis, Minnesota.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Bert: Pre-training of deep bidirectional transformers for language understanding, in: Proceedings of naacL-HLT, Minneapolis, Minnesota

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.636135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.636135Z digest=sha256:8a02120750d2a0b763b9a7cb417efc82526738e6276e2c2bbb6e90e096b9c33a

Observation 4efc0c83-5b3f-4efa-8c3d-00d290f0dc75 · outbound

This paper cites M3DBench: Let's Instruct Large Models with Multi-modal 3D Prompts.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering M3DBench: Let's Instruct Large Models with Multi-modal 3D Prompts

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.739557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.739557Z digest=sha256:446b925b5a6546c9c03eee6ebf41ed0bf0839ab979e06d099cd5d94bb408cc81

Observation 702a733b-b62a-48cb-ae7d-995e5335a206 · outbound

This paper cites an unresolved cited work.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.745097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.745097Z digest=sha256:b8fedffecf41fa3b064b51d2082da6ce83565a51c3bf9378539d9105f94ba51c

Observation 40f574b8-9214-47c8-8f89-6483e5c67ee5 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries, in: Text summarization branches out, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Rouge: A package for automatic evaluation of summaries, in: Text summarization branches out, pp

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.750006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.750006Z digest=sha256:28532a6a0c01f8aab0497c43c656c9a0ca23ae120f352e475d101345c2f47124

Observation 0c5781b4-0858-4eef-a3b7-35e8c34d5fcf · outbound

This paper cites Multi-modal Situated Reasoning in 3D Scenes.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Multi-modal Situated Reasoning in 3D Scenes

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.755264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.755264Z digest=sha256:f0c7bbb639c2793d672ca7a6dccc7ca9e19700d00802fd96b31a9b3136be3882

Observation 63e21f29-94c9-4ed7-9d9f-8ebe23edadb0 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, in: International conference on machine learning, PMLR.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, in: International conference on machine learning, PMLR

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.734334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.734334Z digest=sha256:079cd71610839b8bfbe7009079105c05aa2c2cc73e57a36fc1800a2ea4ae31cf

Observation 4801777a-71ac-4238-8e32-34dbc3de2986 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows, in: Proceedings of the IEEE/CVF international conference on computer vision, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Swin transformer: Hierarchical vision transformer using shifted windows, in: Proceedings of the IEEE/CVF international conference on computer vision, pp

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.765345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.765345Z digest=sha256:2768cde71f496d52c3437cc87f230d46cd9f0854a3b2afc99845769ce96e3403

Observation 7bbd9236-7f49-4e6b-b2a1-c2aebc5ba508 · outbound

This paper cites an unresolved cited work.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.770451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.770451Z digest=sha256:7f52719df7b82a3a181c4a0d073f60d9a1b96828e0f511c75db374e276c9e048

Observation 54a80ed2-517c-40e7-a7bb-3a3fe0f2c389 · outbound

This paper cites Point-voxelcnnforefficient 3ddeeplearning.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Point-voxelcnnforefficient 3ddeeplearning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.781178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.781178Z digest=sha256:dffb37d300c60b6afbac9d376369c49959cc9efd56ede09922ad904e8d9fc525

Observation 40bb24eb-7b9a-4deb-a3ba-dd06e63e1eca · outbound

This paper cites Group-free 3d object detection via transformers, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Group-free 3d object detection via transformers, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.786244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.786244Z digest=sha256:b825491b8582a9f46458d390edcdccbef6cd46ed96fc0e90f6fedeff123240fd

Observation e4100ba6-880d-487a-a321-ca67317f75d0 · outbound

This paper cites From Screens to Scenes: A Survey of Embodied AI in Healthcare.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering From Screens to Scenes: A Survey of Embodied AI in Healthcare

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.760631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.760631Z digest=sha256:76395397c729c60a2ba81a855351baeaf43d144952d2522d4b55d06f74590f77

Observation f1183ca7-278f-4839-9be5-5bc703fcc3fc · outbound

This paper cites Scalable 3d captioning with pretrained models.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Scalable 3d captioning with pretrained models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.795774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.795774Z digest=sha256:65fb8e592f179096da481274d69f629c95077ff90ed2a2e3ae7969f69e1d30d9

Observation 8c2750c8-b066-42ba-af1e-1a379520b0bc · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering SQA3D: Situated Question Answering in 3D Scenes

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.800718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.800718Z digest=sha256:5d651935979e713c647d1617c38af9faa1dadb5502aee7b88403eef3c047376a

Observation 59fcaa02-5cbd-4f67-b3f0-788ad8c2333f · outbound

This paper cites 11976– 11986.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering 11976– 11986

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.775560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.775560Z digest=sha256:859d487417f529bc039b759ea21e4b59dcc20a4e8c4edb3583724aca03b23ad5

Observation f86f2a20-77d6-49fb-960e-298a4b99d70a · outbound

This paper cites Situational awareness matters in 3d vision language reasoning, in: Proceedings of the IEEE/CVF ConferenceonComputerVisionandPatternRecognition,pp.13678– 13688.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Situational awareness matters in 3d vision language reasoning, in: Proceedings of the IEEE/CVF ConferenceonComputerVisionandPatternRecognition,pp.13678– 13688

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.348999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:17.831521Z digest=sha256:7c5b2cf9e836ffe7fb93952cea649a1e4ca8f7bd13fb7868f369fab40282f8c7

Observation 458a1ac0-4749-4bc8-9af4-d951fffb7016 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning, in: Findings of the Association for Computational Linguistics: ACL 2022, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Chartqa: A benchmark for question answering about charts with visual and logical reasoning, in: Findings of the Association for Computational Linguistics: ACL 2022, pp

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.330227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:17.884997Z digest=sha256:6f1ce962f56a1251147751e03211465145cbdac0724923fe24e53e123003cc5e

Observation 0959a972-4ed6-413d-bc01-042f2d96cec7 · outbound

This paper cites Transformer- based vision-language alignment for robot navigation and question answering.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Transformer- based vision-language alignment for robot navigation and question answering

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.791284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.791284Z digest=sha256:9192fb3070dcc1584489e84835e38947b163c72f992b163116be37a25f098097

Observation 161797f3-3e26-4a7d-b6f0-f21d92395d3c · outbound

This paper cites Situated language understandingat25milesperhour,in:Proceedingsofthe15thAnnual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL), pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Situated language understandingat25milesperhour,in:Proceedingsofthe15thAnnual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL), pp

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.293287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:17.919386Z digest=sha256:d9558d75c7011fd8fc190f48a82588423fd7304a87a4cf2426ea94d5864c683d

Observation 3fe083a6-c902-4670-babc-ed24d338dcd8 · outbound

This paper cites Bridging the gap between 2d and 3d visual question answering: A fusion approach for 3d vqa, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Bridging the gap between 2d and 3d visual question answering: A fusion approach for 3d vqa, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.276344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:17.924731Z digest=sha256:5a3583833d746f6c0564002e8dd47fbb95aca1f94fce47a9b1f51ef36adce000

Observation 2bc2d738-ed83-4694-8d2c-5464d99d1fcb · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Openeqa: Embodied question answering in the era of foundation models, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.364328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:17.819884Z digest=sha256:98cab8ebfb2a7307beb0502b12995a53a2cfde36c2216d2a9813ad92dff9e539

Observation a0793ba7-5424-492a-ac86-c466215232d2 · outbound

This paper cites Bleu: a method forautomaticevaluationofmachinetranslation,in:Proceedingsofthe 40thannualmeetingoftheAssociationforComputationalLinguistics, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Bleu: a method forautomaticevaluationofmachinetranslation,in:Proceedingsofthe 40thannualmeetingoftheAssociationforComputationalLinguistics, pp

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.260466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:17.936144Z digest=sha256:e1837a651cc767d77b090ed896c4dea6c8d0b5abe60604bdbceb247deb79d21e

Observation c0c8011a-85ff-4af6-9bd6-df05eabcd90f · outbound

This paper cites Clip-guided vision-language pre- training for question answering in 3d scenes, in: Proceedings of the IEEE/CVFConferenceonComputerVisionandPatternRecognition, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Clip-guided vision-language pre- training for question answering in 3d scenes, in: Proceedings of the IEEE/CVFConferenceonComputerVisionandPatternRecognition, pp

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.243933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:17.941160Z digest=sha256:6fe9a7c3c5186de765d9b829625dea2385dc0d92aed5303e26c268543918a17d

Observation d2c4bb3e-b282-4c9c-8eaa-cb4e24f55608 · outbound

This paper cites Wordnet: a lexical database for english.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Wordnet: a lexical database for english

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.310515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:17.913380Z digest=sha256:0bb6f96e8f346895328c0939eb16f2a838b4ed543f28fcf7928a3897c306207d

Observation 6cabdd2f-da26-41df-8c74-3617643202a7 · outbound

This paper cites Glove: Global vec- tors for word representation, in: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Glove: Global vec- tors for word representation, in: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pp

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.205906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:17.953075Z digest=sha256:cd79e9b9e5222e2bb5fc5e7d64aca7471d3552a98655d34161feed0d520f82b3

Observation c81878b0-94d5-4b26-9ebd-405478650924 · outbound

This paper cites 9277–9286.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering 9277–9286

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.188705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:17.959269Z digest=sha256:ed7934a2db85d0bec97b86ffaf0c377c8ef0ac5c0f45010df5514033bad1df3a

Observation 60c0cdfb-23a2-41a8-afea-2df919149701 · outbound

This paper cites Frozen Transformers in Language Models Are Effective Visual Encoder Layers.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Frozen Transformers in Language Models Are Effective Visual Encoder Layers

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.930017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.930017Z digest=sha256:b0bf0754ddaa452d1c543fc624a56449142d61a26355241c9267e1e1b3f76a21

Observation a06c324c-05a8-43c0-8876-31210bbddbed · outbound

This paper cites Pointnet++: Deep hierarchicalfeaturelearningonpointsetsinametricspace.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Pointnet++: Deep hierarchicalfeaturelearningonpointsetsinametricspace

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.152709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:17.970363Z digest=sha256:faaa13d2cce90c33d4ffe4373eb94cf4413a762bac377ef0c3e2178e2b6c32f0

Observation f777bd15-7c0f-461f-97d4-6520f1f124ec · outbound

This paper cites Langsplat: 3d language gaussian splatting, in: Proceedings of the IEEE/CVF ConferenceonComputerVisionandPatternRecognition,pp.20051– 20060.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Langsplat: 3d language gaussian splatting, in: Proceedings of the IEEE/CVF ConferenceonComputerVisionandPatternRecognition,pp.20051– 20060

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.134834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:17.976372Z digest=sha256:8456a58304d9048bbeb6097e0c2698b8f8d7ce0ed3008c027c815b16bdcb8b1a

Observation 6bb3b8da-7d2d-4d8d-b048-9d2d876d8558 · outbound

This paper cites Openscene:3dsceneunderstandingwith open vocabularies, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Openscene:3dsceneunderstandingwith open vocabularies, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.224915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:17.946976Z digest=sha256:ea1eda3b66b37adef59ae1549e69bdb27fc0b0692a98b0ad19a8373c66817719

Observation 18bdeaa3-da6b-4b76-8730-ef89a431cd02 · outbound

This paper cites Language models are unsupervised multitask learners.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Language models are unsupervised multitask learners

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.105079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:17.989114Z digest=sha256:e307c550b1a101ac43c9f2bfaaeac577648f022319ed69cda676b402781e4a09

Observation c5c2937b-0a21-466a-88a8-a1403b85f529 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.087066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:17.995346Z digest=sha256:372afc577f12a4a67d0fe2e27ed4a928847b79363e9450675775aa13244b7f65

Observation 0f197cae-0cf6-4be4-b553-dbd398825c44 · outbound

This paper cites Pointnet:Deeplearning on point sets for 3d classification and segmentation, in: Proceedings Z.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Pointnet:Deeplearning on point sets for 3d classification and segmentation, in: Proceedings Z

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.171580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:17.964885Z digest=sha256:01187ba83d84781414000b5403bb8bed2f1ffc7d105b07532e3a63297afb11b5

Observation 08ec80d4-085b-405b-bdca-2042072d1aae · outbound

This paper cites Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:18.008160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:18.008160Z digest=sha256:235475662056d91633f174c4bc6510d5e33c28e2b48220820f27808f8a4d7d2a

Observation c05988c6-3d54-4554-9422-911d308c0889 · outbound

This paper cites 3d is here: Point cloud library (pcl), in: 2011 IEEE international conference on robotics and automation, IEEE.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering 3d is here: Point cloud library (pcl), in: 2011 IEEE international conference on robotics and automation, IEEE

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.070008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:18.013925Z digest=sha256:e605234dfee90406e44c0a6f13aac6dda44fd0eca1882e8e5eb0df0444e4fa09

Observation 1751947b-f9a5-4835-af09-afd17b371734 · outbound

This paper cites Learning transferable visual models from natural language supervision, in: International conference on machine learning, PMLR.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Learning transferable visual models from natural language supervision, in: International conference on machine learning, PMLR

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.983140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.983140Z digest=sha256:acf5639975e768975bcb294e4c809db1bab2fa87fb436d0ea4dbf99f9ff64361

Observation 54c087d1-90e1-4404-993d-7eeb62a14e26 · outbound

This paper cites Mask3d: Mask transformer for 3d semantic instance segmentation, in: 2023 IEEE International Conference on Robotics and Automation (ICRA), IEEE.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Mask3d: Mask transformer for 3d semantic instance segmentation, in: 2023 IEEE International Conference on Robotics and Automation (ICRA), IEEE

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.035134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:18.023660Z digest=sha256:e53c913c2a0d95e09d5ae3088a12b2c932f37241b847ea1b25d77f45f2052888

Observation 59110341-08da-491c-90e9-4c67baeea08c · outbound

This paper cites Evaluating Zero-Shot GPT-4V Performance on 3D Visual Question Answering Benchmarks.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Evaluating Zero-Shot GPT-4V Performance on 3D Visual Question Answering Benchmarks

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-08-09T19:24:19.013183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:18.028898Z digest=sha256:17ac173d17b708fda3e9f14d37b6d051d219a04cf8392d47f742d12371d56ab9

Observation aaf700ab-c215-47b7-a9ea-6ff33ed1bbec · outbound

This paper cites SQuAD: 100,000+ questions for machine comprehension of text, in: Su, J., Duh, K., Carreras, X.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering SQuAD: 100,000+ questions for machine comprehension of text, in: Su, J., Duh, K., Carreras, X

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:18.001184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:18.001184Z digest=sha256:def53f3ac3794a7e683a4e3c607a7d4c620a552b4f5fc2085777a82633a7ce76

Observation 49020635-2f2e-4c86-92e0-833503882ef7 · outbound

This paper cites an unresolved cited work.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-09T19:24:20.017994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:18.038721Z digest=sha256:083f7a2687f4df246c66e1f151f32a0dcc75ba08e691e3f07f9b3b689e4801e0

Observation 82b63eac-17bb-4ff0-82be-79266cba391c · outbound

This paper cites an unresolved cited work.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-09T19:24:19.984296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:18.047963Z digest=sha256:049afc3dfd1905eb228c3e2fcb9f747d62357dc65e6f578691c0b011856d2bef

Observation 6ac02b77-880a-418d-8219-196a4c375a74 · outbound

This paper cites Habitat: A platform for embodied ai research, in: Proceedings of the IEEE/CVF international conference on computer vision, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Habitat: A platform for embodied ai research, in: Proceedings of the IEEE/CVF international conference on computer vision, pp

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.052660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:18.019065Z digest=sha256:9f0885336df7906bef9bf70052af8047a270b5f0cf30411107c73786dcb63290

Observation fb935da2-e5d3-44ab-8b97-834e69656197 · outbound

This paper cites an unresolved cited work.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-09T19:24:19.951331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:18.063033Z digest=sha256:a523fd8f9ccb8ce64926d48a8bbf48406a60ab46107df72e7a04341ca2c77533

Observation 05ecb402-ab6f-4dff-8938-1ae7fd2ae615 · outbound

This paper cites an unresolved cited work.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Unresolved cited work

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:18.074040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:18.074040Z digest=sha256:446d18b75e2a81b11d070b1cde21299b92425ff731b70163a35d46aeff716a4c

Observation 602e8709-d571-49b0-93e2-02867534d7f0 · outbound

This paper cites Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:18.033660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:18.033660Z digest=sha256:f8a5ad3dfc3dfd0707b5621a10280acc884f3f70939ac4dac0d02a06e612adf4

Observation d0126e34-0730-4c3b-9603-a12ed31c2a4c · outbound

This paper cites Attention is all you need.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Attention is all you need

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:18.089500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:18.089500Z digest=sha256:72bc425089961f60b9a0b0e3115c9977a7d291c8f9132f974c93f2a47b10dc5a

Observation 0888d035-6604-4ecd-ba50-113077d7e7d9 · outbound

This paper cites Cider: Consensus-based image description evaluation, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Cider: Consensus-based image description evaluation, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:19.890016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:18.095281Z digest=sha256:32bdb2970252d09afe390cd27997b91a52f6115279c2520b831177393dbaef9f

Observation 5202b772-2e1e-416b-b31c-85050eb81cdb · outbound

This paper cites Rio: 3d object instance re-localization in changing indoor environ- ments, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Rio: 3d object instance re-localization in changing indoor environ- ments, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:19.868550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:18.100362Z digest=sha256:29a37b69fa12ff18d37aaf492c9a668e9aaf292f3632b3fd154d403d6790a6d2

Observation b8add8e9-b150-47c8-8cde-4b3d2dd41e20 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:18.105248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:18.105248Z digest=sha256:48f67ba4799220291816bc8025bc2b7a017937f18fca38641e62287839a5e849

Observation b79f58dc-148a-4d0d-82f8-4206a4d8aa86 · outbound

This paper cites Space3D-Bench: Spatial 3D Question Answering Benchmark.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Space3D-Bench: Spatial 3D Question Answering Benchmark

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:18.057476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:18.057476Z digest=sha256:2fee162dcabf7130b85e3721e62783a27d1b359dcb36ec8eb83dfa5094c81910

Observation 85db0a83-c175-46f9-8e5a-bf7773b43ba2 · outbound

This paper cites Medical vqa, in: Visual Question Answering: From Theory to Application.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Medical vqa, in: Visual Question Answering: From Theory to Application

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:19.831073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:18.116148Z digest=sha256:019efeab5c15b6eb866b9352c10604a8ecf9799d3458fff6d94f2d55a7bcdbf6

Observation 07c8ad4f-d80a-4aa7-9143-b8bfe8035493 · outbound

This paper cites SplatTalk: 3D VQA with Gaussian Splatting.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering SplatTalk: 3D VQA with Gaussian Splatting

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:18.068674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:18.068674Z digest=sha256:2d844a74d176793528637728ba7df23f722fbd238e913d43f3d5723bcb20d7f2

Observation 69a842f5-e186-46ca-b09c-26fa2642b265 · outbound

This paper cites an unresolved cited work.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering Unresolved cited work

Reference 99

Resolution
unresolved
raw_fallback, observed 2026-08-09T19:24:19.811776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:18.126915Z digest=sha256:05fbd24e62c73b4abf9c2fd7be67cafe8caf4dbba28c0097e3007eecc1be4906

Observation 465d05e9-243e-4786-acc3-b1f04e9ea7a9 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering LLaMA: Open and Efficient Foundation Language Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:18.079030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:18.079030Z digest=sha256:d856359441765ae59ed4a865fdf1922e1d5a1cec6470ceca6544ba9d290b24b9

Pith citing papers

Observation dfa2a3c7-da90-4b0f-a8a3-343e810bad04 · inbound

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts cites this paper.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:14.341180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:14.341180Z digest=sha256:867d556e1bad9f00a40d7793f42afb09bf333544e9f283f5bedf24653e79c7d5

Observation 3943a55d-6f90-4551-beb5-39eca2afddb8 · inbound

POMA-3D: The Point Map Way to 3D Scene Understanding cites this paper.

POMA-3D: The Point Map Way to 3D Scene Understanding Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.610268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:27:27.347592Z digest=sha256:4a801e6e953c2210eb0acf3e522d89014d7cc9d858ddd58fc44eb0f5b1fd8e3f

Observation 1633c144-0855-4c72-932e-4ee62c5586ca · inbound

RGB-Pointmap Pretraining for Unified 3D Scene Understanding cites this paper.

RGB-Pointmap Pretraining for Unified 3D Scene Understanding Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T21:23:17.935190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T21:19:49.421653Z digest=sha256:e4ec7f4608fbbd9db7c9bbe2f34a597d132d5be5844f1fd9f6e426a0cfd78418