Pith. sign in

Paper Citation Record · LEDGER

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

As of 9 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 2 inbound Pith citation observations for arXiv:2507.04047.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04047 v2

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:02:33.215747Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:54:57.184897Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T19:03:08.599602Z

Reference resolution

96 of 96 outbound references displayed

  • verified exact2
  • verified fuzzy59
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e2a31111-e2e4-4890-a72a-2b5fe7745999 · outbound

This paper cites Scanents3d: Exploit- ing phrase-to-3d-object correspondences for improved visio- linguistic models in 3d scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanents3d: Exploit- ing phrase-to-3d-object correspondences for improved visio- linguistic models in 3d scenes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.156912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.156912Z digest=sha256:9337ce74ab6883e2f20fd0a43f27a2126f65053fd12fdd8613142cac20789e37

Observation 7f2de7b2-1a14-431d-a5c8-30bbdd9a6843 · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.251850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.251850Z digest=sha256:cc544cd44b3ffce0ab1ed66e26b941fe89955d8d1a8321005b98dac9d85f4ae5

Observation d4f365fa-c59c-42b1-9e52-546e5375435c · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanqa: 3d question answering for spatial scene understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.317126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.317126Z digest=sha256:71dcd5895f54f4241a3458b56c0a8058988d4516f2a877d223095e25a9346ea3

Observation c49bc4df-6ccb-4a9e-9bbe-fd1b50ea2390 · outbound

This paper cites Do as i can, not as i say: Grounding language in robotic affordances.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Do as i can, not as i say: Grounding language in robotic affordances

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.434833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.434833Z digest=sha256:97a6be05e14ce828bdf53db389433cd8e4f6349f158f6d5bf4892481e76346d7

Observation bb7ccfdb-041b-411f-993c-23f9995f6668 · outbound

This paper cites 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.497301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.497301Z digest=sha256:6d45208f628a6c9bb1f289e7f376788e19c8e28efb23620150fd5697e1d25b25

Observation f9dddfbd-8cbb-48ea-b292-57527f9aadad · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Emerg- ing properties in self-supervised vision transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.605674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.605674Z digest=sha256:f388daddd9908d808f3b2f16155bad5b29bf101d650e1307d5523eab674ed940

Observation deda0b58-fc06-49a7-b970-3e239931f8ae · outbound

This paper cites Object goal naviga- tion using goal-oriented semantic exploration.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Object goal naviga- tion using goal-oriented semantic exploration

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.679473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.679473Z digest=sha256:62b0978bf54e976e98996a9020d4c5b98904ba2d970ba2eb9681c8c35d0bab28

Observation b823c8a9-85dc-4284-8432-c24334a7769b · outbound

This paper cites Object goal navi- gation using goal-oriented semantic exploration.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Object goal navi- gation using goal-oriented semantic exploration

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.763679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.763679Z digest=sha256:1cea821d31e9edf08e83b835908a3128a3093d9c9f6ed9ecd1724f52027cf4d6

Observation 86802e41-5707-4c97-87d5-a88d5deb6b24 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natu- ral language.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanrefer: 3d object localization in rgb-d scans using natu- ral language

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.842833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.842833Z digest=sha256:ffac61175c30fe076281821828fb6b5b8fd3da665725e6751a65e5bf3388c5e2

Observation 06333299-344d-4207-b617-3a79b9607d6c · outbound

This paper cites Language conditioned spatial relation reasoning for 3d object grounding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Language conditioned spatial relation reasoning for 3d object grounding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.924326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.924326Z digest=sha256:05cc48082ccd252023f1351de0aa1c282f1d4a078ae81345a839bcf5c5c196c6

Observation 778803a5-506f-493f-99d9-949b414b8d32 · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.014240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.014240Z digest=sha256:7a905f34b6e78d989a8487b628ef847a3401569dd4286826ec28486fa3af6472

Observation e4822169-e117-457f-8f5d-242e6d091ec3 · outbound

This paper cites Scan2cap: Context-aware dense captioning in rgb- d scans.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scan2cap: Context-aware dense captioning in rgb- d scans

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.057605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.057605Z digest=sha256:8a357c1ccbc1a3495199aef1acf86f7f6934087e4686f43a763fe20aada51183

Observation 26c200ce-8196-49a5-a404-ede974461053 · outbound

This paper cites Unit3d: A unified trans- former for 3d dense captioning and visual grounding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Unit3d: A unified trans- former for 3d dense captioning and visual grounding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.194145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.194145Z digest=sha256:0808e445309154058646c4ee432c5f5b352815696358cc79c27ff2a676a5a1bf

Observation 96a42d0e-704e-4d5e-b86c-f82703300d84 · outbound

This paper cites Schwing, Alexan- der Kirillov, and Rohit Girdhar.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Schwing, Alexan- der Kirillov, and Rohit Girdhar

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.284849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.284849Z digest=sha256:75dae74652869cc89c173dcd93467207b82bca8a5cf94896c4729ee3615ae334

Observation ed1e8638-0b8a-4a26-9984-d39ecfa84001 · outbound

This paper cites 4d spatio-temporal convnets: Minkowski convolutional neural networks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 4d spatio-temporal convnets: Minkowski convolutional neural networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.391830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.391830Z digest=sha256:1b9306a11cfd9a1ec783fb85a5678c027a4ac9e86759e369333844f8b43bd713

Observation 17ccbf1e-189e-4f93-b408-cafceb94bcaf · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.507880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.507880Z digest=sha256:569a275b44fd11297cbd021c4f51af8a56ec19c777191909ce73f7bc46d2b577

Observation 23dd4bc8-824e-48b8-99bd-72233555fe15 · outbound

This paper cites Embodied question answer- ing.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Embodied question answer- ing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.662553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.605207Z digest=sha256:fa1f7b4fda31b86797cddc08f387a4e634ebdad9114163b33723205033dbeb29

Observation fec0ec8e-b139-451f-b6b8-91dcf8ba2993 · outbound

This paper cites A survey of embodied ai: From simulators to research tasks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation A survey of embodied ai: From simulators to research tasks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.648665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.730308Z digest=sha256:35d5c6306569c181c017121ca5f8e87dfd93ec09928e7d325f5e21cedd065853

Observation 6db28803-3cf2-4251-ad77-654dcd228b67 · outbound

This paper cites The One RING: a Robotic Indoor Navigation Generalist.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation The One RING: a Robotic Indoor Navigation Generalist

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.828420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.828420Z digest=sha256:765dc33117c075ce8649056fb41ff885bec3e62bf71c445e0d7a6d7252f9077a

Observation 4a984db5-a1cb-4e85-842d-0a3a593d6b52 · outbound

This paper cites Spoc: Imitating short- est paths in simulation enables effective navigation and ma- nipulation in the real world.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Spoc: Imitating short- est paths in simulation enables effective navigation and ma- nipulation in the real world

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.633572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.859253Z digest=sha256:ba6bba2253aab8828e3a32ca980a1cfab8782e8a977cb00d4f135ab83ee44c4c

Observation 713006fa-6b05-414d-a629-451cd2717320 · outbound

This paper cites Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.864163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.864163Z digest=sha256:408e9e026ae568bafa3417b94be5ec6da4c28e61c4757677dc3485046f2393d5

Observation 07a13e23-4b26-4210-91eb-1174c07e9df7 · outbound

This paper cites Efficient graph-based image segmentation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Efficient graph-based image segmentation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.618261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.869619Z digest=sha256:4860d499350494f40a8d19c52b14e6d542ec6fbd214ff50f11ddd519ba2cac91

Observation 1a3fd2f0-a809-4758-ad4d-a2be5e0be2cc · outbound

This paper cites Cows on pasture: Base- lines and benchmarks for language-driven zero-shot object navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Cows on pasture: Base- lines and benchmarks for language-driven zero-shot object navigation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.604356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.873766Z digest=sha256:fd12c352a3d0e4cf2a54611795c0cf6cf2d65f7893b675542818a70eb31ff98b

Observation 8c68ce8f-f0a4-4e70-9234-da915752a457 · outbound

This paper cites Scaling open-vocabulary image segmentation with image- level labels.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scaling open-vocabulary image segmentation with image- level labels

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.589762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.878891Z digest=sha256:d0819d0efc7b747d32885158d00fb3f07facfca256f887b1595003baf665d28c

Observation 45f432e4-e104-4efe-9322-2a706db8a7fa · outbound

This paper cites Conceptgraphs: Open-vocabulary 3d scene graphs for per- ception and planning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Conceptgraphs: Open-vocabulary 3d scene graphs for per- ception and planning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.575555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.883132Z digest=sha256:600c9b4c6bb6207ec67ea8f1e26dab374bf64e04b643eb79742906849e1c2019

Observation adf93ba6-943c-4359-bfef-d23cc8551d51 · outbound

This paper cites Viewrefer: Grasp the multi-view knowledge for 3d visual grounding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Viewrefer: Grasp the multi-view knowledge for 3d visual grounding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.561287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.887401Z digest=sha256:b1b0dfa5c669af088887aaa28dff1e570652203dfc07b39df209b23955df13ad

Observation 7d587f8b-2923-40f4-8826-daaabe53f5d4 · outbound

This paper cites Transrefer3d: Entity-and- relation aware transformer for fine-grained 3d visual ground- ing.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Transrefer3d: Entity-and- relation aware transformer for fine-grained 3d visual ground- ing

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.546556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.891844Z digest=sha256:3d7a1533762c08fe50a8d2fb65495c7ae138fde07f07bad49e1eaadd071527ce

Observation 0a0a1021-5c64-4d71-8425-ea1db88b305a · outbound

This paper cites Vln bert: A recurrent vision- and-language bert for navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Vln bert: A recurrent vision- and-language bert for navigation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.531842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.896182Z digest=sha256:99755d7e8bb074b0c7f885d85145c6cb148130825be8bbeddb48c8d9bfbcebab

Observation d4c652b4-9733-408a-8d2e-c5eb0468f7ee · outbound

This paper cites 3d-llm: In- jecting the 3d world into large language models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3d-llm: In- jecting the 3d world into large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.515695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.901425Z digest=sha256:64a9cc13673fbbe3f4433d17d7a285f9d0f45cb1cc0342a76c61bc9e68ecfce6

Observation 7f29a1b4-1c1e-4eb9-955e-9b9381b77dcd · outbound

This paper cites A real-time occupancy map from multiple video streams.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation A real-time occupancy map from multiple video streams

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.498323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.906117Z digest=sha256:5e8d8bf02ac36050aed5d09eb209648f58b5ce93245355a8c7d65af4de374083

Observation 9f43c2d6-7681-47f4-b4a6-54b61de930e7 · outbound

This paper cites An embodied generalist agent in 3d world.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation An embodied generalist agent in 3d world

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.481002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.910518Z digest=sha256:db491eecef5ed446ec602d7ff1269f83cef88fdca21eb320d5c6bc250c3b6a75

Observation 4c5594b4-3ee3-422a-b71b-7dba27b78f5f · outbound

This paper cites Language models as zero-shot planners: Ex- tracting actionable knowledge for embodied agents.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Language models as zero-shot planners: Ex- tracting actionable knowledge for embodied agents

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.463894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.914903Z digest=sha256:c8505a9394674b86d362290de529b07a96aebc43232e740c0b214ac7c5778c2f

Observation 143deb32-0c82-488b-81f8-15730725a8d9 · outbound

This paper cites V oxposer: Composable 3d value maps for robotic manipulation with language models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation V oxposer: Composable 3d value maps for robotic manipulation with language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.448269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.919522Z digest=sha256:6bacb3e154285eb66275cb83587ff592fc596322ed7633148fd9425ee2e3e500

Observation 39bbe7cb-7eaf-46c9-a293-620806b5320b · outbound

This paper cites Bottom up top down detection transform- ers for language grounding in images and point clouds.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Bottom up top down detection transform- ers for language grounding in images and point clouds

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.433025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.924413Z digest=sha256:ea6684de08105a31d545a20f67815e36a73bcb9c0497cca09344a6738a0cd758

Observation 8f216072-19fc-4d8f-9511-ee0f109fc652 · outbound

This paper cites Conceptfusion: Open-set multi- modal 3d mapping.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Conceptfusion: Open-set multi- modal 3d mapping

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.417777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.929001Z digest=sha256:54d44d6abcf15f1449bb245c9ca83efb7e3887ec44973234aa731e3cacaa6542

Observation e52ef426-8946-4d81-856e-63b42e949343 · outbound

This paper cites Sceneverse: Scaling 3d vision-language learning for grounded scene understanding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sceneverse: Scaling 3d vision-language learning for grounded scene understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.402368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.934381Z digest=sha256:ae936828ce8bd5665a00e5b382e3e4e6159c5f01ed325186692be1097c229923

Observation b46b61ae-463b-42ff-a9ca-487841e009f3 · outbound

This paper cites Goat-bench: A benchmark for multi-modal lifelong navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Goat-bench: A benchmark for multi-modal lifelong navigation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.386201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.939241Z digest=sha256:6dbff5b6549a213b5b2c41864bc5d8a730c33f372dcfcfa98ea69e39f293a386

Observation 9eeaf62f-5c89-4be2-be2a-02d3f7d04fb8 · outbound

This paper cites Realfred: An em- bodied instruction following benchmark in photo-realistic environments.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Realfred: An em- bodied instruction following benchmark in photo-realistic environments

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.371553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.943797Z digest=sha256:8a2ca42b3ef81486671178a788b2d72811463a25ca17197f2f6a7805fdb71fc9

Observation 30f6df15-b03d-4f0c-b060-014a2e9d8e01 · outbound

This paper cites Segment anything.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Segment anything

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.357025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.948191Z digest=sha256:bbf62900363bfb0644c9733c6773defe223a78a6644487d09ea885b39c14c7b1

Observation b1919572-cd83-449d-826b-e48b8505acfe · outbound

This paper cites UniCLIP: Unified Framework for Contrastive Language-Image Pre-training.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation UniCLIP: Unified Framework for Contrastive Language-Image Pre-training

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:02:33.577190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.952478Z digest=sha256:72b3557b8897cf0b1c89873ccdaa4c9f4e5add40f67d5118da5d5f49388f3bc1

Observation 8207f357-d3b8-407e-a437-96f5b9a07aa8 · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Less is more: Clipbert for video-and-language learning via sparse sampling

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.342337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.956776Z digest=sha256:875608a0650be683d8002efab5b3a2a90f195e5bc2b68f7b722db3b26612e2ab

Observation ad92fa97-2683-4468-9993-289fb62236ff · outbound

This paper cites BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.960840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.960840Z digest=sha256:32b9e222388ac36effe64f4b985ff559f916a904b4ab44f94db038af501f3930

Observation 7d582015-4573-4d72-90f1-53a9abbd8ed0 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.966191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.966191Z digest=sha256:5213e6d6f3b89921a771aaff6896ce4acc3826418b3f02d93b3c7c44911fb9aa

Observation 2a1c475a-e27f-4406-9282-799d56d74131 · outbound

This paper cites Panoptic segformer: Delving deeper into panoptic segmen- tation with transformers.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Panoptic segformer: Delving deeper into panoptic segmen- tation with transformers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.328116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.971408Z digest=sha256:e14a3bb3d5c74f313455cd94e103b60358ec220870c49074ecce62d2e6710494

Observation 53c90110-e900-4bf6-b726-6d8daf6d9d1e · outbound

This paper cites Code as policies: Language model programs for embodied con- trol.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Code as policies: Language model programs for embodied con- trol

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.312598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.975605Z digest=sha256:b831e0279d37415a0b624d8b436a6fca2f987730d9ebc3013c9bff4125d5028c

Observation b986deef-cc20-4681-a7cf-88509af4eabd · outbound

This paper cites NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.980063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.980063Z digest=sha256:7750643e5fb197d62add79b6f721a76e5960afd2ca5599a9a23a64250c55018e

Observation 2d1902b8-80a2-4ddf-90eb-aa605a81c199 · outbound

This paper cites Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.984728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.984728Z digest=sha256:4fa877d63c041f578215d9a29c1625a3ed8d66f2831c1f30851dd1d13bfd3f30

Observation c545085f-88f9-4e78-bb8c-50e4337486a6 · outbound

This paper cites Sqa3d: Situated question answering in 3d scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sqa3d: Situated question answering in 3d scenes

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.296444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.989899Z digest=sha256:1bf99e1968293ac709a3177cca13fb44fa3fa08ba3f65515bc221eedb43511ea

Observation 41653663-ad6f-4d35-8548-2f546cc1f18e · outbound

This paper cites Zson: Zero-shot object-goal navigation using multimodal goal embeddings.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Zson: Zero-shot object-goal navigation using multimodal goal embeddings

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.280449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.994814Z digest=sha256:03251ae29f35c8696b5c7b93cc153b1e4720a61cf2c189f113da9cbd995306e5

Observation e621b6db-fe2c-4a9d-8df8-dbfe3a847b75 · outbound

This paper cites Openeqa: Embodied question answering in the era of foun- dation models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Openeqa: Embodied question answering in the era of foun- dation models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.211271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:32.999861Z digest=sha256:545894cb21fe44549146c26d0f1fc4665a72d289454fed9ba428d9f8a782d4ef

Observation 9b824185-5025-4395-aca8-665e709e9b42 · outbound

This paper cites Openeqa: Embodied question answering in the era of foun- dation models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Openeqa: Embodied question answering in the era of foun- dation models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.152079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.005435Z digest=sha256:61e0636d507b53034a1909785dc5a7090ebcce82b544dac1b71b63786ca81815

Observation bb3fbf5b-e554-44e3-9818-b6a3b50163c3 · outbound

This paper cites Spatial memory.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Spatial memory

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.114012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.010136Z digest=sha256:0e7ee70e9fe9bbc7f2cbfafee42380a1c55d94fc3328d027d8c69d5dde24750d

Observation 6da5b11a-36a1-404a-b98c-a3b75be6c88e · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation DINOv2: Learning Robust Visual Features without Supervision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.014670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.014670Z digest=sha256:a276b8e9c75ee5c7489bfa53acb6e686da6115c38ebb4bbc7c22a504dd9f6b1b

Observation 729ec64a-1e53-4c9e-af84-67004af85111 · outbound

This paper cites Teach: Task-driven embodied agents that chat.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Teach: Task-driven embodied agents that chat

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.099440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.019013Z digest=sha256:06c65834e8870f4f56cd679db887c974a80f1364d9f0628dda34226ee0465ddf

Observation 91e13b98-5db9-4a06-b856-b57512660a26 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Openscene: 3d scene understanding with open vocabularies

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.082872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.023557Z digest=sha256:14faf803747cad128363f7d19e94ca15c107fbcd432f2c699376d2e1f6543665

Observation 0bb53d3a-b4fb-4ce6-a6bd-76e809f22e22 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Learn- ing transferable visual models from natural language super- vision

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.067457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.028017Z digest=sha256:3da881ff056a31b7e1f7f76a73fba3f6faf0cff95d12b187a2df764c2e6b6614

Observation 8337e663-f1b0-4a05-9243-ec3752d2a8ee · outbound

This paper cites Pirlnav: Pretraining with imitation and rl finetuning for objectnav.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Pirlnav: Pretraining with imitation and rl finetuning for objectnav

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.051867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.032520Z digest=sha256:880116f445a4fe598b73a73198c9e94b5440ad3c4f2c4a47cfe5f44eae674263

Observation 738a7c9a-cb3b-4a08-b6fb-3f9a641194f3 · outbound

This paper cites Sayplan: Ground- ing large language models using 3d scene graphs for scalable task planning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sayplan: Ground- ing large language models using 3d scene graphs for scalable task planning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.036765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.037520Z digest=sha256:42432bf55ab0ac20298273aec40870c4f7f8a80d6423eaf251a612a5ec127d23

Observation 841e3bdb-4be8-4196-9e45-3ab0b52d98ca · outbound

This paper cites Explore until Confident: Efficient Exploration for Embodied Question Answering.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Explore until Confident: Efficient Exploration for Embodied Question Answering

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.042822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.042822Z digest=sha256:c299cd6a210fdafd538054c9f0ef64a50d4b37c190d3c8ed9e6725c211518949

Observation 2bd1b137-7252-436e-90e0-50c9dee394da · outbound

This paper cites Language- grounded indoor 3d semantic segmentation in the wild.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Language- grounded indoor 3d semantic segmentation in the wild

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.019320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.047240Z digest=sha256:82f24a23a760bdec9380438a01afea67cdff190a7a32d71ce989289fb867d453

Observation acc76002-445d-4d6a-a6a7-d62271306133 · outbound

This paper cites Habitat: A plat- form for embodied ai research.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Habitat: A plat- form for embodied ai research

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.002845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.051322Z digest=sha256:53f587811d1f2544862c3d95b2c3e325fcb2281b338d0edd5dc399787e300af0

Observation 7bc365d5-0260-4142-b6f4-cb24fdbc7afd · outbound

This paper cites Proximal Policy Optimization Algorithms.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Proximal Policy Optimization Algorithms

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.055692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.055692Z digest=sha256:36325a6051d28bcf6845ead0d6d45c43e9c6c16724e4dc6db3837e8c3bd51651

Observation 46baae52-902b-4d3a-bb2d-173f6c84d29a · outbound

This paper cites Mask3d: Mask trans- former for 3d semantic instance segmentation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Mask3d: Mask trans- former for 3d semantic instance segmentation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.987323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.059795Z digest=sha256:4851f960fdbfbf154b6c75aa9d716979ca813cccbb01ca7e4bcee146cd2443eb

Observation 0a686420-7872-4d1c-a62c-54e40f39f5ea · outbound

This paper cites Alfred: A benchmark for interpreting grounded instructions for everyday tasks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Alfred: A benchmark for interpreting grounded instructions for everyday tasks

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.971219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.064643Z digest=sha256:161c43b2ea9fe4c41ba4fac3480f03edc41ef7162f2c1f7aea7d8b7643f726ab

Observation 170add76-540e-4418-8d40-8792d0b32b8c · outbound

This paper cites Llm-planner: Few-shot grounded planning for embodied agents with large language models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Llm-planner: Few-shot grounded planning for embodied agents with large language models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.954645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.069207Z digest=sha256:9055c919e5b52de7f6d01c7037a345d5eeec7f1834344cb86b286cc7b6628427

Observation 9bf74e8c-1c66-4fcf-9d3e-71c08d19150b · outbound

This paper cites Habitat 2.0: Training home assistants to rearrange their habitat.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Habitat 2.0: Training home assistants to rearrange their habitat

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.073653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.073653Z digest=sha256:af4ad68e10616fb1e649c0c8606728bc934301cebfb2ba7e28ee5f6226da86a9

Observation c9530d43-f3e2-4889-b87e-9045a8cd09d1 · outbound

This paper cites Sumner, Marc Pollefeys, Federico Tombari, and Francis Engelmann.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sumner, Marc Pollefeys, Federico Tombari, and Francis Engelmann

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.927417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.078435Z digest=sha256:72450718d4a4a5191edf013f76b36a5f8ad60c4c6eb83eeb87ee0ee47c61974e

Observation c9361746-d768-4873-a170-e937b6194819 · outbound

This paper cites Rio: 3d object instance re- localization in changing indoor environments.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Rio: 3d object instance re- localization in changing indoor environments

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.911980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.083330Z digest=sha256:567b2a29555664a31b68857ed6d32a45996bb055506c49236c3c6334bb6745f6

Observation c3958155-876a-4424-a8c8-c14914faa310 · outbound

This paper cites Embodiedscan: A holistic multi- modal 3d perception suite towards embodied ai.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Embodiedscan: A holistic multi- modal 3d perception suite towards embodied ai

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.895603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.088684Z digest=sha256:f90733f6822a391fed03fc451d5211833275cc4f92e8567b6516868c2abf4b78

Observation 9840311b-3a6c-44c1-b628-fe68e2d051ac · outbound

This paper cites DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.094358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.094358Z digest=sha256:48acde52031a99f8ae1aa437659ec001d272b7b93640551e0bc0766c57115c20

Observation d855c8e7-24ce-4c92-bf00-1cde8f9bce5c · outbound

This paper cites Ver: Scaling on- policy rl leads to the emergence of navigation in embodied rearrangement.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Ver: Scaling on- policy rl leads to the emergence of navigation in embodied rearrangement

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.880023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.099421Z digest=sha256:76879ff0c8ee8cc76a73a602e2c40fda2a5dcb4586a27b75c32801284987b87b

Observation c4dbf192-0e51-4d22-9090-66021d4bdc9b · outbound

This paper cites Scenegraphfusion: Incre- mental 3d scene graph prediction from rgb-d sequences.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scenegraphfusion: Incre- mental 3d scene graph prediction from rgb-d sequences

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.865248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.103822Z digest=sha256:6be39011b8437af0290c4e204e32a4f14f652c684f1af4391ec085281877822c

Observation 73cebb80-6c31-4bd7-be82-84128aef8574 · outbound

This paper cites EmbodiedSAM: Online Segment Any 3D Thing in Real Time.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation EmbodiedSAM: Online Segment Any 3D Thing in Real Time

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.108579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.108579Z digest=sha256:11ec6f96fa241a717fc6e623296abf3cef9436ab182007a2e8e321fbb74b46b9

Observation 8f204022-07ac-431f-b824-1607b07efda7 · outbound

This paper cites Habitat-matterport 3d semantics dataset.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Habitat-matterport 3d semantics dataset

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.850264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.113772Z digest=sha256:3fb7880e0918b9489d1e3c9be49926c81873d9fea550b4cf98c798deeb61db37

Observation 66540fd2-5ebc-42da-8a41-49a9ceb41a84 · outbound

This paper cites Frontier-based exploration using multiple robots.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Frontier-based exploration using multiple robots

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.835273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.118819Z digest=sha256:7650db47d505c2722801fc327b3dd453067b212cb0b21cc11d38478c14cd6c45

Observation 1525a073-ef2b-44c6-bb44-cf0d9203795b · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.123605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.123605Z digest=sha256:47ab79df374b84629531ddcd3adb5c78be7451bff881cb136e64415ccbfadebf

Observation 6b28ee52-6bb1-45e5-a211-0a0961a5e8fc · outbound

This paper cites 3d-mem: 3d scene memory for embodied exploration and reasoning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3d-mem: 3d scene memory for embodied exploration and reasoning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.819523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.128531Z digest=sha256:9ce31566042c746e35149476db4a296567ac28968727b43be340be7969d09ad3

Observation f7902776-3426-4812-ba1d-835c3e02c367 · outbound

This paper cites Vlfm: Vision-language frontier maps for zero-shot semantic navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Vlfm: Vision-language frontier maps for zero-shot semantic navigation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.804490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.133341Z digest=sha256:5279a77192546a961a64196c428755155a991a9f075cb9bfb87f31496e8bbb6a

Observation e846a639-0e12-4c7f-96d5-1bab80731938 · outbound

This paper cites Hm3d-ovon: A dataset and bench- mark for open-vocabulary object goal navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Hm3d-ovon: A dataset and bench- mark for open-vocabulary object goal navigation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.788045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.138214Z digest=sha256:212869d7793bfb851a410709f0f21a963994393f79b94be7177ee8b0150c5fdc

Observation 0e5784ad-8c37-4039-8d60-94a6b91f6d0e · outbound

This paper cites Frontier semantic exploration for visual target navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Frontier semantic exploration for visual target navigation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.770186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.142761Z digest=sha256:38d772f437365428b8b65fbc0937701ca43fffe3c28633d44d0c55a4ced922c7

Observation aae9a2f4-215d-4950-8a5f-1f54cf5c022c · outbound

This paper cites Instancere- fer: Cooperative holistic understanding for visual ground- ing on point clouds through instance multi-level contextual referring.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Instancere- fer: Cooperative holistic understanding for visual ground- ing on point clouds through instance multi-level contextual referring

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.754549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.147218Z digest=sha256:80499c58e06b610fd90285e9dd099e2ba4fcc826626d49f6db8ae0204a8e38d9

Observation 2609d05e-91ce-4d80-b285-5d697180c3b5 · outbound

This paper cites PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.151309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.151309Z digest=sha256:8713da562a3617d23993d2bddf0d303aa82b32f60ee9ba3a1358999997b97b59

Observation 583e3afd-8e7f-47ef-9722-8da0e9f8a503 · outbound

This paper cites Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.155998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.155998Z digest=sha256:1936ab24fe7d68733923508fedc2a2869c64831105d61524e97c3470f81ff011

Observation 6063f092-66ec-41ba-84c4-cb0f6f1045a0 · outbound

This paper cites Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:02:33.334560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.160598Z digest=sha256:99e42ed2784086a18f012378f3f952cc8bd56a28dbc5ea91f9e1d9eda9b96e9f

Observation 5f9780c8-10d5-4986-8f9a-4d97c8147ca9 · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d objects.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Multi3drefer: Grounding text description to multiple 3d objects

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.740367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.165673Z digest=sha256:ee2b4fc19d88d00a1ff0023bbbc53e2f93737a76f62464158dbf371edd5b6d64

Observation f7906a49-1d19-4666-ad00-b2842a1db9ca · outbound

This paper cites Microsoft kinect sensor and its effect.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Microsoft kinect sensor and its effect

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.723322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.170543Z digest=sha256:80d254ea34d321680f13ae1c708b4223d14c05f19abab846e6e065aa6dd14597

Observation f7fba5c5-b1d6-49a2-a672-f6d67140bc4e · outbound

This paper cites Task-oriented Sequential Grounding and Navigation in 3D Scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Task-oriented Sequential Grounding and Navigation in 3D Scenes

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.174687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.174687Z digest=sha256:fc61f11beb046131f70eb9e684d21cfbe7a61d14912f9f675a8cbaf400bae715

Observation b0f133cb-4108-4947-b1fc-836c269d090e · outbound

This paper cites To- wards explainable 3d grounded visual question answering: A new benchmark and strong baseline.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation To- wards explainable 3d grounded visual question answering: A new benchmark and strong baseline

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.708124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.179080Z digest=sha256:ba6c71504cc1185030d6a6d6f7c388e66d498e02834c5d42aeb96e3756921b11

Observation a10528fa-1096-4e34-adf2-8d1bfbaa1104 · outbound

This paper cites Fast Segment Anything.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Fast Segment Anything

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.183289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.183289Z digest=sha256:b566c6f84511ae183d14fb3a21dfd8408807271b59c1e2ecf6fceb190e0398cb

Observation c0dbc718-15b3-4fef-ba94-993a9eaf9604 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.187976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.187976Z digest=sha256:32378972877ab64eeee86e38ff09ebc75d987fc3c4391845fbd4082b9d6ff7fe

Observation 83388a1b-427e-42dd-a5f5-0aa425a01d7c · outbound

This paper cites Dual memory units with uncertainty regulation for weakly supervised video anomaly detection.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Dual memory units with uncertainty regulation for weakly supervised video anomaly detection

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.690761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.193178Z digest=sha256:40c438aecbf3ddc2ffde20775b53bad9ab40e96be661feb0e94c2a966569802c

Observation 96ec4986-1805-49e8-ad9c-1def13a40b0b · outbound

This paper cites Scanreason: Empowering 3d visual grounding with reasoning capabilities.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanreason: Empowering 3d visual grounding with reasoning capabilities

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.676105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.197976Z digest=sha256:9eef04a87545899fa64cecae13382ad2ae2461498e7570f29e3932bcc60755a0

Observation d4465aff-60cc-4572-b653-dfdb711f5ee1 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.660133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.202562Z digest=sha256:b7517dde2a7f0e46faaf2f64144855a9ad38441f974d7b640781553277e75263

Observation 225f2ccb-d3ce-4433-9dbc-f8934a0e1169 · outbound

This paper cites Unifying 3d vision-language understanding via prompt- able queries.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Unifying 3d vision-language understanding via prompt- able queries

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.644142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.206561Z digest=sha256:40374007242449c889ab5150e1b76128941348618a80ac0e2621cbc43d3dd03d

Observation 17a85cb1-9423-4a72-994e-4fc5374ca733 · outbound

This paper cites TANGO: Training-free Embodied AI Agents for Open-world Tasks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation TANGO: Training-free Embodied AI Agents for Open-world Tasks

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.210895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.210895Z digest=sha256:cb2dba2e866ca12a722511b43261b7ddbfcac2faa3360acf3aae0954362d4636

Observation be143c56-3127-4de0-84c2-0274233b3c85 · outbound

This paper cites Generalized decoding for pixel, image, and language.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Generalized decoding for pixel, image, and language

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.628398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:02:33.215747Z digest=sha256:f60b98bc427feaac59ac543197e626b3121495db7bd25314d1431804565d1dff

Pith citing papers

Observation 8b12b765-c17c-4532-8932-0c2f5ffe3a7d · inbound

SPG: Style-Prompting Guidance for Style-Specific Content Creation cites this paper.

SPG: Style-Prompting Guidance for Style-Specific Content Creation Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:57.184897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:57.184897Z digest=sha256:af1619e9e5b482b92afac2dc6f0b7de63294af050694a96234eb666aa0afd7ee

Observation 7363683b-1a7a-41e1-8f77-a918c1574094 · inbound

FSUNav: A Cerebrum-Cerebellum Architecture for Fast, Safe, and Universal Zero-Shot Goal-Oriented Navigation cites this paper.

FSUNav: A Cerebrum-Cerebellum Architecture for Fast, Safe, and Universal Zero-Shot Goal-Oriented Navigation Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:03:08.600996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T18:58:48.779518Z digest=sha256:783903208c7445d46197bb709ec796732db2b4542b99de33e01d8e8516d030f3