Pith. sign in

Paper Citation Record · LEDGER

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

As of 8 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 2 inbound Pith citation observations for arXiv:2507.04047.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04047 v2

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:02:33.215747Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:54:57.184897Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T19:03:08.599602Z

Reference resolution

96 of 96 outbound references displayed

  • verified exact2
  • verified fuzzy59
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e2a31111-e2e4-4890-a72a-2b5fe7745999 · outbound

This paper cites Scanents3d: Exploit- ing phrase-to-3d-object correspondences for improved visio- linguistic models in 3d scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanents3d: Exploit- ing phrase-to-3d-object correspondences for improved visio- linguistic models in 3d scenes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.156912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.156912Z digest=sha256:9337ce74ab6883e2f20fd0a43f27a2126f65053fd12fdd8613142cac20789e37

Observation 7f2de7b2-1a14-431d-a5c8-30bbdd9a6843 · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.251850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.251850Z digest=sha256:cc544cd44b3ffce0ab1ed66e26b941fe89955d8d1a8321005b98dac9d85f4ae5

Observation d4f365fa-c59c-42b1-9e52-546e5375435c · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanqa: 3d question answering for spatial scene understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.317126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.317126Z digest=sha256:71dcd5895f54f4241a3458b56c0a8058988d4516f2a877d223095e25a9346ea3

Observation c49bc4df-6ccb-4a9e-9bbe-fd1b50ea2390 · outbound

This paper cites Do as i can, not as i say: Grounding language in robotic affordances.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Do as i can, not as i say: Grounding language in robotic affordances

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.434833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.434833Z digest=sha256:97a6be05e14ce828bdf53db389433cd8e4f6349f158f6d5bf4892481e76346d7

Observation bb7ccfdb-041b-411f-993c-23f9995f6668 · outbound

This paper cites 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.497301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.497301Z digest=sha256:6d45208f628a6c9bb1f289e7f376788e19c8e28efb23620150fd5697e1d25b25

Observation f9dddfbd-8cbb-48ea-b292-57527f9aadad · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Emerg- ing properties in self-supervised vision transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.605674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.605674Z digest=sha256:f388daddd9908d808f3b2f16155bad5b29bf101d650e1307d5523eab674ed940

Observation deda0b58-fc06-49a7-b970-3e239931f8ae · outbound

This paper cites Object goal naviga- tion using goal-oriented semantic exploration.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Object goal naviga- tion using goal-oriented semantic exploration

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.679473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.679473Z digest=sha256:62b0978bf54e976e98996a9020d4c5b98904ba2d970ba2eb9681c8c35d0bab28

Observation b823c8a9-85dc-4284-8432-c24334a7769b · outbound

This paper cites Object goal navi- gation using goal-oriented semantic exploration.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Object goal navi- gation using goal-oriented semantic exploration

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.763679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.763679Z digest=sha256:1cea821d31e9edf08e83b835908a3128a3093d9c9f6ed9ecd1724f52027cf4d6

Observation 86802e41-5707-4c97-87d5-a88d5deb6b24 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natu- ral language.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanrefer: 3d object localization in rgb-d scans using natu- ral language

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.842833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.842833Z digest=sha256:ffac61175c30fe076281821828fb6b5b8fd3da665725e6751a65e5bf3388c5e2

Observation 06333299-344d-4207-b617-3a79b9607d6c · outbound

This paper cites Language conditioned spatial relation reasoning for 3d object grounding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Language conditioned spatial relation reasoning for 3d object grounding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:31.924326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:31.924326Z digest=sha256:05cc48082ccd252023f1351de0aa1c282f1d4a078ae81345a839bcf5c5c196c6

Observation 778803a5-506f-493f-99d9-949b414b8d32 · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.014240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.014240Z digest=sha256:7a905f34b6e78d989a8487b628ef847a3401569dd4286826ec28486fa3af6472

Observation e4822169-e117-457f-8f5d-242e6d091ec3 · outbound

This paper cites Scan2cap: Context-aware dense captioning in rgb- d scans.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scan2cap: Context-aware dense captioning in rgb- d scans

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.057605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.057605Z digest=sha256:8a357c1ccbc1a3495199aef1acf86f7f6934087e4686f43a763fe20aada51183

Observation 26c200ce-8196-49a5-a404-ede974461053 · outbound

This paper cites Unit3d: A unified trans- former for 3d dense captioning and visual grounding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Unit3d: A unified trans- former for 3d dense captioning and visual grounding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.194145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.194145Z digest=sha256:0808e445309154058646c4ee432c5f5b352815696358cc79c27ff2a676a5a1bf

Observation 96a42d0e-704e-4d5e-b86c-f82703300d84 · outbound

This paper cites Schwing, Alexan- der Kirillov, and Rohit Girdhar.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Schwing, Alexan- der Kirillov, and Rohit Girdhar

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.284849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.284849Z digest=sha256:75dae74652869cc89c173dcd93467207b82bca8a5cf94896c4729ee3615ae334

Observation ed1e8638-0b8a-4a26-9984-d39ecfa84001 · outbound

This paper cites 4d spatio-temporal convnets: Minkowski convolutional neural networks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 4d spatio-temporal convnets: Minkowski convolutional neural networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.391830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.391830Z digest=sha256:1b9306a11cfd9a1ec783fb85a5678c027a4ac9e86759e369333844f8b43bd713

Observation 17ccbf1e-189e-4f93-b408-cafceb94bcaf · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.507880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.507880Z digest=sha256:569a275b44fd11297cbd021c4f51af8a56ec19c777191909ce73f7bc46d2b577

Observation 23dd4bc8-824e-48b8-99bd-72233555fe15 · outbound

This paper cites Embodied question answer- ing.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Embodied question answer- ing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.662553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.605207Z digest=sha256:337780257304ad935f4246edb6ff0055d62cea57aedc70d9cc72dd82a492040a

Observation fec0ec8e-b139-451f-b6b8-91dcf8ba2993 · outbound

This paper cites A survey of embodied ai: From simulators to research tasks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation A survey of embodied ai: From simulators to research tasks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.648665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.730308Z digest=sha256:45ca6cb955869f3ceaec6ea9280cf598cdafc72a4abbc0f406039784b1cefedb

Observation 6db28803-3cf2-4251-ad77-654dcd228b67 · outbound

This paper cites The One RING: a Robotic Indoor Navigation Generalist.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation The One RING: a Robotic Indoor Navigation Generalist

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.828420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.828420Z digest=sha256:765dc33117c075ce8649056fb41ff885bec3e62bf71c445e0d7a6d7252f9077a

Observation 4a984db5-a1cb-4e85-842d-0a3a593d6b52 · outbound

This paper cites Spoc: Imitating short- est paths in simulation enables effective navigation and ma- nipulation in the real world.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Spoc: Imitating short- est paths in simulation enables effective navigation and ma- nipulation in the real world

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.633572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.859253Z digest=sha256:9709cd2652e32f68444ee528bc4a8982448f849bfad202863a17ecb8236d82fc

Observation 713006fa-6b05-414d-a629-451cd2717320 · outbound

This paper cites Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.864163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.864163Z digest=sha256:408e9e026ae568bafa3417b94be5ec6da4c28e61c4757677dc3485046f2393d5

Observation 07a13e23-4b26-4210-91eb-1174c07e9df7 · outbound

This paper cites Efficient graph-based image segmentation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Efficient graph-based image segmentation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.618261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.869619Z digest=sha256:f78607403ebb420f260ebfc084f9b3770f82e44fa4a50c16295620d82de9e3e4

Observation 1a3fd2f0-a809-4758-ad4d-a2be5e0be2cc · outbound

This paper cites Cows on pasture: Base- lines and benchmarks for language-driven zero-shot object navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Cows on pasture: Base- lines and benchmarks for language-driven zero-shot object navigation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.604356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.873766Z digest=sha256:0784be243609f952de99125dc9918e5ad4e0ff55e63709c84b358e1d89e49ccb

Observation 8c68ce8f-f0a4-4e70-9234-da915752a457 · outbound

This paper cites Scaling open-vocabulary image segmentation with image- level labels.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scaling open-vocabulary image segmentation with image- level labels

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.589762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.878891Z digest=sha256:9d6bea2d0e11a6618018df40df9ac1e328c8eb4ed1f38ed8632efd737cabfdbc

Observation 45f432e4-e104-4efe-9322-2a706db8a7fa · outbound

This paper cites Conceptgraphs: Open-vocabulary 3d scene graphs for per- ception and planning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Conceptgraphs: Open-vocabulary 3d scene graphs for per- ception and planning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.575555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.883132Z digest=sha256:6d2aae8293ce100fcd801449b056f7d73d531b68b443c1bea523673c2f792ae4

Observation adf93ba6-943c-4359-bfef-d23cc8551d51 · outbound

This paper cites Viewrefer: Grasp the multi-view knowledge for 3d visual grounding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Viewrefer: Grasp the multi-view knowledge for 3d visual grounding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.561287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.887401Z digest=sha256:af92ed834f7215acc51e2e8f7225f622e0b6d5d4d23d3465eeea5b1fd3309f23

Observation 7d587f8b-2923-40f4-8826-daaabe53f5d4 · outbound

This paper cites Transrefer3d: Entity-and- relation aware transformer for fine-grained 3d visual ground- ing.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Transrefer3d: Entity-and- relation aware transformer for fine-grained 3d visual ground- ing

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.546556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.891844Z digest=sha256:f1908db6ee5fda7758b23461b98809f86252ef3ef1c483ed515b02c9cf7c785e

Observation 0a0a1021-5c64-4d71-8425-ea1db88b305a · outbound

This paper cites Vln bert: A recurrent vision- and-language bert for navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Vln bert: A recurrent vision- and-language bert for navigation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.531842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.896182Z digest=sha256:6689aa3d36f3d8163beb7f8d414a5339ec4d8fa2d40d3d3f6fd8be02c785f4af

Observation d4c652b4-9733-408a-8d2e-c5eb0468f7ee · outbound

This paper cites 3d-llm: In- jecting the 3d world into large language models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3d-llm: In- jecting the 3d world into large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.515695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.901425Z digest=sha256:a3d54c8ac40918bf7c00cba1853377fa8396668f2896dea9e37500d795a4eb9d

Observation 7f29a1b4-1c1e-4eb9-955e-9b9381b77dcd · outbound

This paper cites A real-time occupancy map from multiple video streams.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation A real-time occupancy map from multiple video streams

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.498323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.906117Z digest=sha256:259063f066d60c261449b97bca16a59d214c436bc9960c6ac0a9ce128a23ae65

Observation 9f43c2d6-7681-47f4-b4a6-54b61de930e7 · outbound

This paper cites An embodied generalist agent in 3d world.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation An embodied generalist agent in 3d world

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.481002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.910518Z digest=sha256:ffd821e35cbf5df577e2683363f474d5521fd0eb12289d18e5402f25b519332a

Observation 4c5594b4-3ee3-422a-b71b-7dba27b78f5f · outbound

This paper cites Language models as zero-shot planners: Ex- tracting actionable knowledge for embodied agents.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Language models as zero-shot planners: Ex- tracting actionable knowledge for embodied agents

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.463894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.914903Z digest=sha256:69ed10bc179e3bd2cdc0f5f422711c9495b9350d5c5df4eace348cdbe188dd0a

Observation 143deb32-0c82-488b-81f8-15730725a8d9 · outbound

This paper cites V oxposer: Composable 3d value maps for robotic manipulation with language models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation V oxposer: Composable 3d value maps for robotic manipulation with language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.448269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.919522Z digest=sha256:9bda2949f80320314a7614db98179670d01bf3b76a16a704fa05dcc0d37fe6dd

Observation 39bbe7cb-7eaf-46c9-a293-620806b5320b · outbound

This paper cites Bottom up top down detection transform- ers for language grounding in images and point clouds.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Bottom up top down detection transform- ers for language grounding in images and point clouds

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.433025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.924413Z digest=sha256:24c86e0be129c04a3a62c91346a3e6af694b169d9cccb9278737590b6bd5eabd

Observation 8f216072-19fc-4d8f-9511-ee0f109fc652 · outbound

This paper cites Conceptfusion: Open-set multi- modal 3d mapping.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Conceptfusion: Open-set multi- modal 3d mapping

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.417777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.929001Z digest=sha256:ad10362eb8ce148cb29abfdfc021805a78d9c20b1e86b15888433f4423358389

Observation e52ef426-8946-4d81-856e-63b42e949343 · outbound

This paper cites Sceneverse: Scaling 3d vision-language learning for grounded scene understanding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sceneverse: Scaling 3d vision-language learning for grounded scene understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.402368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.934381Z digest=sha256:743ec98b3bd2ca0241baa79c3ed994a1e652282a64b3c673456f5c4c4b9f4a27

Observation b46b61ae-463b-42ff-a9ca-487841e009f3 · outbound

This paper cites Goat-bench: A benchmark for multi-modal lifelong navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Goat-bench: A benchmark for multi-modal lifelong navigation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.386201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.939241Z digest=sha256:084555c218cbd3dcf3afe0c0ab9c01a8f899d98746f12e70cd6ddf871e42c232

Observation 9eeaf62f-5c89-4be2-be2a-02d3f7d04fb8 · outbound

This paper cites Realfred: An em- bodied instruction following benchmark in photo-realistic environments.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Realfred: An em- bodied instruction following benchmark in photo-realistic environments

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.371553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.943797Z digest=sha256:e53d96f5e011644538b31281c4bcdff50c34c32c4119bf5ff50b075b63939692

Observation 30f6df15-b03d-4f0c-b060-014a2e9d8e01 · outbound

This paper cites Segment anything.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Segment anything

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.357025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.948191Z digest=sha256:ce15a9ff466fd00ea0da5ff2b76ef287e520027f763b8f70d5b4e78960de3f19

Observation b1919572-cd83-449d-826b-e48b8505acfe · outbound

This paper cites UniCLIP: Unified Framework for Contrastive Language-Image Pre-training.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation UniCLIP: Unified Framework for Contrastive Language-Image Pre-training

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:02:33.577190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.952478Z digest=sha256:7ccd66771a7c3ce221b79db1b77e5716ffb5684fb98820251579af6018a88aaa

Observation 8207f357-d3b8-407e-a437-96f5b9a07aa8 · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Less is more: Clipbert for video-and-language learning via sparse sampling

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.342337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.956776Z digest=sha256:201e66e5547efb4595b8e5a6f683875eda1863229d5fb20527f0c5a7f9c27d93

Observation ad92fa97-2683-4468-9993-289fb62236ff · outbound

This paper cites BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.960840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.960840Z digest=sha256:32b9e222388ac36effe64f4b985ff559f916a904b4ab44f94db038af501f3930

Observation 7d582015-4573-4d72-90f1-53a9abbd8ed0 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.966191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.966191Z digest=sha256:5213e6d6f3b89921a771aaff6896ce4acc3826418b3f02d93b3c7c44911fb9aa

Observation 2a1c475a-e27f-4406-9282-799d56d74131 · outbound

This paper cites Panoptic segformer: Delving deeper into panoptic segmen- tation with transformers.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Panoptic segformer: Delving deeper into panoptic segmen- tation with transformers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.328116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.971408Z digest=sha256:df40aaa750efa6cfce6b4e81b1e3ea3e837696a2bc93113100a456d0cad31361

Observation 53c90110-e900-4bf6-b726-6d8daf6d9d1e · outbound

This paper cites Code as policies: Language model programs for embodied con- trol.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Code as policies: Language model programs for embodied con- trol

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.312598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.975605Z digest=sha256:f27b5e40ff161d17f4204d89ad3b6eb1c26f128da12edeb2e5bec19ae1174c4c

Observation b986deef-cc20-4681-a7cf-88509af4eabd · outbound

This paper cites NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.980063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.980063Z digest=sha256:7750643e5fb197d62add79b6f721a76e5960afd2ca5599a9a23a64250c55018e

Observation 2d1902b8-80a2-4ddf-90eb-aa605a81c199 · outbound

This paper cites Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:32.984728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:32.984728Z digest=sha256:4fa877d63c041f578215d9a29c1625a3ed8d66f2831c1f30851dd1d13bfd3f30

Observation c545085f-88f9-4e78-bb8c-50e4337486a6 · outbound

This paper cites Sqa3d: Situated question answering in 3d scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sqa3d: Situated question answering in 3d scenes

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.296444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.989899Z digest=sha256:075b3f43cd796d10117159d01b78a5c58dbdfc58c4180f0edba4d1a23a0241a3

Observation 41653663-ad6f-4d35-8548-2f546cc1f18e · outbound

This paper cites Zson: Zero-shot object-goal navigation using multimodal goal embeddings.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Zson: Zero-shot object-goal navigation using multimodal goal embeddings

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.280449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.994814Z digest=sha256:f377cb46cfb96419c944c2312989fd3d67295f0ad304334e461e31f7a450dbb8

Observation e621b6db-fe2c-4a9d-8df8-dbfe3a847b75 · outbound

This paper cites Openeqa: Embodied question answering in the era of foun- dation models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Openeqa: Embodied question answering in the era of foun- dation models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.211271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:32.999861Z digest=sha256:aa76b9b683677597a62cd890c2e8d7abba8776322384ed723ebd6cfb801e35ea

Observation 9b824185-5025-4395-aca8-665e709e9b42 · outbound

This paper cites Openeqa: Embodied question answering in the era of foun- dation models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Openeqa: Embodied question answering in the era of foun- dation models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.152079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.005435Z digest=sha256:8f40d23be34c4cb9cda1ae9dffdf282c37cae8ead536d1389dbcc63f218ab58b

Observation bb3fbf5b-e554-44e3-9818-b6a3b50163c3 · outbound

This paper cites Spatial memory.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Spatial memory

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.114012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.010136Z digest=sha256:da3c1fe1501a966599998d0738fd32eba7ce9262f26a630c62cc231350988382

Observation 6da5b11a-36a1-404a-b98c-a3b75be6c88e · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation DINOv2: Learning Robust Visual Features without Supervision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.014670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.014670Z digest=sha256:3e271cb1a96b38ff33be024002b3389bf5553d39fadb00d6621dcd174fd99c6a

Observation 729ec64a-1e53-4c9e-af84-67004af85111 · outbound

This paper cites Teach: Task-driven embodied agents that chat.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Teach: Task-driven embodied agents that chat

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.099440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.019013Z digest=sha256:652eaaee79cf7863e8c1795b144095fc77904f2bcf08f04356ba082e4de07bdf

Observation 91e13b98-5db9-4a06-b856-b57512660a26 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Openscene: 3d scene understanding with open vocabularies

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.082872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.023557Z digest=sha256:e2ce92cc4c3dd4ac6146a94bc822e0d0d1f2833cca48eb716398ee0402ac2778

Observation 0bb53d3a-b4fb-4ce6-a6bd-76e809f22e22 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Learn- ing transferable visual models from natural language super- vision

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.067457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.028017Z digest=sha256:caf08d127b0947d2554dc4cde21db6167fb6e47c4f864b56ffad75e0db6333d1

Observation 8337e663-f1b0-4a05-9243-ec3752d2a8ee · outbound

This paper cites Pirlnav: Pretraining with imitation and rl finetuning for objectnav.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Pirlnav: Pretraining with imitation and rl finetuning for objectnav

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.051867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.032520Z digest=sha256:fcdf547bc9567cea0ee952c892e2ae86f8ee9df329066bf3fce40d7de21da3a8

Observation 738a7c9a-cb3b-4a08-b6fb-3f9a641194f3 · outbound

This paper cites Sayplan: Ground- ing large language models using 3d scene graphs for scalable task planning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sayplan: Ground- ing large language models using 3d scene graphs for scalable task planning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.036765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.037520Z digest=sha256:ec51d001c94f4a26f4a2976b0b34b2b2357361792fcb20e0951615d3dbd507b0

Observation 841e3bdb-4be8-4196-9e45-3ab0b52d98ca · outbound

This paper cites Explore until Confident: Efficient Exploration for Embodied Question Answering.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Explore until Confident: Efficient Exploration for Embodied Question Answering

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.042822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.042822Z digest=sha256:c299cd6a210fdafd538054c9f0ef64a50d4b37c190d3c8ed9e6725c211518949

Observation 2bd1b137-7252-436e-90e0-50c9dee394da · outbound

This paper cites Language- grounded indoor 3d semantic segmentation in the wild.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Language- grounded indoor 3d semantic segmentation in the wild

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.019320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.047240Z digest=sha256:a80ad0346171999f4b0041ddc5c1150cf9c0c06d3d3e0b5aedc7103346625026

Observation acc76002-445d-4d6a-a6a7-d62271306133 · outbound

This paper cites Habitat: A plat- form for embodied ai research.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Habitat: A plat- form for embodied ai research

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:34.002845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.051322Z digest=sha256:f721890abe3625d78a472800dfc546af88915c1a1556cd8392944501a4c9d2eb

Observation 7bc365d5-0260-4142-b6f4-cb24fdbc7afd · outbound

This paper cites Proximal Policy Optimization Algorithms.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Proximal Policy Optimization Algorithms

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.055692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.055692Z digest=sha256:36325a6051d28bcf6845ead0d6d45c43e9c6c16724e4dc6db3837e8c3bd51651

Observation 46baae52-902b-4d3a-bb2d-173f6c84d29a · outbound

This paper cites Mask3d: Mask trans- former for 3d semantic instance segmentation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Mask3d: Mask trans- former for 3d semantic instance segmentation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.987323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.059795Z digest=sha256:efef40439efa35aac2a830bdbe299846bf10c476d7dda7f66369c5431751a4d2

Observation 0a686420-7872-4d1c-a62c-54e40f39f5ea · outbound

This paper cites Alfred: A benchmark for interpreting grounded instructions for everyday tasks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Alfred: A benchmark for interpreting grounded instructions for everyday tasks

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.971219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.064643Z digest=sha256:9ace543f7a66efa965492ebb510b4b5c6ba2721550a72877211c4c7083b0647a

Observation 170add76-540e-4418-8d40-8792d0b32b8c · outbound

This paper cites Llm-planner: Few-shot grounded planning for embodied agents with large language models.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Llm-planner: Few-shot grounded planning for embodied agents with large language models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.954645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.069207Z digest=sha256:672d82182a4273e0cd225bf1b4146dc357d43593c6c4befe1442a23e771771b8

Observation 9bf74e8c-1c66-4fcf-9d3e-71c08d19150b · outbound

This paper cites Habitat 2.0: Training home assistants to rearrange their habitat.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Habitat 2.0: Training home assistants to rearrange their habitat

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.073653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.073653Z digest=sha256:af4ad68e10616fb1e649c0c8606728bc934301cebfb2ba7e28ee5f6226da86a9

Observation c9530d43-f3e2-4889-b87e-9045a8cd09d1 · outbound

This paper cites Sumner, Marc Pollefeys, Federico Tombari, and Francis Engelmann.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Sumner, Marc Pollefeys, Federico Tombari, and Francis Engelmann

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.927417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.078435Z digest=sha256:1d14908378eb9baa0629413564eacd222e4c6b20d4a4669683ee451cfbe2e5b6

Observation c9361746-d768-4873-a170-e937b6194819 · outbound

This paper cites Rio: 3d object instance re- localization in changing indoor environments.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Rio: 3d object instance re- localization in changing indoor environments

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.911980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.083330Z digest=sha256:41dc201969af2f5f0a1848199433e17dd999bd0bad38e70de51b016b3233de98

Observation c3958155-876a-4424-a8c8-c14914faa310 · outbound

This paper cites Embodiedscan: A holistic multi- modal 3d perception suite towards embodied ai.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Embodiedscan: A holistic multi- modal 3d perception suite towards embodied ai

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.895603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.088684Z digest=sha256:4c06ec47c959feb7a948cb8e8f2e8a2a86f8dfef70f79cb73439aebd5b332a5b

Observation 9840311b-3a6c-44c1-b628-fe68e2d051ac · outbound

This paper cites DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.094358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.094358Z digest=sha256:48acde52031a99f8ae1aa437659ec001d272b7b93640551e0bc0766c57115c20

Observation d855c8e7-24ce-4c92-bf00-1cde8f9bce5c · outbound

This paper cites Ver: Scaling on- policy rl leads to the emergence of navigation in embodied rearrangement.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Ver: Scaling on- policy rl leads to the emergence of navigation in embodied rearrangement

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.880023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.099421Z digest=sha256:ef7bf9b9acc7dc38fb1b2a8bbe7503e66a594533db8ff3e0d898a818638f9058

Observation c4dbf192-0e51-4d22-9090-66021d4bdc9b · outbound

This paper cites Scenegraphfusion: Incre- mental 3d scene graph prediction from rgb-d sequences.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scenegraphfusion: Incre- mental 3d scene graph prediction from rgb-d sequences

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.865248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.103822Z digest=sha256:283e746fb1b222744d03da0197cbeef8204fe6d3b95035b4a6bce729512f6955

Observation 73cebb80-6c31-4bd7-be82-84128aef8574 · outbound

This paper cites EmbodiedSAM: Online Segment Any 3D Thing in Real Time.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation EmbodiedSAM: Online Segment Any 3D Thing in Real Time

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.108579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.108579Z digest=sha256:11ec6f96fa241a717fc6e623296abf3cef9436ab182007a2e8e321fbb74b46b9

Observation 8f204022-07ac-431f-b824-1607b07efda7 · outbound

This paper cites Habitat-matterport 3d semantics dataset.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Habitat-matterport 3d semantics dataset

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.850264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.113772Z digest=sha256:3010b2afb0f0f67278de4c9d2ddf9f9e0b3febe365f8ddf83e46f877131fbbde

Observation 66540fd2-5ebc-42da-8a41-49a9ceb41a84 · outbound

This paper cites Frontier-based exploration using multiple robots.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Frontier-based exploration using multiple robots

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.835273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.118819Z digest=sha256:cdbacabf4e20d93707fde224c5acdc599f3ddc2dc2fedf30a226c923a14562a8

Observation 1525a073-ef2b-44c6-bb44-cf0d9203795b · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.123605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.123605Z digest=sha256:47ab79df374b84629531ddcd3adb5c78be7451bff881cb136e64415ccbfadebf

Observation 6b28ee52-6bb1-45e5-a211-0a0961a5e8fc · outbound

This paper cites 3d-mem: 3d scene memory for embodied exploration and reasoning.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3d-mem: 3d scene memory for embodied exploration and reasoning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.819523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.128531Z digest=sha256:9ab96e2227e9ea2e976ebf08410b68a94140c532dda2caf158d7cbce7277e55a

Observation f7902776-3426-4812-ba1d-835c3e02c367 · outbound

This paper cites Vlfm: Vision-language frontier maps for zero-shot semantic navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Vlfm: Vision-language frontier maps for zero-shot semantic navigation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.804490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.133341Z digest=sha256:a2c6187f284b57fcaba45029dc25275114846898604ae79228eb2e911b15d864

Observation e846a639-0e12-4c7f-96d5-1bab80731938 · outbound

This paper cites Hm3d-ovon: A dataset and bench- mark for open-vocabulary object goal navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Hm3d-ovon: A dataset and bench- mark for open-vocabulary object goal navigation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.788045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.138214Z digest=sha256:0c65de2da7b41eb7fa5222e4e85d887da684ffe543b2f58b44060ea3e53a55fe

Observation 0e5784ad-8c37-4039-8d60-94a6b91f6d0e · outbound

This paper cites Frontier semantic exploration for visual target navigation.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Frontier semantic exploration for visual target navigation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.770186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.142761Z digest=sha256:2a3d8c0b03fdde41ab3064833ab76ad629a5d6e319cc3d0e9b2785e25ce0443e

Observation aae9a2f4-215d-4950-8a5f-1f54cf5c022c · outbound

This paper cites Instancere- fer: Cooperative holistic understanding for visual ground- ing on point clouds through instance multi-level contextual referring.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Instancere- fer: Cooperative holistic understanding for visual ground- ing on point clouds through instance multi-level contextual referring

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.754549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.147218Z digest=sha256:a3bed1444ffb86af1084bd8f9bc72bac03398fc4038e50a6a427dfc10f524871

Observation 2609d05e-91ce-4d80-b285-5d697180c3b5 · outbound

This paper cites PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.151309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.151309Z digest=sha256:8713da562a3617d23993d2bddf0d303aa82b32f60ee9ba3a1358999997b97b59

Observation 583e3afd-8e7f-47ef-9722-8da0e9f8a503 · outbound

This paper cites Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.155998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.155998Z digest=sha256:1936ab24fe7d68733923508fedc2a2869c64831105d61524e97c3470f81ff011

Observation 6063f092-66ec-41ba-84c4-cb0f6f1045a0 · outbound

This paper cites Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:02:33.334560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.160598Z digest=sha256:8dee864f87d3e37279c65960da5d22d8635cf2cda25b80b6cddfc5975f551647

Observation 5f9780c8-10d5-4986-8f9a-4d97c8147ca9 · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d objects.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Multi3drefer: Grounding text description to multiple 3d objects

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.740367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.165673Z digest=sha256:37b15e3cc51ce3a2c77f37ba6057317d9079f75edd2ca428f4e2881f4f6919f3

Observation f7906a49-1d19-4666-ad00-b2842a1db9ca · outbound

This paper cites Microsoft kinect sensor and its effect.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Microsoft kinect sensor and its effect

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.723322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.170543Z digest=sha256:ea6980ec603bbaf7ad371e215e365c170204cee1584e106ec749c8874b1ae40b

Observation f7fba5c5-b1d6-49a2-a672-f6d67140bc4e · outbound

This paper cites Task-oriented Sequential Grounding and Navigation in 3D Scenes.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Task-oriented Sequential Grounding and Navigation in 3D Scenes

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.174687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.174687Z digest=sha256:eb622ab93a594babc5c335496c6bc9c34972faca2a848569ae99ebfd2f11aabf

Observation b0f133cb-4108-4947-b1fc-836c269d090e · outbound

This paper cites To- wards explainable 3d grounded visual question answering: A new benchmark and strong baseline.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation To- wards explainable 3d grounded visual question answering: A new benchmark and strong baseline

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.708124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.179080Z digest=sha256:eecf877c6e0dc212706ca512c5c5711da69318f5b3e68b257df9a6981ae6dabe

Observation a10528fa-1096-4e34-adf2-8d1bfbaa1104 · outbound

This paper cites Fast Segment Anything.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Fast Segment Anything

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.183289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.183289Z digest=sha256:b566c6f84511ae183d14fb3a21dfd8408807271b59c1e2ecf6fceb190e0398cb

Observation c0dbc718-15b3-4fef-ba94-993a9eaf9604 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.187976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.187976Z digest=sha256:32378972877ab64eeee86e38ff09ebc75d987fc3c4391845fbd4082b9d6ff7fe

Observation 83388a1b-427e-42dd-a5f5-0aa425a01d7c · outbound

This paper cites Dual memory units with uncertainty regulation for weakly supervised video anomaly detection.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Dual memory units with uncertainty regulation for weakly supervised video anomaly detection

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.690761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.193178Z digest=sha256:6bb123773c8c033f0c172ff54fa041db1538a3601998a951537ec69d28a153d7

Observation 96ec4986-1805-49e8-ad9c-1def13a40b0b · outbound

This paper cites Scanreason: Empowering 3d visual grounding with reasoning capabilities.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Scanreason: Empowering 3d visual grounding with reasoning capabilities

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.676105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.197976Z digest=sha256:6ad85bb1333df746841ffd69375615709553858f134771db1f3213bc528890fb

Observation d4465aff-60cc-4572-b653-dfdb711f5ee1 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.660133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.202562Z digest=sha256:036cb1c206bf0c85b73049b32ecb6408b15cfb30b8ec422c49a9d1df968f4609

Observation 225f2ccb-d3ce-4433-9dbc-f8934a0e1169 · outbound

This paper cites Unifying 3d vision-language understanding via prompt- able queries.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Unifying 3d vision-language understanding via prompt- able queries

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.644142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.206561Z digest=sha256:b83e3e22e23665bbc07d82ac36ac519b28e571fa58dcbfd0d5bd14f5a1c93e8c

Observation 17a85cb1-9423-4a72-994e-4fc5374ca733 · outbound

This paper cites TANGO: Training-free Embodied AI Agents for Open-world Tasks.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation TANGO: Training-free Embodied AI Agents for Open-world Tasks

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.210895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.210895Z digest=sha256:cb2dba2e866ca12a722511b43261b7ddbfcac2faa3360acf3aae0954362d4636

Observation be143c56-3127-4de0-84c2-0274233b3c85 · outbound

This paper cites Generalized decoding for pixel, image, and language.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation Generalized decoding for pixel, image, and language

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:02:33.628398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:02:33.215747Z digest=sha256:58616d38b3f949ff35061a79fdffad78c93b4a858783a1551a6f483d461059c7

Pith citing papers

Observation 8b12b765-c17c-4532-8932-0c2f5ffe3a7d · inbound

SPG: Style-Prompting Guidance for Style-Specific Content Creation cites this paper.

SPG: Style-Prompting Guidance for Style-Specific Content Creation Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:57.184897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:57.184897Z digest=sha256:af1619e9e5b482b92afac2dc6f0b7de63294af050694a96234eb666aa0afd7ee

Observation 7363683b-1a7a-41e1-8f77-a918c1574094 · inbound

FSUNav: A Cerebrum-Cerebellum Architecture for Fast, Safe, and Universal Zero-Shot Goal-Oriented Navigation cites this paper.

FSUNav: A Cerebrum-Cerebellum Architecture for Fast, Safe, and Universal Zero-Shot Goal-Oriented Navigation Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:03:08.600996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T18:58:48.779518Z digest=sha256:8566feb4db45bb6cb06103a548748e8524c293699423cb02c212762a5db75617