Pith. sign in

Paper Citation Record · LEDGER

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields

As of 11 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 0 inbound Pith citation observations for arXiv:2506.23352.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23352 v1

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:49:48.230001Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

96 of 96 outbound references displayed

  • verified exact2
  • verified fuzzy64
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9a2a456f-48b1-461b-9bab-b5febf5309c0 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Scanqa: 3d question answering for spatial scene understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.054070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.054070Z digest=sha256:20eef025317b9f7932f7f8eb14f6f8b9da226ce70002a8b6f01908fe23bc5064

Observation 7bba69c6-6d78-4fd5-89d6-9cdd55d2ada1 · outbound

This paper cites Qwen2.5-VL Technical Report.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.094181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.094181Z digest=sha256:2dae8a6e19cabb2a2ce802bea0f9635d80b450fddd60b66020ee1a3bf08a880e

Observation 3922716b-968a-46df-9b06-41652b632605 · outbound

This paper cites Henriques, Andrew Zisserman, and Andrea Vedaldi.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Henriques, Andrew Zisserman, and Andrea Vedaldi

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.194838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.194838Z digest=sha256:5d87aad587a3b4a0249682461922ff928bdce1d0a69cadd2019482ce261ebaf4

Observation 05a9f140-72af-403e-8111-ef8f7d6c539b · outbound

This paper cites Blumer, Qingx- uan Chen, and Francis Engelmann.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Blumer, Qingx- uan Chen, and Francis Engelmann

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.256991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.256991Z digest=sha256:f5991aa5b7ed68ae54ecc7db7ee78dca48cf85727dab5edc333c4d853b797f9c

Observation b093582a-470c-4d28-b589-ce745a4c8e3c · outbound

This paper cites A persistent spatial semantic representation for high-level natural language instruction execution.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields A persistent spatial semantic representation for high-level natural language instruction execution

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.392137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.392137Z digest=sha256:04581f1ffe39fa2f1733d1c6f52283a155138d649bee7ce8bfe2e352817c53d9

Observation 71ca8a59-5ff8-47a7-9292-8244e40f89d9 · outbound

This paper cites Prompt-rsvqa: Prompt- ing visual context to a language model for remote sensing visual question answering.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Prompt-rsvqa: Prompt- ing visual context to a language model for remote sensing visual question answering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.499765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.499765Z digest=sha256:6067d6d6d915265d25c9ff8b87276320fdaf4673fb910415f8ce7a2ef2f212fb

Observation dee5502b-394d-4b8a-9d94-53bcfd086c67 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natu- ral language.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Scanrefer: 3d object localization in rgb-d scans using natu- ral language

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.634508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.634508Z digest=sha256:b4664d3f33a5af5bb925f6d76334dc1bd969de2f7bb76ebb9ab9f34a071dd321

Observation 2d18a5ce-27c7-4499-9d5a-6992720f8660 · outbound

This paper cites Panoptic vision-language feature fields.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Panoptic vision-language feature fields

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.745471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.745471Z digest=sha256:a980c22fb12e1fa774436c1d78d9f23c70dc2f7193aab67587f0efe7e2670fff

Observation e45035e5-1c0c-469c-a16e-8289272ccb08 · outbound

This paper cites Stylecity: Large-scale 3d urban scenes stylization.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Stylecity: Large-scale 3d urban scenes stylization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.839057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.839057Z digest=sha256:fadfca25a2015bfd10a638dc2be867b0c5730874dca28a26d84576216f2a644b

Observation d71297e6-2235-430c-96e6-bf273376ff1e · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.929873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.929873Z digest=sha256:dc09927a09538d0682bff4aa40c01bb7ead1e2a458adfd9226174bdc5668a7e2

Observation 00fe7b97-3e74-4516-ab3c-9da16c510e9c · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.012679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.012679Z digest=sha256:a6f85fcd21ccf4a063898262252d861d583b0d77e305541f8123f0d1739161e1

Observation da1aa099-20a3-46e0-a396-41eb294e5e9e · outbound

This paper cites LayoutGPT: Compositional Visual Planning and Generation with Large Language Models.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields LayoutGPT: Compositional Visual Planning and Generation with Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.132312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.132312Z digest=sha256:a278abc46cd96fe2b31a5d08f5a53a8b20426e3b7ee1259c4eef473f3d0fee73

Observation 1999571b-cd9a-407c-bc5f-ccfc1b13af7e · outbound

This paper cites Dynamic 3d gaussian fields for urban areas.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Dynamic 3d gaussian fields for urban areas

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.190261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.190261Z digest=sha256:2b720415abfe88051b9e877af18262db4e941d29a2acafb9642794101c903346

Observation 70eb2b3b-63eb-4257-a5a6-a8e64cbdd79a · outbound

This paper cites Ue4-nerf:neural radiance field for real-time rendering of large-scale scene.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Ue4-nerf:neural radiance field for real-time rendering of large-scale scene

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.253760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.253760Z digest=sha256:66007f4bbb3ece303346c1018913fefc56fdba4f9ad7819226e6e28955449eaf

Observation a5fa8581-1343-47b9-8e7c-a86977f9e80d · outbound

This paper cites StreetSurf: Extending Multi-view Implicit Surface Reconstruction to Street Views.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields StreetSurf: Extending Multi-view Implicit Surface Reconstruction to Street Views

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.307473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.307473Z digest=sha256:20316b327c899d703b170245c72256b16ec9ecd8b5f6504a94ad296aa1973c9d

Observation 36a3ba48-3582-43fb-8f84-603d4d76c0a2 · outbound

This paper cites Visual program- ming: Compositional visual reasoning without training.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Visual program- ming: Compositional visual reasoning without training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.378479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.378479Z digest=sha256:ffdf180edc03cf0dbca4328fac8da7ea94efa6f09de6b85fdeb41f2506a64ee4

Observation 15cf77ae-37cc-48a1-9156-f7661f6bd688 · outbound

This paper cites Pigeon: Predicting image geolocations.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Pigeon: Predicting image geolocations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.437098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.437098Z digest=sha256:6e21ee413371b674f5db6ee7084fae50731e50c8dd5338ac0fc824342252eced

Observation a4e7e80a-1a68-44a6-a34d-90d8b1bbb3b8 · outbound

This paper cites Dragon: Drone and ground gaussian splatting for 3d building reconstruction.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Dragon: Drone and ground gaussian splatting for 3d building reconstruction

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.490242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.490242Z digest=sha256:9ad97550e7e265f7dcc958eeed8c7133c3171795f3281459b1d9e6431b73281e

Observation eb5e3a97-aac3-425d-a4cc-cc6dec241d00 · outbound

This paper cites 3d-llm: Inject- ing the 3d world into large language models.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3d-llm: Inject- ing the 3d world into large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.533958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.533958Z digest=sha256:8ee670c87f6b88b32403f19e630399b8fe7f5c01a9caa8e601b41a93dc05c63f

Observation 3edb0cda-0e2e-45c5-b5c1-bc48b5709486 · outbound

This paper cites RSGPT: A Remote Sensing Vision Language Model and Benchmark.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.574175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.574175Z digest=sha256:5afee2d3ae7b8896994b177be541f8db1eaaefa9bbf456670d351e1e01307032

Observation 57381278-8928-4fa2-bd2c-0c0df261d55d · outbound

This paper cites TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.614807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.614807Z digest=sha256:36d3af917a13c5ab53a64572080612e3de6b0c25b5a043e9f27c5458ffe68aa3

Observation cacaa869-fdd5-498b-aa6c-f7c052d768b2 · outbound

This paper cites GraspSplats: Efficient Manipulation with 3D Feature Splatting.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields GraspSplats: Efficient Manipulation with 3D Feature Splatting

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.660608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.660608Z digest=sha256:d00c5f757f64498194cd3cd009ad4161614c93b4e36be12f3e6fc69cf9d9517a

Observation da349638-f488-43c6-9293-b7e9a72dbcaa · outbound

This paper cites Fastlgs: Speeding up lan- guage embedded gaussians with feature grid mapping.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Fastlgs: Speeding up lan- guage embedded gaussians with feature grid mapping

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:50:00.199756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:43.729855Z digest=sha256:cba86d1ea5805b6c92878bb4a9cc3bd74a92484660b556b31f08b5949b98b757

Observation 4cc0bab2-8de2-4525-a191-6972b1f1a3ef · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3d gaussian splatting for real-time radiance field rendering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:50:00.069090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:43.778265Z digest=sha256:1aace15a3dfc445c79c5737d435bb14a1306ea353d5d5382b2d1ae50700360a5

Observation 5079a02f-46e7-4f71-96b1-0ae8264c004f · outbound

This paper cites A hierarchical 3d gaussian representation for real-time ren- dering of very large datasets.ACM Transactions on Graphics (TOG), 43(4), 2024.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields A hierarchical 3d gaussian representation for real-time ren- dering of very large datasets.ACM Transactions on Graphics (TOG), 43(4), 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.894873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:43.817893Z digest=sha256:5683b1fc9a25c43b6a9ebcbafb851e0356ac83eb0ff3ee920022e6f4aad8cd6f

Observation f03dc12f-8688-4d4f-b78a-3b4d89abee4c · outbound

This paper cites LERF: Language embed- ded radiance fields.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields LERF: Language embed- ded radiance fields

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.718873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:43.885600Z digest=sha256:2d7b8506920e2a4ebf28efdaa664a4b9448980c7f9cadb7b9bd5305cd11a1006

Observation 4d690336-278a-482b-b22a-6b0373ac9ea7 · outbound

This paper cites Lobell, and Ste- fano Ermon.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Lobell, and Ste- fano Ermon

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.590614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:43.952849Z digest=sha256:3ef6f30a42b0e18559fbb2ad43a1d93dfe4d88b3f156da7dd3f7540d3369017b

Observation 77bafc28-a913-4c26-8aad-80d6f1db3910 · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.476659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:44.000320Z digest=sha256:fc3640628e29f779d5aa1c9e65438d5a07ce89703f5ccd6f574613748766756f

Observation 512e23fd-0770-4804-8f56-d514e6f22f0c · outbound

This paper cites Segment any- thing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Segment any- thing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.303634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:44.042944Z digest=sha256:b342ca74178d8abc7aada3580d5d29298f7b4f95aad80bf866186216bc491c60

Observation 383964ea-5fb3-40a4-b019-7b71dbe8a3ed · outbound

This paper cites SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:44.101737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:44.101737Z digest=sha256:d71a9e2e0e51665278a53220c1f8371eaaa482055bb6b4c0a7e515ae78326963

Observation db4b74fd-35ac-4f0d-96c4-308024254d38 · outbound

This paper cites Decomposing nerf for editing via feature field dis- tillation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Decomposing nerf for editing via feature field dis- tillation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.166663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:44.136124Z digest=sha256:f390f965eeea865a9045b79d93e43cb3ca95101bc902cd607a2ff53d8fb47672

Observation 4594861b-9cfa-4255-9ca0-32b88981eb59 · outbound

This paper cites Text2pos: Text-to-point-cloud cross-modal localiza- tion.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Text2pos: Text-to-point-cloud cross-modal localiza- tion

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.064991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:44.170793Z digest=sha256:f7b74cb56c08d6cfa4e1809806f0f78d5c7c8a2bda446f12fb91c8d1d765bc1f

Observation 8eb4957f-fbb9-41b1-8285-94555bb82c3d · outbound

This paper cites Geochat: Grounded large vision-language model for remote sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Geochat: Grounded large vision-language model for remote sensing

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.903803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:44.222774Z digest=sha256:691c09cbce1a2244abe7982416f81983d3635632495e07ed6d7dcaf792b7aa45

Observation 7b1fd120-5535-440d-8b72-83d94a97fa55 · outbound

This paper cites NeRF-XL: Scaling nerfs with multiple GPUs.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields NeRF-XL: Scaling nerfs with multiple GPUs

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.773105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:44.267005Z digest=sha256:414027302ee6542d276ee8246e22a9c7b42982a02ed2388331af07815e6a1d73

Observation 51b0f3b3-7f83-4852-bc1d-694b87713b73 · outbound

This paper cites Matrixcity: A large-scale city dataset for city-scale neural rendering and beyond.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Matrixcity: A large-scale city dataset for city-scale neural rendering and beyond

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.613797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:44.309914Z digest=sha256:c4d00432656de3db036b7641548189b0ab7330837785c089df1c7fe7859abe91

Observation a5a46439-664f-40d8-a915-3f451efd8976 · outbound

This paper cites Vastgaussian: Vast 3d gaus- sians for large scene reconstruction.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Vastgaussian: Vast 3d gaus- sians for large scene reconstruction

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.486658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:44.345633Z digest=sha256:9cd9df75db1c8cf0bb65601fa30d16e499c2be17e02bf9f50b631967f1927726

Observation 11471b57-9a6a-442e-adb6-c54a6d7a3237 · outbound

This paper cites Capturing, reconstructing, and simulating: the urbanscene3d dataset.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Capturing, reconstructing, and simulating: the urbanscene3d dataset

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.345746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:44.377705Z digest=sha256:9c6e665f843a1a70201de5ddde604770720b66172683d0c226c741983a7fd67d

Observation 2ffa59ba-1675-421a-89c6-c32d5bd75650 · outbound

This paper cites Re- moteclip: A vision language foundation model for remote sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Re- moteclip: A vision language foundation model for remote sensing

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.217038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:44.441781Z digest=sha256:1b4509f9350acd727854293bb9f23aca0b1bf79dbe247b3a719dd337cc960413

Observation 321362df-eaa4-42f8-ad62-66f8dfe04287 · outbound

This paper cites Visual instruction tuning.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Visual instruction tuning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.049953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:44.492129Z digest=sha256:14be3aec2ceaa6c0dfb1b64e402b4232cacfa99166b1965e5feea990cb423150

Observation 7da5c1eb-1eb3-48c3-abb9-73bcdc066f97 · outbound

This paper cites Citygaus- sian: Real-time high-quality large-scale scene rendering with gaussians.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Citygaus- sian: Real-time high-quality large-scale scene rendering with gaussians

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.943733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:44.536960Z digest=sha256:e05cc4422a1a560f16d1bff0dbe16bc11e8f48b75f6b54488d4e4dfd429cb756

Observation c5b516e5-64a8-4bbc-9865-6fbfc9fb13b8 · outbound

This paper cites Citygaussianv2: Efficient and geometri- cally accurate reconstruction for large-scale scenes, 2024.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Citygaussianv2: Efficient and geometri- cally accurate reconstruction for large-scale scenes, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.732635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:44.579863Z digest=sha256:414aba69dbb3df21aa8a143ace9a3bf3b9e06e006abf9cfaf41977bd97029f6a

Observation 64464ef6-4f58-491e-a64e-c1aea177f006 · outbound

This paper cites Citygaussianv2: Efficient and geometri- cally accurate reconstruction for large-scale scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Citygaussianv2: Efficient and geometri- cally accurate reconstruction for large-scale scenes

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.486653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:44.633032Z digest=sha256:d82d8a5bbb0cd0e18781c111bbc3822d4af8a2a505b38d8040998879d445746f

Observation a39d028d-c282-40a8-841b-91b5017f9523 · outbound

This paper cites Chameleon: Plug-and-play compositional reasoning with large language models.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Chameleon: Plug-and-play compositional reasoning with large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.326550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:44.673401Z digest=sha256:02e301ee956ec0d9a013b4f494adb75684915eda1f9a3269c7b58b08e5937086

Observation 411a2e79-1115-4057-b457-d2753476d1ce · outbound

This paper cites Exploring models and data for remote sensing im- age caption generation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Exploring models and data for remote sensing im- age caption generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.153451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:44.722564Z digest=sha256:188e6869957cd8ade3ceb178eaf299b5add7cd66dfc05173efc580ed3d583b05

Observation 42f8f1c8-c23b-4d8e-bec0-fae4bce9c14e · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:44.779870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:44.779870Z digest=sha256:3755eb38ca952c8bc0041bcb6d24c469cc1617b7cb9e03939932aa558238b5ae

Observation 1c23294a-ff1c-43ac-9251-0d118650d631 · outbound

This paper cites A multiscale grouping transformer with clip latents for re- mote sensing image captioning.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields A multiscale grouping transformer with clip latents for re- mote sensing image captioning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.999396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:44.839366Z digest=sha256:3ebada86d98dd4a9cd52477cf280f86b07287d351ab5426ab781c3b2b6a17968

Observation 8692614f-bbd4-425f-b298-23576a5f2c93 · outbound

This paper cites Llama 3.2 connect 2024: Vision on the edge and mo- bile devices.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Llama 3.2 connect 2024: Vision on the edge and mo- bile devices

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.865133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:44.886549Z digest=sha256:1794c18845465b7c7d6fd19db4ed5a274a961cdb2cb6326b635c9dc538f53f01

Observation 968bc7a2-0782-4099-b787-0fc6d70410f8 · outbound

This paper cites Srinivasan, Matthew Tancik, Jonathan T.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Srinivasan, Matthew Tancik, Jonathan T

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.632184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:44.962424Z digest=sha256:32fa7651e708cfb8971a8b665c6e494283d5e3e90307e796818a8fceca0b9a24

Observation dadfd791-b5fc-4459-a5cf-e26c4aa29d90 · outbound

This paper cites Cityrefer: Geography-aware 3d visual grounding dataset on city-scale point cloud data.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Cityrefer: Geography-aware 3d visual grounding dataset on city-scale point cloud data

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.486953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:45.026632Z digest=sha256:dda22533dfbda76ee78a60338c2f3fadaf6922203025e8ec89087cd4f073cdfb

Observation c84971c3-8ab1-4f8c-98e6-f143ba7b7feb · outbound

This paper cites Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.346668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:45.078369Z digest=sha256:b829d69c8d54ce9b00225eef9b432155d7a117db7f565e43659a9c14f35185a2

Observation a6f53a50-2fbb-4fa8-9bb5-82087eb5ffff · outbound

This paper cites Hello gpt-4o.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Hello gpt-4o

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.242315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:45.171325Z digest=sha256:d762ae044cd49d821fc071aea66270a5b55e75c751c2710810343591d5b0de1b

Observation a7d2eac8-8edb-49c0-9d47-9a7287cd4e45 · outbound

This paper cites Vhm: Versatile and honest vision lan- guage model for remote sensing image analysis.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Vhm: Versatile and honest vision lan- guage model for remote sensing image analysis

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.086960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:45.242656Z digest=sha256:3088f39e202c71a44f9679c25edb7aef142148db4d686539f78655ca978d712b

Observation ede4e48d-4a1f-44e3-8d07-8c976f4a375d · outbound

This paper cites Langsplat: 3d language gaussian splatting.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Langsplat: 3d language gaussian splatting

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.894524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:45.317392Z digest=sha256:6c698218436dca187c8435631f8beb864779e83cce35e2f6da27c17759041064

Observation 597031c2-4287-4768-9ec9-9c81916f7abe · outbound

This paper cites Deep semantic understanding of high resolution remote sensing image.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Deep semantic understanding of high resolution remote sensing image

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.736156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:45.454471Z digest=sha256:93df91891327a23dec91bd7cb29b4767869b70efeed2eac7abc1a4d33fec747d

Observation c47a1362-4105-4afd-b80a-93a9f0db17c5 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Learn- ing transferable visual models from natural language super- vision

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.596461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:45.537922Z digest=sha256:980a6228a21f44dd2b76b0032d5865dc5bd9ab771a64127346c2d659431bd677

Observation 12ebebed-4dce-4e81-9e20-ce61af8ba3db · outbound

This paper cites Derf: Decom- posed radiance fields.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Derf: Decom- posed radiance fields

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.445680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:45.650119Z digest=sha256:0a02b8c9d343176aebf877257a229b3c31437c89e53b6234e8eedda4961e2d65

Observation 72561cb0-d64b-4d2e-9e0d-559e86cabb15 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Toolformer: Language models can teach themselves to use tools

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.249562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:45.703626Z digest=sha256:93cab4d29a402a045766e16d2a8c97c237ba1281085cdbef2c13a8e11045c0f1

Observation c64e465f-305b-429a-8906-0612a0ab092e · outbound

This paper cites CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:45.758919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:45.758919Z digest=sha256:b7c7d08db9f9b4fa196bfe7114bd08a551b1489bfb240377426c9a62e6742be0

Observation 1cca6068-bdbe-4e95-b372-49d2d1cc4e29 · outbound

This paper cites Language embedded 3d gaussians for open- vocabulary scene understanding.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Language embedded 3d gaussians for open- vocabulary scene understanding

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.086891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:45.808762Z digest=sha256:6f875c0149913960d683333a5416467fd20cad1cdf30306f5b7bf7247f3cfe0c

Observation 753f4515-7c61-4625-8f72-df9435fd83cc · outbound

This paper cites Real-time view synthesis for large scenes with millions of square meters.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Real-time view synthesis for large scenes with millions of square meters

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.947992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:45.867334Z digest=sha256:df9948183fb5d83c4849d4dc80c6d92c217f5911b247f0aa90bf2753569e5a2b

Observation 86280683-e13f-4125-a85d-f36cb77c5652 · outbound

This paper cites City-on-web: Real-time neural rendering of large- scale scenes on the web.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields City-on-web: Real-time neural rendering of large- scale scenes on the web

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.815656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:45.905136Z digest=sha256:b44411d5b1cd19f802bae886720eea25efe5391e251b7710c4ef637d00854b42

Observation b9c0206e-f8e7-4604-82a5-16898e815aeb · outbound

This paper cites De- composing 3d scenes into objects via unsupervised volume segmentation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields De- composing 3d scenes into objects via unsupervised volume segmentation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.633027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:46.023299Z digest=sha256:837e677d1584b2f8814372af9fe650fc4f3e56b20c3ed878992d56d65004fd38

Observation 700651a4-cb5f-4101-b462-ca96833df9cb · outbound

This paper cites Modular visual question answering via code generation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Modular visual question answering via code generation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.454110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:46.069829Z digest=sha256:1539a1c2e387e76facf90d44b10faea1ccee1bece1d60ad0bfc52921fb359b9f

Observation dc431695-f1e0-4f6f-a2d7-2bbbf9863b43 · outbound

This paper cites 3d ques- tion answering for city scene understanding.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3d ques- tion answering for city scene understanding

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.292203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:46.122863Z digest=sha256:36839e95f3e29e7e2a8da714de3aa73eb316d595c6fcdf133176db484950a0e4

Observation a116152e-ee2d-47bc-a708-0b273c10d0b5 · outbound

This paper cites Visual grounding in remote sensing images.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Visual grounding in remote sensing images

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.192127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:46.179689Z digest=sha256:e9620102485db85a0db3a88e011f32503497f50bd889fe009c43f483fce1563d

Observation 89523449-652f-4b7d-93aa-f2de235e1a3b · outbound

This paper cites ViperGPT: Visual inference via python execution for reasoning.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields ViperGPT: Visual inference via python execution for reasoning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.059622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:46.235926Z digest=sha256:0beebbc791ef5f248453571306abaf3038d350448f95a7c85155ea7908b50a7e

Observation d3fd2ac5-28eb-43fd-947f-dc05078d97f3 · outbound

This paper cites Srinivasan, Jonathan T.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Srinivasan, Jonathan T

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.881502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:46.293702Z digest=sha256:cfdfddbc037a7d34c9c1f8a1e8e76ad70b87deacdef947d4e7a398f465ffa920

Observation 57e570d9-5479-4183-9948-0f315a385603 · outbound

This paper cites Crs-diff: Controllable generative re- mote sensing foundation model.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Crs-diff: Controllable generative re- mote sensing foundation model

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.729228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:46.366132Z digest=sha256:8e7341318fc2f87cf4601ea8bef2a999547ea87f2fa9f3b9bf9741aa25a3476c

Observation e963ca73-2e31-42fd-b422-939edfc5e32f · outbound

This paper cites MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:46.416409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:46.416409Z digest=sha256:95d3d6e7297927538734077ed75176a67cd66db924531e248db042d5c5c2f8d8

Observation 43ac9ec0-b7ff-427a-a3de-96a9fb84d01e · outbound

This paper cites Mega-nerf: Scalable construction of large-scale nerfs for virtual fly-throughs.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Mega-nerf: Scalable construction of large-scale nerfs for virtual fly-throughs

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.611663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:46.462758Z digest=sha256:72c1f3cc932588b8b49c4fe40a2e0b1b07dd94e1f72327bc3158ee21e523344f

Observation 66328b1d-9627-491e-b845-e8701b4ae3d7 · outbound

This paper cites Geoclip: Clip-inspired alignment between locations and im- ages for effective worldwide geo-localization.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Geoclip: Clip-inspired alignment between locations and im- ages for effective worldwide geo-localization

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.524417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:46.542012Z digest=sha256:f476db0421d46e49769cfb63d2fcf56f8b39dc93c48e51f734f7565a1b6d4bac

Observation 08dcae4e-331b-4fde-a88e-c931cd680623 · outbound

This paper cites Skyscript: A large and semanti- cally diverse vision-language dataset for remote sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Skyscript: A large and semanti- cally diverse vision-language dataset for remote sensing

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.281261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:46.608758Z digest=sha256:7ccd3076104c41b5ea465232040ae92c30bf0f970465172ec389cc0d2f6eae5e

Observation 96d58b29-89ee-43af-b61c-3fb6d54e869a · outbound

This paper cites Text2loc: 3d point cloud localization from natural language.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Text2loc: 3d point cloud localization from natural language

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:52.994039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:46.661207Z digest=sha256:8cbccec81e73441039a0032e47131f9b5e895dbb3189c493c9e751c91d9565c8

Observation e6fdab54-f947-4033-8866-5f2f7e532a5c · outbound

This paper cites Bungeenerf: Progressive neural radiance field for extreme multi-scale scene rendering.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Bungeenerf: Progressive neural radiance field for extreme multi-scale scene rendering

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:52.782372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:46.717819Z digest=sha256:00368c21d4c6ed8975e0975ad2225e989359a0cb9034c1a1451b7f43a007144d

Observation 74ea19dd-d55c-4610-96da-75d4db71c88a · outbound

This paper cites Citydreamer: Compositional generative model of unbounded 3D cities.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Citydreamer: Compositional generative model of unbounded 3D cities

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:52.490533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:46.764874Z digest=sha256:3c16f3ec2b52ea0020646a4574807e51074b2f162b7ae7bb8ff77777f9523f7f

Observation 7cb1dfe4-9caa-4660-8e0f-322ee862b86d · outbound

This paper cites Generative Gaussian Splatting for Unbounded 3D City Generation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Generative Gaussian Splatting for Unbounded 3D City Generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:46.818968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:46.818968Z digest=sha256:4c9c8cb501fe7b8c09f560f042fe076f4c7b5d27a98f95aab6a9a4184488844c

Observation 8ad4195c-3cd5-4c51-9efb-b631cd4f3f37 · outbound

This paper cites Grid-guided neural radiance fields for large urban scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Grid-guided neural radiance fields for large urban scenes

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:52.184167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:46.895366Z digest=sha256:e5125f5d138b87abe1cb6ee8d6160cbaffd20c4034cd4afebe338cef09adc0b0

Observation b760559c-52cb-4306-a908-3e67e5e05eeb · outbound

This paper cites Pointllm: Empowering large language models to understand point clouds.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Pointllm: Empowering large language models to understand point clouds

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:51.883576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:46.938093Z digest=sha256:f021805e694737ab2364052c7ca883f565c2e11065ef3d38a4067880bd569ce6

Observation e3dd9d04-2419-4c03-a6a6-a96cfc73e4c1 · outbound

This paper cites Addressclip: Empowering vision-language models for city-wide image address localization.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Addressclip: Empowering vision-language models for city-wide image address localization

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:51.674509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:47.105618Z digest=sha256:dd29e1b389369ee185c0d0c63f1de2ecb1fc76819996afea151d64ad32539017

Observation c5b83c66-8f0f-4dff-983b-606d80e0b1ee · outbound

This paper cites Unisim: A neural closed-loop sensor simulator.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Unisim: A neural closed-loop sensor simulator

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:51.420687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:47.217718Z digest=sha256:1aebc0ce03f91a2d267b8e63d99375e7fb15aec4cb801e05e92b55b65ab2cd1b

Observation 8a8f6348-3a1f-4b59-8b62-811d51aae34f · outbound

This paper cites Scannet++: A high-fidelity dataset of 3d indoor scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Scannet++: A high-fidelity dataset of 3d indoor scenes

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:51.161037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:47.345334Z digest=sha256:d3c97149e0a94bc6324574a4e630e2f86397f0597036d158321255f6ced07423

Observation 928e6764-3973-4ad1-bf7b-d15f99d7e2e9 · outbound

This paper cites Dogs: Distributed-oriented gaus- sian splatting for large-scale 3d reconstruction via gaussian consensus.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Dogs: Distributed-oriented gaus- sian splatting for large-scale 3d reconstruction via gaussian consensus

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:50.891902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:47.442586Z digest=sha256:1899c37e22f068f7c7d27ee944c0b74ea486722787d428b46383d1dac11e117c

Observation c06fb373-00e8-4647-aca5-bfeb0b7331a2 · outbound

This paper cites PreSight: Enhancing Autonomous Vehicle Perception with City-Scale NeRF Priors.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields PreSight: Enhancing Autonomous Vehicle Perception with City-Scale NeRF Priors

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:49:48.524762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:47.509275Z digest=sha256:8eb6969a0a0484a2eb55c9b4446a7025657cb14f2a9e05cbb0164b5127fda312

Observation 07bb5d59-1330-4260-90f3-f1ce37af3872 · outbound

This paper cites Exploring a fine-grained multiscale method for cross-modal remote sensing image re- trieval.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Exploring a fine-grained multiscale method for cross-modal remote sensing image re- trieval

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:50.727057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:47.551158Z digest=sha256:320045ae0029b6460ef3ecbd5cc085a1503873b61ea0348f317b19ed076575cb

Observation c8b06b90-3f35-40f9-bfbd-76a741394e33 · outbound

This paper cites Garfield++: Reinforced gaussian ra- diance fields for large-scale 3d scene reconstruction, 2024.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Garfield++: Reinforced gaussian ra- diance fields for large-scale 3d scene reconstruction, 2024

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:50.489234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:47.607078Z digest=sha256:27a9fdd596c507c0e8a7f49d08aa16cb5fb2c690399f6dc4901452b03f241711

Observation 1444202d-ec7b-4a38-9b77-da8efe02732a · outbound

This paper cites 3DitScene: Editing any scene via language-guided disentan- gled gaussian splatting.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3DitScene: Editing any scene via language-guided disentan- gled gaussian splatting

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:50.181200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:47.663922Z digest=sha256:781be2648acd065fc4edc8d656fc8efd13f064026bfacae2f6a81b91e6a40a0b

Observation 5d435d0e-5c3e-4e41-bda2-42231a60e7fa · outbound

This paper cites Earthgpt: A universal multi-modal large lan- guage model for multi-sensor image comprehension in re- mote sensing domain.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Earthgpt: A universal multi-modal large lan- guage model for multi-sensor image comprehension in re- mote sensing domain

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.989182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:47.719996Z digest=sha256:d63eeb797f201c476415551efab0ba974c321dd2e94759edaaf0ea0fef83c082

Observation cb0f6bdb-ae78-497e-8085-d4022ea2b322 · outbound

This paper cites EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:47.792628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:47.792628Z digest=sha256:9d83663bfbe230acb0d64620d7101692d4f195fdd11b03bc13565ff5e59c13b6

Observation a6fde006-503c-4e13-a313-7da9fb1a1243 · outbound

This paper cites Efficient Large-scale Scene Representation with a Hybrid of High-resolution Grid and Plane Features.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Efficient Large-scale Scene Representation with a Hybrid of High-resolution Grid and Plane Features

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:49:48.374875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:47.847955Z digest=sha256:2bb64bb9dbdf1c49be08fd6ecd8ed0b9515c113f6262f038dc11994e60c56a18

Observation a7ef6112-0d56-40b4-8e46-add371eadc78 · outbound

This paper cites Aerial lifting: Neural urban semantic and building instance lifting from aerial imagery.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Aerial lifting: Neural urban semantic and building instance lifting from aerial imagery

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.770435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:47.900648Z digest=sha256:79e4104c160b0a3e39ea0e25a9d32d86cbe846bc935f7e9d988e9fcacae746a0

Observation de21f140-582b-4c75-9827-6bbac2003409 · outbound

This paper cites Rs5m and georsclip: A large-scale vision- language dataset and a large vision-language model for remote sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Rs5m and georsclip: A large-scale vision- language dataset and a large vision-language model for remote sensing

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.525950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:47.958787Z digest=sha256:9887cee554c7ccb24b526bea9246a753c0443473fd6fb5b4ba1bc43e43d7a732

Observation 861b8f96-6aef-48a5-9807-82467594e760 · outbound

This paper cites Mutual Attention Inception Network for Remote Sensing Visual Question Answering.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Mutual Attention Inception Network for Remote Sensing Visual Question Answering

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.261423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:48.037601Z digest=sha256:ea00330aab7d5441c3bc349662ef7a083963a4ec9121b6a52b86afd13bea06d8

Observation e9afdf7b-ffbb-44f9-8cac-136746af8ab1 · outbound

This paper cites Drivinggaussian: Composite gaussian splatting for surrounding dynamic au- tonomous driving scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Drivinggaussian: Composite gaussian splatting for surrounding dynamic au- tonomous driving scenes

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.088711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:48.120675Z digest=sha256:0dfb8774aab80ef46beae7c90585a4dcbc588b34ce8b1b1dcf8a6446f7f5cf46

Observation 067fb573-df17-4484-9755-f991eb3f2950 · outbound

This paper cites Towards vision- language geo-foundation models: A survey.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Towards vision- language geo-foundation models: A survey

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:48.162222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:48.162222Z digest=sha256:cf759a7b73bde8313e9a650554708a2367d3f5d6b2bec9cb80f46dfb12a54319

Observation 665b8c4d-6538-4db8-9d99-a80c8c9f3ed5 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:48.941651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:48.227457Z digest=sha256:9b93a8d8ed03e8a977853cb4618498b5b96c08f9e1f97ea0593d802e22088494

Observation 0eddabe3-4fb5-4160-a229-5338d77d636e · outbound

This paper cites 'yes' if {ANSWER1} < {ANSWER2} else 'no'.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 'yes' if {ANSWER1} < {ANSWER2} else 'no'

Reference 2023

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T21:49:48.805064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:49:48.230001Z digest=sha256:a24ed94eff00cb26710e66efc7640346da46f6d52f96a1312ba85a9905566a85

Pith citing papers

No inbound Pith citation observations are available.