Pith. sign in

Paper Citation Record · LEDGER

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields

As of 8 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 0 inbound Pith citation observations for arXiv:2506.23352.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23352 v1

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:49:48.230001Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

96 of 96 outbound references displayed

  • verified exact2
  • verified fuzzy64
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9a2a456f-48b1-461b-9bab-b5febf5309c0 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Scanqa: 3d question answering for spatial scene understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.054070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.054070Z digest=sha256:b4cf88bb5cbe15cf411121114ff9ba0b8e72a55f6a1643ae2891a0d7880e443f

Observation 7bba69c6-6d78-4fd5-89d6-9cdd55d2ada1 · outbound

This paper cites Qwen2.5-VL Technical Report.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.094181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.094181Z digest=sha256:c1c3487c5eb23b6e1131c968a77aad3af1ef44d38034d15e41e1ff689ea724d1

Observation 3922716b-968a-46df-9b06-41652b632605 · outbound

This paper cites Henriques, Andrew Zisserman, and Andrea Vedaldi.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Henriques, Andrew Zisserman, and Andrea Vedaldi

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.194838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.194838Z digest=sha256:3e048391ce2f15d00b9ac6287f031e7fa75bad29db4af2c7cb61c52c451ef001

Observation 05a9f140-72af-403e-8111-ef8f7d6c539b · outbound

This paper cites Blumer, Qingx- uan Chen, and Francis Engelmann.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Blumer, Qingx- uan Chen, and Francis Engelmann

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.256991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.256991Z digest=sha256:1ef4c8f8b3fc2c788c47790e75b090fae3c41ce33b2581e4dfe828d9b126c37a

Observation b093582a-470c-4d28-b589-ce745a4c8e3c · outbound

This paper cites A persistent spatial semantic representation for high-level natural language instruction execution.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields A persistent spatial semantic representation for high-level natural language instruction execution

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.392137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.392137Z digest=sha256:3924a82a349d2ea8b6d62db0eae9325cf438ce6dabb1489fc598444577e67443

Observation 71ca8a59-5ff8-47a7-9292-8244e40f89d9 · outbound

This paper cites Prompt-rsvqa: Prompt- ing visual context to a language model for remote sensing visual question answering.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Prompt-rsvqa: Prompt- ing visual context to a language model for remote sensing visual question answering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.499765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.499765Z digest=sha256:69c49f34eddaa698d7f3129dfc3cdf9312724f716ebd4a57713bff51da6ba9b3

Observation dee5502b-394d-4b8a-9d94-53bcfd086c67 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natu- ral language.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Scanrefer: 3d object localization in rgb-d scans using natu- ral language

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.634508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.634508Z digest=sha256:22442dc252531d7eb81c309587f829b68194ca0b9d0f8c4c8ad8071caeabcbd3

Observation 2d18a5ce-27c7-4499-9d5a-6992720f8660 · outbound

This paper cites Panoptic vision-language feature fields.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Panoptic vision-language feature fields

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.745471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.745471Z digest=sha256:6693c6c9d51d8313a74cb1278582609d6dd40cfde43105a38d3e8b3eba070803

Observation e45035e5-1c0c-469c-a16e-8289272ccb08 · outbound

This paper cites Stylecity: Large-scale 3d urban scenes stylization.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Stylecity: Large-scale 3d urban scenes stylization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.839057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.839057Z digest=sha256:d6a55185ee73172e0f21984ffdb3d8abc800dd2640de225f2d8bbe537a261716

Observation d71297e6-2235-430c-96e6-bf273376ff1e · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.929873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.929873Z digest=sha256:f3ebc5077491b63db631ec0f9bb661435a2ac267023caac55656ec2a84f4e798

Observation 00fe7b97-3e74-4516-ab3c-9da16c510e9c · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.012679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.012679Z digest=sha256:c02bd50299a234367acce9a5e88af39f2c5110844f8e97eda994ff1239c3847f

Observation da1aa099-20a3-46e0-a396-41eb294e5e9e · outbound

This paper cites LayoutGPT: Compositional Visual Planning and Generation with Large Language Models.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields LayoutGPT: Compositional Visual Planning and Generation with Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.132312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.132312Z digest=sha256:618def62bd69893ed925a857edf3d2f5c43b49f6984828bf73848db504150165

Observation 1999571b-cd9a-407c-bc5f-ccfc1b13af7e · outbound

This paper cites Dynamic 3d gaussian fields for urban areas.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Dynamic 3d gaussian fields for urban areas

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.190261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.190261Z digest=sha256:13c471aab96a9a8146f84f91b105adcce73b83715730470a9e80e4d4e40e2f02

Observation 70eb2b3b-63eb-4257-a5a6-a8e64cbdd79a · outbound

This paper cites Ue4-nerf:neural radiance field for real-time rendering of large-scale scene.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Ue4-nerf:neural radiance field for real-time rendering of large-scale scene

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.253760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.253760Z digest=sha256:0aa11825f42a4bdd27e2b02da6a11d4f19984103b32e035d1c560950e4b69812

Observation a5fa8581-1343-47b9-8e7c-a86977f9e80d · outbound

This paper cites StreetSurf: Extending Multi-view Implicit Surface Reconstruction to Street Views.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields StreetSurf: Extending Multi-view Implicit Surface Reconstruction to Street Views

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.307473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.307473Z digest=sha256:a9aa4dbcb860d2a6ecadc6534e8c1683efc6e7351b130026ae92f4280eea9043

Observation 36a3ba48-3582-43fb-8f84-603d4d76c0a2 · outbound

This paper cites Visual program- ming: Compositional visual reasoning without training.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Visual program- ming: Compositional visual reasoning without training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.378479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.378479Z digest=sha256:3d01f9a97695285a6cd74d02a99ed5fb8064bd835d08b44c5eb283779b9f2012

Observation 15cf77ae-37cc-48a1-9156-f7661f6bd688 · outbound

This paper cites Pigeon: Predicting image geolocations.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Pigeon: Predicting image geolocations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.437098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.437098Z digest=sha256:6d03486aef576db9809371030ff50016be8e550fca517d20e018707f6af9bd59

Observation a4e7e80a-1a68-44a6-a34d-90d8b1bbb3b8 · outbound

This paper cites Dragon: Drone and ground gaussian splatting for 3d building reconstruction.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Dragon: Drone and ground gaussian splatting for 3d building reconstruction

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.490242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.490242Z digest=sha256:a301015849fc22ce63f5c7e4350b4e3e8a71282a82c2bac8f14584957a893a2e

Observation eb5e3a97-aac3-425d-a4cc-cc6dec241d00 · outbound

This paper cites 3d-llm: Inject- ing the 3d world into large language models.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3d-llm: Inject- ing the 3d world into large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.533958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.533958Z digest=sha256:6cb09e71037cd3040f4b57e47b48855f7f93682bf0d6b711fab5559606c29f0e

Observation 3edb0cda-0e2e-45c5-b5c1-bc48b5709486 · outbound

This paper cites RSGPT: A Remote Sensing Vision Language Model and Benchmark.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.574175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.574175Z digest=sha256:c7501f2daade0013d78612ef27556c893aab16fbceaa92fa43196735d7f57a07

Observation 57381278-8928-4fa2-bd2c-0c0df261d55d · outbound

This paper cites TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.614807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.614807Z digest=sha256:612d4c0705173d76ee3405b2ae1907e62e4a21f172865882b46fc13e23750f01

Observation cacaa869-fdd5-498b-aa6c-f7c052d768b2 · outbound

This paper cites GraspSplats: Efficient Manipulation with 3D Feature Splatting.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields GraspSplats: Efficient Manipulation with 3D Feature Splatting

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.660608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.660608Z digest=sha256:e3e447f59446110d01ce5c6223595739655eaf5b717e6db3491a0021c53bfb2e

Observation da349638-f488-43c6-9293-b7e9a72dbcaa · outbound

This paper cites Fastlgs: Speeding up lan- guage embedded gaussians with feature grid mapping.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Fastlgs: Speeding up lan- guage embedded gaussians with feature grid mapping

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:50:00.199756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:43.729855Z digest=sha256:c80f04a72c0cbf0d0cd0bdf7faef881ef78d606776e8c74924cce8bbf0d54213

Observation 4cc0bab2-8de2-4525-a191-6972b1f1a3ef · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3d gaussian splatting for real-time radiance field rendering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:50:00.069090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:43.778265Z digest=sha256:fe23c37c93252dc4c8747ab8fd8f477a499117d42bab2907720cf0c34bf76a94

Observation 5079a02f-46e7-4f71-96b1-0ae8264c004f · outbound

This paper cites A hierarchical 3d gaussian representation for real-time ren- dering of very large datasets.ACM Transactions on Graphics (TOG), 43(4), 2024.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields A hierarchical 3d gaussian representation for real-time ren- dering of very large datasets.ACM Transactions on Graphics (TOG), 43(4), 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.894873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:43.817893Z digest=sha256:d880723f0ca3665aef2a891d96d62d76d51f1e5dc74076c35d3f13b3a6c2128d

Observation f03dc12f-8688-4d4f-b78a-3b4d89abee4c · outbound

This paper cites LERF: Language embed- ded radiance fields.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields LERF: Language embed- ded radiance fields

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.718873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:43.885600Z digest=sha256:102814babc6c7c91a1e1f98ecb037cbbefc1f3b0c58399aac0e6fe13c9fe5f3f

Observation 4d690336-278a-482b-b22a-6b0373ac9ea7 · outbound

This paper cites Lobell, and Ste- fano Ermon.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Lobell, and Ste- fano Ermon

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.590614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:43.952849Z digest=sha256:cfe19261b1bcbd7578fa19eb54b349d444d1736e0bcc9f1cad65531396628652

Observation 77bafc28-a913-4c26-8aad-80d6f1db3910 · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.476659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:44.000320Z digest=sha256:eea5c3af7962acf5edd5e230c49ce6a0b4f45705a2636f8d558abbe4d847b5e0

Observation 512e23fd-0770-4804-8f56-d514e6f22f0c · outbound

This paper cites Segment any- thing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Segment any- thing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.303634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:44.042944Z digest=sha256:b7090c48e26829b13e695f78990124aed56602e0e5056444ffd9e3cd399e43c3

Observation 383964ea-5fb3-40a4-b019-7b71dbe8a3ed · outbound

This paper cites SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:44.101737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:44.101737Z digest=sha256:e4dcd67c25a869ffe186c9b53639be734f8eb8771c635480afd6a3dec8ee2875

Observation db4b74fd-35ac-4f0d-96c4-308024254d38 · outbound

This paper cites Decomposing nerf for editing via feature field dis- tillation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Decomposing nerf for editing via feature field dis- tillation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.166663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:44.136124Z digest=sha256:0176d199780fcd4928bcb20769dda17159c58cebc76b318ba6021f7073f36447

Observation 4594861b-9cfa-4255-9ca0-32b88981eb59 · outbound

This paper cites Text2pos: Text-to-point-cloud cross-modal localiza- tion.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Text2pos: Text-to-point-cloud cross-modal localiza- tion

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.064991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:44.170793Z digest=sha256:4aa467706eae7ffad4cebeae24604c944aa5c969ac2b9148fc33a231c48011df

Observation 8eb4957f-fbb9-41b1-8285-94555bb82c3d · outbound

This paper cites Geochat: Grounded large vision-language model for remote sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Geochat: Grounded large vision-language model for remote sensing

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.903803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:44.222774Z digest=sha256:632a928ae0af115fda7c2530b0490719a77d993e6478715cd1cf9c96f1055adf

Observation 7b1fd120-5535-440d-8b72-83d94a97fa55 · outbound

This paper cites NeRF-XL: Scaling nerfs with multiple GPUs.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields NeRF-XL: Scaling nerfs with multiple GPUs

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.773105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:44.267005Z digest=sha256:e7017831c44aca887f6ebc1a0115c9fc22d1a530672d49fc0d114ca24af69343

Observation 51b0f3b3-7f83-4852-bc1d-694b87713b73 · outbound

This paper cites Matrixcity: A large-scale city dataset for city-scale neural rendering and beyond.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Matrixcity: A large-scale city dataset for city-scale neural rendering and beyond

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.613797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:44.309914Z digest=sha256:79c2b8ac3cc4c958d1d83a2f897c2d510b4691ed9b117f5a8fb1350f6a2a666d

Observation a5a46439-664f-40d8-a915-3f451efd8976 · outbound

This paper cites Vastgaussian: Vast 3d gaus- sians for large scene reconstruction.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Vastgaussian: Vast 3d gaus- sians for large scene reconstruction

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.486658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:44.345633Z digest=sha256:53c4b79156e6a747e137b19b78e429de5a68f7f238a3bbb074bc10800e021d24

Observation 11471b57-9a6a-442e-adb6-c54a6d7a3237 · outbound

This paper cites Capturing, reconstructing, and simulating: the urbanscene3d dataset.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Capturing, reconstructing, and simulating: the urbanscene3d dataset

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.345746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:44.377705Z digest=sha256:734b0c56b68660751e2bc93ad084b9830887d886f5df93dac485787f60e87e83

Observation 2ffa59ba-1675-421a-89c6-c32d5bd75650 · outbound

This paper cites Re- moteclip: A vision language foundation model for remote sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Re- moteclip: A vision language foundation model for remote sensing

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.217038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:44.441781Z digest=sha256:5acec9ba0f1e41794ac0dd225c8916ee56f489d16521666b2ba8d726a8e63fdd

Observation 321362df-eaa4-42f8-ad62-66f8dfe04287 · outbound

This paper cites Visual instruction tuning.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Visual instruction tuning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.049953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:44.492129Z digest=sha256:faeb74f711228af210bb9e3e51e6cd60f6c236cced58f4d9e31f3a0cf773674c

Observation 7da5c1eb-1eb3-48c3-abb9-73bcdc066f97 · outbound

This paper cites Citygaus- sian: Real-time high-quality large-scale scene rendering with gaussians.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Citygaus- sian: Real-time high-quality large-scale scene rendering with gaussians

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.943733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:44.536960Z digest=sha256:56979b0addda0fce170146f77f1a4ee75a5cbf183572c0c10c0642406ac49ba7

Observation c5b516e5-64a8-4bbc-9865-6fbfc9fb13b8 · outbound

This paper cites Citygaussianv2: Efficient and geometri- cally accurate reconstruction for large-scale scenes, 2024.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Citygaussianv2: Efficient and geometri- cally accurate reconstruction for large-scale scenes, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.732635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:44.579863Z digest=sha256:8005bb8edb2312e513215d13993860040c05ead34aff5dc74039efb0184c27fa

Observation 64464ef6-4f58-491e-a64e-c1aea177f006 · outbound

This paper cites Citygaussianv2: Efficient and geometri- cally accurate reconstruction for large-scale scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Citygaussianv2: Efficient and geometri- cally accurate reconstruction for large-scale scenes

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.486653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:44.633032Z digest=sha256:a54e711778d2db9787505b1e5ff3da8554b155b1b423577087b13c6fae69d73d

Observation a39d028d-c282-40a8-841b-91b5017f9523 · outbound

This paper cites Chameleon: Plug-and-play compositional reasoning with large language models.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Chameleon: Plug-and-play compositional reasoning with large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.326550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:44.673401Z digest=sha256:1a294ce622740fa4aecb598a95f8eb4b41fb7a6bed86683898f03d218c200243

Observation 411a2e79-1115-4057-b457-d2753476d1ce · outbound

This paper cites Exploring models and data for remote sensing im- age caption generation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Exploring models and data for remote sensing im- age caption generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.153451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:44.722564Z digest=sha256:92abfe5d758b94575d9cc8621269194a00d8958013c52dc941b39707d991b575

Observation 42f8f1c8-c23b-4d8e-bec0-fae4bce9c14e · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:44.779870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:44.779870Z digest=sha256:d512d26b44a887c49a9799d36cdc0726bdc9d17fb0b88660e3d4a98b04d2c4ae

Observation 1c23294a-ff1c-43ac-9251-0d118650d631 · outbound

This paper cites A multiscale grouping transformer with clip latents for re- mote sensing image captioning.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields A multiscale grouping transformer with clip latents for re- mote sensing image captioning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.999396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:44.839366Z digest=sha256:5f1c24da33fe7a964b8cba628b4e0feb43bf5537262b53c02b224b2f817823f1

Observation 8692614f-bbd4-425f-b298-23576a5f2c93 · outbound

This paper cites Llama 3.2 connect 2024: Vision on the edge and mo- bile devices.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Llama 3.2 connect 2024: Vision on the edge and mo- bile devices

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.865133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:44.886549Z digest=sha256:9bab25a24e90628839e59e44a15454f16b5f2c100b367518ea81148dce272c9d

Observation 968bc7a2-0782-4099-b787-0fc6d70410f8 · outbound

This paper cites Srinivasan, Matthew Tancik, Jonathan T.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Srinivasan, Matthew Tancik, Jonathan T

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.632184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:44.962424Z digest=sha256:6d9dd7cc005f4eea8a89bd8ec1b9f1015ea93710bd90aecfde99cb260b994c41

Observation dadfd791-b5fc-4459-a5cf-e26c4aa29d90 · outbound

This paper cites Cityrefer: Geography-aware 3d visual grounding dataset on city-scale point cloud data.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Cityrefer: Geography-aware 3d visual grounding dataset on city-scale point cloud data

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.486953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:45.026632Z digest=sha256:ae8b3b0eb6e55d8bee9a058597b2a4c0f9287c5c05dbbd5fe78c091c3a8e99c6

Observation c84971c3-8ab1-4f8c-98e6-f143ba7b7feb · outbound

This paper cites Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.346668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:45.078369Z digest=sha256:19487001e4846628d33f9e6707c9d14b6a33ccb76b1b5620624c8db4a10a58fc

Observation a6f53a50-2fbb-4fa8-9bb5-82087eb5ffff · outbound

This paper cites Hello gpt-4o.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Hello gpt-4o

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.242315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:45.171325Z digest=sha256:8e5a9a5361aaeaabe6c246660192dc6e23ac7e15e2806e8a9d4a8ec290599cc1

Observation a7d2eac8-8edb-49c0-9d47-9a7287cd4e45 · outbound

This paper cites Vhm: Versatile and honest vision lan- guage model for remote sensing image analysis.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Vhm: Versatile and honest vision lan- guage model for remote sensing image analysis

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.086960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:45.242656Z digest=sha256:029b16ce1fa649f3e181b0c76b05397ea116921995cacb12ad0f855e2301e27d

Observation ede4e48d-4a1f-44e3-8d07-8c976f4a375d · outbound

This paper cites Langsplat: 3d language gaussian splatting.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Langsplat: 3d language gaussian splatting

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.894524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:45.317392Z digest=sha256:1d55207e0c90344394d464a7d73c7309138fc2bf5ff7ce84ccd150e0a4251cdd

Observation 597031c2-4287-4768-9ec9-9c81916f7abe · outbound

This paper cites Deep semantic understanding of high resolution remote sensing image.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Deep semantic understanding of high resolution remote sensing image

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.736156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:45.454471Z digest=sha256:4731b4b5de711190c16eb170349d115d2239d1b347c907e4283a047b55fc5557

Observation c47a1362-4105-4afd-b80a-93a9f0db17c5 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Learn- ing transferable visual models from natural language super- vision

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.596461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:45.537922Z digest=sha256:8c4e2eaaa909ec1a6c1b50f31c613c39f0a8552aa62b66b40289b506c88fe02d

Observation 12ebebed-4dce-4e81-9e20-ce61af8ba3db · outbound

This paper cites Derf: Decom- posed radiance fields.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Derf: Decom- posed radiance fields

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.445680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:45.650119Z digest=sha256:31da3f2a9716e4c91ce1e5a59a5fe8e847a72f087efbb4c5de6c58b6a10695fc

Observation 72561cb0-d64b-4d2e-9e0d-559e86cabb15 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Toolformer: Language models can teach themselves to use tools

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.249562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:45.703626Z digest=sha256:edd98b5bda1803677c573e3a14a9d0ad925ef31fb6ef074b484c51d158eac4e8

Observation c64e465f-305b-429a-8906-0612a0ab092e · outbound

This paper cites CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:45.758919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:45.758919Z digest=sha256:aa4d9912f591a1fa4afe71ffb91fa2f6131f561aa935f266392dac1d76a0f7f5

Observation 1cca6068-bdbe-4e95-b372-49d2d1cc4e29 · outbound

This paper cites Language embedded 3d gaussians for open- vocabulary scene understanding.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Language embedded 3d gaussians for open- vocabulary scene understanding

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.086891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:45.808762Z digest=sha256:68470e1df41713bf6efceeb3c999f8d0721850f9becd645b30bc0d6dc79dc07b

Observation 753f4515-7c61-4625-8f72-df9435fd83cc · outbound

This paper cites Real-time view synthesis for large scenes with millions of square meters.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Real-time view synthesis for large scenes with millions of square meters

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.947992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:45.867334Z digest=sha256:368edf36fb407802fe6a184d68cf899c9f0905318405c5b8f08b36f92f7505eb

Observation 86280683-e13f-4125-a85d-f36cb77c5652 · outbound

This paper cites City-on-web: Real-time neural rendering of large- scale scenes on the web.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields City-on-web: Real-time neural rendering of large- scale scenes on the web

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.815656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:45.905136Z digest=sha256:ac6bc81aa000ad7d6ec68c199bab8cadf02387833bdabff214bb660ec481ee08

Observation b9c0206e-f8e7-4604-82a5-16898e815aeb · outbound

This paper cites De- composing 3d scenes into objects via unsupervised volume segmentation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields De- composing 3d scenes into objects via unsupervised volume segmentation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.633027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:46.023299Z digest=sha256:08c85168335e0d3bcd95624af8c078903756da95f608f87aa5dc9a9c86bc7b1a

Observation 700651a4-cb5f-4101-b462-ca96833df9cb · outbound

This paper cites Modular visual question answering via code generation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Modular visual question answering via code generation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.454110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:46.069829Z digest=sha256:49ebd7bf432f91558dee8ba3cfa2360376399660c07944b2edcf2953e1c3b126

Observation dc431695-f1e0-4f6f-a2d7-2bbbf9863b43 · outbound

This paper cites 3d ques- tion answering for city scene understanding.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3d ques- tion answering for city scene understanding

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.292203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:46.122863Z digest=sha256:d4e5ba950c080586e4ad3f1c8fc90a36df6d8586b917536cdb94128226b38979

Observation a116152e-ee2d-47bc-a708-0b273c10d0b5 · outbound

This paper cites Visual grounding in remote sensing images.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Visual grounding in remote sensing images

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.192127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:46.179689Z digest=sha256:6d3a2f905647b7e4ef50126a1eedab71fedd3c28063d3203cab89f26446d847b

Observation 89523449-652f-4b7d-93aa-f2de235e1a3b · outbound

This paper cites ViperGPT: Visual inference via python execution for reasoning.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields ViperGPT: Visual inference via python execution for reasoning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.059622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:46.235926Z digest=sha256:d2fbff66fb72cb094b032f8f0d3c45058b2f4765b9ea4b724dcf6ea7c76ab768

Observation d3fd2ac5-28eb-43fd-947f-dc05078d97f3 · outbound

This paper cites Srinivasan, Jonathan T.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Srinivasan, Jonathan T

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.881502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:46.293702Z digest=sha256:5e90cedd8a57adfb4aa265b627a3f0d8be0a18eb533a324e12c5f5d2d15aa7f6

Observation 57e570d9-5479-4183-9948-0f315a385603 · outbound

This paper cites Crs-diff: Controllable generative re- mote sensing foundation model.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Crs-diff: Controllable generative re- mote sensing foundation model

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.729228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:46.366132Z digest=sha256:a59d8b61c832843e4e1eb1f44fef7962cc1b60f8c4f2cbf47023c8eb54bbe82d

Observation e963ca73-2e31-42fd-b422-939edfc5e32f · outbound

This paper cites MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:46.416409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:46.416409Z digest=sha256:8fdef3e5c0d5ecf970ccc095394649082476cf43bde00ff45148f1fa4944efcc

Observation 43ac9ec0-b7ff-427a-a3de-96a9fb84d01e · outbound

This paper cites Mega-nerf: Scalable construction of large-scale nerfs for virtual fly-throughs.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Mega-nerf: Scalable construction of large-scale nerfs for virtual fly-throughs

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.611663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:46.462758Z digest=sha256:95813a5363f29d334d3583c6a4ed1b79def7ca0e50f0dd1a5e05354ee2beead9

Observation 66328b1d-9627-491e-b845-e8701b4ae3d7 · outbound

This paper cites Geoclip: Clip-inspired alignment between locations and im- ages for effective worldwide geo-localization.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Geoclip: Clip-inspired alignment between locations and im- ages for effective worldwide geo-localization

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.524417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:46.542012Z digest=sha256:19488feec4be74aed3ed50992ebbc9e79cb38d16daedd4316049835d47d7831e

Observation 08dcae4e-331b-4fde-a88e-c931cd680623 · outbound

This paper cites Skyscript: A large and semanti- cally diverse vision-language dataset for remote sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Skyscript: A large and semanti- cally diverse vision-language dataset for remote sensing

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.281261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:46.608758Z digest=sha256:245a19921166f9282f2fb1ab5769cc91f8b34bbee9968472f9bc800587c21975

Observation 96d58b29-89ee-43af-b61c-3fb6d54e869a · outbound

This paper cites Text2loc: 3d point cloud localization from natural language.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Text2loc: 3d point cloud localization from natural language

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:52.994039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:46.661207Z digest=sha256:929866db62acd126b91e2c2c9f1ee323390545fe2bcb5f420dc290a655db089b

Observation e6fdab54-f947-4033-8866-5f2f7e532a5c · outbound

This paper cites Bungeenerf: Progressive neural radiance field for extreme multi-scale scene rendering.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Bungeenerf: Progressive neural radiance field for extreme multi-scale scene rendering

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:52.782372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:46.717819Z digest=sha256:09ca4a45e4051cd45fdab66f59a530a8a4b7b8bb58bf0413fc74b370d3988b1a

Observation 74ea19dd-d55c-4610-96da-75d4db71c88a · outbound

This paper cites Citydreamer: Compositional generative model of unbounded 3D cities.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Citydreamer: Compositional generative model of unbounded 3D cities

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:52.490533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:46.764874Z digest=sha256:18c8301a8a7e5ea816eaf5ab977c2428deffa3dceb9e5d5b1d5019b9e14c79a6

Observation 7cb1dfe4-9caa-4660-8e0f-322ee862b86d · outbound

This paper cites Generative Gaussian Splatting for Unbounded 3D City Generation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Generative Gaussian Splatting for Unbounded 3D City Generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:46.818968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:46.818968Z digest=sha256:d07408443f3ef5c003dd21c2e223db054fdbdb78124e9669bb142fda53c76823

Observation 8ad4195c-3cd5-4c51-9efb-b631cd4f3f37 · outbound

This paper cites Grid-guided neural radiance fields for large urban scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Grid-guided neural radiance fields for large urban scenes

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:52.184167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:46.895366Z digest=sha256:abcfc51ccf91856afac523621c9572ac890dab84b2dce2bb6e2c61c99b5cfebe

Observation b760559c-52cb-4306-a908-3e67e5e05eeb · outbound

This paper cites Pointllm: Empowering large language models to understand point clouds.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Pointllm: Empowering large language models to understand point clouds

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:51.883576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:46.938093Z digest=sha256:52373553386d0a3baef519d1f171984bca944f332544e232dcab31f7d83c7403

Observation e3dd9d04-2419-4c03-a6a6-a96cfc73e4c1 · outbound

This paper cites Addressclip: Empowering vision-language models for city-wide image address localization.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Addressclip: Empowering vision-language models for city-wide image address localization

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:51.674509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:47.105618Z digest=sha256:d3d4b0be775223e8663e5e997acda7e8b297ca261560a3bd35de44de62f667ba

Observation c5b83c66-8f0f-4dff-983b-606d80e0b1ee · outbound

This paper cites Unisim: A neural closed-loop sensor simulator.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Unisim: A neural closed-loop sensor simulator

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:51.420687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:47.217718Z digest=sha256:dd24ab7185bbe43bbc75b68af6dc24b437f8c410b35fee85426efc127d0f704c

Observation 8a8f6348-3a1f-4b59-8b62-811d51aae34f · outbound

This paper cites Scannet++: A high-fidelity dataset of 3d indoor scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Scannet++: A high-fidelity dataset of 3d indoor scenes

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:51.161037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:47.345334Z digest=sha256:e458f05215b0e1351bf6da7c5fec1ff9ffda807cfc62c541fac0b087591424e4

Observation 928e6764-3973-4ad1-bf7b-d15f99d7e2e9 · outbound

This paper cites Dogs: Distributed-oriented gaus- sian splatting for large-scale 3d reconstruction via gaussian consensus.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Dogs: Distributed-oriented gaus- sian splatting for large-scale 3d reconstruction via gaussian consensus

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:50.891902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:47.442586Z digest=sha256:549e04d0fd84cd1203e9de4651ae371f07ea11c076e788b1509896e46b37f975

Observation c06fb373-00e8-4647-aca5-bfeb0b7331a2 · outbound

This paper cites PreSight: Enhancing Autonomous Vehicle Perception with City-Scale NeRF Priors.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields PreSight: Enhancing Autonomous Vehicle Perception with City-Scale NeRF Priors

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:49:48.524762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:47.509275Z digest=sha256:bbc7694451bed51f521d4f2ed7402fdd2ed3fcc7beb3e52cd977fde72df2146d

Observation 07bb5d59-1330-4260-90f3-f1ce37af3872 · outbound

This paper cites Exploring a fine-grained multiscale method for cross-modal remote sensing image re- trieval.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Exploring a fine-grained multiscale method for cross-modal remote sensing image re- trieval

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:50.727057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:47.551158Z digest=sha256:92374cd6ba9e636a40e9b8a69f672d21a9f1da4b290ada10baac308d8b5be00d

Observation c8b06b90-3f35-40f9-bfbd-76a741394e33 · outbound

This paper cites Garfield++: Reinforced gaussian ra- diance fields for large-scale 3d scene reconstruction, 2024.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Garfield++: Reinforced gaussian ra- diance fields for large-scale 3d scene reconstruction, 2024

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:50.489234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:47.607078Z digest=sha256:3af8cb03e49995276edd06e415a631def63abcc08dd961382d986f1de40d9a13

Observation 1444202d-ec7b-4a38-9b77-da8efe02732a · outbound

This paper cites 3DitScene: Editing any scene via language-guided disentan- gled gaussian splatting.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3DitScene: Editing any scene via language-guided disentan- gled gaussian splatting

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:50.181200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:47.663922Z digest=sha256:922a850f7d834ca7fed3b5c1e14f040c51385b7afef6dc75ecea82b8342ed6a8

Observation 5d435d0e-5c3e-4e41-bda2-42231a60e7fa · outbound

This paper cites Earthgpt: A universal multi-modal large lan- guage model for multi-sensor image comprehension in re- mote sensing domain.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Earthgpt: A universal multi-modal large lan- guage model for multi-sensor image comprehension in re- mote sensing domain

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.989182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:47.719996Z digest=sha256:7f545ea05e675c83040dec15a43d8ee69be2b6278cbaa1b456e1cdd3fe8217b7

Observation cb0f6bdb-ae78-497e-8085-d4022ea2b322 · outbound

This paper cites EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:47.792628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:47.792628Z digest=sha256:b4fcc5cd3d8d816ba0c36d9622fd54364e231fef7a274dfc45796142a0caf694

Observation a6fde006-503c-4e13-a313-7da9fb1a1243 · outbound

This paper cites Efficient Large-scale Scene Representation with a Hybrid of High-resolution Grid and Plane Features.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Efficient Large-scale Scene Representation with a Hybrid of High-resolution Grid and Plane Features

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:49:48.374875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:47.847955Z digest=sha256:e314de92a359c4c5a23c66f3cb963d65da4ac9d377240dc19145d96b07fee487

Observation a7ef6112-0d56-40b4-8e46-add371eadc78 · outbound

This paper cites Aerial lifting: Neural urban semantic and building instance lifting from aerial imagery.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Aerial lifting: Neural urban semantic and building instance lifting from aerial imagery

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.770435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:47.900648Z digest=sha256:3e0029f590b5520d79b20a00388a1959cf6b857cdb493cd3494f5e4e68360db8

Observation de21f140-582b-4c75-9827-6bbac2003409 · outbound

This paper cites Rs5m and georsclip: A large-scale vision- language dataset and a large vision-language model for remote sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Rs5m and georsclip: A large-scale vision- language dataset and a large vision-language model for remote sensing

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.525950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:47.958787Z digest=sha256:4e1954674510c391aeb3009e51c71025b4d14abd9e219cfa62285df4bf1be7bd

Observation 861b8f96-6aef-48a5-9807-82467594e760 · outbound

This paper cites Mutual Attention Inception Network for Remote Sensing Visual Question Answering.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Mutual Attention Inception Network for Remote Sensing Visual Question Answering

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.261423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:48.037601Z digest=sha256:7285b995d315e08fc0d66f4fe30bbd85e002f85ef31fb4ac827e1979e802790b

Observation e9afdf7b-ffbb-44f9-8cac-136746af8ab1 · outbound

This paper cites Drivinggaussian: Composite gaussian splatting for surrounding dynamic au- tonomous driving scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Drivinggaussian: Composite gaussian splatting for surrounding dynamic au- tonomous driving scenes

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.088711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:48.120675Z digest=sha256:f70c9ea37d0e80c85cf8e99438945b8f42bea2a24e1459c16cac49bcc80b2575

Observation 067fb573-df17-4484-9755-f991eb3f2950 · outbound

This paper cites Towards vision- language geo-foundation models: A survey.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Towards vision- language geo-foundation models: A survey

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:48.162222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:48.162222Z digest=sha256:ec18183cf3d0cd28a8bb8473331818bf2cd10a8bc47fafc37fb33049434c2cb4

Observation 665b8c4d-6538-4db8-9d99-a80c8c9f3ed5 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:48.941651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:48.227457Z digest=sha256:b2dc083129cd0191ff7a994ad78b704ea64a1f31e86253d51d982191f242c905

Observation 0eddabe3-4fb5-4160-a229-5338d77d636e · outbound

This paper cites 'yes' if {ANSWER1} < {ANSWER2} else 'no'.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 'yes' if {ANSWER1} < {ANSWER2} else 'no'

Reference 2023

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T21:49:48.805064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:49:48.230001Z digest=sha256:a480c0a0dddbf46be015ea4d6f5aa8b21c936d29c68e640ebb444956fc918755

Pith citing papers

No inbound Pith citation observations are available.