REVIEW 4 cited by
GPT4GEO: How a Language Model Sees the World's Geography
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) have shown remarkable capabilities across a broad range of tasks involving question answering and the generation of coherent text and code. Comprehensively understanding the strengths and weaknesses of LLMs is beneficial for safety, downstream applications and improving performance. In this work, we investigate the degree to which GPT-4 has acquired factual geographic knowledge and is capable of using this knowledge for interpretative reasoning, which is especially important for applications that involve geographic data, such as geospatial analysis, supply chain management, and disaster response. To this end, we design and conduct a series of diverse experiments, starting from factual tasks such as location, distance and elevation estimation to more complex questions such as generating country outlines and travel networks, route finding under constraints and supply chain analysis. We provide a broad characterisation of what GPT-4 (without plugins or Internet access) knows about the world, highlighting both potentially surprising capabilities but also limitations.
Forward citations
Cited by 4 Pith papers
-
MapStory: Prototyping Editable Map Animations with LLM Agents
Natural language scripts can be turned into editable, geospatially grounded map animations through MapStory's dual-agent LLM architecture.
-
Towards Interpretable Geo-localization: a Concept-Aware Global Image-GPS Alignment Framework
A concept bottleneck projecting images and GPS into a subspace of geographic concepts improves GeoCLIP from 10.8 to 13.2 percent top-1 km accuracy on Im2GPS3k and adds semantic explanations.
-
Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning
A new benchmark called GEOHALUBENCH measures how often LLMs invent, omit, or confuse real-world places and relations, and a dynamic-beta KTO method reduces these errors on the benchmark.
-
The World As Large Language Models See It: Exploring the reliability of LLMs in representing geographical features
GPT-4o and Gemini 2.0 Flash approximate the geography of Austria but show systematic biases in coordinates and elevations and frequent errors in assigning federal states.
Discussion (0). Sign in to comment.