Pith. sign in

Paper Citation Record · LEDGER

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images

As of 18 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2508.21565.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21565 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:43:52.244483Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy40
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa73e3e2-8e30-475d-92a4-f9dd7715655b · outbound

This paper cites Google street view: Capturing the world at street level.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Google street view: Capturing the world at street level

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.663051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:51.819293Z digest=sha256:35e59a36d0b41b228724a65797ee34c6b7a18bad1a678e3c30b8383ad69bacf4

Observation 3aab4094-9239-493d-864c-2edac5b47858 · outbound

This paper cites Evaluation methods for landscapes with greenery.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Evaluation methods for landscapes with greenery

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.639234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:51.827290Z digest=sha256:9a4172736275a3ee116242d3ee4c66c8f4114119287d51058b4e15570e250178

Observation a1fb3897-2151-4dca-947c-80252575d21d · outbound

This paper cites Browning, Jiaying Dong, Kuiran Zhang, Shuai Yuan, H¨useyin Ertan ˙Inan, Olivia McAnirlin, Dani T.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Browning, Jiaying Dong, Kuiran Zhang, Shuai Yuan, H¨useyin Ertan ˙Inan, Olivia McAnirlin, Dani T

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.612154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:51.847823Z digest=sha256:ac87855811cfd327217d4598bdc19f32d8757250d91fa91cc767664ef3c4352e

Observation e8d66e39-f41c-437c-a787-4cba57e8efd7 · outbound

This paper cites The green window view index: automated multi-source visi- bility analysis for a multi-scale assessment of green window views.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images The green window view index: automated multi-source visi- bility analysis for a multi-scale assessment of green window views

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.589324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:51.859841Z digest=sha256:2bdd65be20de30a46828fb0f18d642dc569dde1238063bb7d333f29b85575362

Observation 74d718cc-de8b-481c-9e96-cbea460508ac · outbound

This paper cites Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:51.870945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:51.870945Z digest=sha256:1312213df1d259d3a8981a2433ea13903e24b761310b4f70119ccd646c5c28de

Observation afb504ce-f976-4375-bcc5-24f1e60f4a60 · outbound

This paper cites End-to- end object detection with transformers.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images End-to- end object detection with transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.546217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:51.879017Z digest=sha256:64898a99d97632d2348a93c445e89d8b2b0c147102bf9e6c46e6c02de055adc4

Observation c147cb24-a1c3-4ea8-bd4f-47b197d61d61 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.521387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:51.886291Z digest=sha256:abad927c954e2d110bdbb14954853a5119bb8e82e9c9f9852a93d43d7ce5f2b6

Observation a1ea861b-0d50-4ccb-a644-ed1ec0c13e3a · outbound

This paper cites Evaluating implied urban nature vital- ity in san francisco: An interdisciplinary approach combining census data, street view images, and social media analysis.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Evaluating implied urban nature vital- ity in san francisco: An interdisciplinary approach combining census data, street view images, and social media analysis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.500260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:51.893409Z digest=sha256:c9074fe37c11ca8572aaf8a6d65a9bbf208babde845ac87228e425374888becc

Observation 9251068b-ef49-40f2-95f1-c7f4d905fe27 · outbound

This paper cites Visual chain- of-thought prompting for knowledge-based visual reasoning.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Visual chain- of-thought prompting for knowledge-based visual reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.479241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:51.899935Z digest=sha256:f3b7591a92690c5a150500c9a6443e8a1843431ec4804b6df081c26d542534e7

Observation 034e278e-a75c-4733-bb7b-0bc140cb65fd · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images The cityscapes dataset for semantic urban scene understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.450500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:51.907753Z digest=sha256:d5b8bb4e3a42016a6838ca9d402b7a9ef3b83625da633b3ef99d2eeafcdbe44c

Observation 8ac9be47-7346-48d5-963b-985d87dbaeb4 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.421217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:51.916567Z digest=sha256:1c6f7cc0786c981ec6dcfcf0e9029f20b24347969a8f33f7957cff66330c2c8f

Observation 30e2affe-ff1b-4eac-af09-77c67577f38a · outbound

This paper cites Feeling Nature: Measuring perceptions of biophilia across global biomes using visual AI,.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Feeling Nature: Measuring perceptions of biophilia across global biomes using visual AI,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.401105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:51.922767Z digest=sha256:ae75eacd597ee19aab315dcf95bc77b88e1547170ec9c370125a0877ba89c90e

Observation 59eb0f1b-c05a-4a86-b9c0-572a2d45cd4d · outbound

This paper cites an unresolved cited work.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:43:53.380042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:51.931046Z digest=sha256:ed3b815527a1a750b8b9568a0bcf09ed95ad5a576694d499ff961721655c0acc

Observation 896a4ee2-26e0-48dc-a808-025e6a3038c5 · outbound

This paper cites The processing of negation and polarity: An overview.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images The processing of negation and polarity: An overview

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.339939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:51.938652Z digest=sha256:e9ad7d1de49b2775f7671231f805ee6337ef432a69c7df064c1a7753b7531fe0

Observation 1e2386f4-122c-4e00-96f1-806d80b07ccd · outbound

This paper cites Vqa-lol: Visual question answering under the lens of logic.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Vqa-lol: Visual question answering under the lens of logic

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.318735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:51.952079Z digest=sha256:7b3d8727ba7f0e092b3c64369fa862f4c68bfae799ff1c831cb2e2780b98063f

Observation e4a42176-215c-4c1e-9e75-0355d00db18d · outbound

This paper cites an unresolved cited work.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:43:53.299469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:51.963718Z digest=sha256:e605bbd59307ba1c0e934606486311a8cae6df809fd4dbfca8635b675e273485

Observation a27f872f-ccec-46c1-8cf2-84ec2b934bd4 · outbound

This paper cites an unresolved cited work.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:43:53.278062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:51.969332Z digest=sha256:dee459fe9ca5443abbe9aeabb1682b21cb71aa3d10aec644f0f36abe5ea569fb

Observation 8f65b9f8-92b4-47b9-a2ed-7820c2fceb60 · outbound

This paper cites Natural Adversarial Examples.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Natural Adversarial Examples

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:51.975849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:51.975849Z digest=sha256:9ea421536d15aa70737b9a195a486bac5d7f8f79c6cd40327e91630c9b4ed8d7

Observation ad873a86-b2d6-4301-afae-49cd838a606f · outbound

This paper cites Which street is hotter? street morphol- ogy may hold clues -thermal environment mapping based on street view imagery.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Which street is hotter? street morphol- ogy may hold clues -thermal environment mapping based on street view imagery

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.250774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:51.984863Z digest=sha256:9805f283c6cc998f82bd52be56ec7c621ef98e9815bba719e06af707b5b296d8

Observation c20955f7-af37-40ca-b428-c52899b1d746 · outbound

This paper cites Hudson and Christopher D.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Hudson and Christopher D

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.222186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:51.993458Z digest=sha256:f99b44799ceab0201902f0582f06a795410016946ed228e2c5d64be0b6ac99c9

Observation ef39f770-dd50-4905-b031-0f75a8d8f645 · outbound

This paper cites Lawrence Zitnick, and Ross Girshick.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Lawrence Zitnick, and Ross Girshick

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.194527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.001857Z digest=sha256:2522929461d5916059f859d395c3e66cbf1b80ba2b208843abe2437dd992c189

Observation a516bb51-ec3e-4461-9433-bf4296dd86a9 · outbound

This paper cites Negated and Misprimed Probes for Pretrained Language Models: Birds Can Talk, But Cannot Fly.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Negated and Misprimed Probes for Pretrained Language Models: Birds Can Talk, But Cannot Fly

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.008408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.008408Z digest=sha256:3c8162e5629cedc8932531ab5d89dc9323fa228ada4161cba15552e239230dfd

Observation f1b6b3ae-0fe0-4dbc-a7a4-73c76d0983dc · outbound

This paper cites Large language models are zero-shot reasoners.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Large language models are zero-shot reasoners

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.162260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.015411Z digest=sha256:a6cb4ddb948b432ea49ca37ca58596b4f0e3ad8908fc92ef72778e84b709623b

Observation 321bfdfc-88b5-43fd-9f77-bad28e8087ba · outbound

This paper cites Understanding counterfactuality: A review of experimental evidence for the dual meaning of counterfactuals.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Understanding counterfactuality: A review of experimental evidence for the dual meaning of counterfactuals

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.135549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.021415Z digest=sha256:f9a67d26e4118d08e6b6c4016ca4ada50daf9fb7005ecd1fbc77d58b037173c0

Observation 73401931-cbf7-4d71-ab53-b03723f9f749 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen im- age encoders and large language models.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Blip- 2: Bootstrapping language-image pre-training with frozen im- age encoders and large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.111299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.027458Z digest=sha256:0263778c08f6bc97c2fe5d9daafe8f67ec8edd86ad05fbd52d015aa09aa0d412

Observation 8a271aaf-a8e8-4376-8eda-b2f5c4d56c29 · outbound

This paper cites Assessing street-level urban greenery using google street view and a modified green view index.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Assessing street-level urban greenery using google street view and a modified green view index

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.078958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.034515Z digest=sha256:8a8930ed33bba1c1e13b6d37a81710e6532da14f53a540b4bea3cf13b4a5f127

Observation 73a49552-63b5-4c00-9117-49fdd9df3423 · outbound

This paper cites Generalizing vision-language models to novel domains: A comprehensive survey.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Generalizing vision-language models to novel domains: A comprehensive survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.041500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.041500Z digest=sha256:f38d66f7c1e4cac3136ff026e000ab5f1cfd0790e581c55164513b82155dc3aa

Observation 57ba5a6c-0101-40fd-bde1-f45f0b336c66 · outbound

This paper cites Eyes can deceive: Benchmarking counterfactual rea- soning abilities of multi-modal large language models.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Eyes can deceive: Benchmarking counterfactual rea- soning abilities of multi-modal large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.051160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.048824Z digest=sha256:0e05301023edfd6ebc750e1e516f84d23a293b7799425344fbaf80c5c1778ef9

Observation 3ce008d1-5d7e-4e8a-8b80-0ffb16569066 · outbound

This paper cites Evaluating human perception of building exteriors using street view imagery.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Evaluating human perception of building exteriors using street view imagery

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.020377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.060380Z digest=sha256:100a5a5338c9b33c50dba6d8553cb378ece2be31b47f3f5d61cdd02270acd630

Observation 54854b56-3587-41bf-9ac8-c81ee1c67148 · outbound

This paper cites Improved baselines with visual instruction tuning.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Improved baselines with visual instruction tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.066628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.066628Z digest=sha256:1f54692a5a684dbd956ddaf5983557574efa5fe029f3dc8f110198c413d554ee

Observation cfd2e46e-acbe-4658-83bd-ff45a21775ac · outbound

This paper cites Efficacy of Synthetic Data as a Benchmark.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Efficacy of Synthetic Data as a Benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.076558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.076558Z digest=sha256:d81a45346bceec0dfccaebd10daf9f3655c98d599ab595a95ab0ab311dd7d33f

Observation 5f6f332f-f551-458e-bf6d-e67a7ebae481 · outbound

This paper cites Review of methods used to es- timate the sky view factor in urban street canyons.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Review of methods used to es- timate the sky view factor in urban street canyons

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.972412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.084073Z digest=sha256:2547c8c0ec299184d93c8756678f00d6c731f0109fb474b7b982d50e6f072f65

Observation 4d6efcc9-874f-485c-b432-09bc7612846a · outbound

This paper cites Objective scoring of streetscape walkability related to leisure walking: Statistical modeling approach with semantic segmentation of google street view images.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Objective scoring of streetscape walkability related to leisure walking: Statistical modeling approach with semantic segmentation of google street view images

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.948854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.090882Z digest=sha256:5dba3c5bc7c490ebaa5b640ce808b5d2d26d56f0d656ace3df64f7952a5eefc6

Observation 8d40bd89-9d58-421e-bacc-c3b29d53fe68 · outbound

This paper cites Streetscore - predicting the perceived safety of one mil- lion streetscapes.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Streetscore - predicting the perceived safety of one mil- lion streetscapes

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.920991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.096672Z digest=sha256:ad9451f356b3126f6ab539b10a894d88ed883da65c95d3e565b7bf80a0f91bf0

Observation f7984d46-b145-4026-bdba-89620c64ac7c · outbound

This paper cites The mapillary vistas dataset for semantic understanding of street scenes.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images The mapillary vistas dataset for semantic understanding of street scenes

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.897550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.102021Z digest=sha256:a79f7f74b83a05cd58affc8fb448314aff15549d2a7895a0dfd3f150b09cd062

Observation 714a7014-80c6-4efc-8125-cbfcf2e43b1b · outbound

This paper cites Counterfactual vqa: A cause- effect look at language bias.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Counterfactual vqa: A cause- effect look at language bias

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.875949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.108525Z digest=sha256:3c0b35e3016d09b333c226aa3f7030c1e9eefaf909789ec1ad6550da0c6d6e7d

Observation 565549e6-5b89-4887-9a93-0608d7d0247d · outbound

This paper cites Evaluating the subjective percep- tions of streetscapes using street-view images.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Evaluating the subjective percep- tions of streetscapes using street-view images

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.840854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.117459Z digest=sha256:8b3e41d3a32f048d57f245992b542ed49e8a4fcf450d034c6921205a7724d496

Observation 5c478263-af87-4d41-9a62-0384e917d7c4 · outbound

This paper cites Learning transferable visual models from natural language supervision.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Learning transferable visual models from natural language supervision

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.810090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.124226Z digest=sha256:cdf60d2c4c05d2e62851a87809f8eff75c00e2c7eeca6227085760fd2c63fdf8

Observation f4217137-e637-49d5-baad-d2f8a655fbf9 · outbound

This paper cites Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.783106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.134226Z digest=sha256:330c0e7c63a900a011ae631f84f939dff6d9174752880e9dc3f74fc5ea6bb488

Observation c7917aaf-da19-4f57-91ae-3ff0fdab65e3 · outbound

This paper cites Visual cot: Advancing multi-modal language models with a compre- hensive dataset and benchmark for chain-of-thought reason- ing.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Visual cot: Advancing multi-modal language models with a compre- hensive dataset and benchmark for chain-of-thought reason- ing

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.749086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.145483Z digest=sha256:614258f9bb60796184d5c4ef7a464d5cd15684cdb4a6141dc9a476309e24ef2f

Observation d06398fc-d73b-4ba4-aaa4-61ac3c064d4d · outbound

This paper cites Is synthetic data all we need? benchmarking the robustness of models trained with synthetic images.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Is synthetic data all we need? benchmarking the robustness of models trained with synthetic images

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.724551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.151861Z digest=sha256:175bf7b73e22a49b107dfff90021dedb7b3e8f6e584ae3987080b1dc79235cc3

Observation 314146b7-e38e-4fac-9014-88c5c09a5ab6 · outbound

This paper cites Svensson.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Svensson

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.698831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.157932Z digest=sha256:6afcd8367b4b8ac56ebe9b1859e03627b7a9660d35db36501fadc2636c2a3374

Observation c045f472-8ee5-47ff-8d08-9841f4de7986 · outbound

This paper cites A new benchmark: On the util- ity of synthetic data with blender for bare supervised learn- ing and downstream domain adaptation.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images A new benchmark: On the util- ity of synthetic data with blender for bare supervised learn- ing and downstream domain adaptation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.663049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.165500Z digest=sha256:4ad905790449cb340aece424e7af0a96febe8a0be7d11bfaa5b2e74a9a9b5382

Observation 5ed35ed8-5924-49c6-82f0-ecdeb80ea2b2 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.173641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.173641Z digest=sha256:e853084ff5c3911d0fa59c1573f65b4ceae3c360debe808bff9d0e750f4bdf4e

Observation 31c188cc-dbb3-4a3c-b4c1-5d5c0a230415 · outbound

This paper cites Negation: A Pink Elephant in the Large Language Models' Room?.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Negation: A Pink Elephant in the Large Language Models' Room?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.181805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.181805Z digest=sha256:16979d218467a8732a040b114b6bfe09982733a246245393878ca04e58254cb4

Observation d676c83d-a650-4194-9a1e-d672822df9e4 · outbound

This paper cites Al- varez.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Al- varez

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.630835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.187965Z digest=sha256:0b093f668f1d347755725cd1a033cb3244c4ce8bfd4a3a21067d2dc9b72af2d4

Observation a065ad02-441c-4927-83fa-cdfbcdeeb656 · outbound

This paper cites Self-consistency improves chain of thought reasoning in lan- guage models.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Self-consistency improves chain of thought reasoning in lan- guage models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.598326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.195234Z digest=sha256:744c49e30bec1d3146468a9ec94a0227d8894d0f76f9363315d5f75b701da95d

Observation ff0f101b-8ba0-49a6-833d-513bb6a8efaa · outbound

This paper cites Chi, Quoc V.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Chi, Quoc V

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.572855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.201830Z digest=sha256:9df637f300ee079f310237b0576d080cbe9139ecb63b3529f81e66d806630656

Observation 0802696a-1af0-4e2f-8a0a-665d74fa90cc · outbound

This paper cites Alvarez, and Ping Luo.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Alvarez, and Ping Luo

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.536709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.212828Z digest=sha256:17da0f699a6ef9b52b8f822c06a50400d0e5d8e66d3acb4a8838ff6f0ab7d815

Observation 20365222-5643-4cc0-89c8-6a25d8ace7c8 · outbound

This paper cites Improve vision language model chain-of-thought reasoning, 2024.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Improve vision language model chain-of-thought reasoning, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.514988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.219179Z digest=sha256:ce2bb6137c61bb32fe19e75ed5cadc0790fe07c1399f0dd12868da32ca77ff40

Observation df454ef8-5092-4cbf-81b2-d66f19ccb8fa · outbound

This paper cites NegVQA: Can Vision Language Models Understand Negation?.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images NegVQA: Can Vision Language Models Understand Negation?

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.235796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.235796Z digest=sha256:cd81d2ec9881f961486172c4a3f98f720461948fbdff7391c4bcdb949d135ebb

Observation f516c9c9-7109-428c-9898-a64ab16fc56e · outbound

This paper cites A study on the impact of visible green index and vegetation structures on brain wave change in residential landscape.Urban Forestry & Urban Greening, 64:127299, 2021.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images A study on the impact of visible green index and vegetation structures on brain wave change in residential landscape.Urban Forestry & Urban Greening, 64:127299, 2021

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.491168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:43:52.244483Z digest=sha256:56f4b5c954795d4a2f3f9e5cb1a456730e1247df14ad6e1b7501e4378e4aad23

Pith citing papers

No inbound Pith citation observations are available.