Pith. sign in

Paper Citation Record · LEDGER

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images

As of 18 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2508.21565.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21565 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:43:52.244483Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy40
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa73e3e2-8e30-475d-92a4-f9dd7715655b · outbound

This paper cites Google street view: Capturing the world at street level.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Google street view: Capturing the world at street level

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.663051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:51.819293Z digest=sha256:7417c8d8f95264b2955fff60c472de1f94dccc439e8f3c48e3f2857ed84b9232

Observation 3aab4094-9239-493d-864c-2edac5b47858 · outbound

This paper cites Evaluation methods for landscapes with greenery.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Evaluation methods for landscapes with greenery

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.639234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:51.827290Z digest=sha256:f38de8408fceb16da64e974f16ed065fddd3a2e6ebd22497aeba181dc5a373aa

Observation a1fb3897-2151-4dca-947c-80252575d21d · outbound

This paper cites Browning, Jiaying Dong, Kuiran Zhang, Shuai Yuan, H¨useyin Ertan ˙Inan, Olivia McAnirlin, Dani T.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Browning, Jiaying Dong, Kuiran Zhang, Shuai Yuan, H¨useyin Ertan ˙Inan, Olivia McAnirlin, Dani T

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.612154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:51.847823Z digest=sha256:0187b6755727992d5f670a850df7adccd654f588857bf98f36b7c62cc1446204

Observation e8d66e39-f41c-437c-a787-4cba57e8efd7 · outbound

This paper cites The green window view index: automated multi-source visi- bility analysis for a multi-scale assessment of green window views.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images The green window view index: automated multi-source visi- bility analysis for a multi-scale assessment of green window views

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.589324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:51.859841Z digest=sha256:48ac14f973cf7f4904b311927982a4a598b591d9fe5dcc67e2baf287238d6398

Observation 74d718cc-de8b-481c-9e96-cbea460508ac · outbound

This paper cites Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:51.870945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:51.870945Z digest=sha256:1312213df1d259d3a8981a2433ea13903e24b761310b4f70119ccd646c5c28de

Observation afb504ce-f976-4375-bcc5-24f1e60f4a60 · outbound

This paper cites End-to- end object detection with transformers.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images End-to- end object detection with transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.546217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:51.879017Z digest=sha256:1eca9ffa05842c2f7f1d1025161803742a3838d03a5754679f5b5154407f3926

Observation c147cb24-a1c3-4ea8-bd4f-47b197d61d61 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.521387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:51.886291Z digest=sha256:f8e9cc83859704fc7d9c962da27cefb6bb188576d5a03387cb92f7a746daede8

Observation a1ea861b-0d50-4ccb-a644-ed1ec0c13e3a · outbound

This paper cites Evaluating implied urban nature vital- ity in san francisco: An interdisciplinary approach combining census data, street view images, and social media analysis.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Evaluating implied urban nature vital- ity in san francisco: An interdisciplinary approach combining census data, street view images, and social media analysis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.500260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:51.893409Z digest=sha256:e5fc292edc692a8f12dcace170d08e14bf8811ae0c231789cf22f0810cc4399a

Observation 9251068b-ef49-40f2-95f1-c7f4d905fe27 · outbound

This paper cites Visual chain- of-thought prompting for knowledge-based visual reasoning.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Visual chain- of-thought prompting for knowledge-based visual reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.479241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:51.899935Z digest=sha256:1309d32a6815218fcb4b0a3c22abcb7cde22908baedd6ed7daa7093eeff9aebb

Observation 034e278e-a75c-4733-bb7b-0bc140cb65fd · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images The cityscapes dataset for semantic urban scene understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.450500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:51.907753Z digest=sha256:4c3f4a0586c63dd7727b1aa146cf089ab836eb386deb3ed7dfd696b2f6f450ea

Observation 8ac9be47-7346-48d5-963b-985d87dbaeb4 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.421217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:51.916567Z digest=sha256:9a5558c7ed34ca5dda1a83187e419dea0487c19dbddeb7a8f6e3fe76aa6e0880

Observation 30e2affe-ff1b-4eac-af09-77c67577f38a · outbound

This paper cites Feeling Nature: Measuring perceptions of biophilia across global biomes using visual AI,.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Feeling Nature: Measuring perceptions of biophilia across global biomes using visual AI,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.401105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:51.922767Z digest=sha256:9eec0b55c7392f8cdd2e886ee570002d94981b703726a09ec2df267de8d5a5f3

Observation 59eb0f1b-c05a-4a86-b9c0-572a2d45cd4d · outbound

This paper cites an unresolved cited work.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:43:53.380042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:51.931046Z digest=sha256:f22f7e9b8634afe7d017d629ad357f18a1e7acb77fb264f9aa16026e39d2c305

Observation 896a4ee2-26e0-48dc-a808-025e6a3038c5 · outbound

This paper cites The processing of negation and polarity: An overview.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images The processing of negation and polarity: An overview

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.339939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:51.938652Z digest=sha256:61c37f9199624f039b050d590ac7b348bcfd691f2f4203aa1fbfd866138cf45a

Observation 1e2386f4-122c-4e00-96f1-806d80b07ccd · outbound

This paper cites Vqa-lol: Visual question answering under the lens of logic.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Vqa-lol: Visual question answering under the lens of logic

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.318735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:51.952079Z digest=sha256:9d5e1f3089acd1d74924c4c90747ac66274fab66aba3a4592f0cf49925136b80

Observation e4a42176-215c-4c1e-9e75-0355d00db18d · outbound

This paper cites an unresolved cited work.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:43:53.299469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:51.963718Z digest=sha256:4eb447b6d3cea1c1b10b21457666d74f3a320868f3f7c9da625e5cfd970f6d99

Observation a27f872f-ccec-46c1-8cf2-84ec2b934bd4 · outbound

This paper cites an unresolved cited work.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:43:53.278062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:51.969332Z digest=sha256:cfb25033dfb315ad05cbd90fa14f219ee18e1d62a1ff196fe659b24240b446fd

Observation 8f65b9f8-92b4-47b9-a2ed-7820c2fceb60 · outbound

This paper cites Natural Adversarial Examples.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Natural Adversarial Examples

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:51.975849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:51.975849Z digest=sha256:9ea421536d15aa70737b9a195a486bac5d7f8f79c6cd40327e91630c9b4ed8d7

Observation ad873a86-b2d6-4301-afae-49cd838a606f · outbound

This paper cites Which street is hotter? street morphol- ogy may hold clues -thermal environment mapping based on street view imagery.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Which street is hotter? street morphol- ogy may hold clues -thermal environment mapping based on street view imagery

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.250774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:51.984863Z digest=sha256:ad3175bb7c11db182c892df832c1d663ae910362e0420876bd51cc44af80a708

Observation c20955f7-af37-40ca-b428-c52899b1d746 · outbound

This paper cites Hudson and Christopher D.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Hudson and Christopher D

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.222186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:51.993458Z digest=sha256:d3c4888b29170aa86e0a476334811b29169a09b129f15ec594fdfc1b99d2baff

Observation ef39f770-dd50-4905-b031-0f75a8d8f645 · outbound

This paper cites Lawrence Zitnick, and Ross Girshick.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Lawrence Zitnick, and Ross Girshick

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.194527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.001857Z digest=sha256:0b8b3e957589ed7b777a44b33a7135d9bd805be9ab43012547b15ee8d968975c

Observation a516bb51-ec3e-4461-9433-bf4296dd86a9 · outbound

This paper cites Negated and Misprimed Probes for Pretrained Language Models: Birds Can Talk, But Cannot Fly.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Negated and Misprimed Probes for Pretrained Language Models: Birds Can Talk, But Cannot Fly

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.008408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.008408Z digest=sha256:3c8162e5629cedc8932531ab5d89dc9323fa228ada4161cba15552e239230dfd

Observation f1b6b3ae-0fe0-4dbc-a7a4-73c76d0983dc · outbound

This paper cites Large language models are zero-shot reasoners.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Large language models are zero-shot reasoners

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.162260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.015411Z digest=sha256:fe4c3421a218ab899807b9b1d6163cc6dcadf18ed7d0cab41bb377aec205ff34

Observation 321bfdfc-88b5-43fd-9f77-bad28e8087ba · outbound

This paper cites Understanding counterfactuality: A review of experimental evidence for the dual meaning of counterfactuals.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Understanding counterfactuality: A review of experimental evidence for the dual meaning of counterfactuals

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.135549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.021415Z digest=sha256:43e3f2d2bb0312ec3dfd72f6862ae8e2433fc774f2e038204fadb1da46b6eec5

Observation 73401931-cbf7-4d71-ab53-b03723f9f749 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen im- age encoders and large language models.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Blip- 2: Bootstrapping language-image pre-training with frozen im- age encoders and large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.111299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.027458Z digest=sha256:c103b5202cca1413c82286bc488e1afd3736e336dd55262fedee54e0c0e50c0d

Observation 8a271aaf-a8e8-4376-8eda-b2f5c4d56c29 · outbound

This paper cites Assessing street-level urban greenery using google street view and a modified green view index.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Assessing street-level urban greenery using google street view and a modified green view index

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.078958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.034515Z digest=sha256:705906afdc18686fa8afce57795f361e1fec28e39ae2c69871ff2342bab36b3e

Observation 73a49552-63b5-4c00-9117-49fdd9df3423 · outbound

This paper cites Generalizing vision-language models to novel domains: A comprehensive survey.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Generalizing vision-language models to novel domains: A comprehensive survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.041500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.041500Z digest=sha256:f38d66f7c1e4cac3136ff026e000ab5f1cfd0790e581c55164513b82155dc3aa

Observation 57ba5a6c-0101-40fd-bde1-f45f0b336c66 · outbound

This paper cites Eyes can deceive: Benchmarking counterfactual rea- soning abilities of multi-modal large language models.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Eyes can deceive: Benchmarking counterfactual rea- soning abilities of multi-modal large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.051160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.048824Z digest=sha256:aab297d523ac559ee867435402d41a0592462ca9f7dba5466beb31f597b1e316

Observation 3ce008d1-5d7e-4e8a-8b80-0ffb16569066 · outbound

This paper cites Evaluating human perception of building exteriors using street view imagery.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Evaluating human perception of building exteriors using street view imagery

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.020377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.060380Z digest=sha256:aff025b07277284d9e89bd924964e9c38a8f374dabbe526dea5138d7381c56b7

Observation 54854b56-3587-41bf-9ac8-c81ee1c67148 · outbound

This paper cites Improved baselines with visual instruction tuning.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Improved baselines with visual instruction tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.066628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.066628Z digest=sha256:1f54692a5a684dbd956ddaf5983557574efa5fe029f3dc8f110198c413d554ee

Observation cfd2e46e-acbe-4658-83bd-ff45a21775ac · outbound

This paper cites Efficacy of Synthetic Data as a Benchmark.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Efficacy of Synthetic Data as a Benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.076558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.076558Z digest=sha256:d81a45346bceec0dfccaebd10daf9f3655c98d599ab595a95ab0ab311dd7d33f

Observation 5f6f332f-f551-458e-bf6d-e67a7ebae481 · outbound

This paper cites Review of methods used to es- timate the sky view factor in urban street canyons.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Review of methods used to es- timate the sky view factor in urban street canyons

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.972412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.084073Z digest=sha256:46b91338c0051dab04d86729e47fbbe90a4b53c7e950c55b08d0e057428aabd2

Observation 4d6efcc9-874f-485c-b432-09bc7612846a · outbound

This paper cites Objective scoring of streetscape walkability related to leisure walking: Statistical modeling approach with semantic segmentation of google street view images.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Objective scoring of streetscape walkability related to leisure walking: Statistical modeling approach with semantic segmentation of google street view images

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.948854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.090882Z digest=sha256:d6177768ac5d6b5ec14720c45bead078f842031fa7f978303ac8fdaaf1fc15a6

Observation 8d40bd89-9d58-421e-bacc-c3b29d53fe68 · outbound

This paper cites Streetscore - predicting the perceived safety of one mil- lion streetscapes.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Streetscore - predicting the perceived safety of one mil- lion streetscapes

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.920991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.096672Z digest=sha256:c030df1acae06308a1b5fec25c7d9e06d4a62cf39062abda7ec12063684bdbfd

Observation f7984d46-b145-4026-bdba-89620c64ac7c · outbound

This paper cites The mapillary vistas dataset for semantic understanding of street scenes.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images The mapillary vistas dataset for semantic understanding of street scenes

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.897550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.102021Z digest=sha256:36f1535e08a5772a1054b8f717cca4267eb5a84f7f0951b52414f6bc39b55f89

Observation 714a7014-80c6-4efc-8125-cbfcf2e43b1b · outbound

This paper cites Counterfactual vqa: A cause- effect look at language bias.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Counterfactual vqa: A cause- effect look at language bias

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.875949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.108525Z digest=sha256:f0e6eb0a5b19dae5f012e7f6749a1742e1abd2debce3479ed8e793213da99904

Observation 565549e6-5b89-4887-9a93-0608d7d0247d · outbound

This paper cites Evaluating the subjective percep- tions of streetscapes using street-view images.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Evaluating the subjective percep- tions of streetscapes using street-view images

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.840854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.117459Z digest=sha256:3e2409fda2ba5c79b59cadc2db9a333431f90529eebb4c6426724e12372423c1

Observation 5c478263-af87-4d41-9a62-0384e917d7c4 · outbound

This paper cites Learning transferable visual models from natural language supervision.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Learning transferable visual models from natural language supervision

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.810090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.124226Z digest=sha256:e6763370abec89a595a50ddad14b3366573b7e452f16083025ee54eee14869c5

Observation f4217137-e637-49d5-baad-d2f8a655fbf9 · outbound

This paper cites Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.783106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.134226Z digest=sha256:b9f48d47777ee73bb6578204f68f8fff0ba0c5acdbc190f0314c69d890f24cc8

Observation c7917aaf-da19-4f57-91ae-3ff0fdab65e3 · outbound

This paper cites Visual cot: Advancing multi-modal language models with a compre- hensive dataset and benchmark for chain-of-thought reason- ing.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Visual cot: Advancing multi-modal language models with a compre- hensive dataset and benchmark for chain-of-thought reason- ing

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.749086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.145483Z digest=sha256:02500b55727cd96a501b2e586458450ad0c775628480f2e7a2e39198d7f09157

Observation d06398fc-d73b-4ba4-aaa4-61ac3c064d4d · outbound

This paper cites Is synthetic data all we need? benchmarking the robustness of models trained with synthetic images.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Is synthetic data all we need? benchmarking the robustness of models trained with synthetic images

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.724551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.151861Z digest=sha256:c8a611aa508a0090cb8dd41bcf072afef4ab79ee0c5781b1b811e0d6b55820d7

Observation 314146b7-e38e-4fac-9014-88c5c09a5ab6 · outbound

This paper cites Svensson.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Svensson

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.698831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.157932Z digest=sha256:86853fb4c93e5a636c8197df88b5b2cbc4506e6cbffb0f537cb20441f79d01c4

Observation c045f472-8ee5-47ff-8d08-9841f4de7986 · outbound

This paper cites A new benchmark: On the util- ity of synthetic data with blender for bare supervised learn- ing and downstream domain adaptation.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images A new benchmark: On the util- ity of synthetic data with blender for bare supervised learn- ing and downstream domain adaptation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.663049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.165500Z digest=sha256:37094e166b94ab0a8bb0fc079956db53d640a9f224fea1e6cf1677080e4df98e

Observation 5ed35ed8-5924-49c6-82f0-ecdeb80ea2b2 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.173641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.173641Z digest=sha256:e853084ff5c3911d0fa59c1573f65b4ceae3c360debe808bff9d0e750f4bdf4e

Observation 31c188cc-dbb3-4a3c-b4c1-5d5c0a230415 · outbound

This paper cites Negation: A Pink Elephant in the Large Language Models' Room?.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Negation: A Pink Elephant in the Large Language Models' Room?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.181805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.181805Z digest=sha256:16979d218467a8732a040b114b6bfe09982733a246245393878ca04e58254cb4

Observation d676c83d-a650-4194-9a1e-d672822df9e4 · outbound

This paper cites Al- varez.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Al- varez

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.630835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.187965Z digest=sha256:e0ae0b21d0b0e1cf871f926a48d6968edf5e0ab0dd0aa9a4cefc8a9ffd55b493

Observation a065ad02-441c-4927-83fa-cdfbcdeeb656 · outbound

This paper cites Self-consistency improves chain of thought reasoning in lan- guage models.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Self-consistency improves chain of thought reasoning in lan- guage models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.598326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.195234Z digest=sha256:54602fedb2e2482cedf64b0e008ee84a9c6d40e394a505f8edd4330518d2951d

Observation ff0f101b-8ba0-49a6-833d-513bb6a8efaa · outbound

This paper cites Chi, Quoc V.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Chi, Quoc V

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.572855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.201830Z digest=sha256:9e3855f863e0350f9811e6a9543185826def7c6fc88a7059444faa6898a7e9ca

Observation 0802696a-1af0-4e2f-8a0a-665d74fa90cc · outbound

This paper cites Alvarez, and Ping Luo.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Alvarez, and Ping Luo

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.536709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.212828Z digest=sha256:d563cf5a6b849a675c1837227d30da18db5d992a151d773b5a4919a0d984ab77

Observation 20365222-5643-4cc0-89c8-6a25d8ace7c8 · outbound

This paper cites Improve vision language model chain-of-thought reasoning, 2024.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Improve vision language model chain-of-thought reasoning, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.514988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.219179Z digest=sha256:89db091eae26d6c90a7279ea6d4e234adf743109e2cc8c4e68d6675a2c8d0040

Observation df454ef8-5092-4cbf-81b2-d66f19ccb8fa · outbound

This paper cites NegVQA: Can Vision Language Models Understand Negation?.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images NegVQA: Can Vision Language Models Understand Negation?

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.235796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.235796Z digest=sha256:cd81d2ec9881f961486172c4a3f98f720461948fbdff7391c4bcdb949d135ebb

Observation f516c9c9-7109-428c-9898-a64ab16fc56e · outbound

This paper cites A study on the impact of visible green index and vegetation structures on brain wave change in residential landscape.Urban Forestry & Urban Greening, 64:127299, 2021.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images A study on the impact of visible green index and vegetation structures on brain wave change in residential landscape.Urban Forestry & Urban Greening, 64:127299, 2021

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.491168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T16:43:52.244483Z digest=sha256:a2c88cda076049315f6bcae8e8983cd7f0e94b1c4774a1e4d8b9aecade00f6ec

Pith citing papers

No inbound Pith citation observations are available.