Pith. sign in

Paper Citation Record · LEDGER

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

As of 8 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 6 inbound Pith citation observations for arXiv:2506.01277.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01277 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:52:10.218789Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:12:10.857306Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T23:06:20.239381Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact4
  • verified fuzzy20
  • unresolved14
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5b870ec5-2ecf-4f1f-9453-a203fb34e9ec · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Learning Transferable Visual Models From Natural Language Supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.326032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.326032Z digest=sha256:24bb8e5408f202eb860d0e9aa7bb0ac2d2b01cbeace7a188dea942f55ad5b8f5

Observation 6a14bccb-dcf5-45f6-b2a3-7a98fa9acb10 · outbound

This paper cites an unresolved cited work.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:52:13.349011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:08.362373Z digest=sha256:6b1f28966bafd4fd9fe629c8bb5e456820f49694b5adbbac066e9e39bd15f252

Observation 1141b9f7-2500-4a76-8e86-177824a35b62 · outbound

This paper cites PlaNet - photo geolocation with con- volutional neural networks.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models PlaNet - photo geolocation with con- volutional neural networks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.387126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.387126Z digest=sha256:dde3fb34532eecd80b8e66a03ac1ed71a9096899257c05c4e75b01766647a832

Observation 148ace07-28fb-49c5-85a7-6ba3ad394ec2 · outbound

This paper cites PIGEON: Predicting Image Geolocations.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models PIGEON: Predicting Image Geolocations

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:52:10.869770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:08.433343Z digest=sha256:a1e87ca72d68e9ad594e0bc0e0b6896ec549542e6f5db9c9f920a0dc89036103

Observation fee7286f-50b1-4b48-a43b-13fa9eb5c07f · outbound

This paper cites Image-Based Geolocation Using Large Vision-Language Models.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Image-Based Geolocation Using Large Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.551981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.551981Z digest=sha256:3f8816d71fcf7fb28b5932503bfd8f3a98b38b2ad5b24a8271905ac98cbf0cdf

Observation a2dc9744-dbad-4f52-852a-7bfcb6c5134f · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.601226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.601226Z digest=sha256:b13600dcd0741836bc0eca43538010bb75cf6b6ca72782d42c9d32602e7be8c8

Observation 7b13da83-2661-40bd-b0f2-60b00ee1c4f7 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.630800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.630800Z digest=sha256:7ae6f4376207f046623750130a7bf989368949bf1ad15958fcaf4fe883ebbdf5

Observation 30c03c75-74da-45cd-a91b-d1218652f3af · outbound

This paper cites Neural network ensembles.IEEE Transactions on Pattern Analysis and Machine Intelligence, 12(10):993–1001, 1990.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Neural network ensembles.IEEE Transactions on Pattern Analysis and Machine Intelligence, 12(10):993–1001, 1990

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:13.244819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:08.677012Z digest=sha256:438ec473b61d0433ecd7329974dd7dccfdba8f869d0becde9d4f156eaf4b9cb7

Observation a64f0ec5-2882-4c39-8759-e00d3c24ef5a · outbound

This paper cites Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.724712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.724712Z digest=sha256:bc292ac430bf6a27a6ca511c1b80fc33431dc1a8217398392063151f7528c1f8

Observation 8dcf04f6-b1eb-42bb-9bd8-6608e1952def · outbound

This paper cites Where We Are and What We're Looking At: Query Based Worldwide Image Geo-localization Using Hierarchies and Scenes.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Where We Are and What We're Looking At: Query Based Worldwide Image Geo-localization Using Hierarchies and Scenes

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:52:10.616398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:08.764186Z digest=sha256:14ab780c5b33e0742039b928d8b14f0c521977b6fd9a6b3e599424d05e0a8f8c

Observation eb76507a-c259-486b-a3ce-9e3b118146cc · outbound

This paper cites The Mapil- lary Vistas dataset for semantic understanding of street scenes.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models The Mapil- lary Vistas dataset for semantic understanding of street scenes

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:13.100674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:08.801352Z digest=sha256:2ef44f11b0fedbee10353080ac5d64c2996cced74003ebd3cd85a0aa5771c52c

Observation 25452abf-1357-45ad-88ff-b01d1e19727f · outbound

This paper cites GeoNames geographical database.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models GeoNames geographical database

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.947008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:08.826571Z digest=sha256:577933be1aec98fe2110b80841e8dea76f9b8c61359182b6b1bd1322587e2849

Observation 5169bda4-4f7d-42c3-a9a2-4dcc23558192 · outbound

This paper cites Qwen2.5-VL Technical Report.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Qwen2.5-VL Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:09.047711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.047711Z digest=sha256:448a91e2fb9a0e504b5576a8be1b99ebdda41223b890b1b5e42137ca04878f37

Observation 54ddb3a1-3421-4792-a162-fd35c6d2a87d · outbound

This paper cites Decoupled Weight Decay Regularization.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Decoupled Weight Decay Regularization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:09.081528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.081528Z digest=sha256:0e6fcad940d22fde2c17a12073144b0e14cef773ae9617db86b534506d347f41

Observation 9ad4f2a6-8671-4596-87c9-e9dc77b2b336 · outbound

This paper cites OpenStreetView-5M: The Many Roads to Global Visual Geolocation.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models OpenStreetView-5M: The Many Roads to Global Visual Geolocation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:09.120531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.120531Z digest=sha256:9567bb2a19bd93182198be89f5474991b0bb24f01dbdad1401c4d2d2365e8dfb

Observation 193b1e22-a3e6-49ce-8702-3dc16a604b34 · outbound

This paper cites The Claude 3 model family: Opus, Sonnet, Haiku.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models The Claude 3 model family: Opus, Sonnet, Haiku

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.868679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:09.193032Z digest=sha256:0fe93b876b50e2422e0fae308bcb229b96ffb9ecad7798cf007f81b6ec1f15c9

Observation 7631d502-cc44-4fdb-9cae-f977578e16e5 · outbound

This paper cites Visual Instruction Tuning.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Visual Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:09.261284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.261284Z digest=sha256:6c6740c43e7d92448c438f3d07b18da53c2d873ca834032a4546a0ce9d8b9f54

Observation b8c927a4-7674-4873-acef-447b02dfb975 · outbound

This paper cites Mistral-Small-3.1-24B-Instruct-2503.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Mistral-Small-3.1-24B-Instruct-2503

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.768597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:09.302998Z digest=sha256:69acdad2d02ee333a42d7fcb16c166db1b4215ad797bc17752cbf10be26eeb1e

Observation 00ff3373-43a2-4640-9a60-16d45423daa3 · outbound

This paper cites Revisiting IM2GPS in the Deep Learning Era.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Revisiting IM2GPS in the Deep Learning Era

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:09.371329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.371329Z digest=sha256:07e16c4ef6760dd66c75a49e501e63dd3f7a7939b6305058242f37c17523da6a

Observation fe163f24-9bf9-4a1e-9068-23c0b9512cd1 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:09.515725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.515725Z digest=sha256:b5e74be49412a144b750ea8a2ee80f2150eb96ce52f9ef1ef9e59846430c2bf4

Observation 4c703f5e-d585-4cf8-9785-e4d83a4ec20f · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 24

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:52:09.555930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.555930Z digest=sha256:f00ea3fa262fbdc9d8365e152ba5a4a4b94c7d6c52a5712a35eeb92262e52872

Observation e70fda6e-38f2-403b-a5b8-43f8a2764ae7 · outbound

This paper cites Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:52:10.415712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:09.473416Z digest=sha256:16c67909d13ff83a349d9534c3675cf2a40084ad7698eff5c2454cbf4e472413

Observation 7dbbf803-2ed1-4664-ac0f-00481042e4bb · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.675504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:09.594635Z digest=sha256:0c650a7a4e90b014015f0424e472cc4ecedfe2c416dcccf978a75a5b8d5f7640

Observation 947e2810-6378-40e2-b3cf-419d7b1409a3 · outbound

This paper cites Guidelines: • The answer NA means that the paper has no limitation while the answer No means that the paper has limitations, but those are not discussed in the paper.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper has no limitation while the answer No means that the paper has limitations, but those are not discussed in the paper

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.601524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:09.627183Z digest=sha256:a0082a6522f0c3df68915d1c95691b10be239cdaf272e19005ce4ae4c4f4ae57

Observation 9bcd7b9a-f880-435d-aec0-4ad3a7a11462 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.486517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:09.687952Z digest=sha256:9463d28402234965faf915955a945c4cc0a54dd83bc7ef950048a7b6c50ff5a3

Observation 62811774-990e-4f8c-9518-4fa0e1845b78 · outbound

This paper cites The appendices include hyperparameters for SFT and information about computational resources used.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models The appendices include hyperparameters for SFT and information about computational resources used

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.393269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:09.751346Z digest=sha256:c40ca19f299ff1e052f965d8c424626e075a1b64385b46764d93f363c51b8599

Observation 7dc63ab1-963b-4bcd-b217-e92994f77b42 · outbound

This paper cites Our appendices provide detailed instructions regarding implementation, hyperparameters, and experimental setup to facilitate reproduction.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Our appendices provide detailed instructions regarding implementation, hyperparameters, and experimental setup to facilitate reproduction

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.301578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:09.792357Z digest=sha256:ebbd4cb38eb201955c642053d066e5d58448930db2b17dbeba3a0d0a5c109fbf

Observation 1dba0e3d-4065-44e2-a2ff-01d772224858 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper does not include experiments

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.220821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:09.842854Z digest=sha256:cef9945e2ce8f41ccc4330de5705e9c80007bc929962725829e39c5508443839

Observation 5cc46a9c-3c83-45bd-a47d-a8079eb9fd12 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper does not include experiments

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.110456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:09.878555Z digest=sha256:e3240cf06c2b5fe3689bd8d0161ef64c1d6eeb7d653c6798f1e87a8a7ce54259

Observation 180e1af7-572b-4d03-8655-1beb3d7b8761 · outbound

This paper cites We also report model sizes and memory requirements.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models We also report model sizes and memory requirements

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.988746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:09.888164Z digest=sha256:3fd13e4512db873af13e026965b37913c9e56456f53e72b111ed70a18a29737a

Observation 13c45ece-fa9c-4e7a-836b-b714c48619b8 · outbound

This paper cites We use publicly available datasets, acknowledge relevant prior work, and are transparent about our methodologies.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models We use publicly available datasets, acknowledge relevant prior work, and are transparent about our methodologies

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.886563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:09.925377Z digest=sha256:bbc0197db1a76a806809780cdc362503b714a847f2db9efdef54b58ddf704cbf

Observation 54276147-772a-4883-9c90-95322fe4ce64 · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.755852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:09.978969Z digest=sha256:03f668a2af8f9e940a88b21dacafe016599afbc0836603dc5725ceb9e7a6cd88

Observation bbf48e5d-e190-40dd-9d6b-1ed9331aa26a · outbound

This paper cites 27 Guidelines: • The answer NA means that the paper poses no such risks.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models 27 Guidelines: • The answer NA means that the paper poses no such risks

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.621881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:10.027853Z digest=sha256:6ac5213a0f5a2926d298b454c2d3fc3553af6f5236f3a49df199daa61ebcded4

Observation 3ef73617-d963-4be5-98a5-33ed5ad543e8 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper does not use existing assets

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.507759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:10.068261Z digest=sha256:44f2b6ab19cdb8887e991825af00b54eab3f287573dacd98aa4207d2443b1841

Observation 8125e676-e808-4062-8bd4-3cebc78c0535 · outbound

This paper cites This documentation will be released alongside the dataset.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models This documentation will be released alongside the dataset

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.310523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:10.112687Z digest=sha256:2e14cf705934fbb5e694d5acb96f617325ef712f74caddbc61632ad9f4186931

Observation 04d645b3-ad62-46d9-a66e-d21e141be845 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.158826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:10.143213Z digest=sha256:f8360bf23a4ce80d91705236deef6a58090cb4b0b315b526626481df1a32fbcd

Observation c14f16c0-6101-481f-82fd-6053b2217ae5 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:10.997861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:10.218789Z digest=sha256:5b61ffc30a68cbedab1a5880aefd6cc295ce1ab5d67cd7cfb9bf64a964560b46

Observation 2183609d-e1a8-486a-821c-40d15e3a9070 · outbound

This paper cites G3: An Effective and Adaptive Framework for Worldwide Geolocalization Using Large Multi-Modality Models.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models G3: An Effective and Adaptive Framework for Worldwide Geolocalization Using Large Multi-Modality Models

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:52:10.788245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:08.501464Z digest=sha256:0ab3741ddd87ade6128fc93ba23f6a27045990293f376a94b4028c77fe480172

Observation b42b2fe3-b71a-4bb4-a522-2881740e889e · outbound

This paper cites Gemma 3 Technical Report.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Gemma 3 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.967423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.967423Z digest=sha256:9dc93c6ec0aa2b6b97c861ef7b1dcd2982e5efc013bc80144b8a5140f0fe722d

Pith citing papers

Observation 5da9a71f-b019-499b-aeee-179f85ab3b52 · inbound

A Global-Local Cross-Attention Network for Ultra-high Resolution Remote Sensing Image Semantic Segmentation cites this paper.

A Global-Local Cross-Attention Network for Ultra-high Resolution Remote Sensing Image Semantic Segmentation GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:10.857306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:10.857306Z digest=sha256:532a8c313190f3b1f2d894b7b3e1a2ae2450e822bf14ce1483004243b91231e8

Observation c1b68c6f-60f2-4d40-b263-be1b4624e2e8 · inbound

Building Lightweight Semantic Segmentation Models for Aerial Images Using Dual Relation Distillation cites this paper.

Building Lightweight Semantic Segmentation Models for Aerial Images Using Dual Relation Distillation GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:29.112400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:00:29.112400Z digest=sha256:628bca378c8ac383a947ceaf45f9a4d56aa49f66b1cc5a21143a1de4aa6dfa6d

Observation 0a34a1ef-010f-4b43-8de7-cf7d9ee70961 · inbound

GDGS: 3D Gaussian Splatting Via Geometry-Guided Initialization And Dynamic Density Control cites this paper.

GDGS: 3D Gaussian Splatting Via Geometry-Guided Initialization And Dynamic Density Control GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:24:31.913301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:24:31.913301Z digest=sha256:d49324810f62df7a538e8bc4b3cea10f2f3f91f5b587fb0b020d96a7acb5fc50

Observation 4aab417b-4cd4-4650-9ded-b40c41c22580 · inbound

From Pixels to Places: A Systematic Benchmark for Evaluating Image Geolocalization Ability in Large Language Models cites this paper.

From Pixels to Places: A Systematic Benchmark for Evaluating Image Geolocalization Ability in Large Language Models GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:50:47.271597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T23:49:59.795575Z digest=sha256:29006b3a6ef8d0cd218afa8810e5cd225ca266fd09ca37bbccf209af96133b81

Observation 556d334e-529f-4853-a0ff-ca369f038b59 · inbound

Do VLMs See What Sensors Feel? A Scalable Expert-Guided Design for Wheelchair Accessibility Assessment from Street View cites this paper.

Do VLMs See What Sensors Feel? A Scalable Expert-Guided Design for Wheelchair Accessibility Assessment from Street View GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:20.241225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T14:42:47.815192Z digest=sha256:e34345f36a0ccd2a444360565b99c1c18fcf85543ea09c8139ede484672751e5

Observation 33241b1a-4fed-42af-80bf-13ec998b5cfb · inbound

DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization cites this paper.

DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-30T22:45:19.083071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T22:45:19.083071Z digest=sha256:3eee3768d91a2be5b40f7a195792804758c3d5c058902df710b103efe42aa967