Pith. sign in

Paper Citation Record · LEDGER

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

As of 16 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 6 inbound Pith citation observations for arXiv:2506.01277.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01277 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:52:10.218789Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:12:10.857306Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T23:06:20.239381Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact4
  • verified fuzzy20
  • unresolved14
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5b870ec5-2ecf-4f1f-9453-a203fb34e9ec · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Learning Transferable Visual Models From Natural Language Supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.326032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.326032Z digest=sha256:368f7ffa2ba6faa2502ef3f9cdc655710630dd23ab132694ee95698dd22beb8e

Observation 6a14bccb-dcf5-45f6-b2a3-7a98fa9acb10 · outbound

This paper cites an unresolved cited work.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:52:13.349011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:08.362373Z digest=sha256:aa57ce4f3f1cabea2dc2a7e6191be4737ae33a4dde687c13a910a59ab94954a9

Observation 1141b9f7-2500-4a76-8e86-177824a35b62 · outbound

This paper cites PlaNet - photo geolocation with con- volutional neural networks.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models PlaNet - photo geolocation with con- volutional neural networks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.387126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.387126Z digest=sha256:91fe1c76dfdc2fae4b9cdbf41ffcdb314e1b17131f966e0f36d0b90edfb86be0

Observation 148ace07-28fb-49c5-85a7-6ba3ad394ec2 · outbound

This paper cites PIGEON: Predicting Image Geolocations.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models PIGEON: Predicting Image Geolocations

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:52:10.869770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:08.433343Z digest=sha256:44c7ce2f51c1fba517477c8acdda397564ba3e67501b40ff8d4993d6a176ab3e

Observation fee7286f-50b1-4b48-a43b-13fa9eb5c07f · outbound

This paper cites Image-Based Geolocation Using Large Vision-Language Models.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Image-Based Geolocation Using Large Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.551981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.551981Z digest=sha256:de06e2e1d16e0876fa6642bf7b849ba8cb024ec5bcd3bed20e59478fe57368f0

Observation a2dc9744-dbad-4f52-852a-7bfcb6c5134f · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.601226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.601226Z digest=sha256:44fee9218a1dd389e1aae701636bbb988c699d91926f44f48d4fdeb5bb50f714

Observation 7b13da83-2661-40bd-b0f2-60b00ee1c4f7 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.630800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.630800Z digest=sha256:720d31778236dcff843f5ad670cd7f256852c3372c0002bdb6df9d241089a3dd

Observation 30c03c75-74da-45cd-a91b-d1218652f3af · outbound

This paper cites Neural network ensembles.IEEE Transactions on Pattern Analysis and Machine Intelligence, 12(10):993–1001, 1990.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Neural network ensembles.IEEE Transactions on Pattern Analysis and Machine Intelligence, 12(10):993–1001, 1990

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:13.244819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:08.677012Z digest=sha256:3535576d17e9a15b7882b0c16207d02ee6c72f74332f751ab8967b417028b5aa

Observation a64f0ec5-2882-4c39-8759-e00d3c24ef5a · outbound

This paper cites Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.724712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.724712Z digest=sha256:6b986c12bf432668a1d14f4ef7e443d78c59d073f97e2c6a6bb8219c5c7c200c

Observation 8dcf04f6-b1eb-42bb-9bd8-6608e1952def · outbound

This paper cites Where We Are and What We're Looking At: Query Based Worldwide Image Geo-localization Using Hierarchies and Scenes.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Where We Are and What We're Looking At: Query Based Worldwide Image Geo-localization Using Hierarchies and Scenes

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:52:10.616398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:08.764186Z digest=sha256:38ac48d0c4e7786b7fc696ac9165f82de01bf2331f3db0716c0cc91f67bb86f2

Observation eb76507a-c259-486b-a3ce-9e3b118146cc · outbound

This paper cites The Mapil- lary Vistas dataset for semantic understanding of street scenes.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models The Mapil- lary Vistas dataset for semantic understanding of street scenes

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:13.100674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:08.801352Z digest=sha256:6763c47b8f92e7961eb83148a2e987e26e3abbba8a9fba490f1cf3d5dacfe4d2

Observation 25452abf-1357-45ad-88ff-b01d1e19727f · outbound

This paper cites GeoNames geographical database.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models GeoNames geographical database

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.947008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:08.826571Z digest=sha256:bdc2483fe4f5b77b6fe2e7de7e3a04626b7ceb861ffe04799fed49ec64c41ab5

Observation 5169bda4-4f7d-42c3-a9a2-4dcc23558192 · outbound

This paper cites Qwen2.5-VL Technical Report.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Qwen2.5-VL Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:09.047711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.047711Z digest=sha256:f90605f8620ff3720b3e026a81fe7d34eb7906562cce4acdf0497b288d26130f

Observation 54ddb3a1-3421-4792-a162-fd35c6d2a87d · outbound

This paper cites Decoupled Weight Decay Regularization.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Decoupled Weight Decay Regularization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:09.081528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.081528Z digest=sha256:1ed5885805557c763f3a966ccefabe3c17cd49126219b14ba464e3220df625f2

Observation 9ad4f2a6-8671-4596-87c9-e9dc77b2b336 · outbound

This paper cites OpenStreetView-5M: The Many Roads to Global Visual Geolocation.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models OpenStreetView-5M: The Many Roads to Global Visual Geolocation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:09.120531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.120531Z digest=sha256:ad5dd8530a58fbfd674cd25cee605916233573b5da5e350ec598cb23af9c4a11

Observation 193b1e22-a3e6-49ce-8702-3dc16a604b34 · outbound

This paper cites The Claude 3 model family: Opus, Sonnet, Haiku.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models The Claude 3 model family: Opus, Sonnet, Haiku

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.868679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:09.193032Z digest=sha256:b6c9d28ac835cd3cde8387512991507e76e1c406b8d762f48bdf5996628beadb

Observation 7631d502-cc44-4fdb-9cae-f977578e16e5 · outbound

This paper cites Visual Instruction Tuning.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Visual Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:09.261284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.261284Z digest=sha256:9883be4dac274cf1fc44abb6721bff9e5064f66e5c082f94c28105e2c06d8f56

Observation b8c927a4-7674-4873-acef-447b02dfb975 · outbound

This paper cites Mistral-Small-3.1-24B-Instruct-2503.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Mistral-Small-3.1-24B-Instruct-2503

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.768597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:09.302998Z digest=sha256:c2a83370f8fe82c4bb82fb7d48c88d05676892ecc0aa3f181b96c8dbe458fd1a

Observation 00ff3373-43a2-4640-9a60-16d45423daa3 · outbound

This paper cites Revisiting IM2GPS in the Deep Learning Era.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Revisiting IM2GPS in the Deep Learning Era

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:09.371329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.371329Z digest=sha256:4dc387f09a460a8f03b827c7b242865b83306affde347d1193f401b088b1fe88

Observation fe163f24-9bf9-4a1e-9068-23c0b9512cd1 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:09.515725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.515725Z digest=sha256:ef744b392130d066c80737688d120673c1f2c146db8214193efe7c76eeae421f

Observation 4c703f5e-d585-4cf8-9785-e4d83a4ec20f · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 24

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:52:09.555930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.555930Z digest=sha256:690a66ea23cc5cf83bff62e6eda58e99b3fbf01ba5207a97f574c086504bc4d5

Observation e70fda6e-38f2-403b-a5b8-43f8a2764ae7 · outbound

This paper cites Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:52:10.415712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:09.473416Z digest=sha256:d5bafcac0b72cbdbaae5d18be51c18f3c87b4130564c0e894e8618c04bf78610

Observation 7dbbf803-2ed1-4664-ac0f-00481042e4bb · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.675504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:09.594635Z digest=sha256:b103398e142d1b8ef3bd87f3a394cc872b24e01c632b63176f1b0d9d2a0ac727

Observation 947e2810-6378-40e2-b3cf-419d7b1409a3 · outbound

This paper cites Guidelines: • The answer NA means that the paper has no limitation while the answer No means that the paper has limitations, but those are not discussed in the paper.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper has no limitation while the answer No means that the paper has limitations, but those are not discussed in the paper

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.601524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:09.627183Z digest=sha256:c4cf668f35625f4bdf92e3616592a69b2807ff01d4badcf30f04cdc89c9780ff

Observation 9bcd7b9a-f880-435d-aec0-4ad3a7a11462 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.486517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:09.687952Z digest=sha256:d9218cceab08b03790fa61bd543ca426650ed16644013c53e4022866483ad508

Observation 62811774-990e-4f8c-9518-4fa0e1845b78 · outbound

This paper cites The appendices include hyperparameters for SFT and information about computational resources used.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models The appendices include hyperparameters for SFT and information about computational resources used

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.393269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:09.751346Z digest=sha256:edfe7201c69ce85d976f289290b05d88fc93bf37dcb276d8ec3216fc5767dd11

Observation 7dc63ab1-963b-4bcd-b217-e92994f77b42 · outbound

This paper cites Our appendices provide detailed instructions regarding implementation, hyperparameters, and experimental setup to facilitate reproduction.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Our appendices provide detailed instructions regarding implementation, hyperparameters, and experimental setup to facilitate reproduction

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.301578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:09.792357Z digest=sha256:d116eab05d1eef66eb9c7d8321dabe3e565423d8179a8514b37efeada9ca09f5

Observation 1dba0e3d-4065-44e2-a2ff-01d772224858 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper does not include experiments

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.220821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:09.842854Z digest=sha256:44b5add7da024e4a175ba8987abd66fa0b89b7382abb318f4e362b2ad0e4c327

Observation 5cc46a9c-3c83-45bd-a47d-a8079eb9fd12 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper does not include experiments

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.110456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:09.878555Z digest=sha256:70756da75b723ccd93e35d4e0b0d08e6329334e99d7c79c28ac9579ac19fd6a9

Observation 180e1af7-572b-4d03-8655-1beb3d7b8761 · outbound

This paper cites We also report model sizes and memory requirements.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models We also report model sizes and memory requirements

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.988746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:09.888164Z digest=sha256:ca517e27e943c891e8f42fd1e12ecd9a4ffc60c641fa09867a45cd7d3e09133f

Observation 13c45ece-fa9c-4e7a-836b-b714c48619b8 · outbound

This paper cites We use publicly available datasets, acknowledge relevant prior work, and are transparent about our methodologies.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models We use publicly available datasets, acknowledge relevant prior work, and are transparent about our methodologies

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.886563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:09.925377Z digest=sha256:1271452b7c14084c22e5f13e8cf3e5da834581ae47c056ae904fb20c83361bb8

Observation 54276147-772a-4883-9c90-95322fe4ce64 · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.755852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:09.978969Z digest=sha256:845dbd5d737589cb417b2854ea6bb9725a4f7e0bb9a942dc516b34b66c425a7b

Observation bbf48e5d-e190-40dd-9d6b-1ed9331aa26a · outbound

This paper cites 27 Guidelines: • The answer NA means that the paper poses no such risks.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models 27 Guidelines: • The answer NA means that the paper poses no such risks

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.621881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:10.027853Z digest=sha256:be9a03a871d972dc11a4fa39d97bf5d2ff379a1d148881e6c2d6f8b03ca7a448

Observation 3ef73617-d963-4be5-98a5-33ed5ad543e8 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper does not use existing assets

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.507759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:10.068261Z digest=sha256:386ceb0b40b53131f1bcd027eec386c4513b24681151282a8fb8ddac40489209

Observation 8125e676-e808-4062-8bd4-3cebc78c0535 · outbound

This paper cites This documentation will be released alongside the dataset.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models This documentation will be released alongside the dataset

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.310523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:10.112687Z digest=sha256:2bc1dfece9cff0863e0efaee8e935d213c18f6669f077db43fa8c4ff209fbe79

Observation 04d645b3-ad62-46d9-a66e-d21e141be845 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.158826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:10.143213Z digest=sha256:b1bfcc7e44c9053f914980678e806fa856bfe3ecacddd33f75e36fe87816afa5

Observation c14f16c0-6101-481f-82fd-6053b2217ae5 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:10.997861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:10.218789Z digest=sha256:2925b04af670b825e0ef5db9bb0a236180b21b1d9aafcaad07969e56d521ff4a

Observation 2183609d-e1a8-486a-821c-40d15e3a9070 · outbound

This paper cites G3: An Effective and Adaptive Framework for Worldwide Geolocalization Using Large Multi-Modality Models.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models G3: An Effective and Adaptive Framework for Worldwide Geolocalization Using Large Multi-Modality Models

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:52:10.788245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:52:08.501464Z digest=sha256:7aa8f32ab4ae4a010187281671c154d80759e5b4b3072badd5734b55325415e7

Observation b42b2fe3-b71a-4bb4-a522-2881740e889e · outbound

This paper cites Gemma 3 Technical Report.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Gemma 3 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.967423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.967423Z digest=sha256:fbdaa03b48054ce2c5b30462a2f509a6eff2b46a1da7be419942df26b730d1ce

Pith citing papers

Observation 5da9a71f-b019-499b-aeee-179f85ab3b52 · inbound

A Global-Local Cross-Attention Network for Ultra-high Resolution Remote Sensing Image Semantic Segmentation cites this paper.

A Global-Local Cross-Attention Network for Ultra-high Resolution Remote Sensing Image Semantic Segmentation GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:10.857306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:10.857306Z digest=sha256:40d8279fd01e7c5c8d0df7080994485fb9758230380efc8af359ebf06feb2141

Observation c1b68c6f-60f2-4d40-b263-be1b4624e2e8 · inbound

Building Lightweight Semantic Segmentation Models for Aerial Images Using Dual Relation Distillation cites this paper.

Building Lightweight Semantic Segmentation Models for Aerial Images Using Dual Relation Distillation GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:29.112400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:00:29.112400Z digest=sha256:e91facf5641de61b12eeefc68e2ce3b2ca8a0bba73136f839854a43fa2b36a04

Observation 0a34a1ef-010f-4b43-8de7-cf7d9ee70961 · inbound

GDGS: 3D Gaussian Splatting Via Geometry-Guided Initialization And Dynamic Density Control cites this paper.

GDGS: 3D Gaussian Splatting Via Geometry-Guided Initialization And Dynamic Density Control GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:24:31.913301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:24:31.913301Z digest=sha256:0835e6f5e6571e7fe9ad766f225706542bd3cc9f23df81a9c72598c366510d92

Observation 4aab417b-4cd4-4650-9ded-b40c41c22580 · inbound

From Pixels to Places: A Systematic Benchmark for Evaluating Image Geolocalization Ability in Large Language Models cites this paper.

From Pixels to Places: A Systematic Benchmark for Evaluating Image Geolocalization Ability in Large Language Models GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:50:47.271597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T23:49:59.795575Z digest=sha256:d4b5e2ac4543942081db7df71ff6b992a24a1bb95a3eae8549471509e75d7ea5

Observation 556d334e-529f-4853-a0ff-ca369f038b59 · inbound

Do VLMs See What Sensors Feel? A Scalable Expert-Guided Design for Wheelchair Accessibility Assessment from Street View cites this paper.

Do VLMs See What Sensors Feel? A Scalable Expert-Guided Design for Wheelchair Accessibility Assessment from Street View GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:20.241225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T14:42:47.815192Z digest=sha256:65f20f3c3c48594cce6259566eedea9ba2a076208ae4ea6393dc064c5e22d56d

Observation 33241b1a-4fed-42af-80bf-13ec998b5cfb · inbound

DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization cites this paper.

DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-30T22:45:19.083071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T22:45:19.083071Z digest=sha256:15d7b25758fcb53bce4ae95b848f846d388d64455ae22e5420db64065b19c4ea