Pith. sign in

Paper Citation Record · LEDGER

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time

As of 11 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2509.02129.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.02129 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:57:17.255005Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf1d124e-84b7-41ae-8bdc-4010ad07f9d6 · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:57:17.877913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:57:16.316998Z digest=sha256:0abd1e5b4a8a848a88769e93897875d2de3a28560450017e6aa058a038a2d185

Observation f5110ec7-93fd-4714-b01c-f7e3064378ef · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:57:17.864379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:57:16.355902Z digest=sha256:96203d7bca46b03138511a0750ad8dd81a7e1ade7926ba82460207421ae5459d

Observation fd67a794-0360-486f-a8e0-5680d2dda27f · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:57:17.850435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:57:16.455330Z digest=sha256:9ccd59f69acbe6cb5012a6d915d99877377aa45e7b9eeff7e13c817c708d6571

Observation 5990be6f-6b82-403f-a1be-9c77f2b4f35c · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:57:17.836818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:57:16.519091Z digest=sha256:2880eefcbbcd72b6dbce32296e15923cca97b2d22bec5115ebe23670216173ab

Observation e029c155-342d-4b6c-a22b-0cc2dc634b09 · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T11:57:16.596950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:57:16.596950Z digest=sha256:f16b9c9ac997f744a0fee2ca55bb1c1567b447612564738eab0540656a0d9719

Observation 87999208-3bcf-46a6-90d6-3171d90122ce · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:57:17.822687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:57:16.683160Z digest=sha256:7dab5e28283b7045857af7ba082a9a69c53c1148d14a88022c9836a77a307b9b

Observation 14285457-87e1-4d05-a671-243df4b6e228 · outbound

This paper cites Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T11:57:16.716582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:57:16.716582Z digest=sha256:0a53ed7c94832044c10abc0c693309002d5b300dabfc062e1079a2071b984893

Observation 0cff251c-719f-4a14-9b30-ce93e8ae5b13 · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:57:17.808279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:57:16.800663Z digest=sha256:b836415480a3ce6215c3e6cdcbc7f12e640e67e0c463074eaa8f326d6e86ea7f

Observation bcf3f1b3-a73c-42a9-be4c-295b3440c589 · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:57:17.794198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:57:16.955019Z digest=sha256:7ea1964eb7b325f992d34eeeecbef373525610371aba9e9269db290c7e407c03

Observation 2b2e8ff4-0239-490e-9170-fb08871dc77f · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:57:17.779771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:57:17.116224Z digest=sha256:987f9e144424111ef02b86e6c94bf7c2ce8372658441591ecc321a8a34a14b07

Observation a17a134d-2444-482f-89be-826632a77ee2 · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:57:17.764029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:57:17.120543Z digest=sha256:d9b2bf41e76901a28fc5871e47a8f633b1f83f17db4feb2912769ae33d3da5a4

Observation 6b92c121-f72f-463a-919a-55f1ef69da0f · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T11:57:17.185085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:57:17.185085Z digest=sha256:60566a4974208c2be41c2fb7238276ccec247b7a7e7c14b9934f023988e6ff72

Observation 30b25eca-efca-4649-a6c8-d395812f596d · outbound

This paper cites Tell Me Where You Are: Multimodal LLMs Meet Place Recognition.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Tell Me Where You Are: Multimodal LLMs Meet Place Recognition

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-05T11:57:17.435287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:57:17.189338Z digest=sha256:2b422d38404730169ce746386e80639eb15d214d26dd69851f2312427b6e40ae

Observation e23b4d26-4498-4661-bff8-d1c5facd7bdb · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T11:57:17.194612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:57:17.194612Z digest=sha256:0b323cf0352e366695a2b1e7d7e7e03dccd0a4546daf440ccc190fe658a7b78e

Observation c266c17b-4bbe-4c14-a8b9-b450801b6029 · outbound

This paper cites TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T11:57:17.198660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:57:17.198660Z digest=sha256:dbf9a48cd49338440195d084216e401c7f2e9daa442d24a6c9682ed9a7cfc10d

Observation 509ee550-866c-4896-8ebe-960f1fe8c9d4 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T11:57:17.203309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:57:17.203309Z digest=sha256:2030277ad4b2cec70be8ad0417fb50eaa1d24bc900d48ba6e37144db70b3d513

Observation 04a8acd0-7b26-4efa-854d-20f1bf1a13b3 · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:57:17.749532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:57:17.208210Z digest=sha256:a77bef1a06a37abe6e4ccefbf61e11147de7471a4a2505db9a4c49c861c06c7a

Observation f9a92daa-1f23-4d7f-ac87-7567485d0efa · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:57:17.734489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:57:17.212218Z digest=sha256:f260e9d4187e7c97c3c13dfc3f15a0591fcfa53f59a85c54e74973ddf38128f4

Observation 65273ceb-6a4d-4eb7-ab27-4ff2ade810b1 · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:57:17.720097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:57:17.216064Z digest=sha256:14df69436efae90fd2cabed3cf468b172549e6df85f12e84aa54a0b3b80a916a

Observation 324a6c31-86b4-420a-888f-0b02322b2b8e · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:57:17.705310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:57:17.220188Z digest=sha256:da4f984dfd74747f0b676bd81effb1b165e1b8ee519dbdb795d3ec482a29596f

Observation 1d75041c-d04c-407a-a219-1a13e8945aee · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T11:57:17.224316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:57:17.224316Z digest=sha256:0120c675534f808452723576b97a0edce8866f5bd76f9145637d53bbea70e108

Observation 60d893ea-1597-4763-9791-91d1b4f62f40 · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:57:17.680668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:57:17.228965Z digest=sha256:b2c2b8162356aa00c093ba5d17ec11206da8fceda04dcd67ec380c0d9c80ca5a

Observation 40a982e5-ef36-4df9-9b57-8794f38db75d · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:57:17.666062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:57:17.232985Z digest=sha256:2ea0132ce541feb3e580a963ed55c367aec578d1b921414269afadec78cf2739

Observation 420fffb0-6471-4d2c-9441-8ffda85c88f6 · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:57:17.651568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:57:17.237103Z digest=sha256:caf6826f5ea318456970911265119d006c861e354d983510be54415ed5888f50

Observation 1e339af1-e62b-49c2-8f8b-6aa90c937897 · outbound

This paper cites an unresolved cited work.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:57:17.636106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:57:17.241091Z digest=sha256:d636c59c58b0f7d03e9a29a2847c92ebdbbec0fda22db4843a36f22a0c85519b

Observation ee72edad-a407-4f03-a9d5-32a6ed516ff4 · outbound

This paper cites NAVIG: Natural Language-guided Analysis with Vision Language Models for Image Geo-localization.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time NAVIG: Natural Language-guided Analysis with Vision Language Models for Image Geo-localization

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-05T11:57:17.299591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:57:17.245721Z digest=sha256:fea09a0c1b1bcf5347e2251a36dd8fd3600d1fcddfe10e36160de757331b6e32

Observation 3f7f82c3-67a0-4d0c-8d48-1ca3b9fa2d63 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time , " * write output.state after.block = add.period write newline

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T11:57:17.250154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:57:17.250154Z digest=sha256:f3abd58deeef4475b5b1b4f3b1085a8799168f2b88510a76f45a5d7962212a9d

Observation d131aa0b-05ae-48fe-a02e-1a6a066da91a · outbound

This paper cites write newline.

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time write newline

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T11:57:17.255005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:57:17.255005Z digest=sha256:c65e05340a9c629225fccd8a7eeab067c8bc6d2bbb2e063155ea2b2f15e9b6eb

Pith citing papers

No inbound Pith citation observations are available.