Pith. sign in

Paper Citation Record · LEDGER

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models

As of 17 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2607.08665.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08665 v2

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T06:32:41.000897Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 576f57c0-9821-48f7-9182-1a6230794c0c · outbound

This paper cites Measuring and improving the energy efficiency of large language models inference,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models Measuring and improving the energy efficiency of large language models inference,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:713ebcb7c64e48d90d2632e3b20fbf05c3520fe48a2436e0ca4533037c7e5ba5

Observation 7b4cf943-e29b-4372-a73a-361faee2469d · outbound

This paper cites Ma- chine learning (ML)-centric resource management in cloud computing: A review and future directions,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models Ma- chine learning (ML)-centric resource management in cloud computing: A review and future directions,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:dd37437ab8ab5ce3960e7b7bcdc3d2ec9791f2cac0e704f35a80b1d13b6ce00a

Observation eee247f3-449d-4aea-b3b2-a16db847b78a · outbound

This paper cites Survey of different large language model architectures: Trends, benchmarks, and challenges,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models Survey of different large language model architectures: Trends, benchmarks, and challenges,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:a86bad11b07ac2a202d5795eca77abcf9ed381f03e313f74342340d19b37c45c

Observation 921fc337-0a3f-48fc-b0f1-7c47315d52e8 · outbound

This paper cites RouterBench: A Benchmark for Multi-LLM Routing System.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models RouterBench: A Benchmark for Multi-LLM Routing System

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:1bb6afce7ecb17c3ac75cf8d04d76d27c7fedceb757a54f0fed0f683e5d9c1fb

Observation e8eb072b-c19a-4218-abdf-cad0a96810ea · outbound

This paper cites LLMRouterBench: A massive benchmark and unified framework for LLM routing,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models LLMRouterBench: A massive benchmark and unified framework for LLM routing,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:53c9556d3582786516cb42f3f1e92444f2825d2444c2dd4c8e7a258b522e2f5d

Observation 807b6aa1-de10-4888-9b88-9d842a148a1e · outbound

This paper cites FrugalGPT: How to use large language models while reducing cost and improving performance,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models FrugalGPT: How to use large language models while reducing cost and improving performance,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:af1b0c67abe9b44848881f2868297d36615e7db6f2bba27220bf9d1509cbd0f7

Observation a1d0569f-f4f0-4d59-ae18-5a323a0dd236 · outbound

This paper cites RouteLLM: Learning to route LLMs from preference data,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models RouteLLM: Learning to route LLMs from preference data,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:fe4531635b5520d621ff287f4be719e029b8635e10af47669b38005992fff5da

Observation 70536be9-9386-4a3f-b144-6f56d5847066 · outbound

This paper cites How Much of the Routing Gap Is Real? Decomposing the Router-to-Oracle Gap into Reproducible Specialist Advantage and Single-Draw Label Noise.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models How Much of the Routing Gap Is Real? Decomposing the Router-to-Oracle Gap into Reproducible Specialist Advantage and Single-Draw Label Noise

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:462be001a7f31577679fd7a9f4dc1d41977b6c9da5bc20a5e48d94654c5bb468

Observation aef4235d-5949-4466-9ed0-a708776dc4c6 · outbound

This paper cites Enhancing LLM reasoning capabilities through brokered multi-expert reflection,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models Enhancing LLM reasoning capabilities through brokered multi-expert reflection,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:c8849ad888cdd9787c13d68424181231381dbdd7ae2db1bba68055fd31b0f922

Observation a9e1a039-b477-4eef-881f-4c6f69cda090 · outbound

This paper cites Classifier cascades and trees for minimizing feature evaluation cost,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models Classifier cascades and trees for minimizing feature evaluation cost,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:eb69179f8a18eead8f21b5fefcc5bac444895b45f28698db6885f5485d3112e0

Observation 67ccf748-9c97-4f2a-b0f8-d7e562aeb212 · outbound

This paper cites Classification with a reject option using a hinge loss,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models Classification with a reject option using a hinge loss,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:e730c5776fd31274ce38a6875b4ccc34d1c436a8bb3e65e812c257814c4deebc

Observation 483ceade-1a1a-4805-a8d9-905c4de52069 · outbound

This paper cites Machine learning with a reject option: A survey,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models Machine learning with a reject option: A survey,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:5efe8e17301a438d1cebdaf778f20cce0937987af97cfd6cde0d61a2944b7651

Observation f85aeb4b-d9ea-422f-817d-f25be22df0b6 · outbound

This paper cites QoS-aware web service rec- ommendation by collaborative filtering,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models QoS-aware web service rec- ommendation by collaborative filtering,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:befd919c71176c71114acf9265f10180505093f315516abc9049ca513646e090

Observation 8433d7b1-4ff8-48ab-8beb-cb53e0b05219 · outbound

This paper cites On the impact of deep neural network calibration on adaptive edge offloading for image classification,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models On the impact of deep neural network calibration on adaptive edge offloading for image classification,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:8ea404c376af54f94387aed6cb83deb7f9dba899267761e9fe6ae7476c33dcf3

Observation 955714da-735f-40a5-b2dd-be7200ac859d · outbound

This paper cites Deep neural networks meet computation offloading in mobile edge networks: Applications, taxonomy, and open issues,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models Deep neural networks meet computation offloading in mobile edge networks: Applications, taxonomy, and open issues,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:3509cc819e25f0a83b007c1ab9d3d6ebefc366d58fa3442de0bbf70316353d71

Observation ed124d94-19a9-4a78-9b39-c0b7915bc23f · outbound

This paper cites Edge-AI: A systematic review on architectures, applications, and challenges,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models Edge-AI: A systematic review on architectures, applications, and challenges,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:a917f7627754748e34fb168ff79993e0281a6821bafe280dd3abb2a4ca6f9715

Observation 66b65895-def7-4dcd-8bf9-fbf77435f3be · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:fbc256cb955b10c4c426f555fe401af8471cdee0ed2bab9bba4cd40f30b672e1

Observation e7701a12-a62e-4a97-9832-3057cded2032 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:3bbd2c4caa0f5a24ff41d3950f247373fb82743841afebfe15ed2d1fdd61ce42

Observation 1e919b69-159d-42d3-bca3-4750e8b2c1af · outbound

This paper cites Ensemble-based uncertainty quantification for reliable large language model classification in social data applications,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models Ensemble-based uncertainty quantification for reliable large language model classification in social data applications,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:53df29fe72176f98b588ba5bd8e7c366a6ee9950a0a1207d5c34fedb95b7fa4d

Observation 11c8c7dd-b43e-4970-9362-f215dfb2c6e7 · outbound

This paper cites RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:05f3e3b12aa73e1398521ac08d89c30e878deab8280a8b98dab80fc2c9ba7b3a

Observation 0bbafed3-edad-46fb-b8a1-f6604e1f8d34 · outbound

This paper cites Beyond the leaderboard: A survey of the science of evaluation, benchmarking, and methodologies for large language models,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models Beyond the leaderboard: A survey of the science of evaluation, benchmarking, and methodologies for large language models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:43c0fe0686ebff47c685cd623ac14f3ccc0a20b9a0fa664530e9732291057bd4

Observation 2a2a6509-aaa9-4b9d-b67a-69398fdaeabb · outbound

This paper cites State of what art? A call for multi-prompt LLM eval- uation,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models State of what art? A call for multi-prompt LLM eval- uation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:81719e7ea781fcd0c6cc3d8807a3cc37db1bd0f56a1975feb3d09de0f4101915

Observation ba11c932-e4f8-41ee-bdc0-a2167b52e9cf · outbound

This paper cites Make large language models efficient: A review,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models Make large language models efficient: A review,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:93b3a2250ac9c9ac3e5500ef47789444988cda877ab38d6237e3ab1f20ad11b1

Observation 86a4003f-2ff6-4d73-b246-0d3dd21d3327 · outbound

This paper cites BrownoutServe: SLO-aware inference serving under bursty workloads for MoE-based LLMs,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models BrownoutServe: SLO-aware inference serving under bursty workloads for MoE-based LLMs,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:dce5b9abb965b162267abbaf7c994fde31621e87dce571073f2bc065f32cc1ca

Observation 5a426239-8a7a-4a53-8c3a-1995a6bdfb66 · outbound

This paper cites iGniter: Interference-aware GPU resource provisioning for predictable DNN inference in the cloud,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models iGniter: Interference-aware GPU resource provisioning for predictable DNN inference in the cloud,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:7268cba7a5df5d50b6a623e470dba963f24233745b86cb57ff35e058c13b6d3e

Observation effe0497-df2f-4ab8-9e8f-20631ace38d0 · outbound

This paper cites Best arm identification in multi-armed bandits,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models Best arm identification in multi-armed bandits,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:ea8d04ed0e0e61d1cf912ed8f91cd1ab6877fabcf122344fb8fdd292660b707c

Observation 9a8cadeb-b0e1-4799-90f8-f82d8747f15f · outbound

This paper cites On the complexity of best- arm identification in multi-armed bandit models,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models On the complexity of best- arm identification in multi-armed bandit models,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:beba9067ffb5e225a7bf3e3f288cef7520195febd0502d4a9554a2b2fe807859

Observation 5086d72a-818b-4525-a4b0-f6b192f0030a · outbound

This paper cites How can we know when language models know? On the calibration of language models for question answering,.

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models How can we know when language models know? On the calibration of language models for question answering,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T06:32:41.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:32:41.000897Z digest=sha256:80819c212ab1c8728a5dbc4f0bdf4bff1d095cc58224f28c75d5618881f163d6

Pith citing papers

No inbound Pith citation observations are available.