Pith. sign in

Paper Citation Record · LEDGER

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World

As of 15 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2608.08239.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08239 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:18:54.165282Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact3
  • verified fuzzy1
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d36c36fa-dbd6-42eb-8e77-ef705d33ba8e · outbound

This paper cites Switchcraft: AI Model Router for Agentic Tool Calling.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Switchcraft: AI Model Router for Agentic Tool Calling

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:18:54.674282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:18:54.095105Z digest=sha256:e2f90ea7e8c07246746da602eeca2664c98980122bbae5b780557ab7d6a97693

Observation e5690e71-0c1a-4b56-83af-97dd95198134 · outbound

This paper cites RouterBench: A Benchmark for Multi-LLM Routing System.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World RouterBench: A Benchmark for Multi-LLM Routing System

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.111313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.111313Z digest=sha256:24942d95abd24233e6c4eec11044e71ab1bc1774994c4543b2f2de31481b616c

Observation acd73c12-76f2-472b-96cb-804c6265671b · outbound

This paper cites RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.116775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.116775Z digest=sha256:406933d24ae5bbf7b670175644a30c834a4e99d0702861876ce6758e454e4304

Observation e8f3a029-e178-4d69-88eb-09b24bd3b99e · outbound

This paper cites Universal Model Routing for Efficient LLM Inference.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Universal Model Routing for Efficient LLM Inference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.121405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.121405Z digest=sha256:da430e1c730fa23c0a4d605341bdfb244921fc7910bfd37bd50566f16f0d9414

Observation 90abd128-3dec-4217-9494-d7aa174c556d · outbound

This paper cites Routerarena: An open platform for comprehensive comparison of llm routers.arXiv preprint arXiv:2510.00202,.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Routerarena: An open platform for comprehensive comparison of llm routers.arXiv preprint arXiv:2510.00202,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.135081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.135081Z digest=sha256:093df7751e133abb4cc592b74a721948b88654b6fa8c204e96983ec51288cab7

Observation 9c2ced0c-2fde-4a61-b8ec-754313d4dd25 · outbound

This paper cites Odar: Principled adaptive routing for llm reasoning via active inference.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Odar: Principled adaptive routing for llm reasoning via active inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.139092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.139092Z digest=sha256:5ad1cf75bef5c801931a23c0cb92faa932dec02263249666fa4dffdc5587a9fe

Observation b5cfd66b-ea8c-41a7-9d32-a05164ad4746 · outbound

This paper cites RouteLLM: Learning to Route LLMs with Preference Data.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World RouteLLM: Learning to Route LLMs with Preference Data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.143606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.143606Z digest=sha256:99478c576ad7e16246fa85cb89837a9bb71fe7f9ef2d88a7a5fb0da7c63c649e

Observation 5b8145ba-adce-4712-9048-a90929237e80 · outbound

This paper cites Route to Reason: Adaptive Routing for LLM and Reasoning Strategy Selection.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Route to Reason: Adaptive Routing for LLM and Reasoning Strategy Selection

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.149064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.149064Z digest=sha256:b368c8ca81e0a64fa156b1d7e2762a8142d85b71636a5b3fc4d6ba87f5efcf8e

Observation 13092041-523f-4119-9c43-b263d655017b · outbound

This paper cites Qwen3 Technical Report.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Qwen3 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.159441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.159441Z digest=sha256:52e495771339c0c40e5bc231e983e86801e8538f0b978730f8425910300bd948

Observation 80955fb7-5ad3-486f-902c-838badcf76d9 · outbound

This paper cites TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.165282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.165282Z digest=sha256:7b89ea144acc64669bfea082bb3b93149b1c3bc87fbc3800796ffee98fb0d8da

Observation 025e749e-fa89-470a-84e5-2907cecea79d · outbound

This paper cites Step-level Optimization for Efficient Computer-use Agents.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Step-level Optimization for Efficient Computer-use Agents

Reference 2016

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:18:54.240259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:18:54.154484Z digest=sha256:178db6f4f7286ee620eac840e6904092742a685f8358413eece3e53be0918c80

Observation 042ff5c3-d06e-4ba3-928b-45ff557256b5 · outbound

This paper cites RouteJudge: An Open Platform for Reproducible and Preference-Aware LLM Routing.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World RouteJudge: An Open Platform for Reproducible and Preference-Aware LLM Routing

Reference 2023

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:18:54.573778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:18:54.126620Z digest=sha256:231afaa3485e7e5c974ae75a2c16061a5e2b0c97b5108e307391d5faf092a4d5

Observation bd504762-ced5-400c-9089-1abeb2024d29 · outbound

This paper cites FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.106043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.106043Z digest=sha256:1a9b74ddd1310175e7eb25111c6f29d8d2df2b8506e5a6f8e62a5d75c604356e

Observation c823f036-d0cf-4726-b02c-791109e2997b · outbound

This paper cites Session-aware agentic routing: Continuity- aware model selection for long-horizon llm agents.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World Session-aware agentic routing: Continuity- aware model selection for long-horizon llm agents

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:18:54.690524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:18:54.131034Z digest=sha256:d065e2e7b6757cd2349b1e25a56f98eeae7d6fc36bfed11bd151df43925ad59d

Observation 4c00d2ec-c5f4-474b-b92c-311bbc0ed06e · outbound

This paper cites AutoMix: Automatically Mixing Language Models.

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World AutoMix: Automatically Mixing Language Models

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-12T00:18:54.100605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:18:54.100605Z digest=sha256:7db78523dae3737aa784b672989389c49cf1ef06d8bb0ddf6b10ef53abd5e729

Pith citing papers

No inbound Pith citation observations are available.