Pith. sign in

Paper Citation Record · LEDGER

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing

As of 10 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 10 inbound Pith citation observations for arXiv:2502.02743.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02743 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:22:29.015064Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:10:53.675879Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T02:25:19.685963Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact4
  • verified fuzzy5
  • unresolved31
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 930e4926-5eb3-4d39-a6a4-8b94aac96986 · outbound

This paper cites Contextualize Me -- The Case for Context in Reinforcement Learning.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Contextualize Me -- The Case for Context in Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.831842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.831842Z digest=sha256:2a07a5c14465e4196cd0db5cf3ba7defb35f2a7db7ddcaebd6c0648770695ba1

Observation ab1e82c0-2cf7-45cf-9f26-609349832f23 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.841954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.841954Z digest=sha256:811bc120d747f7a1fdbad16209312ff3590715a01c7416355caaf8d9f7932a78

Observation f08221ad-75b5-4b95-8eb3-3c5aae85adb7 · outbound

This paper cites A Brief Review of Hypernetworks in Deep Learning.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing A Brief Review of Hypernetworks in Deep Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.847310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.847310Z digest=sha256:4ee2cac76e96c6cdc1d1c78fcb4b3810b697af482437b4440929b5e12a5cef98

Observation 1e5c45ba-38d1-4767-8ae1-eba029549a1a · outbound

This paper cites FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.852947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.852947Z digest=sha256:140ebc066d7dac061d3fe6f338d3e3885832dcfdb6859477118cce86b9b8e27a

Observation 68fd614a-c29e-46c9-92dd-bf214f6c1da0 · outbound

This paper cites Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.858028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.858028Z digest=sha256:4e41c9077a125194854a6155819663ce6cc6e8486ec875cd103c3bb4f44a0998

Observation 8979fb9c-e9a3-44e2-8f3d-7f39bafda339 · outbound

This paper cites Zero-Shot Reinforcement Learning via Function Encoders.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Zero-Shot Reinforcement Learning via Function Encoders

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.878476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.878476Z digest=sha256:d12df38cea50873588859af1afe8ebd29c1e334b770e0d24f724b124af756e81

Observation c8758b8f-16a6-4368-9ef1-576f4dbd4189 · outbound

This paper cites Generalization to New Actions in Reinforcement Learning.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Generalization to New Actions in Reinforcement Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-09T11:22:29.388179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:22:28.883218Z digest=sha256:4a05470b08c9baf8984881a2f795e42e2a4e6e3d68d86de580eb84193082be82

Observation 8caa511e-12f4-4ef1-917a-a22af59ca594 · outbound

This paper cites LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.893446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.893446Z digest=sha256:304a5c33dfdabb91f12ec1786aaefc2bb4f8e6fbecd60d53b74a0d1e534c2906

Observation 8ff53baf-f196-48b2-a4fc-3e5b8ca63220 · outbound

This paper cites A Survey Analyzing Generalization in Deep Reinforcement Learning.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing A Survey Analyzing Generalization in Deep Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.903294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.903294Z digest=sha256:c42407da09eacd1ed9b8318ef83ec73af127a52f80deb905be68738e9be9e182

Observation e845115f-68b9-46d6-903c-d5b9f72ebc25 · outbound

This paper cites Network Randomization: A Simple Technique for Generalization in Deep Reinforcement Learning.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Network Randomization: A Simple Technique for Generalization in Deep Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.908205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.908205Z digest=sha256:3b843214d0db5344c667985d4681bc8091776ccb69971227d2de8c8e7adf0e31

Observation e6887e6a-0877-46d1-8f8a-f439804b5877 · outbound

This paper cites Holistic Evaluation of Language Models.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Holistic Evaluation of Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.913059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.913059Z digest=sha256:a3cc0da610ea4023b4eaaf5b106317af06259edb1438c691f8c2d4ba3b5cf1ce

Observation 7118ebd6-14d2-4dce-838f-52385c04b611 · outbound

This paper cites Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.917976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.917976Z digest=sha256:9b3df607ca7e1bc5ed79e668bed0a93c5a467ff1d8423089ec39f30e87a651f4

Observation 861a997a-073b-43a2-94a6-107bbd0f4d3e · outbound

This paper cites Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.922626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.922626Z digest=sha256:938d82b6e9c4e157f58663851c1a53d5388a8b5c53e4313a47b506a30623b1d0

Observation 56a7bc7c-ac5f-41a9-9305-3a4ae2b2ba86 · outbound

This paper cites AutoMix: Automatically Mixing Language Models.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing AutoMix: Automatically Mixing Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.927497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.927497Z digest=sha256:212f08c4a8543329d7befb8ec46f4449bb392af752153fcee4d6c7a57ef0b90f

Observation 6b368b8f-db85-42d1-b243-d9cd52195ce3 · outbound

This paper cites MetaLLM: A High-performant and Cost-efficient Dynamic Framework for Wrapping LLMs.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing MetaLLM: A High-performant and Cost-efficient Dynamic Framework for Wrapping LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.932266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.932266Z digest=sha256:df0d7182d8cf7514449eacf1cf9079767483a353eeacdada91b802b5affc80e5

Observation 4bc06a7a-4b64-47ae-956f-70ef1f0c2b6e · outbound

This paper cites RouteLLM: Learning to Route LLMs with Preference Data.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing RouteLLM: Learning to Route LLMs with Preference Data

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.937025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.937025Z digest=sha256:0190b96af52dfcc33f88f68ba69b2787fa8a9869a01df79ffe825163d54bf599

Observation 781a6906-2980-41b7-abcd-10f70612d8fc · outbound

This paper cites Policy gradient approaches for multi- objective sequential decision making.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Policy gradient approaches for multi- objective sequential decision making

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:22:29.670077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:22:28.941947Z digest=sha256:2421037237d4f743adbecb41f31d10d96ddf34bdd7e3b34d4ef09f8d1058efed

Observation ddd3671f-54e5-4db2-8dbc-cc9743eaef93 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Code Llama: Open Foundation Models for Code

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.951454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.951454Z digest=sha256:a3fcb830a4927e501b87138750bdf478f68846c3845ce72412a229b0d01a8e33

Observation 34e4d833-4bb7-4021-99dd-8bb6a123b9c0 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.956335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.956335Z digest=sha256:060bd3b80499a7d5b91ebfaebca06c5a0f73f6084f2b15b27836fe353cc85e02

Observation 1c83dd08-bb53-46d1-9d34-4656e4372d80 · outbound

This paper cites Large Language Model Routing with Benchmark Datasets.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Large Language Model Routing with Benchmark Datasets

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.965591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.965591Z digest=sha256:db0739fd632af4c0f2861c7f0fbae0a9ac01727c8e8ec2bea98ffc03d3403892

Observation 97cda4cc-c528-4627-90f9-479cb6c9a8dc · outbound

This paper cites Learning Pareto Set for Multi-Objective Continuous Robot Control.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Learning Pareto Set for Multi-Objective Continuous Robot Control

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-09T11:22:29.145635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:22:28.970295Z digest=sha256:174e2c8ffbac65f4b005955fcdae439e0f45b50424e1bdc148d51309fe4bd46b

Observation 7ac6d6da-8e32-4b34-a0dd-9d1b9d623bdb · outbound

This paper cites Learning Invariances for Policy Generalization.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Learning Invariances for Policy Generalization

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-09T11:22:29.123383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:22:28.974763Z digest=sha256:ea76667ebebc227a2cba31c14b814ce8b549959a023d90b8973f099cd8f3f1f5

Observation 5c4ad380-aff4-4c88-8cf5-3c28898c07a3 · outbound

This paper cites Learning Invariant Representations for Reinforcement Learning without Reconstruction.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Learning Invariant Representations for Reinforcement Learning without Reconstruction

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.984591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.984591Z digest=sha256:37a8a94f6b3eaea350f0e8b342b78c4c60bda0b5e9fe539f3943dc922eedd674

Observation 02f2f55a-38e9-40d5-bf87-24070dd32974 · outbound

This paper cites mixup: Beyond Empirical Risk Minimization.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing mixup: Beyond Empirical Risk Minimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.989984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.989984Z digest=sha256:ff666ba5efda45a0b4b92fefe55f84da813b3706cc63b7b52ec80fae41406524

Observation c214eea0-a91b-4855-8df1-705d464f4073 · outbound

This paper cites Theoretical Analysis A.1.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Theoretical Analysis A.1

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:22:29.656086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:22:28.999316Z digest=sha256:a1481e66a754667ec2d3c461d67a82ec789cab4706ec1f9ace2cfb2c8f7b57d0

Observation 8c20daa2-5818-4673-b3e8-828e12b16eaf · outbound

This paper cites With the calibrated evaluation scores ¯p on a prompt x and a user preference vector ω, the routing action is determined by ˆa = arg maxk∈{k1,k2} ωT [¯pk, −ck].

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing With the calibrated evaluation scores ¯p on a prompt x and a user preference vector ω, the routing action is determined by ˆa = arg maxk∈{k1,k2} ωT [¯pk, −ck]

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:22:29.639945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:22:29.004445Z digest=sha256:9dd5ff6ff045a75ad37b14dee94d72920f2c902e53ca3f432c50666e9d39f028

Observation 0dd9f0e1-a8bf-40d0-8d7e-9155cee582bc · outbound

This paper cites To account for varying user preferences, we evaluate RouteLLM using a range of different thresholds.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing To account for varying user preferences, we evaluate RouteLLM using a range of different thresholds

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:22:29.608747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:22:29.015064Z digest=sha256:16d84e12c7be17c6fd3c280db9377300159dde13766066434e133730595b2d7a

Observation 37ef9fe1-ced8-4d73-a939-e44e61d7c86a · outbound

This paper cites The estimated cost of invoking the models for processing 1M input tokens and generating 1M output tokens.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing The estimated cost of invoking the models for processing 1M input tokens and generating 1M output tokens

Reference 128

Resolution
malformed identifier
raw_fallback, observed 2026-08-09T11:22:29.624214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:22:29.009937Z digest=sha256:3f49e60224125e01854de207d8e6793e8484ab86509cd6dd8d1ae52899f58e25

Observation ec076800-faf5-4fb0-86b8-f4e561d1f171 · outbound

This paper cites Generalization and Regularization in DQN.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Generalization and Regularization in DQN

Reference 1967

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.868513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.868513Z digest=sha256:b1994d4262481e122786b5370ac26b5b7bf0cac4fa4b953be2c5d214f3994bcb

Observation 345d3ba7-16d1-4ae9-87e7-68e07391531d · outbound

This paper cites and Doshi-Velez, F.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing and Doshi-Velez, F

Reference 1987

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:22:29.686218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:22:28.898237Z digest=sha256:1371983610b95ab0316d54f60a99244351ad96cdf5cb402417b7b622fd15cce3

Observation 82e5d072-7412-4b14-85bf-29d062e5d1ac · outbound

This paper cites Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.946558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.946558Z digest=sha256:edd531c976f181fe3dfc373ffb6450686eae79c088505d75187045d450884c30

Observation be890cb1-f802-4b56-bfe0-83c21c0d636d · outbound

This paper cites Tryage: Real-time, intelligent Routing of User Prompts to Large Language Models.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Tryage: Real-time, intelligent Routing of User Prompts to Large Language Models

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.873531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.873531Z digest=sha256:9f753189183afc998de2bd605400c1e7b817d57a2f2dccdfff640ac1f0aafb79

Observation 0747d94e-aab5-475b-b68e-592ed4edb145 · outbound

This paper cites Fusing Models with Complementary Expertise.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Fusing Models with Complementary Expertise

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.979953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.979953Z digest=sha256:c9536188b2ccd7f4f832ae8873752187c50339a2cfa7dcc639181800212a47d9

Observation 5eb8273e-0ad3-4fce-95c9-6fe0f4f378a9 · outbound

This paper cites Proximal Policy Optimization Algorithms.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Proximal Policy Optimization Algorithms

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.961052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.961052Z digest=sha256:c75548dd3528f5bea5bcde768e4f2840121dc4e2232ec11cb6374b8324dfe44e

Observation 9050406b-4c85-4912-ab30-6fd5499245ce · outbound

This paper cites Generalization of Reinforcement Learning with Policy-Aware Adversarial Data Augmentation.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Generalization of Reinforcement Learning with Policy-Aware Adversarial Data Augmentation

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.994652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.994652Z digest=sha256:967df411d2a9c6d30800a1433b26cc6c6433cfb4db590601466e35a29d138f91

Observation c2891f28-fd3c-481a-bd48-daa7ea2de030 · outbound

This paper cites GPT-4 Technical Report.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing GPT-4 Technical Report

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.813999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.813999Z digest=sha256:64c27d5dd7c1fb936140d2f6607fc2317321d887bcbcdc535b957537f342f266

Observation 20a156f6-5bdc-4bd0-88c5-e00dcfde39b3 · outbound

This paper cites Mixtral of Experts.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Mixtral of Experts

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.888358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.888358Z digest=sha256:1355599b6a68cf2a7ed39416852ba735bb3e042b780ce8083d1ffaf2ac2a97ac

Observation f57a25b6-fd3f-4e7a-b931-8f65302520c6 · outbound

This paper cites PD-MORL: Preference-Driven Multi-Objective Reinforcement Learning Algorithm.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing PD-MORL: Preference-Driven Multi-Objective Reinforcement Learning Algorithm

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.826130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.826130Z digest=sha256:fc7cfccd0de7833042ec002e9eccb11cb3e571b87f36446c3e7c57d8c6285647

Observation 405d56ad-0241-441e-9b31-ee974e60cd36 · outbound

This paper cites A Survey on Practical Applications of Multi-Armed and Contextual Bandits.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing A Survey on Practical Applications of Multi-Armed and Contextual Bandits

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.836851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.836851Z digest=sha256:96c431ec879af54187f0ffc5a757c6c8612b0a790567112d502ed40dff1c79c1

Observation fcce90a7-c71b-49de-b617-8e48fbfcdadf · outbound

This paper cites Contrastive Behavioral Similarity Embeddings for Generalization in Reinforcement Learning.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Contrastive Behavioral Similarity Embeddings for Generalization in Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-09T11:22:28.820492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:22:28.820492Z digest=sha256:5f2f809c9957eb11f75c2938814d7dc3bbbbf1ac037209a6a8d365a8fc2d1d31

Observation 4ea73457-752d-4539-9bd2-390d37b57dfe · outbound

This paper cites Robust Reinforcement Learning through Efficient Adversarial Herding.

LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Robust Reinforcement Learning through Efficient Adversarial Herding

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-09T11:22:29.455816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:22:28.863181Z digest=sha256:dd92a350a99b8cec1071ec25aab5cd17a57159ae69d91d8567a4098a4fb58759

Pith citing papers

Observation f5830e30-ac5e-48ce-b8a8-6c0b04578c76 · inbound

Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems cites this paper.

Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T19:10:53.675879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:10:53.675879Z digest=sha256:fb2a6d0c72a5168a922308cae6612656ea679a2171d9d588f6b4f09d6d748dc2

Observation 5f175dcd-2a18-4a06-9aff-522c06734b71 · inbound

Universal Model Routing for Efficient LLM Inference cites this paper.

Universal Model Routing for Efficient LLM Inference LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T23:48:01.031696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:48:01.031696Z digest=sha256:52368a11003b4d61375f4ddddaa976cc5a172cad5243e132816e6711c69b7321

Observation defab73f-003f-4515-9b8d-a9ca3b81fcb0 · inbound

Harnessing Multiple Large Language Models: A Survey on LLM Ensemble cites this paper.

Harnessing Multiple Large Language Models: A Survey on LLM Ensemble LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:25:19.688915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T02:22:28.649071Z digest=sha256:653dac9bf69f8106125a68f47dda3935b47961c1cc902bdace550255bb4349d7

Observation d272b4e4-cca4-4c1e-a726-0fba4a709c93 · inbound

Rethinking Predictive Modeling for LLM Routing: When Simple kNN Beats Complex Learned Routers cites this paper.

Rethinking Predictive Modeling for LLM Routing: When Simple kNN Beats Complex Learned Routers LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:14:57.469183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T15:13:28.927880Z digest=sha256:0f219a79ada292f140fda89b900d0582fb0d874b3c1493c9f781790ec93b6f80

Observation bd38a840-cdc3-457b-8a7b-8697f654ae00 · inbound

Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges cites this paper.

Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing

Reference 269

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:48.858500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:48.858500Z digest=sha256:8ffa462d791c624736591764b1c06346a760164f08a475608b229f6c9425c652

Observation 91f9e57d-68fb-4b1a-95a0-0e34774dda55 · inbound

KRONE: Scalable LLM-Augmented Log Anomaly Detection via Hierarchical Abstraction cites this paper.

KRONE: Scalable LLM-Augmented Log Anomaly Detection via Hierarchical Abstraction LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:00:42.944524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T06:59:23.030671Z digest=sha256:282c5e0a7dbcb55c5e1a4df0757bf4b1ffcd0d625c1d6fab9e899d329bd73916

Observation 510e1c15-fd98-4489-b63e-5fe7d5dfbad7 · inbound

POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving cites this paper.

POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:53:04.247674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:49:43.418576Z digest=sha256:91fdce0705774f3d6d85dbcdabcebec62729e703bb460e75473b782960f25a44

Observation 1756a5e4-ded6-4a2f-98bb-76b7beae0e24 · inbound

Replacing Parameters with Preferences: Federated Alignment of Heterogeneous Vision-Language Models cites this paper.

Replacing Parameters with Preferences: Federated Alignment of Heterogeneous Vision-Language Models LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:41:30.884663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-07T04:07:58.031100Z digest=sha256:1df11fa14979a3199ab2767dbd43aa4c77809fc7fcd55fa97a9c1a03bb2a2603

Observation f8254bf1-e6a1-4fe3-bc5a-24ca5b1d208b · inbound

SCATE: Learning to Supervise Coding Agents for Cost-Effective Test Generation cites this paper.

SCATE: Learning to Supervise Coding Agents for Cost-Effective Test Generation LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-13T01:15:01.090791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:15:01.090791Z digest=sha256:b8f08aff7c0e806b477eae30fa6f591e15abbdd1ead270a07cbad1982fdfc834

Observation dc175f25-7347-46ce-9df0-4ad2217144d1 · inbound

PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference cites this paper.

PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T10:14:14.184268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:14:14.184268Z digest=sha256:5c2c8eb94d1f2d105034c7eb29712cf79044aab9f395d2f49355e8a544f240e4