Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T11:22:29.015064Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 10 inbound Pith citation observations for arXiv:2502.02743.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T11:22:29.015064Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T19:10:53.675879Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-23T02:25:19.685963Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 930e4926-5eb3-4d39-a6a4-8b94aac96986 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Contextualize Me -- The Case for Context in Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab1e82c0-2cf7-45cf-9f26-609349832f23 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing RT-1: Robotics Transformer for Real-World Control at Scale
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f08221ad-75b5-4b95-8eb3-3c5aae85adb7 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing A Brief Review of Hypernetworks in Deep Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e5c45ba-38d1-4767-8ae1-eba029549a1a · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68fd614a-c29e-46c9-92dd-bf214f6c1da0 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8979fb9c-e9a3-44e2-8f3d-7f39bafda339 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Zero-Shot Reinforcement Learning via Function Encoders
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8758b8f-16a6-4368-9ef1-576f4dbd4189 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Generalization to New Actions in Reinforcement Learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8caa511e-12f4-4ef1-917a-a22af59ca594 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ff53baf-f196-48b2-a4fc-3e5b8ca63220 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing A Survey Analyzing Generalization in Deep Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e845115f-68b9-46d6-903c-d5b9f72ebc25 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Network Randomization: A Simple Technique for Generalization in Deep Reinforcement Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6887e6a-0877-46d1-8f8a-f439804b5877 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Holistic Evaluation of Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7118ebd6-14d2-4dce-838f-52385c04b611 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 861a997a-073b-43a2-94a6-107bbd0f4d3e · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56a7bc7c-ac5f-41a9-9305-3a4ae2b2ba86 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing AutoMix: Automatically Mixing Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b368b8f-db85-42d1-b243-d9cd52195ce3 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing MetaLLM: A High-performant and Cost-efficient Dynamic Framework for Wrapping LLMs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bc06a7a-4b64-47ae-956f-70ef1f0c2b6e · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing RouteLLM: Learning to Route LLMs with Preference Data
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 781a6906-2980-41b7-abcd-10f70612d8fc · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Policy gradient approaches for multi- objective sequential decision making
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ddd3671f-54e5-4db2-8dbc-cc9743eaef93 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Code Llama: Open Foundation Models for Code
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34e4d833-4bb7-4021-99dd-8bb6a123b9c0 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c83dd08-bb53-46d1-9d34-4656e4372d80 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Large Language Model Routing with Benchmark Datasets
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97cda4cc-c528-4627-90f9-479cb6c9a8dc · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Learning Pareto Set for Multi-Objective Continuous Robot Control
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7ac6d6da-8e32-4b34-a0dd-9d1b9d623bdb · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Learning Invariances for Policy Generalization
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5c4ad380-aff4-4c88-8cf5-3c28898c07a3 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Learning Invariant Representations for Reinforcement Learning without Reconstruction
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02f2f55a-38e9-40d5-bf87-24070dd32974 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing mixup: Beyond Empirical Risk Minimization
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c214eea0-a91b-4855-8df1-705d464f4073 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Theoretical Analysis A.1
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8c20daa2-5818-4673-b3e8-828e12b16eaf · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing With the calibrated evaluation scores ¯p on a prompt x and a user preference vector ω, the routing action is determined by ˆa = arg maxk∈{k1,k2} ωT [¯pk, −ck]
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0dd9f0e1-a8bf-40d0-8d7e-9155cee582bc · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing To account for varying user preferences, we evaluate RouteLLM using a range of different thresholds
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 37ef9fe1-ced8-4d73-a939-e44e61d7c86a · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing The estimated cost of invoking the models for processing 1M input tokens and generating 1M output tokens
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ec076800-faf5-4fb0-86b8-f4e561d1f171 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Generalization and Regularization in DQN
Reference 1967
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 345d3ba7-16d1-4ae9-87e7-68e07391531d · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing and Doshi-Velez, F
Reference 1987
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 82e5d072-7412-4b14-85bf-29d062e5d1ac · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection
Reference 1999
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be890cb1-f802-4b56-bfe0-83c21c0d636d · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Tryage: Real-time, intelligent Routing of User Prompts to Large Language Models
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0747d94e-aab5-475b-b68e-592ed4edb145 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Fusing Models with Complementary Expertise
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eb8273e-0ad3-4fce-95c9-6fe0f4f378a9 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Proximal Policy Optimization Algorithms
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9050406b-4c85-4912-ab30-6fd5499245ce · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Generalization of Reinforcement Learning with Policy-Aware Adversarial Data Augmentation
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2891f28-fd3c-481a-bd48-daa7ea2de030 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing GPT-4 Technical Report
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20a156f6-5bdc-4bd0-88c5-e00dcfde39b3 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Mixtral of Experts
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f57a25b6-fd3f-4e7a-b931-8f65302520c6 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing PD-MORL: Preference-Driven Multi-Objective Reinforcement Learning Algorithm
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 405d56ad-0241-441e-9b31-ee974e60cd36 · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing A Survey on Practical Applications of Multi-Armed and Contextual Bandits
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcce90a7-c71b-49de-b617-8e48fbfcdadf · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Contrastive Behavioral Similarity Embeddings for Generalization in Reinforcement Learning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ea73457-752d-4539-9bd2-390d37b57dfe · outbound
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing Robust Reinforcement Learning through Efficient Adversarial Herding
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f5830e30-ac5e-48ce-b8a8-6c0b04578c76 · inbound
Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f175dcd-2a18-4a06-9aff-522c06734b71 · inbound
Universal Model Routing for Efficient LLM Inference LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation defab73f-003f-4515-9b8d-a9ca3b81fcb0 · inbound
Harnessing Multiple Large Language Models: A Survey on LLM Ensemble LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d272b4e4-cca4-4c1e-a726-0fba4a709c93 · inbound
Rethinking Predictive Modeling for LLM Routing: When Simple kNN Beats Complex Learned Routers LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bd38a840-cdc3-457b-8a7b-8697f654ae00 · inbound
Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing
Reference 269
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91f9e57d-68fb-4b1a-95a0-0e34774dda55 · inbound
KRONE: Scalable LLM-Augmented Log Anomaly Detection via Hierarchical Abstraction LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 510e1c15-fd98-4489-b63e-5fe7d5dfbad7 · inbound
POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1756a5e4-ded6-4a2f-98bb-76b7beae0e24 · inbound
Replacing Parameters with Preferences: Federated Alignment of Heterogeneous Vision-Language Models LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f8254bf1-e6a1-4fe3-bc5a-24ca5b1d208b · inbound
SCATE: Learning to Supervise Coding Agents for Cost-Effective Test Generation LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc175f25-7347-46ce-9df0-4ad2217144d1 · inbound
PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.