REVIEW 5 major objections 6 minor 2 cited by
BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
T0 review · 5 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Routing queries to cheap models with best-of-n sampling can cut LLM serving cost by up to 60% while keeping response quality within 0.8% of GPT-4o.
desk verdict Genuinely new combination of routing and best-of-n, but the headline cost-quality claim is only measured under the same reward model used for training; needs independent validation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the match-probability multi-head router plus the proxy reward model. The router uses a shared DeBERTa-v3-small backbone that encodes the query once, and $K$ by $N$ lightweight heads that each output $p_{k,n}(q)$, the probability that best-of-n responses from model $k$ match or beat a single GPT-4o response; it is trained on labels $y_n(q)$ that are ground-truth indicators of that event. The proxy reward model is a DeBERTa-v3-large fine-tuned on worst/median/best pairs from 20 responses per training query with the logistic pairwise ranking loss, and it selects the best of the sampled responses at inference. Algorithm 1 combines these pieces: predict match probabilities, filter by threshold t, estimate costs from input and output token prices and average output lengths, and return the highest-proxy-score response from the cheapest valid combination.
What would settle it
Take held-out queries, generate n=1,5,20 responses from a cheap model, score them with the proxy model and with a more authoritative judge such as human raters or LLM-as-judge, and compare the proxy-best response's judge score against a random sample's judge score; if the gap does not rise with n or turns negative on a nontrivial fraction of queries, the proxy-ranking premise fails. Separately, count how often combinations predicted above threshold t actually beat a single GPT-4o response on a set with known outcomes; systematic overprediction would break the reported cost-quality trade-off on other query distributions.
Extended reading notes
Core claim
The central claim is that the cost-quality ordering of LLMs is not fixed: a small model combined with best-of-n sampling sits on a new frontier that a router can exploit. Concretely, BEST-Route presents an algorithm that, for each query, predicts a match probability for every (cheap model, sample count) pair against a powerful reference model, filters to pairs whose predicted probability clears a threshold t, and executes the cheapest of those pairs; if no pair clears t, it uses the reference model once. Cost is estimated from token prices and average output length, and the best response is selected by a lightweight proxy reward model trained with a pairwise ranking loss. The paper reports experiments where this achieves up to 60% cost reduction with only a 0.8% quality drop measured by armoRM on the in-distribution test set, and a 1.59% drop on the out-of-distribution MT-Bench set, with negative quality drops on coding queries when a specialized model is added to the pool.
Load-bearing premise
The load-bearing premise is that the proxy reward model ranks the sampled responses in the same order as the true quality score on the queries the router serves; if that ordering is wrong, best-of-n sampling no longer improves quality and the router's cost savings vanish.
Editorial extensions
If this is right
- A serving system can keep its model portfolio unchanged and still reduce cost by letting the router decide how many responses to sample from each cheap model.
- The threshold t gives operators a tunable cost-quality dial: lower thresholds favor cheap combinations, higher thresholds protect quality by falling back to the reference model.
- Router overhead is small enough for real-time use: at n=20, match-probability prediction takes 0.04s and best-of-n scoring adds 0.58s, roughly 18.7x faster than the fastest local model evaluated.
- Specialized cheaper models can be mixed into the pool without changing the routing algorithm; adding Codestral-22b on coding queries produced better-than-GPT-4o quality at 20% lower cost.
- The approach carries over to out-of-distribution data and other quality metrics, with MT-Bench quality drop of 1.59% at 60% cost reduction.
Reading between the lines
- Because the router separates a shared query encoder from per-(model, n) heads, adding a new cheap model may only require training new lightweight heads; the paper does not demonstrate this transfer, but the architecture invites it.
- The threshold rule assumes predicted match probabilities are calibrated well enough that crossing t means the cheap combination is genuinely as good as the reference; a deployment should track realized match rates to detect silent miscalibration.
- The same proxy-reward best-of-n mechanism could be applied as a test-time compute policy for a single large model, deciding per query how many samples to draw; the paper studies it only in the routing setting.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes BEST-Route, a routing framework that combines a multi-head router with best-of-n sampling. For each query, the router predicts, for each small model and sample count n, the probability that best-of-n sampling from that model matches the quality of a single GPT-4o response; it then selects the cheapest (model, n) combination whose predicted match probability exceeds a threshold t, falling back to GPT-4o when no combination qualifies. The selected responses are ranked by a fine-tuned proxy reward model, which is itself trained on pairwise preferences derived from armoRM scores. Experiments on a new 10K-example dataset and on MT-Bench report up to 60% cost reduction with less than 1% armoRM score drop, and additional BLEU/ROUGE results. The paper also analyzes latency overhead, model usage before/after adding a specialized coding model, and cost-estimation error.
Significance. If the reported results hold, BEST-Route would be a practically useful contribution: it demonstrates that adaptive best-of-n sampling can extend the Pareto frontier of cost versus quality in LLM serving, and the multi-head router design is a sensible way to avoid training K×N independent routers. The paper also ships a concrete system design, detailed cost model, and a new evaluation dataset, which are valuable assets. However, the headline claim is currently supported mainly by armoRM, which is also the training reward for every learned component; the independent BLEU/ROUGE evidence shows much larger quality drops. The missing reporting of the threshold tuning procedure and the absence of error bars make the quantitative claims difficult to interpret. The central idea is defensible, but the evidence needs to be strengthened before the stated conclusions are established.
major comments (5)
- [Section 3.3, Section 4.1, Section 4.2, Table 1] The evaluation metric armoRM is also the training reward for both the proxy reward model and the router. armoRM is used as R_GT to construct the pairwise training set for R_proxy (Section 4.1), as R_GT in the router label y_n(q) (Section 4.2, Eq. 3), and as the response-quality metric in Tables 1–3. The headline claim of '0.8% quality drop' at 60% cost reduction therefore measures closeness to the training objective, not quality as judged by an independent criterion. The paper's own Table 4 shows that under BLEU and ROUGE the same operating point has 18.07% and 21.97% drops, respectively. To support the claimed '<1% performance drop,' the authors should provide an independent evaluation (for example, human judgments, task accuracy, or an out-of-training reward model) and should present the BLEU/ROUGE results as primary evidence alongside armoRM rather than as a secondary robustness check.
- [Section 4.2, Eq. (3)] The label y_n(q) is defined as a probability, but the paper never specifies how this probability is estimated during training. It is not stated whether a single reference response is sampled per query, whether the event R_GT(s*_small) >= R_GT(s_ref) is treated as a binary label, whether multiple responses are averaged, or how ties are handled. Without a concrete estimator, the router training procedure is not reproducible. Please specify the exact construction of the training labels, including the number of reference samples and any smoothing or averaging used.
- [Algorithm 1, Section 5.2, Tables 1–4] The match probability threshold t is a free parameter, and its values are never reported. The cost-reduction operating points of 10%, 20%, 40%, and 60% in Tables 1–4 appear to be produced by adjusting t, but the paper does not state how t is chosen, whether it is tuned on the validation set, or what values correspond to each row. This is load-bearing because the claimed cost-quality trade-off is a function of a fitted operating point; without reporting the selection procedure, the comparison to baselines that do not have an equivalent tunable parameter is not apples-to-apples. Please report the threshold values and the validation-based selection protocol.
- [Tables 1–4] All reported cost-reduction and quality-drop numbers are single point estimates with no error bars, confidence intervals, or significance tests. The test set has only 1K examples, and differences such as 0.19% vs. 0.63% at 10% cost reduction in Table 1 may be within noise. The authors should report means and variances over multiple seeds or bootstrap resamples, at least for the main tables.
- [Section 5.5, Table 3, Table 4] The claim that BEST-Route shows 'robustness under distribution shifts and generalizability to alternative quality metrics' is overstated. Table 3 shows that on MT-Bench the 60% cost-reduction point has a 1.59% armoRM drop, which is already larger than the '<1%' headline; Table 4 shows much larger BLEU/ROUGE drops. The MT-Bench result is still an in-training-metric evaluation, since armoRM is the training reward, and BLEU/ROUGE are lexical overlap metrics with known weak correlation with human judgment, as the paper itself notes in Section 3.3. Please temper the generalization claim and provide an evaluation that does not rely on the training reward.
minor comments (6)
- [Appendix A.1] The word 'specialized' is misspelled as 'specilized' in the description of Codestral-22b.
- [Section 4.1, Figure 2] The y-axis label of Figure 2 reads 'Avg. armoRM score ( )' with an empty parenthetical; please complete the label or remove the empty parentheses.
- [Section 4.2] The symbol n is used both as a specific sample count (e.g., in Eq. 2) and as the maximum sample count in Algorithm 1. Please use distinct notation, such as N_max, to avoid confusion.
- [Section 4.2, Eq. (3)] The notation y_n(q) for the training label and p_n(q) for the predicted probability is confusingly similar; consider using a different symbol for the label, such as l_n(q).
- [Appendix C] The armoRM case study is a single anecdotal example. It is fine as an illustration, but it should not be presented as validation; a sentence clarifying its illustrative status would help.
- [Section 5.1] The text says 'Codes will be released upon acceptance of this work' while the code is also listed with a GitHub URL; please clarify the current availability status.
Circularity Check
BEST-Route's headline '<1% performance drop' is measured with armoRM, the same reward model used to generate training labels for both the proxy reward model and the router; the closed loop makes the claim metric-specific, although the routing algorithm itself is not definitionally circular.
-
self definitional
[Section 3.3; Section 4.1; Section 4.2, Eq. (3); Table 1]
"We assess response quality using armoRM scores Wang et al. [2024a] ... To construct the training set P, we generate n = 20 sample responses S = {s1(q), s2(q), . . . , s20(q)} for each training query q and compute RGT(s(q)) using the armoRM score ... we generate a label yn(q) = P r[RGT(s∗small) ≥ RGT(sref)] (3)"
The response-quality metric for the headline claim (armoRM, Sections 3.3 and 5) is identical to the ground-truth reward R_GT used to train the proxy reward model (Section 4.1) and to form the router's match-probability labels (Eq. 3). Thus the router is trained to predict exactly the event (small-model best-of-n armoRM score >= reference armoRM score) whose realized frequency is then reported as 'performance drop' in Table 1. With the proxy also trained on armoRM pairs, every learned component is optimized against the same score used for evaluation, so the under-1% drop at 60% cost reduction is a within-objective result rather than an independent quality measurement.
full rationale
The routing algorithm itself is self-contained: Algorithm 1 combines a trained proxy reward model, a multi-head router, and a cost model, and the mechanics of the derivation do not assume the conclusion. The reported cost savings are computed from API prices and held-out test data, and no load-bearing self-citation chain is present; the cited prior work (e.g., Ding et al. 2024) only motivates query-difficulty variation and is not used to justify the main result. However, the central quality claim is not independently validated: armoRM serves as both the evaluation metric and the supervision signal for every learned component (the proxy reward model's ranking pairs and the router's match-probability labels). This is a closed optimization loop, so the under-1% performance drop at 60% cost reduction is partially forced by construction, because the system is explicitly trained to preserve armoRM and then evaluated on armoRM. The paper's own Table 4 shows BLEU/ROUGE drops an order of magnitude larger (18.07% and 21.97% at 60% cost reduction), indicating the headline is metric-specific rather than a general quality guarantee. This is a partial circularity in the evaluation metric, not a definitional equivalence of the derivation, so the score is 4 rather than higher.
Assumptions & free parameters
free parameters (3)
- match probability threshold t =
not reported
- maximum sample count N =
20 (proxy training), 5 (router experiments)
- average output length per model =
training-split averages
assumptions (4)
- domain assumption Proxy reward model preserves the ground-truth ranking of responses
- domain assumption armoRM is a reliable proxy for human-judged response quality
- domain assumption Query text alone can predict the match probability y_n(q)
- domain assumption Average training output length predicts per-query cost
Cite this review
Pith. "Pith review of BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute." pith.science (2026). https://pith.science/paper/SQI556FQ
@misc{pith2026250622716,
author = {Pith},
title = {Pith review of: BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute},
year = {2026},
howpublished = {\url{https://pith.science/paper/SQI556FQ}},
note = {Machine review of arXiv:2506.22716}
}
read the original abstract
Large language models (LLMs) are powerful tools but are often expensive to deploy at scale. LLM query routing mitigates this by dynamically assigning queries to models of varying cost and quality to obtain a desired trade-off. Prior query routing approaches generate only one response from the selected model and a single response from a small (inexpensive) model was often not good enough to beat a response from a large (expensive) model due to which they end up overusing the large model and missing out on potential cost savings. However, it is well known that for small models, generating multiple responses and selecting the best can enhance quality while remaining cheaper than a single large-model response. We leverage this idea to propose BEST-Route, a novel routing framework that chooses a model and the number of responses to sample from it based on query difficulty and the quality thresholds. Experiments on real-world datasets demonstrate that our method reduces costs by up to 60% with less than 1% performance drop.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 2 Pith papers
-
Latency and Token-Aware Test-Time Compute
A learned per-query router selects both the inference-scaling method and its compute budget to balance accuracy, token use, and latency, outperforming static strategies on math reasoning.
-
SLMs as Multi-Agent Routers: A Progressive SFT and Reinforcement Learning Approach
A 0.6B router trained by SFT+RL on retrieval-quality rewards reaches 0.771 NDCG@10 across 11 agents, beating intent-prompted LLMs and cutting latency by 82%.
Reference graph
Works this paper leans on
- [1]
- [2]
-
[3]
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell. On the dangers of stochastic parrots : can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness , Accountability , and Transparency , pages 610--623, Virtual Event Canada, Mar. 2021. ACM. ISBN 978-1-4503-8309-7. doi:10.1145/3442188.3445922. URL https://dl.acm.org/do...
arXiv 2021
-
[4]
A global analysis of metrics used for measuring performance in natural language processing
K. Blagec, G. Dorffner, M. Moradi, S. Ott, and M. Samwald. A global analysis of metrics used for measuring performance in natural language processing. arXiv preprint arXiv:2204.11574, 2022
work page Pith review arXiv 2022
- [5]
-
[6]
L. Chen, M. Zaharia, and J. Zou. Frugalgpt: How to use large language models while reducing cost and improving performance. arXiv preprint arXiv:2305.05176, 2023
arXiv 2023
-
[7]
L. Chen, J. Q. Davis, B. Hanin, P. Bailis, I. Stoica, M. Zaharia, and J. Zou. Are more llm calls all you need? towards scaling laws of compound inference systems. arXiv preprint arXiv:2403.02419, 2024 a
arXiv 2024
-
[8]
Z. Chen, A. May, R. Svirschevski, Y. Huang, M. Ryabinin, Z. Jia, and B. Chen. Sequoia: Scalable, robust, and hardware-aware speculative decoding. arXiv preprint arXiv:2402.12374, 2024 b
arXiv 2024
Show all 59 references
-
[9]
Chuang, W
Y.-S. Chuang, W. Fang, S.-W. Li, W.-t. Yih, and J. Glass. Expand, rerank, and retrieve: Query reranking for open-domain question answering. In Findings of the Association for Computational Linguistics: ACL 2023, pages 12131--12147, 2023
2023
-
[10]
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. R \'e . Flashattention: Fast and memory-efficient exact attention with io-awareness. Advances in Neural Information Processing Systems, 35: 0 16344--16359, 2022
2022
-
[11]
D. Ding, S. Amer-Yahia, and L. Lakshmanan. On efficient approximate queries over machine learning models. Proceedings of the VLDB Endowment, 16 0 (4): 0 918--931, 2022
2022
-
[12]
D. Ding, A. Mallick, C. Wang, R. Sim, S. Mukherjee, V. R \"u hle, L. V. Lakshmanan, and A. H. Awadallah. Hybrid llm: Cost-efficient and quality-aware query routing. In The Twelfth International Conference on Learning Representations, 2024
2024
-
[13]
D. Ding, B. Xu, and L. V. Lakshmanan. Occam: Towards cost-efficient and accuracy-aware classification inference. In The Thirteenth International Conference on Learning Representations, 2025
2025
-
[14]
Dubey, A
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024
2024 arXiv
-
[15]
Elsken, J
T. Elsken, J. H. Metzen, and F. Hutter. Neural architecture search: A survey. The Journal of Machine Learning Research, 20 0 (1): 0 1997--2017, 2019
1997
-
[16]
L. Gui, C. G \^a rbacea, and V. Veitch. Bonbon alignment for large language models and the sweetness of best-of-n sampling. arXiv preprint arXiv:2406.00832, 2024
2024 arXiv
-
[17]
Gupta, H
N. Gupta, H. Narasimhan, W. Jitkrittum, A. S. Rawat, A. K. Menon, and S. Kumar. Language model cascades: Token-level uncertainty and beyond. arXiv preprint arXiv:2404.10136, 2024
2024 arXiv
-
[18]
Hassibi, D
B. Hassibi, D. G. Stork, and G. J. Wolff. Optimal brain surgeon and general network pruning. In IEEE international conference on neural networks, pages 293--299. IEEE, 1993
1993
-
[19]
P. He, X. Liu, J. Gao, and W. Chen. Deberta: Decoding-enhanced bert with disentangled attention. arXiv preprint arXiv:2006.03654, 2020
2006 arXiv
-
[20]
Hinton, O
G. Hinton, O. Vinyals, J. Dean, et al. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2 0 (7), 2015
2015 arXiv
-
[21]
Hugging face inference api
HuggingFace. Hugging face inference api. https://huggingface.co/inference-api
-
[22]
Jacob, S
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko. Quantization and training of neural networks for efficient integer-arithmetic-only inference. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2704--2...
2018
-
[23]
Jiang, X
D. Jiang, X. Ren, and B. Y. Lin. Llm-blender: Ensembling large language models with pairwise ranking and generative fusion. arXiv preprint arXiv:2306.02561, 2023
2023 arXiv
-
[24]
Jinnai, T
Y. Jinnai, T. Morimura, K. Ariu, and K. Abe. Regularized best-of-n sampling to mitigate reward hacking for language model alignment. arXiv preprint arXiv:2404.01054, 2024
2024 arXiv
-
[25]
A. Kag, I. Fedorov, A. Gangrade, P. Whatmough, and V. Saligrama. Efficient edge inference by selective query. In The Eleventh International Conference on Learning Representations, 2022
2022
-
[26]
S. Kim, K. Mangalam, S. Moon, J. Malik, M. W. Mahoney, A. Gholami, and K. Keutzer. Speculative decoding with big little decoder. In Thirty-seventh Conference on Neural Information Processing Systems, 2023
2023
-
[27]
Lambert, V
N. Lambert, V. Pyatkin, J. Morrison, L. Miranda, B. Y. Lin, K. Chandu, N. Dziri, S. Kumar, T. Zick, Y. Choi, et al. Rewardbench: Evaluating reward models for language modeling. arXiv preprint arXiv:2403.13787, 2024
2024 arXiv
-
[28]
LeCun, J
Y. LeCun, J. Denker, and S. Solla. Optimal brain damage. Advances in neural information processing systems, 2, 1989
1989
-
[29]
Leviathan, M
Y. Leviathan, M. Kalman, and Y. Matias. Fast inference from transformers via speculative decoding. In International Conference on Machine Learning, pages 19274--19286. PMLR, 2023
2023
-
[30]
C.-Y. Lin. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74--81, 2004
2004
-
[31]
K. Lu, H. Yuan, R. Lin, J. Lin, Z. Yuan, C. Zhou, and J. Zhou. Routing to the expert: Efficient reward-guided ensemble of large language models. arXiv preprint arXiv:2311.08692, 2023
2023 arXiv
-
[32]
Mavromatis, P
C. Mavromatis, P. Karypis, and G. Karypis. Pack of llms: Model fusion at test-time via perplexity optimization. arXiv preprint arXiv:2404.11531, 2024
2024 arXiv
-
[33]
Nakano, J
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, et al. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332, 2021
2021 arXiv
-
[34]
Narasimhan, W
H. Narasimhan, W. Jitkrittum, A. S. Rawat, S. Kim, N. Gupta, A. K. Menon, and S. Kumar. Faster cascades via speculative decoding. arXiv preprint arXiv:2405.19261, 2024
2024 arXiv
-
[35]
I. Ong, A. Almahairi, V. Wu, W.-L. Chiang, T. Wu, J. E. Gonzalez, M. W. Kadous, and I. Stoica. Routellm: Learning to route llms with preference data. arXiv preprint arXiv:2406.18665, 2024
2024 arXiv
-
[36]
OpenAI. Chatgpt. https://chat.openai.com/, a
-
[37]
Openai platform
OpenAI. Openai platform. https://platform.openai.com/overview, b
-
[38]
Ouyang, J
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35: 0 27730--27744, 2022
2022
-
[39]
Papineni, S
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311--318, 2002
2002
-
[40]
S akota, M
M. S akota, M. Peyrard, and R. West. Fly-swat or cannon? cost-effective language model choice via meta-modeling. In Proceedings of the 17th ACM International Conference on Web Search and Data Mining, pages 606--615, 2024
2024
-
[41]
Shnitzer, A
T. Shnitzer, A. Ou, M. Silva, K. Soule, Y. Sun, J. Solomon, N. Thompson, and M. Yurochkin. Large language model routing with benchmark datasets. arXiv preprint arXiv:2309.15789, 2023
2023 arXiv
-
[42]
Skalse, N
J. Skalse, N. Howe, D. Krasheninnikov, and D. Krueger. Defining and characterizing reward gaming. Advances in Neural Information Processing Systems, 35: 0 9460--9471, 2022
2022
-
[43]
Snell, J
C. Snell, J. Lee, K. Xu, and A. Kumar. Scaling llm test-time compute optimally can be more effective than scaling model parameters. arXiv preprint arXiv:2408.03314, 2024
2024 arXiv
-
[44]
Srivatsa, K
K. Srivatsa, K. K. Maurya, and E. Kochmar. Harnessing the power of multiple minds: Lessons learned from llm routing. arXiv preprint arXiv:2405.00467, 2024
2024 arXiv
-
[45]
Stiennon, L
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33: 0 3008--3021, 2020
2020
-
[46]
Treviso, J.-U
M. Treviso, J.-U. Lee, T. Ji, B. v. Aken, Q. Cao, M. R. Ciosici, M. Hassid, K. Heafield, S. Hooker, C. Raffel, et al. Efficient methods for natural language processing: A survey. Transactions of the Association for Computational Linguistics, 11: 0 826--860, 2023
2023
-
[47]
Urban, K
G. Urban, K. J. Geras, S. E. Kahou, O. Aslan, S. Wang, R. Caruana, A. Mohamed, M. Philipose, and M. Richardson. Do deep convolutional nets really need to be deep and convolutional? arXiv preprint arXiv:1603.05691, 2016
2016 arXiv
-
[48]
Vanhoucke, A
V. Vanhoucke, A. Senior, and M. Z. Mao. Improving the speed of neural networks on cpus. 2011
2011
-
[49]
H. Wang, W. Xiong, T. Xie, H. Zhao, and T. Zhang. Interpretable preferences via multi-objective reward modeling and mixture-of-experts. arXiv preprint arXiv:2406.12845, 2024 a
2024 arXiv
-
[50]
J. Wang, J. Wang, B. Athiwaratkun, C. Zhang, and J. Zou. Mixture-of-agents enhances large language model capabilities. arXiv preprint arXiv:2406.04692, 2024 b
2024 arXiv
-
[51]
X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations, 2023
2023
-
[52]
G.-I. Yu, J. S. Jeong, G.-W. Kim, S. Kim, and B.-G. Chun. Orca: A distributed serving system for \ Transformer-Based \ generative models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22), pages 521--538, 2022
2022
-
[53]
M. Yue, J. Zhao, M. Zhang, L. Du, and Z. Yao. Large language model cascades with mixture of thoughts representations for cost-efficient reasoning. arXiv preprint arXiv:2310.03094, 2023
2023 arXiv
-
[54]
Zhang, J
S. Zhang, J. Zhang, D. Ding, M. H. Garcia, A. Mallick, D. Madrigal, M. Xia, V. R \"u hle, Q. Wu, and C. Wang. Ecoact: Economic agent determines when to register what action. arXiv preprint arXiv:2411.01643, 2024 a
2024 arXiv
-
[55]
Zhang, J
S. Zhang, J. Zhang, J. Liu, L. Song, C. Wang, R. Krishna, and Q. Wu. Offline training of language model agents with functions as learnable weights. In Forty-first International Conference on Machine Learning, 2024 b
2024
-
[56]
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al. A survey of large language models. arXiv preprint arXiv:2303.18223, 2023
2023 arXiv
-
[57]
Zheng, W.-L
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685, 2023
2023 arXiv
-
[58]
Zoph and Q
B. Zoph and Q. V. Le. Neural architecture search with reinforcement learning. arXiv preprint arXiv:1611.01578, 2016
2016 arXiv
-
[59]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.