Pith. sign in

Paper Citation Record · LEDGER

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

As of 14 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 6 inbound Pith citation observations for arXiv:2506.06579.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06579 v1

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:56:57.070193Z

measured 106 of 106 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:10:53.372941Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 104 outbound references displayed

  • verified exact3
  • verified fuzzy13
  • unresolved84
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 18d0f5d8-50b3-4731-9b87-671596bc7b41 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.676293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.676293Z digest=sha256:ea1950818b8aa16a0bc09468d310285d114d8e6ca8bada6935c52f6fd4ba18f9

Observation 81b4d2ff-f2a3-422a-9949-b79583aae3c4 · outbound

This paper cites Gpt-3: What is it good for,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Gpt-3: What is it good for,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.681421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.681421Z digest=sha256:3d01386bbd9583567150a70e15609125a086334dedf67bb6a36c560da6695eca

Observation 05ee312a-aeb1-4c58-afe7-62044f8759ee · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.685256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.685256Z digest=sha256:4e6b0d06d5e28794eba5d87b50fe200352b690af67959a02244390f328a758ed

Observation b3796b1a-f0bc-4e92-a908-bdbe0898d0e7 · outbound

This paper cites Mobile Edge Intelligence for Large Language Models: A Contemporary Survey.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Mobile Edge Intelligence for Large Language Models: A Contemporary Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.689467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.689467Z digest=sha256:61dfc7b386b1337bc67ec7233115b8fe3b9ae9e72ac85227355305eec06e551b

Observation 46ff31e9-dd1b-46b2-96a3-52481f9c4c92 · outbound

This paper cites Faster and Lighter LLMs: A Survey on Current Challenges and Way Forward.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Faster and Lighter LLMs: A Survey on Current Challenges and Way Forward

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.693866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.693866Z digest=sha256:67e9391694b583a762e729724a788a58cfec02522aafed7675c5b13e83fd925a

Observation 33ce2435-327e-49b8-953f-09acee7540f3 · outbound

This paper cites A Survey of Resource-efficient LLM and Multimodal Foundation Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A Survey of Resource-efficient LLM and Multimodal Foundation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.698136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.698136Z digest=sha256:3f0bdbdb4c6ffb2bb6757d871ab0cbe544e723082cb5ab96ab9cc8ed4e706ebd

Observation 9fab6bad-1562-4300-b2ef-c1200ed5900d · outbound

This paper cites Efficient compressing and tuning methods for large language models: A systematic literature review,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Efficient compressing and tuning methods for large language models: A systematic literature review,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.703061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.703061Z digest=sha256:ac49b866af13466179b352d13a03e820aced73734352885163a2efee36b48a2f

Observation 61647dc1-bbbe-496c-bfb1-6c23b4989b34 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.706799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.706799Z digest=sha256:e472ed205a163a6728417848780c062c2a0c071df6d215ddd36db0fffc4d978f

Observation 3e1f6fc4-ce73-4bd6-988b-01bc1c177bbd · outbound

This paper cites Evaluation of the phi-3-mini slm for identification of texts related to medicine, health, and sports injuries,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Evaluation of the phi-3-mini slm for identification of texts related to medicine, health, and sports injuries,

Reference 9

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:56:57.788517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:56:56.710904Z digest=sha256:b239f10575518881f3bdc6791f46c148498b3a782d2a9fe5c332e986e50ef920

Observation e5184609-c34b-4f89-b9d8-454290dbc810 · outbound

This paper cites Harnessing moderate- sized language models for reliable patient data deidentification in emergency department records: Algorithm development, validation, and implementation study,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Harnessing moderate- sized language models for reliable patient data deidentification in emergency department records: Algorithm development, validation, and implementation study,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.714951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.714951Z digest=sha256:e846f441fa3be49dcff07df625944c257d27869d9afb9f140051be745da00e75

Observation ef61aed4-6660-40d4-a88f-3ed439004692 · outbound

This paper cites Overview of small language models in practice,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Overview of small language models in practice,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.718751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.718751Z digest=sha256:ba0a8a1b343a3c0182e8f85f3fe1dad914f18df36cc088e6328720e292629fc8

Observation 25ca84e2-fe07-4ddb-9284-b90423a8e8de · outbound

This paper cites MetaLLM: A High-performant and Cost-efficient Dynamic Framework for Wrapping LLMs.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques MetaLLM: A High-performant and Cost-efficient Dynamic Framework for Wrapping LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.722420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.722420Z digest=sha256:21b0ad13c5fd169508cba8a3f4c2f09978aa5f4a12728e205c711ac06cf479a9

Observation 5cf3d3c2-0771-478c-a869-9735cf428956 · outbound

This paper cites EcoAssistant: Using LLM Assistant More Affordably and Accurately.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques EcoAssistant: Using LLM Assistant More Affordably and Accurately

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.726464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.726464Z digest=sha256:c4fd43a2e3eb6a9ddce4de502190d7a9f6b6031616b31fb3d5f32f1c4090982d

Observation ebb02bfe-3add-451a-adfe-b3185ef308e7 · outbound

This paper cites FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.730282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.730282Z digest=sha256:b46fe34616a40653bc1e400b7274f6e24043eebc836dda472f2cc42ac847ed82

Observation 1adff952-db0e-4c63-9f30-5560021ae50e · outbound

This paper cites Harnessing the Power of Multiple Minds: Lessons Learned from LLM Routing.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Harnessing the Power of Multiple Minds: Lessons Learned from LLM Routing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.734663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.734663Z digest=sha256:e0423cf0a7d2142fc9a41048911aa0990567cf988d6f22af41ad6b2f1224ed87

Observation c53fac9d-a16e-49b1-9c65-099db0fc7b9b · outbound

This paper cites Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.738323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.738323Z digest=sha256:1c4c37cc098c51ddabb9e7c00c6c770d2693f3b290d4d797810f05b8b2d3bbdf

Observation e13a777d-61db-46c8-bd3e-3388bac034aa · outbound

This paper cites RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.742261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.742261Z digest=sha256:c29720c799115dda0d3f87c754fdf4edf40e9af6ee1b1526d336ca4b5342315f

Observation 109d4bb6-2963-4afc-a98d-8f293c457236 · outbound

This paper cites RouterBench: A Benchmark for Multi-LLM Routing System.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques RouterBench: A Benchmark for Multi-LLM Routing System

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.746355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.746355Z digest=sha256:5b789a2c06226e249e184c702190965d14331747721db0ba46ffad358d451550

Observation e14414f2-d12c-4f9d-85c1-466be764662a · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A Survey on Efficient Inference for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.750392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.750392Z digest=sha256:da931bcb58a963e5627092a88a31b71fde1c4a552bce667f46c7203a3c7723f5

Observation 5e8ed8fa-1544-40d7-b95d-3d65bd8a07dd · outbound

This paper cites Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.754256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.754256Z digest=sha256:4d307ed28561c93d79d96164c2ad6579e19bebf476e2d4422a52063f7daaaf68

Observation 0c823151-99af-4d27-9d87-0b6c60cb9593 · outbound

This paper cites Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.758394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.758394Z digest=sha256:f21a29f29886e3bbb8fb203899128b1a85492b4f8ec0f8756c7b524ba8d3cbf8

Observation 23b5ca1b-7de2-48db-b85d-84b7783cced2 · outbound

This paper cites LLM Inference Serving: Survey of Recent Advances and Opportunities.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques LLM Inference Serving: Survey of Recent Advances and Opportunities

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.762293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.762293Z digest=sha256:09d451585f8da8ea4b36cd1dfd06e5070468a40861415f1768e8c297c3b32194

Observation 8a93df01-9df9-49f3-a6fd-8babc8be4a46 · outbound

This paper cites Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.766216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.766216Z digest=sha256:2ad47e1d48ed1a67bcad948cdf8d05fc531a2b22d777fba68710407b6dce89fb

Observation 1d8ad172-a5ce-46a3-ad55-399e14c4ea5e · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.770100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.770100Z digest=sha256:90a40fa0b9c20fb8a97f202ceb5cdf4a552941c05c7868a498542ff955144672

Observation 5a9e363e-7f67-4f63-9e44-a51864a34603 · outbound

This paper cites The Efficiency Spectrum of Large Language Models: An Algorithmic Survey.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques The Efficiency Spectrum of Large Language Models: An Algorithmic Survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.774041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.774041Z digest=sha256:0fbf5789b7f66622205572e5a68f5128498b59b0a95afe00cdcf2552f5b4743b

Observation eb037d63-9b43-4127-a192-d9fc29db83ea · outbound

This paper cites Harnessing Multiple Large Language Models: A Survey on LLM Ensemble.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Harnessing Multiple Large Language Models: A Survey on LLM Ensemble

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.778198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.778198Z digest=sha256:f325a0592b22d0cd581b8521c4f0cfc3c5b2305437e1d0704d924599a7e9aebe

Observation 6039e81e-d05b-4f0a-ae3f-a52ab2991376 · outbound

This paper cites A Survey on Collaborative Mechanisms Between Large and Small Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A Survey on Collaborative Mechanisms Between Large and Small Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.786194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.786194Z digest=sha256:fef9245792a3e1bed604747d2cddf9d818ef8c09381fd4ae770b477b24020bd4

Observation 5de78550-fd1b-4748-9cb0-1c7e64efeb56 · outbound

This paper cites A Survey on Mixture of Experts in Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A Survey on Mixture of Experts in Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.789923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.789923Z digest=sha256:9781b42faab4e399a193e23b12f7bbd7e65ddbde5c3971a4f9db2e20dfa823cf

Observation 3577678f-29a4-4934-b91a-f0345d91bd84 · outbound

This paper cites Attention is all you need,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Attention is all you need,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.793800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.793800Z digest=sha256:c21f4a4df5c888dc2795299e921ca2418bf5837f7816544cc4acffd5fbc5808b

Observation 0bbbec51-4dd0-4c8b-a035-89a8f7eb3720 · outbound

This paper cites Improving text classification with transformer,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Improving text classification with transformer,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.797365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.797365Z digest=sha256:b97b2498d0f2011ca285bbf3fec67a9a331e8728ebd53219bab6717c9ece6cb2

Observation 978ff6f8-5e7e-437d-be5b-5011412a4460 · outbound

This paper cites Scaling Neural Machine Translation.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Scaling Neural Machine Translation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.800726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.800726Z digest=sha256:e7c79f56157638b075164f92e98d8eadff98bd798cd03aa39b1f4396e283ff1c

Observation be4655b7-aea5-4364-b2f2-dd5ae28b5a34 · outbound

This paper cites Block-skim: Efficient question answering for transformer,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Block-skim: Efficient question answering for transformer,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.804693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.804693Z digest=sha256:2570941c579db0f9d67e7cbea68ce97cd69e3a0872665341c4fd10330a89c847

Observation b2660a75-897a-4d33-bb3f-35ccf3443b7f · outbound

This paper cites An attentive survey of attention models,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques An attentive survey of attention models,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.808452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.808452Z digest=sha256:834086e99e6ec601562ec103b17491e96fabc3eb952a2ad9d5041eba0196067b

Observation c6c1e688-a417-42a8-b552-c22a9386b9f3 · outbound

This paper cites Attention, please! a survey of neural attention models in deep learning,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Attention, please! a survey of neural attention models in deep learning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.811824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.811824Z digest=sha256:80a84c58e0aa33c6907baf2b1bf6cd153567dc158aa6dda11e078bc5f506e563

Observation 6710cb2b-7a60-47f6-bbc8-cc316fb1a6b0 · outbound

This paper cites Show and tell: A neural image caption generator,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Show and tell: A neural image caption generator,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.815532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.815532Z digest=sha256:4b0a923180ea068341391d481da80275972083aae94804189880b046ef27b6c2

Observation 3ae31ecd-3f5e-4c04-b5b3-662fbbc0d25d · outbound

This paper cites Deep learning,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Deep learning,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.819461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.819461Z digest=sha256:a05a310e2c98567ed48e635f2f4de4ed4613ff1713f6fda64aac22037435bfad

Observation 19dd7d3c-cf54-4f03-9153-8cece3bdd3da · outbound

This paper cites Deep learning,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Deep learning,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.823248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.823248Z digest=sha256:51015dbf3b663c866ba8c157483d062efa9600cd25e30bc2e69c879c227b39c4

Observation ed2ed15a-de27-4b25-a446-5ba96e48fe8b · outbound

This paper cites Long short-term memory,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Long short-term memory,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.827223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.827223Z digest=sha256:7936de309fc921b9f040c71cb0395f0b3e9a906197a9c165426933d28457dca1

Observation 91a31ed1-48ed-4645-aca7-f84d249d3ae8 · outbound

This paper cites Towards Expert-Level Medical Question Answering with Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Towards Expert-Level Medical Question Answering with Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.830717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.830717Z digest=sha256:879856306e2e1a943607c55fb8d51249edff4554aec29d4a8fc2d566b8ce1c47

Observation 61ad6437-42e6-421f-a37d-6989b2f3066d · outbound

This paper cites Training data-efficient image transformers & distillation through attention,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Training data-efficient image transformers & distillation through attention,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.834768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.834768Z digest=sha256:61e43bab79f3bfa3ed7fbce7c62ce7170eb41b6b6db075c6cba5e46e81a02aec

Observation 026b1338-c47b-486d-8f6f-1b19e7c41a99 · outbound

This paper cites SplitLoRA: A Split Parameter-Efficient Fine-Tuning Framework for Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques SplitLoRA: A Split Parameter-Efficient Fine-Tuning Framework for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.838942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.838942Z digest=sha256:6bbfe8022e48b99dbb033e514c2aec54a502e298bfb351d429baadb1505a61eb

Observation 17d69191-bd0e-4a14-bc3c-c4a18dc13238 · outbound

This paper cites ALBERT: A Lite BERT for Self-supervised Learning of Language Representations.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.843455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.843455Z digest=sha256:ab5b9b81330b899513c220e96df43300ea2f8495c1eccddd80753e5c56fb595c

Observation 022ef8eb-2881-47dd-8cf3-a1b946ccdd12 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.847663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.847663Z digest=sha256:be37bd06010b2c2a7d2ebe3f405723457a24c551b7014abab959d004a62bb116

Observation f37dcbe8-5666-47ea-9135-c094306895a8 · outbound

This paper cites Gpt-4 is here: what scientists think,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Gpt-4 is here: what scientists think,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.851308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.851308Z digest=sha256:bb06a9c7d2cb4c07a5f58cdbbfdf98a364f17331a8a88a44e5dc29626d903ffb

Observation 7d9c7891-7d51-4166-a825-b38212f263f3 · outbound

This paper cites Language mod- els are few-shot learners,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Language mod- els are few-shot learners,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.854998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.854998Z digest=sha256:02cbc9e62235a4bd017fcf05360fb437a9a8da6072dee8808a4efe50274f1329

Observation 5296042b-7465-4dcb-acdb-81f0fe42a09e · outbound

This paper cites Multimodal deep learning.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Multimodal deep learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.858659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.858659Z digest=sha256:d216f444881e1ae407deaddc23cff0b69ee744bf74e9bf1e533fddf2aaa73563

Observation 22078948-1812-47f7-aa95-707bf650affd · outbound

This paper cites Multimodal machine learning: A survey and taxonomy,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Multimodal machine learning: A survey and taxonomy,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.862350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.862350Z digest=sha256:21f0519276be9f359db2c6d4bc4e6798c2af6ee099f6cbd6d7c6fa93ad082766

Observation 81b85cd3-5be3-4683-a409-c9a9637d77f2 · outbound

This paper cites Multimodal transformer for unaligned multimodal language sequences,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Multimodal transformer for unaligned multimodal language sequences,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.866015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.866015Z digest=sha256:9e0762fe23acb1d744ba882f7572faa12b7044dcf6990ceb12dcd3d6f81ba209

Observation c9922b25-8d4c-44a1-a3b3-5ced6bfeb68f · outbound

This paper cites GPT-4 Technical Report.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques GPT-4 Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.869757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.869757Z digest=sha256:acd63b6d0dfe2e59d71520ac005eb926aba720f74caab906993de28cc3cdfb8d

Observation afed7e32-5d07-4dbe-837e-0300bfc2d845 · outbound

This paper cites Crossner: Evaluating cross-domain named entity recognition,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Crossner: Evaluating cross-domain named entity recognition,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.873536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.873536Z digest=sha256:369d1c991ea6476d39fb9b775b8e3adfc185320bb95d1d6c0c4abf22ecf87681

Observation 770a5b30-5cab-4377-97c3-9a1bac24c3d7 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Learning transferable visual models from natural language supervision,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.877137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.877137Z digest=sha256:7209b0ecbde7b7f09928f697cea85e0b12907373c4fa2ef19f1e59f155d2e46d

Observation 087f0a8a-8a25-44d6-bc85-4db173ac2277 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques On the Opportunities and Risks of Foundation Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.880936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.880936Z digest=sha256:6b495abdedf1a513d9a88d8125212071283c2edf24a9d3acd11989a1a7af98b6

Observation bade9839-b308-4cfe-bf10-2a8dc2d847d5 · outbound

This paper cites Xtext language engineering for everyone,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Xtext language engineering for everyone,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.885110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.885110Z digest=sha256:d0f8df63293dc0a8b9940a105045562768d17dbb47e57332ccd3d38371b3ddcb

Observation ec150ed6-967c-40c8-a0fa-2959ee49851e · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Align before fuse: Vision and language representation learning with momentum distillation,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.888633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.888633Z digest=sha256:99093b8dee14d5d1ce44cc5a41b3174a17b15cca8d76c925ab55366796b5a9cf

Observation 1a5c4180-e441-4f2a-9a48-1367d5f4289d · outbound

This paper cites Training language models to follow instructions with human feedback,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Training language models to follow instructions with human feedback,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.892388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.892388Z digest=sha256:62074d3c1256f400207b7b46dc6f6c87c12f1264b3195467e0be7c76a79cdf89

Observation 3b22685c-af14-40c0-8907-133b90e7a683 · outbound

This paper cites Hymba: A Hybrid-head Architecture for Small Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Hymba: A Hybrid-head Architecture for Small Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.896718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.896718Z digest=sha256:9e30b31659a5f46e93830857a6514d03736755c309178b303bfc2d3c8ced6ec1

Observation b3a4de7e-c2fd-4b95-a2d4-19d0257026f4 · outbound

This paper cites A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.900912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.900912Z digest=sha256:4175a0e071355f2b5e5527324f71af0249dc455b172407418edb556eb1c63310

Observation b66c68af-5e7d-4b6a-b49d-f98e1538b848 · outbound

This paper cites Small Language Models: Survey, Measurements, and Insights.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Small Language Models: Survey, Measurements, and Insights

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.904953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.904953Z digest=sha256:4c58ad91e1a6cdab93da0f37bb756c892d5ecdc499cbea7d8ad27270f6058c65

Observation ec4eb572-b1a4-4129-b4ec-aee6e84b0c0c · outbound

This paper cites Hallucination is Inevitable: An Innate Limitation of Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Hallucination is Inevitable: An Innate Limitation of Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.908957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.908957Z digest=sha256:ace8b9efcff02d4600aa28ebadc06fa7c14bf26f22f11294d8fb68cc6016467b

Observation d1c0e700-7c5e-4bb3-ad41-ee82a0f844d6 · outbound

This paper cites A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:58.074869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:56:56.912852Z digest=sha256:a9ad241460dd0d4420c7db1a678f0235eb8e8b55fa406d502deb050842c6bedb

Observation 0edc626f-8814-43c8-b43b-70c8fc73aec6 · outbound

This paper cites Towards trustworthy llms: a review on debiasing and dehallucinating in large language models,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Towards trustworthy llms: a review on debiasing and dehallucinating in large language models,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:58.062327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:56:56.916756Z digest=sha256:922a6222be9959a9c9526473877ad7b2411cd8917b0285bfb68959f329cdaf41

Observation c7ad77f5-909c-45ae-9d6b-93fce6f8cad5 · outbound

This paper cites Sometimes painful but certainly promising: Feasibility and trade-offs of language model inference at the edge,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Sometimes painful but certainly promising: Feasibility and trade-offs of language model inference at the edge,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.920457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.920457Z digest=sha256:5f9ef23a84a60e66ac2251135837a6400066f9c0c4c662a649449b64cd792e64

Observation 5b153ab8-f3a6-4034-be29-cb7af28c0ff2 · outbound

This paper cites Tryage: Real-time, intelligent Routing of User Prompts to Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Tryage: Real-time, intelligent Routing of User Prompts to Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.924187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.924187Z digest=sha256:7234feefd9066088359947b1b7e11604325bc112a07810eaa64a7dfcf1d236c6

Observation 2763b83c-da63-41cd-9755-02619710b949 · outbound

This paper cites Fly-swat or cannon? cost- effective language model choice via meta-modeling,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Fly-swat or cannon? cost- effective language model choice via meta-modeling,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:58.050512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:56:56.928196Z digest=sha256:33a2545d183fbe28c94c7790106b4c38645d34ad5bb476b8eb742e89c8ed616c

Observation 3eb1df40-ec5d-4c25-9ffe-fb80b44b4a74 · outbound

This paper cites Routoo: Learning to Route to Large Language Models Effectively.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Routoo: Learning to Route to Large Language Models Effectively

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.932292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.932292Z digest=sha256:3d6ad6c231fe5b6204c68a8921750c05e9af05df60d7cba0f82a7b8857c82a0f

Observation 369e619b-111c-4d37-9083-b7527010f245 · outbound

This paper cites Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.936401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.936401Z digest=sha256:eca147752df18a5131f577431fea52197060647e072b8d4af01bcbc763a6cb53

Observation df27739d-3704-40bb-8686-deff835a0c54 · outbound

This paper cites OptLLM: Optimal Assignment of Queries to Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques OptLLM: Optimal Assignment of Queries to Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.940328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.940328Z digest=sha256:62f1b7fee883446901f83f0e65587c0940eb0065859741a081d9d15315500824

Observation 0c9c37a1-40eb-474b-8550-5c81a28095c2 · outbound

This paper cites Chatgpt,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Chatgpt,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:58.039550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:56:56.944254Z digest=sha256:314ddcdc742f79f18f5e639e3a2335c74e155ee6648083aa3fb46057d7ff38ff

Observation 42d3dc4a-2a64-41c9-99a2-6b0499547c30 · outbound

This paper cites Amazon titan in amazon bedrock,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Amazon titan in amazon bedrock,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:58.028097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:56:56.947839Z digest=sha256:39e4b6f8aa04263c0bac71d48c23d92293896f156b0e1e6d82554ad3a2342082

Observation f221f8e4-7d12-4957-84a7-c3924a8fdfe8 · outbound

This paper cites an unresolved cited work.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:56:58.017464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:56:56.951394Z digest=sha256:b63ca7c0e81e16e518cbfc2b7a9a9d2783f1b6a9aa03d80ec62aae9baa48bb5b

Observation b1854396-ee05-40c7-bc01-4155a4feb7b4 · outbound

This paper cites an unresolved cited work.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:56:58.007120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:56:56.955137Z digest=sha256:a10421ab1b4f51d4098feef91e9c04cf0d2a04ec5540c4bbd9ce2e5844afb80c

Observation 03951d37-536d-408e-baeb-00a54b6df901 · outbound

This paper cites RouteLLM: Learning to Route LLMs with Preference Data.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques RouteLLM: Learning to Route LLMs with Preference Data

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.958761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.958761Z digest=sha256:00c87d2813c0211d53c3f0c83fe446a8a6db80166e1b1dabc4e81d5e19cc80f6

Observation ddbcf273-44eb-4338-99b8-695ce6e04f59 · outbound

This paper cites Cache & Distil: Optimising API Calls to Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Cache & Distil: Optimising API Calls to Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.962993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.962993Z digest=sha256:90db7f2f7f7b5348a8b672ebb4e3e8de401b6e4f5731dc305a9b7d17f749a095

Observation fd6524a2-e721-4cb8-8606-52316e8eb167 · outbound

This paper cites AutoMix: Automatically Mixing Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques AutoMix: Automatically Mixing Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.966738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.966738Z digest=sha256:1cb705bad7ceca667f86b54538243480ff0cbcd64944792b2286ab1aca4a8aef

Observation 6547ef9b-67bb-4c49-a87f-454fe5b60ac1 · outbound

This paper cites Efficient Hybrid Inference for LLMs: Reward-Based Token Modelling with Selective Cloud Assistance.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Efficient Hybrid Inference for LLMs: Reward-Based Token Modelling with Selective Cloud Assistance

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.970762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.970762Z digest=sha256:4e6da796c64ae9ec482c435ed1bc3f8be15558832298bfb0723100e1d5103ade

Observation 41f0d6a2-dacd-4685-9ca5-421c3cbd558c · outbound

This paper cites Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.974688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.974688Z digest=sha256:279e71218e7367239e6f2156d18fc8c873afefb1c09e4270119ba4a369ebd3f4

Observation 39d584d9-4d79-42ad-a834-f6714631ce83 · outbound

This paper cites Query by committee,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Query by committee,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.978614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.978614Z digest=sha256:7c6c787750e6c4c9eabe644777b8902874f47f23414e438dc70aca838b78d825

Observation bc603551-9d84-4c2e-8cea-a0e8cf2a4c24 · outbound

This paper cites A review on edge large language models: Design, execution, and applications,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques A review on edge large language models: Design, execution, and applications,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.990897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:56:56.982111Z digest=sha256:02b1e983fe356328ffa87271c1a94e23b257202510d894deebbc02e15d977ace

Observation 92755cfd-9dbb-452b-be1b-aa346eb0f90f · outbound

This paper cites On-Device Language Models: A Comprehensive Review.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques On-Device Language Models: A Comprehensive Review

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.985768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.985768Z digest=sha256:a3f386c1393f5c6486a60aaa1db4e25835b286da88346895e8665a610e223b00

Observation b897ca20-8400-4e57-94b6-ba9e268f6016 · outbound

This paper cites LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.989615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.989615Z digest=sha256:3f830cc9898982e4b8c299be0d5daf5dfcdd3bb3250f2970352169b7f2c20992

Observation 48890071-a8b5-4073-b2c3-b0e37f358712 · outbound

This paper cites Research.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Research

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.980441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:56:56.993763Z digest=sha256:967d7820b46193cb1b9c991a91d0efe76d252f03064318fbf80e41bd8dc4268c

Observation 6d0c202c-d3d9-4b8f-aace-aa73d8d9d901 · outbound

This paper cites an unresolved cited work.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:56:57.970312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:56:56.997453Z digest=sha256:8e7c3f57fa62b6ecaea1e9399f7eba086c8d5df90aaf389cc4f1bfa2c011bc93

Observation b853f9ce-101b-432a-bc3d-cfbee91bce19 · outbound

This paper cites Edge-first language model inference: Mod- els, metrics, and tradeoffs,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Edge-first language model inference: Mod- els, metrics, and tradeoffs,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.959881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:56:57.001152Z digest=sha256:3f780729d19be034d698a19697b49e0a67aa9cbb8851376eb2478057d9bb8bb7

Observation 47546da5-a959-4dd8-ae59-e3497242d1d4 · outbound

This paper cites UltraEval: A Lightweight Platform for Flexible and Comprehensive Evaluation for LLMs.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques UltraEval: A Lightweight Platform for Flexible and Comprehensive Evaluation for LLMs

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.004691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.004691Z digest=sha256:6db1c0b7363a32e0736e87d530e9a56240d5dff13c8ccce4a5dad849452d18fa

Observation f2707b7c-9c27-42ad-9a28-db59fb148ffb · outbound

This paper cites Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.008858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.008858Z digest=sha256:729e3cda44fe77d8790d074adef7fdb4199d0c5f0d3a6bab0513a464e8f85055

Observation f8fee592-837e-4d70-81fa-7623c76188d6 · outbound

This paper cites MM-LLMs: Recent Advances in MultiModal Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques MM-LLMs: Recent Advances in MultiModal Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.013028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.013028Z digest=sha256:f8379d2e824cb7534999bf92b0689b523663559f1fc67c3632d2212f9b271e1c

Observation 0a64c309-4bed-4ea1-9647-9e54130324de · outbound

This paper cites MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:56:57.208416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:56:57.016882Z digest=sha256:a5452f2c7eb470867e1f5a937de2c3d43d811d6df94c82a7bf0cc700026cbf85

Observation 0d5ee2e9-8d16-4ba5-9f93-528ee93467ee · outbound

This paper cites Multimodal large language models in health care: applications, challenges, and future outlook,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Multimodal large language models in health care: applications, challenges, and future outlook,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.949512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:56:57.020774Z digest=sha256:5c66628ac0616be188922b9a9493ca134dd1693922b7cfa25622b743eff2d574

Observation 37db8c79-d54a-49f6-9a13-f52a9e9359da · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.939194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:56:57.024727Z digest=sha256:214fbf88c1d323a3fe2b01d9cb5bd25388771bc84d5725a2366ee5607fcafc5f

Observation 8356649b-3d40-4c5e-8c39-f0310af58151 · outbound

This paper cites Lexical complexity predic- tion: An overview,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Lexical complexity predic- tion: An overview,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.927904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:56:57.028587Z digest=sha256:69bed6cd70511ab1d45ea194bff85d956b3b71179e80e15308e2757213d7e196

Observation 45265c53-9cc2-4c58-ad15-7189185fc78c · outbound

This paper cites Collaborative cross- modal fusion with large language model for recommendation,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Collaborative cross- modal fusion with large language model for recommendation,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.916891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:56:57.032260Z digest=sha256:398f9cb1154b27773959ae0261a8d7caeec809ffc353b3a967d135eccd895f33

Observation 639313a1-8624-4d86-bd63-0c21eb2fcc4b · outbound

This paper cites Communication-Efficient Distributed On-Device LLM Inference Over Wireless Networks.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Communication-Efficient Distributed On-Device LLM Inference Over Wireless Networks

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.036167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.036167Z digest=sha256:76be39e64c82a5332eac000636bf812d50c2c9fc53c76e0eedabba20363dddc8

Observation f9cf4371-264d-46b4-9bf7-326b519ac0a1 · outbound

This paper cites Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.040284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.040284Z digest=sha256:a7099f3b55d0a6affb59eed21fb2b56a448f035bb023019d2f18ad5eb8f276ce

Observation 5fa027e9-05a7-4355-bc66-3c02c9e9534c · outbound

This paper cites Privacy-preserved LLM Cascade via CoT-enhanced Policy Learning.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Privacy-preserved LLM Cascade via CoT-enhanced Policy Learning

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.044447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.044447Z digest=sha256:60c27a20d28c5337b95eccd8a482569e024adf9c68bc00d820c79397c4dc7c12

Observation 19bbad48-bb75-497c-ae97-ff936e166272 · outbound

This paper cites Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.048415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.048415Z digest=sha256:e8ff6184e62adfb374047023819a4ad4943dfeda20e2fbd23104aedd71e01b75

Observation ccf21cc0-c770-4c37-83a8-322785e006a9 · outbound

This paper cites XAI meets LLMs: A Survey of the Relation between Explainable AI and Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques XAI meets LLMs: A Survey of the Relation between Explainable AI and Large Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.052089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.052089Z digest=sha256:156ab72cf2c76be795e1a06ca2e616b9c066a272b9b5d9905d97dae36fa7b9d1

Observation 65439416-3169-4241-99f4-6904b9f83c1c · outbound

This paper cites PickLLM: Context-Aware RL-Assisted Large Language Model Routing.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques PickLLM: Context-Aware RL-Assisted Large Language Model Routing

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.056470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.056470Z digest=sha256:b35f027a5e37bf2ac8cb72caf110ddfc6219b31465edc7fba8125b8dd8281826

Observation 2c26cbc3-633d-40a4-8602-45546a00ee96 · outbound

This paper cites Text Sanitization Beyond Specific Domains: Zero-Shot Redaction & Substitution with Large Language Models.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Text Sanitization Beyond Specific Domains: Zero-Shot Redaction & Substitution with Large Language Models

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:56:57.138941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:56:57.062055Z digest=sha256:4344ba4e5732e886ddfa4e40953f3b0546249d7a65968078cf81b9ea53d78808

Observation f6fcb862-cb94-4226-a586-8dd499f3a109 · outbound

This paper cites Trusted llm inference on the edge with smart contracts,.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Trusted llm inference on the edge with smart contracts,

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:56:57.906148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:56:57.066381Z digest=sha256:d0aff3c31395942e14c955b6473ff78fb2860cda76a8f41ac5c573adbfbb16c9

Observation 055c6e2d-3fc3-4bae-b046-10296140e601 · outbound

This paper cites Encryption-Friendly LLM Architecture.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Encryption-Friendly LLM Architecture

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:57.070193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:57.070193Z digest=sha256:dfaa61eb1f806bb6c5a72cd00cb0cb9c890b13e7375e95619b56b9ec18a30db7

Pith citing papers

Observation 8bd8a001-206d-4097-95ca-4f1690881d27 · inbound

Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems cites this paper.

Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T19:10:53.372941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:10:53.372941Z digest=sha256:d679abae9aa8ca7bae8f702fe93386458abe139614884f39db8b1e09cd98c5bb

Observation 51c91ec7-75e7-407b-bacb-1972153792d3 · inbound

AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives cites this paper.

AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T19:34:25.790205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:34:25.790205Z digest=sha256:2237b963e2bb9c6affd3a10d565b4f820e0b9cd63b04b9af3d366edd08832a92

Observation 744745d6-6957-4b28-bb15-51ab1b51eaec · inbound

When Less is Enough: Efficient Inference via Collaborative Reasoning cites this paper.

When Less is Enough: Efficient Inference via Collaborative Reasoning Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:41:42.531986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-09T19:27:04.267404Z digest=sha256:d76f759ca9d0b393f30051b4db991d42e4a4e696cb52623d68fa1eb47178a071

Observation 34674ce5-e34d-40c4-97ff-f856eab0a7c5 · inbound

SOMA: Efficient Multi-turn LLM Serving via Small Language Model cites this paper.

SOMA: Efficient Multi-turn LLM Serving via Small Language Model Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:42:04.221906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T01:37:27.361504Z digest=sha256:7f91eef5269ec544779bdd667f962e2d1cbe3031c720d15de1d71d6a37a9f59b

Observation 4f495410-efb5-4926-87f1-8dcf3557bdb5 · inbound

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces cites this paper.

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-06-27T22:31:21.605739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T22:22:52.690010Z digest=sha256:0cc8ad2efe1a49fc4def1c379aac13ad70949a0611196aa50277c340c87a0fb1

Observation 4fbda2fd-a0ef-45e3-b4ae-6be656850ee8 · inbound

From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing cites this paper.

From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

Reference 147

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:17:08.806454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T22:54:28.452796Z digest=sha256:31eabfd6835cd28c3c67091055b53c9fe2c6dcc43b08f178155a5a74773a2cc0