Pith. sign in

Paper Citation Record · LEDGER

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts

As of 22 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2508.00234.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.00234 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:22:42.600721Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact4
  • verified fuzzy24
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7efcac1-e101-4a8f-aee0-b89606f8e866 · outbound

This paper cites AIoT smart home via autonomous LLM agents,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts AIoT smart home via autonomous LLM agents,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.090189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.438667Z digest=sha256:4a1ed9e1563e57d860b9053b60b96c0d62d69565a49a6a4676905a4912865014

Observation 5becd7f5-7a18-4a0b-8b99-0a7ff6399765 · outbound

This paper cites Large language models for human- ai co-creation of robotic dance performances,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Large language models for human- ai co-creation of robotic dance performances,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.079977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.442311Z digest=sha256:9a49afa4f932834a73175f01eac74576afa80ff62f01666676e71d9231786905

Observation 66562e3f-b0ec-43cb-95a3-6d327ed826de · outbound

This paper cites EdgeFM: Leveraging foundation model for open-set learning on the edge,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts EdgeFM: Leveraging foundation model for open-set learning on the edge,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.070683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.445752Z digest=sha256:15af0f95b85a1c934b35e15835e61ad1afea445fb20dc7c5b9e4f36db1c57d70

Observation fa1a07e2-b807-4ec9-afdd-780f71e7630e · outbound

This paper cites WDMoE: Wireless Distributed Large Language Models with Mixture of Experts.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts WDMoE: Wireless Distributed Large Language Models with Mixture of Experts

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:22:42.849405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.449489Z digest=sha256:85a0e72c8e7c907fc6bf2bda7d9fa66c988002ab21eacd477aaf9f93a019ad2c

Observation e29751e0-7f8d-4110-a95a-ee4b27e27192 · outbound

This paper cites On Protecting the Data Privacy of Large Language Models (LLMs): A Survey.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts On Protecting the Data Privacy of Large Language Models (LLMs): A Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.453414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.453414Z digest=sha256:ed9e62055c14d4b4a3028a411c56884842f6571f91d6619bd31f92c7fcd06487

Observation 94d46d84-12c1-4b01-9fc2-9714dd1b58c2 · outbound

This paper cites Edge intelligence: Paving the last mile of artificial intelligence with edge computing,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Edge intelligence: Paving the last mile of artificial intelligence with edge computing,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.061734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.457777Z digest=sha256:60cb49bc36af6a34ba75f9ae1914085dee27f52442ac1f3a8a3480c38b8bb043

Observation 057fe7f9-4c30-4db0-bcf3-867bc48a5d07 · outbound

This paper cites Enabling AI-Generated Content (AIGC) Services in Wireless Edge Networks.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Enabling AI-Generated Content (AIGC) Services in Wireless Edge Networks

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:22:42.826912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.461675Z digest=sha256:1205574166bde9483500568ee5468183bea29f868839775de7e748693f79a739

Observation 178e19ba-07ac-4d24-ae51-acba3c2955d2 · outbound

This paper cites Toward Scalable Generative AI via Mixture of Experts in Mobile Edge Networks.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Toward Scalable Generative AI via Mixture of Experts in Mobile Edge Networks

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:22:42.812607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.465930Z digest=sha256:799aa65b303cded59840026cb9e462096b6f3502903ce7ee88a2ce854a8fcc7a

Observation a15f1f24-6d08-4389-8ebc-a4dcd44239c4 · outbound

This paper cites LLM-Blender: Ensembling large language models with pairwise ranking and generative fusion,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts LLM-Blender: Ensembling large language models with pairwise ranking and generative fusion,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.052200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.469627Z digest=sha256:4aaf1675940a1bfb992b0df007da158fefc6d6e5d4bd40898a62ab240d8ee007

Observation c1324dbe-b74d-4485-a844-e9374342ebe8 · outbound

This paper cites Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.473633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.473633Z digest=sha256:402130128c946ff89d03922fca0da322d6dc0de8923e3e5f7abbf9de2abca2cb

Observation 5851814e-fe47-4e49-9f96-335532db538d · outbound

This paper cites Orca: A distributed serving system for Transformer-based generative models,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Orca: A distributed serving system for Transformer-based generative models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.042587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.477410Z digest=sha256:3e30e9561cbc9d03fd157c8eb2bf2c035590c48a9fc8bbf6aa66df1d91851fe5

Observation f57748b3-8b08-4739-8117-89915ee233ad · outbound

This paper cites Efficient memory management for large language model serving with PagedAttention,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Efficient memory management for large language model serving with PagedAttention,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.033027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.480957Z digest=sha256:bd0a0a1444678a17d13da2346ac0cc14213e704c92e7cdf284f137ba988d08ae

Observation d6fb5336-512d-48cd-8483-794d817d97f1 · outbound

This paper cites TensorOpera Router: A Multi-Model Router for Efficient LLM Inference.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts TensorOpera Router: A Multi-Model Router for Efficient LLM Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.484238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.484238Z digest=sha256:3ea448e1b5a49c9de48008e7ff0171ad187ed2de2daf97c5afb03dc0537fd96a

Observation c3c4b001-6546-45ba-b1e2-86507aa5b3cd · outbound

This paper cites Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.487916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.487916Z digest=sha256:015b890be96456f621fd57f0ba0388f05b92cee4c916f4c2312e39ec1aa2a198

Observation 5d215f12-f43b-4d2c-be98-bc4d39c890e6 · outbound

This paper cites Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.491345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.491345Z digest=sha256:50e5ece3cded47df01ba30aec13f04b11e5d86c78ba3173a3537a5c1c4df6777

Observation 9e1a3c60-1439-493e-8f61-0320821726da · outbound

This paper cites RouteLLM: Learning to Route LLMs with Preference Data.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts RouteLLM: Learning to Route LLMs with Preference Data

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.495457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.495457Z digest=sha256:d2233a10adda32958ac995354fb989be8350e8f57852f68b82d3e8f1400c5615

Observation c8f02fcc-0387-4b39-9431-998ed626a0a2 · outbound

This paper cites BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.499001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.499001Z digest=sha256:6c525f41c0b501d504a05d2ee53324e28947540a0afa4b93460cac5e247d1064

Observation 6f8399dc-ff72-4234-bc1e-d302dff8fba4 · outbound

This paper cites Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.502527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.502527Z digest=sha256:3a5761de389e0bd8cf10b3a6281c1aced2330c039272b8fee6dc027f6e44dd40

Observation 5feb3bae-8f4e-4fc5-a454-b3d2bd10255d · outbound

This paper cites S3: Increasing gpu utiliza- tion during generative inference for higher throughput,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts S3: Increasing gpu utiliza- tion during generative inference for higher throughput,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.022410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.506125Z digest=sha256:77d2d1bd70652e33663ba3f4f4fce990d2a24fe0618b4f6d82dcc7191d77f2ca

Observation ff93db18-1f56-45f4-af19-a093d6a6f59e · outbound

This paper cites FlexGen: High-throughput generative inference of large language models with a single gpu,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts FlexGen: High-throughput generative inference of large language models with a single gpu,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.012284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.509785Z digest=sha256:3e47e64c3e65f72bf081623a12f2c61c7f4ec565fcbe2eb524ddc2baf4be452b

Observation 31825be3-8091-4d3c-942e-620e5bc4d5ab · outbound

This paper cites FlashAttention: Fast and memory-efficient exact attention with io-awareness,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts FlashAttention: Fast and memory-efficient exact attention with io-awareness,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:43.002330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.512936Z digest=sha256:d1a1bec41908ff635843707260d6b1cce827067a1d50a3093520e9340328b8a8

Observation 8e83768f-be94-4e1b-9b39-1bfced14ec38 · outbound

This paper cites HeteGen: Efficient heterogeneous parallel inference for large language models on resource-constrained devices,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts HeteGen: Efficient heterogeneous parallel inference for large language models on resource-constrained devices,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.992601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.516096Z digest=sha256:bd65a97a2791178f9928f173d8622d567caefb3feb8b956a4b9389603a6d4163

Observation 6a1aaa9b-af2a-45c9-952d-ed89d8481d74 · outbound

This paper cites ExeGPT: Constraint-aware resource scheduling for LLM inference,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts ExeGPT: Constraint-aware resource scheduling for LLM inference,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.983278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.519086Z digest=sha256:620b8e84ed1986bc5781a307c72de9e376f4b41daca109cf659a96f1aa7f3d03

Observation 120312c5-8335-48a5-8248-6cb87de67a0f · outbound

This paper cites Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.522976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.522976Z digest=sha256:9774f4a95a249d399578086fe123eec990a6fa8a0a0575e793ab80158ca20fba

Observation ed792d59-abd2-4839-bdd1-2093adadd75e · outbound

This paper cites Splitwise: Efficient generative LLM inference using phase splitting,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Splitwise: Efficient generative LLM inference using phase splitting,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.971972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.526269Z digest=sha256:98dd6e85c749b58b7f67e2303d3994bfc441dd9d98e23f1798712ce3313f0262

Observation 99811022-c3ff-47f0-adb3-784c8973e78f · outbound

This paper cites DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.529491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.529491Z digest=sha256:c60dc44960ca02882cbe9a6d844fa6ea4169683a0c1498ecb89828d6cdf978c6

Observation 29ca24c1-7612-4688-bc65-4af530a53ced · outbound

This paper cites Merge, Ensemble, and Cooperate! A Survey on Collaborative Strategies in the Era of Large Language Models.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Merge, Ensemble, and Cooperate! A Survey on Collaborative Strategies in the Era of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.532900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.532900Z digest=sha256:11abae8183e0bc94eaf5cf1fd0f3966fbb98a50b2d9e5dc00b089912cb8844a8

Observation 972eceb0-f50d-4179-990e-02817071252e · outbound

This paper cites Large Language Model Routing with Benchmark Datasets.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Large Language Model Routing with Benchmark Datasets

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.536535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.536535Z digest=sha256:bd2d7aa502c165c865269f463f60b9ba47db7e15b70a00272b7603f9b7d08751

Observation 907191d6-53e4-470b-a88a-964cba56d57b · outbound

This paper cites Octopus v4: Graph of language models.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Octopus v4: Graph of language models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:22:42.699685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.540098Z digest=sha256:f1c4d60c2617fae0b44611d48e7203b55b20e6ce5cba906f2052673690813be8

Observation 959d65b7-5550-4a95-9b9c-bbbd9e40ae62 · outbound

This paper cites GraphRouter: A Graph-based Router for LLM Selections.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts GraphRouter: A Graph-based Router for LLM Selections

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.543528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.543528Z digest=sha256:56f22094707178a00bc450c5894eeca82a00a0c0d018f99279368bbd46616cd0

Observation 6f7900a0-9cbd-40a5-9bfd-a4e84da2a1da · outbound

This paper cites Eagle: Efficient Training-Free Router for Multi-LLM Inference.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Eagle: Efficient Training-Free Router for Multi-LLM Inference

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.547078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.547078Z digest=sha256:4a6f8cd61bdae720cde00861f118265c5d4177be2c55acc0a4d9124aa85d2887

Observation f118e507-38d5-4f06-9c12-fb9cc61baa17 · outbound

This paper cites RouterBench: A Benchmark for Multi-LLM Routing System.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts RouterBench: A Benchmark for Multi-LLM Routing System

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.551092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.551092Z digest=sha256:f08c707c4926ff98d05c6368b0925249627d4bd105626e8fe079161a4b43d24a

Observation 00649949-c493-4f18-b74b-9b6853a6aac3 · outbound

This paper cites Reinforcement learning in dynamic task scheduling: A review,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Reinforcement learning in dynamic task scheduling: A review,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.962721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.554729Z digest=sha256:6a33279d2cc59d7e4735bf1f12a0444059a05d1448ab888b678065a321206fd7

Observation 8ec0fde6-11b0-46de-aa00-7d53d3a016aa · outbound

This paper cites Collaborative learning-based scheduling for kubernetes-oriented edge-cloud network,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Collaborative learning-based scheduling for kubernetes-oriented edge-cloud network,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.953572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.558272Z digest=sha256:cd9cc2b940d46d9164af4939068733d1d07fab4a6c5613f9d72bb72f8b4b20ab

Observation 8f2438b1-a10b-460d-9795-efc97d7fd20e · outbound

This paper cites Clipper: A low-latency online prediction serving system,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Clipper: A low-latency online prediction serving system,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.943669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.561384Z digest=sha256:30111d4b37c53b38133d059abd156513c0298e931e6768b31d1a80bbe881bcb5

Observation 4fc7aa00-d5ce-4947-bc46-6a2ca3ac062c · outbound

This paper cites Tapfinger: Task place- ment and fine-grained resource allocation for edge machine learning,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Tapfinger: Task place- ment and fine-grained resource allocation for edge machine learning,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.933924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.564545Z digest=sha256:7364755510cd842e6eeea01fe80988b4893ffc831a3fe2d70238ca289163824a

Observation 805502f2-61a6-4900-91c7-df0b40ef9c12 · outbound

This paper cites The non- stochastic multiarmed bandit problem,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts The non- stochastic multiarmed bandit problem,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.923358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.567752Z digest=sha256:34c92b0b77d259d5da0a1623ffe1a63231f94b7b0b95d984f6f01d26900aea4a

Observation aeecd7cb-b573-4c6d-86db-3b8f5be55375 · outbound

This paper cites Alpaca: A strong, replicable instruction- following model,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Alpaca: A strong, replicable instruction- following model,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.913571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.570832Z digest=sha256:ea0e6fa82abdcce864b04c86ae7e51bc2b47e63d554ce12dbaf0f01a396ec6eb

Observation 5ee694a6-7663-4955-b1ba-ba5df4dc6475 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.573898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.573898Z digest=sha256:2fbac40a3330e21a07a01c07303fc45592a1da2442a5a7267d2c447f7f1a9a44

Observation e6446bf0-638f-4d86-8d3b-14769f76bdcb · outbound

This paper cites Introducing Mpt-7b: A new standard for open-source, commercially usable LLMs, 2023,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Introducing Mpt-7b: A new standard for open-source, commercially usable LLMs, 2023,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.903359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.577353Z digest=sha256:a726016772596be57050f44611d1b3131e9f09741269221b298b5780e0e877e9

Observation ccc560dd-3b4a-4490-ae35-6b60998217b6 · outbound

This paper cites BERTScore: Evaluating text generation with BERT,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts BERTScore: Evaluating text generation with BERT,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.893031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.580526Z digest=sha256:3a46a850bc7de2e12b118d9f2912b8d4ce86bb338efbcc1d41667efc76e4b78f

Observation 06130ea9-117f-42d2-958d-ebcc97f7adf5 · outbound

This paper cites Soft Actor-Critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Soft Actor-Critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.882580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.583787Z digest=sha256:f89ed4d1df09d85d5652830105f74ec1262836e31aa1104ba9ea9f1459fb2485

Observation 218b2815-e419-46b1-8d47-d6dd83cffc48 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.586945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.586945Z digest=sha256:e4827171b8cdbd37b050db32fd4961db544ad473c883f094f2f643cbd1c674bf

Observation 349b3954-1ab8-4e79-93c3-68053e355c59 · outbound

This paper cites PyTorch: An im- perative style, high-performance deep learning library,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts PyTorch: An im- perative style, high-performance deep learning library,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.872062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.590788Z digest=sha256:d2b7a79ec79727b133bd351e0eeeac509e9b8f786304e0bcc63a2af14f30d0cf

Observation c3e0348b-b300-4653-b8be-43de4843abe8 · outbound

This paper cites TorchRL: A data-driven decision-making library for PyTorch.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts TorchRL: A data-driven decision-making library for PyTorch

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.593937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.593937Z digest=sha256:3cd03630c822693a1853911899d2d931c0ad13eb2875b32301aa477235bba3ec

Observation 744176ad-faf2-4bd9-8275-e68d4e776812 · outbound

This paper cites Fast Graph Representation Learning with PyTorch Geometric.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Fast Graph Representation Learning with PyTorch Geometric

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.597363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.597363Z digest=sha256:5c176fbbba509f46132ec529e9c96f4d8657f3ed309a8ff357e8c4e27a7e4d03

Observation 9f7890c3-0601-4b7c-935e-fca2d8f40366 · outbound

This paper cites BERT: Pre- training of deep bidirectional transformers for language understanding,.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts BERT: Pre- training of deep bidirectional transformers for language understanding,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:22:42.860805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T10:22:42.600721Z digest=sha256:f76d34cdc97ca5d40a804cd4dd9f1c91e713364fd9008d4149c93bb015c648da

Pith citing papers

No inbound Pith citation observations are available.