Pith. sign in

Paper Citation Record · LEDGER

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving

As of 14 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 4 inbound Pith citation observations for arXiv:2512.04013.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.04013 v3

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T18:43:32.575076Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T23:16:46.161659Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T00:35:10.221606Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 67c303cd-ba22-4a7b-86ad-464690d13be6 · outbound

This paper cites Infercept: efficient intercept support for augmented large language model inference.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Infercept: efficient intercept support for augmented large language model inference

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:27.386587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:27.386587Z digest=sha256:9767b77bd2c234e7f31e4501dca348cdbcbe03ef3f87cbcff0595e7c95b8e8e6

Observation ff7a8620-12a5-494a-a017-a9093cbe60bc · outbound

This paper cites Taming {Throughput-Latency} tradeoff in {LLM} inference with {Sarathi-Serve}.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Taming {Throughput-Latency} tradeoff in {LLM} inference with {Sarathi-Serve}

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:27.512933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:27.512933Z digest=sha256:6c1bf8b9284c63048f7a55f65be2d331b042ce01e9ff11001b53476fdd983044

Observation fd4a6a13-d068-4778-b4e5-fc8fa69a3ca3 · outbound

This paper cites Revisiting service level objectives and system level metrics in large language model serving.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Revisiting service level objectives and system level metrics in large language model serving

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:27.628880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:27.628880Z digest=sha256:4f8003519da1a769bc58c4ab401c9a1a7f8ff0f914825142a2e10e667ab1b173

Observation e9316851-8943-49d3-9131-6e569b986d96 · outbound

This paper cites Model context protocol (mcp).

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Model context protocol (mcp)

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:27.736822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:27.736822Z digest=sha256:cd4ff49a95608fecee7ff582c2f804e3b31b975fe9cdf22903b2f2a149ee95d3

Observation f111267d-5bb3-4e45-a2ed-af70dde5d220 · outbound

This paper cites From good to great: Improving math reasoning with tool-augmented interleaf prompting.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving From good to great: Improving math reasoning with tool-augmented interleaf prompting

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:27.865578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:27.865578Z digest=sha256:b30d2de47e57054e7e735e82296a68536d86ddaa8c0dc0dca8bf9079b7d9275c

Observation ff33d1e1-b9f0-4e2a-a8d3-751b46dbf341 · outbound

This paper cites Advancing tool-augmented large language models: Inte- grating insights from errors in inference trees.Advances in Neural Information Processing Systems, 37:106555– 106581, 2024.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Advancing tool-augmented large language models: Inte- grating insights from errors in inference trees.Advances in Neural Information Processing Systems, 37:106555– 106581, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:28.132979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:28.132979Z digest=sha256:e9fe80251082e4314b4c078ac05aad5357c7730874973c3bd0bf392c30e94c04

Observation 2b363ce5-31f8-477c-b737-9ecb9df569cc · outbound

This paper cites ChatCoT: Tool-Augmented Chain-of-Thought Reasoning on Chat-based Large Language Models.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving ChatCoT: Tool-Augmented Chain-of-Thought Reasoning on Chat-based Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:28.220908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:28.220908Z digest=sha256:ef2f3089c7df981ab2a75446c59d386406eba6636a931b08bf58af80d50314eb

Observation 4f543e4a-108e-424c-a310-63199f4d8aa5 · outbound

This paper cites Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625– 630, 2024.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625– 630, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:28.358990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:28.358990Z digest=sha256:03fa749f650ad2fca8eeb17d28ba0573e2f0858c34f6baa91de67397bc44870f

Observation e0bfac4b-a9e2-4e46-b71d-eaf42f765e6d · outbound

This paper cites MCP-Zero: Active Tool Discovery for Autonomous LLM Agents.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving MCP-Zero: Active Tool Discovery for Autonomous LLM Agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:28.515792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:28.515792Z digest=sha256:ea24c104284eabc48fc83a0af3533f41e0d33a83692968f536cb82fe46e082d1

Observation 50c05bb9-0fe4-4866-bb0b-8cde0a358e20 · outbound

This paper cites Efficient llm scheduling by learning to rank.Advances in Neural Information Processing Systems, 37:59006–59029, 2024.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Efficient llm scheduling by learning to rank.Advances in Neural Information Processing Systems, 37:59006–59029, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:28.642831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:28.642831Z digest=sha256:69601ba370593e06726bce52fc4b0b5f8917de2f90f8cb083ace91d663e5edb1

Observation 066f641a-bbf7-4a02-a9f5-7308e1fa448d · outbound

This paper cites Jetcheva, and Hardi Trivedi.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Jetcheva, and Hardi Trivedi

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:28.718262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:28.718262Z digest=sha256:efd5afc9e59ed60e3151a00665109eaddee63f275efe3b6b1e658d350a60ae8a

Observation b7557ef2-f386-4d82-86f8-bcdf36f92130 · outbound

This paper cites Apt-serve: Adaptive request scheduling on hybrid cache for scalable llm inference serving.Proceedings of the ACM on Management of Data, 3(3):1–28, 2025.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Apt-serve: Adaptive request scheduling on hybrid cache for scalable llm inference serving.Proceedings of the ACM on Management of Data, 3(3):1–28, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:28.880808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:28.880808Z digest=sha256:1ebfb068d68c328db9e62d3329ecd8f6f9adb669e3f88819bb5874ac10a237b3

Observation cd7e3ae0-b221-4779-996c-64e3fa541e64 · outbound

This paper cites Asynchronous LLM Function Calling.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Asynchronous LLM Function Calling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:29.008241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:29.008241Z digest=sha256:0ba10dadbe524da4b4468ff6e7b9701bd98f3f2882a73a9f85f53e865aefe319

Observation 03f327b7-3238-4f4f-8c35-82e0b97cb622 · outbound

This paper cites A study on classi- fication based concurrent api calls and optimal model combination for tool augmented llms for ai agent.Sci- entific Reports, 15(1):20579, 2025.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving A study on classi- fication based concurrent api calls and optimal model combination for tool augmented llms for ai agent.Sci- entific Reports, 15(1):20579, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:29.138575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:29.138575Z digest=sha256:44c96b70bef2f07d4f910d2b66b1792ee9f33e77ab502c0117dbb1aa7d217863

Observation cf67f45b-3f9e-4fbf-8534-75b8eedb9a7c · outbound

This paper cites Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings.Advances in neural information processing systems, 36:45870–45894, 2023.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings.Advances in neural information processing systems, 36:45870–45894, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:29.281902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:29.281902Z digest=sha256:611e1a427b8e9a199e56f141e9c1a006fb74adc9b7ade59b7c99587685485fae

Observation 59bc1e72-68c0-48f4-a4f0-72f2ad15655a · outbound

This paper cites Shuffleinfer: Disaggregate llm inference for mixed downstream workloads.ACM Transactions on Architecture and Code Optimization, 2025.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Shuffleinfer: Disaggregate llm inference for mixed downstream workloads.ACM Transactions on Architecture and Code Optimization, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:29.481736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:29.481736Z digest=sha256:543346784a60dafc6afc88b086643e5896c1bb71ef8a8f26b39605157f5bce89

Observation f02c14a8-c891-432f-9f95-b01ee43d2528 · outbound

This paper cites Tightllm: Maximizing throughput for llm inference via adaptive offloading policy.IEEE Transactions on Computers, 2025.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Tightllm: Maximizing throughput for llm inference via adaptive offloading policy.IEEE Transactions on Computers, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:29.608669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:29.608669Z digest=sha256:4e7e8b93a254ecb36e1c31f854c467b63e750112054cbca7ff8ae3eee6188d79

Observation 24de9e02-25b6-469b-82f7-bc4c0ccd6b7b · outbound

This paper cites Accelerating llm serving for multi-turn dialogues with efficient resource management.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Accelerating llm serving for multi-turn dialogues with efficient resource management

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:29.714372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:29.714372Z digest=sha256:410f7dd6e7758839a348246993d57660f08ee5d0758a8e5556bd9c2d42f6d1f8

Observation a8d071d6-57ed-4f8f-8dd5-3ce9ab223eb4 · outbound

This paper cites S3: increasing gpu utilization during generative inference for higher throughput.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving S3: increasing gpu utilization during generative inference for higher throughput

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:29.817105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:29.817105Z digest=sha256:b22b21ead18f56399aedd365d39742331fa22d513043b1c6e0e617ec60af99a3

Observation c8832e6e-db31-4eb7-8d6e-95543d8339ec · outbound

This paper cites Optimizing goodput through sharing for batch analytics with deadlines.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Optimizing goodput through sharing for batch analytics with deadlines

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:29.929131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:29.929131Z digest=sha256:769a7fc57ea3b44a544943f69a324bdfcba4f6cbdea8887481f9414005175c06

Observation 314f4266-187f-4d0c-a260-22b32d2659c0 · outbound

This paper cites Efficient memory manage- ment for large language model serving with pagedatten- tion.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Efficient memory manage- ment for large language model serving with pagedatten- tion

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:30.046652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:30.046652Z digest=sha256:d3d21f2a3559b2990232425d7847b7daaeaf4ad91bb2bbc82c8ba764fcb0547a

Observation fc265919-a298-41fc-9975-27457dc032d4 · outbound

This paper cites GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:30.156406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:30.156406Z digest=sha256:d8ccc7ded34bfd865b9ee2767ea910ae5734814c067aca843a01f7b00f7c9a5c

Observation 17fea0fe-b551-4dd8-a24a-e3a6176f7ac5 · outbound

This paper cites Augmented Language Models: a Survey.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Augmented Language Models: a Survey

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:30.283410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:30.283410Z digest=sha256:fb42fdfc198036867f984861beebade5aeef54156f619e7996691b40b475dd94

Observation c2596146-f0f2-4e60-97d9-32de7bad0821 · outbound

This paper cites Introducing function calling in chatgpt.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Introducing function calling in chatgpt

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:30.405668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:30.405668Z digest=sha256:e29b756be3700f8b472a582cf7f0ccef269c31bac558daf46a20759877a80a48

Observation 63ab32b7-83fa-424d-880e-4ae40400bbf6 · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Splitwise: Efficient generative llm inference using phase splitting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:30.513540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:30.513540Z digest=sha256:249f7feae9ab7428bbb27ade799bccd70425e5c31966c1f0f7a0229526596b56

Observation d56e7b59-8ed3-4411-bdbd-5068a2497485 · outbound

This paper cites WebRL: Training LLM web agents via self-evolving online curriculum reinforcement learning.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving WebRL: Training LLM web agents via self-evolving online curriculum reinforcement learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:30.592405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:30.592405Z digest=sha256:51f7607b43254b7fedbac395cc3b12bc12813c0c3fc83b9e2e82f335cc815d89

Observation be7f0d62-8ea7-4f7c-b5be-da4706076009 · outbound

This paper cites Tool learning with foun- dation models.ACM Computing Surveys, 57(4):1–40, 2024.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Tool learning with foun- dation models.ACM Computing Surveys, 57(4):1–40, 2024

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:30.703365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:30.703365Z digest=sha256:38e8307d4f3b116d12fafc56f657bb6f03d758600f9d8db89e30cdcdf4db93d7

Observation a85e1ebb-75a7-461b-9321-65bb6eb50b70 · outbound

This paper cites ToolLLM: Facilitating large lan- guage models to master 16000+ real-world APIs.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving ToolLLM: Facilitating large lan- guage models to master 16000+ real-world APIs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:30.788942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:30.788942Z digest=sha256:79c24bb82f21f4ed923f826d41dc897468ac8994fe9bf8724bfc470bdb12f1c6

Observation b06a4ee4-7246-419b-afb6-62c7d28dbc26 · outbound

This paper cites Tool- former: Language models can teach themselves to use tools.Advances in Neural Information Processing Sys- tems, 36:68539–68551, 2023.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Tool- former: Language models can teach themselves to use tools.Advances in Neural Information Processing Sys- tems, 36:68539–68551, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:30.897508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:30.897508Z digest=sha256:e83719445d2ad7f06356a799b6139fbd2879505de9a6a3837fb18f758150e4b4

Observation 6650787b-9992-4f2d-9483-c76ed4992d1f · outbound

This paper cites DON’t STOP ME NOW: EMBEDDING BASED SCHEDULING FOR LLMS.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving DON’t STOP ME NOW: EMBEDDING BASED SCHEDULING FOR LLMS

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:31.007239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:31.007239Z digest=sha256:0f6b825c835666562147ed731a8cde84789a45cce55177c3aea91c61327d7b88

Observation c4333ffe-36c9-4718-8fcf-4bca5319d1ed · outbound

This paper cites Fast Inference for Augmented Large Language Models.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Fast Inference for Augmented Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:31.141452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:31.141452Z digest=sha256:f07913f1e936ecd24a360a59158921b42a0e639ba8cc94dd17e3fcafede5c3e3

Observation c11fd6a3-c04c-497f-8dd2-a1f9f578d4e2 · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Flexgen: High-throughput generative inference of large language models with a single gpu

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:31.288578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:31.288578Z digest=sha256:22eb4a1708719c9bd108f92f802c06c424813a322d778c41a952d7750d72d2a7

Observation a57485fd-c763-4d24-aca8-c80aab7ebf60 · outbound

This paper cites DRAGIN: Dynamic Retrieval Augmented Generation based on the Information Needs of Large Language Models.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving DRAGIN: Dynamic Retrieval Augmented Generation based on the Information Needs of Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:31.398999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:31.398999Z digest=sha256:2904286cf26a81a4e7675f69f10269bcbdede1e2060c7ccff4bedca00601f88e

Observation 7aa51e48-48a6-48e3-bb39-5646e1755402 · outbound

This paper cites Wang and A.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Wang and A

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:31.543812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:31.543812Z digest=sha256:819433df199966b10e077f1e0bb41e4afa7e66a0e4203f5883c972a7c7e00897

Observation 82bd4dbd-0214-47f3-b765-0934590afe9b · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Fast Distributed Inference Serving for Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:31.712787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:31.712787Z digest=sha256:461b9e8bc11602148b2af3b0c467a298a3efac324a1ca3ec6462d1f42919ccc8

Observation 656bd126-cf33-4f5a-968f-a9d99d15cf4a · outbound

This paper cites A Toolbox, Not a Hammer -- Multi-TAG: Scaling Math Reasoning with Multi-Tool Aggregation.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving A Toolbox, Not a Hammer -- Multi-TAG: Scaling Math Reasoning with Multi-Tool Aggregation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:31.887361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:31.887361Z digest=sha256:d96db0ae1f2f045e2a578970fd30b44017243932f8a40dc22d89b7bd5f054681

Observation db4ca4a5-0cef-42e8-86a2-46bcad201e81 · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based} generative models.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Orca: A distributed serving system for {Transformer-Based} generative models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:32.048467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:32.048467Z digest=sha256:e9844adbeae60e52a7d3fbb2fa7361cf9160af247c0ad4264521c1e8e98f9297

Observation 21731362-e518-4e15-8f5b-d16505317704 · outbound

This paper cites {SHEPHERD}: Serving {DNNs} in the wild.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving {SHEPHERD}: Serving {DNNs} in the wild

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:32.163793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:32.163793Z digest=sha256:5df1682075d8ecc8fd155a6b32b0caf789d844b6ffa9ea95cd15fc6f88196c8f

Observation bf4d822e-3346-4f3d-8fa1-a1fc2168bb5d · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving OPT: Open Pre-trained Transformer Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:32.279616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:32.279616Z digest=sha256:026db88a61dc9ade432ddd985892d3582ee7d647b90c2d56b24a3b4aef40a7dc

Observation c3093174-0517-4fa4-b929-d3879215e197 · outbound

This paper cites Webpilot: a versatile and au- tonomous multi-agent system for web task execution with strategic exploration.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Webpilot: a versatile and au- tonomous multi-agent system for web task execution with strategic exploration

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:32.371311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:32.371311Z digest=sha256:64cc748d5d9fc5dcc1ef50d9e6a17c2deee048ddfbb06e3168490c0c48d23ab8

Observation 5c9b5d6f-6bb7-494e-afd6-e67a180a207f · outbound

This paper cites Response length perception and sequence scheduling: an llm-empowered llm infer- ence pipeline.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Response length perception and sequence scheduling: an llm-empowered llm infer- ence pipeline

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:32.451871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:32.451871Z digest=sha256:1db50c5d6174fb268019b5a9047f80eac476bf4927af7631163201bc2f638ddf

Observation 57fa7e3a-34fc-44d3-8ed9-f0a0e70fc139 · outbound

This paper cites {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:32.575076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:32.575076Z digest=sha256:2722b599e4722b30d38298de7a2748b30fc5bc564acab807fcca8c73803663db

Observation b8dfccfb-94c0-42b7-b120-56c57d07f7f7 · outbound

This paper cites an unresolved cited work.

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T18:43:27.933278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:43:27.933278Z digest=sha256:8bd8b67d029b5b87cf08c92a867c769fba2d97a665477a771ab3b7a02383f576

Pith citing papers

Observation d157a2be-c120-415b-95a0-a02d4a444cc1 · inbound

Efficient Multi-round LLM Inference over Disaggregated Serving cites this paper.

Efficient Multi-round LLM Inference over Disaggregated Serving AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:46.161659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:46.161659Z digest=sha256:5eec24bba588952fc3611394ce4c6cd2cc285a7b896c5a3ecbde942effa0b544

Observation e9d5cf38-73a0-406b-97d0-83e3602788ed · inbound

Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving cites this paper.

Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-13T22:20:51.838309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:20:51.838309Z digest=sha256:23f9c0453c6a05663cb25ac05c211ea33056a67e94af49119435649199ab39c4

Observation 8f48c869-091a-4c5a-8dbd-9b2d05ec1801 · inbound

Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference cites this paper.

Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-13T02:18:53.109644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T18:27:13.781231Z digest=sha256:cc376e4549cd82f76a8c5bed9df9fdd82af682bacbe570d23c04aa997af60061

Observation 2ba4fdb8-0140-4635-b7dc-25b138f82528 · inbound

Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference cites this paper.

Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-13T02:18:53.109644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-01T00:31:12.965178Z digest=sha256:ff6fbdf4261b4024fbc1b0c2b2251e0c164a97209f31835823d9c5509b3ec32a