REVIEW 4 cited by
Andes: Defining and Enhancing Quality-of-Experience in LLM-Based Text Streaming Services
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Large language models (LLMs) are now at the core of conversational AI services such as real-time translation and chatbots, which provide live user interaction by incrementally streaming text to the user. However, existing LLM serving systems fail to provide good user experience because their optimization metrics are not always aligned with user experience. In this paper, we first introduce and define the notion of Quality-of-Experience (QoE) for text streaming services by considering each user's end-to-end interaction timeline. Based on this, we propose Andes, a QoE-aware LLM serving system that enhances user experience by ensuring that users receive the first token promptly and subsequent tokens at a smooth, digestible pace, even during surge periods. This is enabled by Andes's preemptive request scheduler that dynamically prioritizes requests at the token granularity based on each request's expected QoE gain and GPU resource usage. Our evaluations demonstrate that, compared to state-of-the-art LLM serving systems, Andes improves the average QoE by up to $4.7\times$ given the same GPU resource, or saves up to 61% GPU resources while maintaining the same high QoE.
Forward citations
Cited by 4 Pith papers
-
Leaky Language Models: Stealing Architecture and Inference Optimizations via Per-Token Timing
Per-token generation timing leaks speculative decoding and draft-model context length from Gemini, and recovers layer count and hidden size of Llama-family models with top-5 accuracy up to 65% when both are unknown.
-
GPU-to-Grid: Voltage Regulation via GPU Utilization Control
GPU batch-size control, driven by real-time voltage and latency feedback, can serve as a fast distribution-grid voltage regulation resource.
-
KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows
KVFlow uses workflow-aware eviction priorities and overlapped KV prefetching to cut cache-miss latency in LLM multi-agent serving.
-
ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
ConsumerBench reveals that running multiple generative AI apps concurrently on a consumer GPU causes severe starvation under greedy allocation and wasted capacity under static partitioning, driving the need for SLO-aw...
Discussion (0). Continue with ORCID to comment.