Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:27:53.745929Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 4 inbound Pith citation observations for arXiv:2505.07203.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:27:53.745929Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T12:55:10.063756Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-18T08:51:08.950391Z
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0243f9e1-63e6-4859-ac0a-d6beb68c31c6 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications https: //character.ai/
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 53d40eb0-3707-4757-b408-28f2da28b813 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications [Online; accessed 2025-04-17]
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3bf8fc66-045e-467d-ad8f-671410d9738d · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Taming{Throughput-Latency} tradeoff in{LLM} inference with {Sarathi-Serve}
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bf5283b4-44a0-4f54-b508-f081fe947733 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Pytorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ea0f3eef-bd27-4fbe-99c8-d8ab76bd2752 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications The ai code editor
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b9b9b9d6-3fd9-41da-ae41-a92fbcb18c4f · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Glimpse: Continuous, real-time object recog- nition on mobile devices
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dd654ef5-101c-418e-a198-43a3279fafe2 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Feature engineering for machine learning and data analytics
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e9d28194-8d4e-46e4-8d4e-5690219f1d89 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01c3e016-8daa-4660-b225-d83c3c8a69e6 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Oneadapt: Fast adaptation for deep learning applications via back- propagation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e9c30c07-e086-4ae1-afaa-1bb06b2273dd · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Accmpeg: Optimizing video encoding for accurate video analytics
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fd00cd96-e99b-45a2-86cf-6162d77e606f · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Empowering Many, Biasing a Few: Generalist Credit Scoring through Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4919d590-6d62-4440-ba4f-6481f1d2c410 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications 360brew: A decoder-only foundation model for personalized ranking and recommendation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6044832-a151-4b28-ab8c-a25cf61898ab · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Gpt-3: Its nature, scope, limits, and consequences
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b5c677f-e06e-4c82-a30f-12d19dbcc643 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Github copilot - write code faster
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 87bf785e-f30f-46f5-bde2-8443c4eac836 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Tiresias: A {GPU} cluster manager for distributed deep learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b9c6bfc9-2db1-4ff4-bbd7-c9a99cb3d920 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 070ec353-da31-46bf-9097-ec604a14edf7 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87b1e661-6b83-44e0-b97f-63c38887672c · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Epic: Efficient position-independent context caching for serving large language models, 2024
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation afae9e35-653b-42c7-86ab-913980a668e6 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d26bc59e-c4cb-4528-b497-48d5f1fe3b59 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Gear: An efficient kv cache com- pression recipe for near-lossless generative inference of llm, 2024
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 24550dde-8603-49d5-8f9d-0c92bdc5ac57 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Over-fitting and model tuning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7a2840c0-8f67-4532-a4f5-63002c5bcb22 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Efficient memory management for large language model serving with 13 pagedattention
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a8a825b4-1a37-424a-91a2-8606daec7c64 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Efficient memory management for large language model serving with pagedattention
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3af2354d-a1b5-4add-80e9-ae4b4bc9219a · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Spam-T5: Benchmarking Large Language Models for Few-Shot Email Spam Detection
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4d1eb7f-1ea4-4c94-87a7-cf06ed827ba1 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Depression detection on social media with large language models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40c1f9e7-a35a-4767-85d1-9e452da647be · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Reducto: On-camera filtering for resource-efficient real-time video analytics
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3217ccf5-0339-4387-a970-07e14193a430 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Terapipe: Token-level pipeline parallelism for training large-scale language models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6496282b-7c4e-4f0c-8a51-4d1587ba24f1 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Edge assisted real-time object detection for mobile augmented reality
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4f462cbd-2303-4fe8-a5dd-47995a8a342c · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Gonzalez, Ion Stoica, and Matei Zaharia
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f536499a-cf67-4118-b5c5-3fd2b85d6881 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications FinGPT: Democratizing Internet-scale Data for Financial Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc014c3c-8e22-4e48-902d-c1ad9d1fc420 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Cachegen: Kv cache compression and streaming for fast large language model serving
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3e442f03-9673-4d2a-b0de-4f72bf8e1b4d · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 047d1cb1-4bb2-4c18-8fed-76d619473651 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Github - lmcache/lmcache: Redis for llms
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 07b981e4-390d-4878-af50-bed6259c4cd0 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Chatgpt: Conversational language model
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4dc108da-b91e-4536-91b3-14426a1005dd · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Generative agents: Interac- tive simulacra of human behavior
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f0185cbd-b785-46d2-9b5c-d417a446525c · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Optimus: an efficient dynamic resource scheduler for deep learn- ing clusters
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8c9461ee-d736-4803-9a83-a4178d128321 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Perplexity is a free ai search engine
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9b1acce1-2265-4fe2-bd50-d0ab5d9a0bcd · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3927c8b-6136-4ba6-b91a-264d15d65b3a · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Rah! recsys–assistant–human: A human-centered recommendation framework with llm agents
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 637b7295-4500-4d2f-98b3-98f68cc2d8c1 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Playing games with ais: the limits of gpt-3 and similar large language models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2ef96d4e-921a-48b0-a532-543f457614fc · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Beyond Classification: Financial Reasoning in State-of-the-Art Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3925683a-e76d-4d6b-ae8e-78799ab38ecd · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Enhancing Recommender Systems with Large Language Model Reasoning Graphs
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation becd3254-6802-4aa7-bf8c-d64054083e98 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Chain-of-thought prompt- ing elicits reasoning in large language models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d6ff7978-52d9-447f-a655-bb4f442ad64e · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications A survey on large language models for recommendation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5d1acd8-b8a7-4907-9529-8b63e542db75 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Predict- ing loan default in peer-to-peer lending using narrative data
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 30cf5885-a546-4145-8433-9491ff7869fe · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Auto-GPT for Online Decision Making: Benchmarks and Additional Opinions
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9da14eb-f966-4f74-9245-ba3710c9315f · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Mini- mizing hallucinations and communication costs: Adversarial debate and voting mechanisms in llm-based multi-agents
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0eeba333-fa42-4744-aaf8-95a48fcbdd05 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Cacheblend: Fast large language model serving for rag with cached knowledge fusion
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 899ff7b9-bb18-44ad-bec8-6fb6014d425a · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e751b87d-52b3-4af0-b225-c9c631042b71 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Orca: A distributed serving system for {Transformer-Based} generative models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8e54bc2c-54a8-4dbf-be4f-4961db2a084d · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Towards explaining the effects of data preprocessing on machine learning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 51bae741-9221-4c33-a038-53bcb0c90103 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Caravan: practical online learning of in-network ml models with labeling agents
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation af44b6bd-290f-4c2e-a910-3c0c899563da · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Ll- maaa: Making large language models as active annotators
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 46778d6b-fc90-4c16-93a6-ce1a3260d7e2 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications H2o: Heavy-hitter oracle for efficient generative in- ference of large language models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ed04e13f-d7a1-46ca-a052-8cef1fad50b4 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Mpic: Position-independent multimodal context caching system for efficient mllm serving, 2025
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8488794d-3968-4225-9121-505c1348aa6b · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Gonzalez, Clark Barrett, and Ying Sheng
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82f531d2-ea1b-43f4-931e-94e4302d6649 · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serv- ing
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7929bbf0-80a4-4cab-94d4-14e929b0c38c · outbound
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Unresolved cited work
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb2e10b5-71de-4c19-8415-8f9c15abf136 · inbound
TetriServe: Efficiently Serving Mixed DiT Workloads PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80f2c909-7b31-4f36-b684-7462641c5c4f · inbound
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f8d7267a-7304-4f84-a113-e6452afaff3c · inbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5179647a-9af1-4c01-919e-362ad2d2382b · inbound
Beyond Storage: State as a Runtime Control Problem in Parallel and Distributed Systems PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.