Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T16:59:01.487312Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 1 inbound Pith citation observation for arXiv:2501.12689.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T16:59:01.487312Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T14:42:35.999530Z
A source-named dated measurement, never combined with another source.
Source: cited_works
92 of 92 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cd32cb47-cc50-4020-86d5-2130697c6f13 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching https://developers.google.com/ search/docs/appearance/ai-overviews
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e500c042-14d4-4c90-8ab8-36aa279a2743 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching https://aws.amazon.com/codewhisperer/
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 380085fe-9edc-4b6a-ad44-5a23f39d0a4c · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching https://claude.ai/
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44786cd6-a2bf-470c-a50d-a425c26c6d57 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching https://character.ai/
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f88c1446-2145-4f77-94db-aa0af443544b · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching https://openai.com/index/ introducing-deep-research/
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 096681f9-7a28-426c-a187-75d8f547d723 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching https://api-docs.deepseek.com/guides/ kv_cache
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a6f70cf-8326-48c0-8a96-4e6c24dec579 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching https: //github.com/deepseek-ai/open-infra-index/blob/main/ 202502OpenSourceWeek/day_6_one_more_thing_ deepseekV3R1_inference_system_overview.md
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cf6f2cc-1fcb-4feb-9706-572214791dd1 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching https: //developers.googleblog.com/en/gemini-15-flash-8b-is-now- generally-\available-for-use/
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b45434b-deb8-4bd2-8354-411b2d10f417 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching https://ai.google.dev/gemini-api/docs/ caching?lang=python
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a97d271e-8ce3-4965-9ec3-9fc697263bcf · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching https://github.com/features/copilot/
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0b110c8-ed32-401e-bfef-cb6af0771a32 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching https://www
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1ccee4c-a3ce-4e1a-b78f-9966255d546c · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching https://huggingface.co/spaces/lmarena-ai/chatbot-arena- leaderboard
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2d29a30-d0b2-4492-aed9-dff524a20be2 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching https://huggingface.co/docs/ api-inference/index
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 992e6427-342e-4c0b-8bc3-c95b28dd62ad · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching https://github
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b9ecb950-bb36-4c65-8103-eda5c8aa0075 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching https://microsoft.github.io/msmarco/
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c5bbda78-31ab-4abf-9cdf-c7cbfdd066bc · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching https: //github.com/explosion/spaCy
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e38a789c-d390-4ee9-9638-9b4f0356c303 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching https://www.databricks.com/blog/building-cost-optimized- chatbot-semantic-caching, 2024
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 69b137e8-c30b-4caa-a7d5-36e0f26602a0 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0fe5ba3-8f17-44ec-8617-f9f70b688597 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Analysis of thompson sampling for the multi-armed bandit problem
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 92175fbd-bfe9-4eb6-a4ad-da7e1613a51a · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7123322-f633-448a-a5f5-657ab115c5da · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Gptcache: An open-source semantic cache for llm applica- tions enabling faster answers and cost savings
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 074cb387-ae3a-4d81-bb27-b477f80bad82 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Findings of the 2016 conference on machine translation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f90915e7-10a2-444b-8b22-0d8c5fff2872 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching JAX: compos- able transformations of Python+NumPy programs, 2018
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e71ff12a-149e-4d43-a45e-c90fbf043580 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Language Models are Few-Shot Learners
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1044e1e8-587f-484f-8592-841ad40b297d · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Are more llm calls all you need? towards scaling laws of compound inference systems
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a194a693-786f-4437-9fad-67e3d8be7938 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Safe RLHF: Safe Reinforcement Learning from Human Feedback
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75dec9d3-3c05-46d1-8630-51089c5c7376 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Learning semantic similarity in a continuous space
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1546857c-aac3-4d70-b3a7-7cc1bead340b · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching A Survey on In-context Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a51bbd8-2ba0-4fe2-82bf-c7da66dfb1dc · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Gemini: A Family of Highly Capable Multimodal Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 480858d0-a692-4121-bd7c-9e26947eb758 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching The Llama 3 Herd of Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ba9fc8b-ef05-4783-add3-b425f3128708 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Apple intelligence foundation lan- guage models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d44d57c7-7753-48a5-b1a1-6f5e882e17dc · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching A Theory of Emergent In-Context Learning as Implicit Structure Induction
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6b05c68-830a-4601-8a77-16dc123ccd23 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb5f9fb0-9d97-4ce2-98f3-c9f04f6c252e · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c425dc93-770c-4898-b7cc-6db9de6902b4 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Evaluation of Best-of-N Sampling Strategies for Language Model Alignment
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c5439a9-4d0e-4605-a805-1f7377e26f93 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Active Retrieval Augmented Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dae16513-ac8f-466d-a8b0-8c28aaec925c · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching MegaScale: Scaling large language model training to more than 10,000 GPUs
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e7f5ad9c-8aad-4a38-ad0e-7897ee31dd45 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Billion-scale similarity search with GPUs
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ead06160-c83b-40fe-912f-208322b47337 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Tanh works better with asymmetry
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 391c229e-cd53-430d-aed8-b10dcc8d28ef · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c08c3152-cb8f-4692-b3d7-6337991f75ba · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 946dc1bd-1535-494a-ab91-c7ea8e33cf37 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Efficient memory management for large language model serving with pagedattention
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b3e0f31-9917-4819-9713-88ad72c33937 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Auto-GDA: Automatic Domain Adaptation for Efficient Grounding Verification in Retrieval-Augmented Generation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e7d1cf4a-5f0f-42f2-a50a-6e3b78c0ed52 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Retrieval-augmented generation for knowledge- intensive nlp tasks
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c4163a19-a8e2-496d-b1cf-f4fd42b79230 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Dpsynthe- sizer: differentially private data synthesizer for privacy preserving data sharing
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c38910f0-5b83-440e-9b42-dd41ee577771 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Schapire
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3de0911c-4795-4685-b6fe-85e16e4bce7a · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Gon- zalez, and Ion Stoica
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6f803b86-85e5-468a-aa9f-53d9479639e6 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3537dd21-d42d-472a-83eb-32702659733f · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Openorca: An open dataset of gpt augmented flan reasoning traces
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1ca1e0b8-f16c-4aef-a0bb-dc53bfee0978 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Parrot: Efficient serving of llm-based applications with semantic variable
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e2b46443-4199-4a28-b683-bdf3080bd62e · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Andes: Defining and enhancing quality-of- experience in llm-based text streaming services
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 95a57241-3c22-4ef3-909f-094181b591e5 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching In-context Learning with Retrieved Demonstrations for Language Models: A Survey
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 522b6ef3-5e6b-411e-9d0f-a74d285a6b0c · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Simpo: Simple preference optimization with a reference-free reward
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33dffffc-c17f-42f0-ab03-fb5177ecbbf9 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 035a3950-fb3e-470a-b310-124f02b63984 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching RouteLLM: Learning to Route LLMs with Preference Data
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b70c6d1a-371a-4f0d-a4bc-6fad3bf9fb16 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Training language models to follow instruc- tions with human feedback
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5d3ec71e-b22e-4de4-856e-e8c4bfd0f1ba · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Splitwise: Efficient gen- erative llm inference using phase splitting
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 605cb783-6603-4532-b2bd-4b797d720249 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d477c421-d779-42f1-919d-6a1ad668f1bb · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Modserve: Scal- able and resource-efficient large multimodal model serving
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1204164-7432-4757-a02d-d0e905e17bc4 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2a15ae50-03c3-4a63-8cee-834b174626cb · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching The probabilistic relevance framework: Bm25 and beyond
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6b26c39d-d4b9-496e-a29f-ca20b9578756 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 540059ff-f982-4dc7-b9d3-a703af6d1982 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Gonzalez, and Ion Stoica
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f03d290f-c6ca-42d2-9b55-dfa4dac278ac · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching A statistical interpretation of term specificity and its application in retrieval
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e4714e4c-3b96-4f0f-835b-eb5358916a14 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Hygen: Efficient llm serving via elastic online-offline request co-location
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48874c01-f19b-4c00-985c-05ded19a3a08 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Hashimoto
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95cd782d-7f5b-481f-8aa0-a38814cc11e8 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Gemma 2: Improving Open Language Models at a Practical Size
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 567a4a46-69a7-4919-a334-1efc461420a6 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e129664e-46f7-4ad3-8439-a31b468b8046 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Emergent Abilities of Large Language Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52f8979c-3250-407b-8261-f1ee720fae31 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Fast Distributed Inference Serving for Large Language Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ec0c356-036f-414a-9093-2aff9256d5b1 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching dLoRA: Dynamically orchestrating requests and adapters for LoRA LLM serving
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4e61e9fa-c230-405c-b437-f0c5dc32a5e8 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Why in-context learning models are good few-shot learners? In ICLR, 2025
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7c569c36-6cf0-4a3b-8347-14970b0e2415 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a437b52c-ed8b-4a15-8e53-6e5611e3efb2 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac1ef0f7-dd5d-4b55-9f7f-c44df8a7be25 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Generating Data for Symbolic Language with Large Language Models
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fd80b406-6514-4059-b32c-6213b6ee35f1 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Compositional exemplars for in-context learning
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5d0ae6f1-7cb1-472d-bc84-92625cdb4022 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Orca: A distributed serving system for{Transformer- Based} generative models
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 28518b5b-a337-4f2c-9f10-67ee85eb4b12 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Longrag: A dual-perspective retrieval- augmented generation paradigm for long-context question answering
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 299158f3-1257-4348-870e-49c0a38a585e · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 745952cf-51da-46fe-b11f-2cf26ae7ff9a · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Judging llm-as-a-judge with mt-bench and chatbot arena.NeurIPS, 2023
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 72e24730-f6f0-4451-ac9f-bf749807b0f8 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Gonzalez, Clark Barrett, and Ying Sheng
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 52914cab-38cd-435e-ba70-e6f7090abdba · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b4d1226-6145-435e-961e-ed29311ae7da · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Distillspec: Improving speculative decoding via knowledge distillation, 2024
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0a8c8721-424c-4d62-9124-cb322f5dac56 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching We can bound this with the union bound: 𝑃(ˆ𝑖𝑇 ≠ 1)≤ 𝑁∑︁ 𝑖=2 𝑃(𝜇𝑖 >𝜇1) (2)
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fdef6cda-73fe-4b83-a309-51dda7a6e446 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching We can state this more formally for the number of comparisons,𝑚𝑖(𝑇), for a sufficiently large T: 𝑚𝑖(𝑇)≥ 𝐾 log(𝑇) Δ2 𝑖 (3) where𝐾 is a positive constant
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 29f295ec-864f-4b79-b281-308edef2edff · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Let the em- pirical difference be ˆΔ𝑖(𝑚) = 𝜇1−𝜇𝑖 after𝑚 compar- isons, whose true mean is the utility gap Δ𝑖 =𝑈1−𝑈𝑖
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f823556b-fdd4-4bbf-b6c4-68c904ce3010 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8b09cd53-4a64-4bdd-978b-c201b058ab8d · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Substitut- ing this result back into the union bound from step 1 gives the final bound
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6df4f04f-d3fe-4924-9082-5c8d802e3434 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9152419b-e6e1-4fb5-af5b-d387d722330a · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation aac84641-c28b-4673-86ee-cb524d096197 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bb69b9a2-e6f5-45fd-b009-dbb4adbd88f4 · outbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b0f988cc-6209-4367-8174-1ca3a132931f · inbound
CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems IC-Cache: Efficient Large Language Model Serving via In-context Caching
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.