Pith. sign in

Paper Citation Record · LEDGER

IC-Cache: Efficient Large Language Model Serving via In-context Caching

As of 11 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 1 inbound Pith citation observation for arXiv:2501.12689.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12689 v3

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:59:01.487312Z

measured 93 of 93 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:42:35.999530Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

92 of 92 outbound references displayed

  • verified exact2
  • verified fuzzy36
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd32cb47-cc50-4020-86d5-2130697c6f13 · outbound

This paper cites https://developers.google.com/ search/docs/appearance/ai-overviews.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://developers.google.com/ search/docs/appearance/ai-overviews

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.196339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.196339Z digest=sha256:e8c024b8648abb77d4cf404c1f8f15adef845a71e35f475803e16eb2ee98f642

Observation e500c042-14d4-4c90-8ab8-36aa279a2743 · outbound

This paper cites https://aws.amazon.com/codewhisperer/.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://aws.amazon.com/codewhisperer/

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.200708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.200708Z digest=sha256:79c49e7f16b8eda5cf1a9d0426c236ad2c76d4ac0e4a3668fddd4e124b02bdcb

Observation 380085fe-9edc-4b6a-ad44-5a23f39d0a4c · outbound

This paper cites https://claude.ai/.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://claude.ai/

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.204302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.204302Z digest=sha256:22fcfb133247b993f89050ef3cd08692a3ce5b40a3be1d982cf172a8b045a493

Observation 44786cd6-a2bf-470c-a50d-a425c26c6d57 · outbound

This paper cites https://character.ai/.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://character.ai/

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.209020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.209020Z digest=sha256:5b9440c66e07425c309a29d8b3d924fd80db03c88ae4b01053631828b4cd711f

Observation f88c1446-2145-4f77-94db-aa0af443544b · outbound

This paper cites https://openai.com/index/ introducing-deep-research/.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://openai.com/index/ introducing-deep-research/

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.212859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.212859Z digest=sha256:cb3e6f112676fdbf215f9a9a15c487354e02cedbdbadaf52ffc1a31bb903165e

Observation 096681f9-7a28-426c-a187-75d8f547d723 · outbound

This paper cites https://api-docs.deepseek.com/guides/ kv_cache.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://api-docs.deepseek.com/guides/ kv_cache

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.216531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.216531Z digest=sha256:24bb5d660b32be974832e2e6e500def284f6bd1fcb9b82e7a9b8fcb8b4c05dc3

Observation 3a6f70cf-8326-48c0-8a96-4e6c24dec579 · outbound

This paper cites https: //github.com/deepseek-ai/open-infra-index/blob/main/ 202502OpenSourceWeek/day_6_one_more_thing_ deepseekV3R1_inference_system_overview.md.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https: //github.com/deepseek-ai/open-infra-index/blob/main/ 202502OpenSourceWeek/day_6_one_more_thing_ deepseekV3R1_inference_system_overview.md

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.219494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.219494Z digest=sha256:e63f95f5105bea8869e537b1bfe510a60b73ca94663966663ddc4966cfc4bf4d

Observation 2cf6f2cc-1fcb-4feb-9706-572214791dd1 · outbound

This paper cites https: //developers.googleblog.com/en/gemini-15-flash-8b-is-now- generally-\available-for-use/.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https: //developers.googleblog.com/en/gemini-15-flash-8b-is-now- generally-\available-for-use/

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.222104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.222104Z digest=sha256:e588425afc06a6fb68f8069021145f3ec8e3fcfecd5ef92b760d10897a3b1f1a

Observation 3b45434b-deb8-4bd2-8354-411b2d10f417 · outbound

This paper cites https://ai.google.dev/gemini-api/docs/ caching?lang=python.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://ai.google.dev/gemini-api/docs/ caching?lang=python

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.224912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.224912Z digest=sha256:691363ac0e991c8bb877eca41243887e45ea75eaa7f678bef65ef6fd42403c42

Observation a97d271e-8ce3-4965-9ec3-9fc697263bcf · outbound

This paper cites https://github.com/features/copilot/.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://github.com/features/copilot/

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.227790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.227790Z digest=sha256:f3bebd2676e877fd74ffd6f2d2fc8e82561fbd16610ed5fcf0725efddbaf4dae

Observation b0b110c8-ed32-401e-bfef-cb6af0771a32 · outbound

This paper cites https://www.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://www

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.230372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.230372Z digest=sha256:343fce351601377fc4e10682ee4df2a67e26d74fae43d65248888695cecf6196

Observation a1ccee4c-a3ce-4e1a-b78f-9966255d546c · outbound

This paper cites https://huggingface.co/spaces/lmarena-ai/chatbot-arena- leaderboard.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://huggingface.co/spaces/lmarena-ai/chatbot-arena- leaderboard

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.233061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.233061Z digest=sha256:8f14874f3f94833c149a6e30162b6449b77bc0d9ee1d5a6b063cc58144ee7e5e

Observation d2d29a30-d0b2-4492-aed9-dff524a20be2 · outbound

This paper cites https://huggingface.co/docs/ api-inference/index.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://huggingface.co/docs/ api-inference/index

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.422543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.235646Z digest=sha256:d3ead013abeed3ab236841435595cf93afac67cd4617e159ed583839509d4a99

Observation 992e6427-342e-4c0b-8bc3-c95b28dd62ad · outbound

This paper cites https://github.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://github

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.413250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.238251Z digest=sha256:18bc2d3997fcafed7ea68304b5131faec4ca07b3129f4af19fad3243eccdc713

Observation b9ecb950-bb36-4c65-8103-eda5c8aa0075 · outbound

This paper cites https://microsoft.github.io/msmarco/.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://microsoft.github.io/msmarco/

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.404981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.241257Z digest=sha256:84ffda2bf99d25b2288e804e6b34a171acb6d0e0563ff0eac5353065b4f7c8e0

Observation c5bbda78-31ab-4abf-9cdf-c7cbfdd066bc · outbound

This paper cites https: //github.com/explosion/spaCy.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https: //github.com/explosion/spaCy

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.396331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.243979Z digest=sha256:de9d9d1b3c6b711ff58ae98ca63838f64edb8ede4391e6ec97fa39e0ad40a9d5

Observation e38a789c-d390-4ee9-9638-9b4f0356c303 · outbound

This paper cites https://www.databricks.com/blog/building-cost-optimized- chatbot-semantic-caching, 2024.

IC-Cache: Efficient Large Language Model Serving via In-context Caching https://www.databricks.com/blog/building-cost-optimized- chatbot-semantic-caching, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.387503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.246580Z digest=sha256:1578ab5276dcea1cf5adf0ff414ca74931fa56e02385ee8f3bd18339abfd4ed7

Observation 69b137e8-c30b-4caa-a7d5-36e0f26602a0 · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

IC-Cache: Efficient Large Language Model Serving via In-context Caching SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.249697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.249697Z digest=sha256:deb0d82f0234a4a32b315550bf71e40a127d67cca1a842c544270f307f44f0f9

Observation b0fe5ba3-8f17-44ec-8617-f9f70b688597 · outbound

This paper cites Analysis of thompson sampling for the multi-armed bandit problem.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Analysis of thompson sampling for the multi-armed bandit problem

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.379244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.252860Z digest=sha256:2deea656f6a4baa4e296c4211522020dac5f3f96c495c8187447daeeb8218ca6

Observation 92175fbd-bfe9-4eb6-a4ad-da7e1613a51a · outbound

This paper cites Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.255587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.255587Z digest=sha256:060a35d971164d27197c1477cce7c925be6c3b093bbef2d423e0a4145669103c

Observation c7123322-f633-448a-a5f5-657ab115c5da · outbound

This paper cites Gptcache: An open-source semantic cache for llm applica- tions enabling faster answers and cost savings.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Gptcache: An open-source semantic cache for llm applica- tions enabling faster answers and cost savings

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.370400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.259377Z digest=sha256:4327489ac8f3e7487d633d27654c25b047616070645b2effeab9bade87fdfe49

Observation 074cb387-ae3a-4d81-bb27-b477f80bad82 · outbound

This paper cites Findings of the 2016 conference on machine translation.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Findings of the 2016 conference on machine translation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.361569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.262365Z digest=sha256:a8f2b7e215f0b03c5f4e9c38194b7391702ece7a856989898c1677a8c49047b7

Observation f90915e7-10a2-444b-8b22-0d8c5fff2872 · outbound

This paper cites JAX: compos- able transformations of Python+NumPy programs, 2018.

IC-Cache: Efficient Large Language Model Serving via In-context Caching JAX: compos- able transformations of Python+NumPy programs, 2018

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.351675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.265474Z digest=sha256:b6052a1d17066456c5ee3867d8085dde1fe5a75a161f1cd00676070a5c2f555e

Observation e71ff12a-149e-4d43-a45e-c90fbf043580 · outbound

This paper cites Language Models are Few-Shot Learners.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Language Models are Few-Shot Learners

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.269384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.269384Z digest=sha256:0e4bd795ab5c758a1ebbc7a27f6c71b532614fee41be5f31eb64be2e4008bec9

Observation 1044e1e8-587f-484f-8592-841ad40b297d · outbound

This paper cites Are more llm calls all you need? towards scaling laws of compound inference systems.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Are more llm calls all you need? towards scaling laws of compound inference systems

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.341134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.273733Z digest=sha256:a960ca2bfba7e595e1e7476da0061d424f67ddbf516daf14c88615b9bda64585

Observation a194a693-786f-4437-9fad-67e3d8be7938 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.277250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.277250Z digest=sha256:1dadfedcb217560228d01651c5de45547b70286bd5a94a499704d40b3aa241c0

Observation 75dec9d3-3c05-46d1-8630-51089c5c7376 · outbound

This paper cites Learning semantic similarity in a continuous space.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Learning semantic similarity in a continuous space

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.331126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.280850Z digest=sha256:0e37badebf66a9624e7e4bb075512d2594b5d0678c012573361c51d03a6e65c1

Observation 1546857c-aac3-4d70-b3a7-7cc1bead340b · outbound

This paper cites A Survey on In-context Learning.

IC-Cache: Efficient Large Language Model Serving via In-context Caching A Survey on In-context Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.284457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.284457Z digest=sha256:4b7619b998d139fd1d0d95946e39e4bdb1c24a6052a822ac613906d8322352b0

Observation 1a51bbd8-2ba0-4fe2-82bf-c7da66dfb1dc · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Gemini: A Family of Highly Capable Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.288550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.288550Z digest=sha256:537ee2b68ecbc3e91b78076e367fb6b37dfa3315715786e6ddafee07a1b9ac72

Observation 480858d0-a692-4121-bd7c-9e26947eb758 · outbound

This paper cites The Llama 3 Herd of Models.

IC-Cache: Efficient Large Language Model Serving via In-context Caching The Llama 3 Herd of Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.291516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.291516Z digest=sha256:6a1ee12b6deae568fc67033fbcbbc4a23e629bd5836806c13715406000dfeb1c

Observation 9ba9fc8b-ef05-4783-add3-b425f3128708 · outbound

This paper cites Apple intelligence foundation lan- guage models.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Apple intelligence foundation lan- guage models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.294266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.294266Z digest=sha256:bccaf1d5cde967ad0f44b5bbf5079d9e27b4a5fd9b3697bc13d8607e6ce63a87

Observation d44d57c7-7753-48a5-b1a1-6f5e882e17dc · outbound

This paper cites A Theory of Emergent In-Context Learning as Implicit Structure Induction.

IC-Cache: Efficient Large Language Model Serving via In-context Caching A Theory of Emergent In-Context Learning as Implicit Structure Induction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.296886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.296886Z digest=sha256:5b4ad31f4c0881eeb0234d716253b9e66ac14311c9448f039e592346857eb619

Observation b6b05c68-830a-4601-8a77-16dc123ccd23 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.299994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.299994Z digest=sha256:4e0faf2d92bf91a741c6b67af5ca0652656a83ceca1417799397802b00595492

Observation eb5f9fb0-9d97-4ce2-98f3-c9f04f6c252e · outbound

This paper cites An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4.

IC-Cache: Efficient Large Language Model Serving via In-context Caching An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.302934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.302934Z digest=sha256:91f5c256ec1c90656054945a6f35385bff0d05df374d481da1305320a4e4e0b9

Observation c425dc93-770c-4898-b7cc-6db9de6902b4 · outbound

This paper cites Evaluation of Best-of-N Sampling Strategies for Language Model Alignment.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Evaluation of Best-of-N Sampling Strategies for Language Model Alignment

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.305843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.305843Z digest=sha256:09f0fd198a480928f4c41cd1ff456b210511bafce73c48f7e4fcdd6a1865c6d2

Observation 3c5439a9-4d0e-4605-a805-1f7377e26f93 · outbound

This paper cites Active Retrieval Augmented Generation.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Active Retrieval Augmented Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.308781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.308781Z digest=sha256:cb385e1d2eea777cf38825a864940022a478f7d67e2a15b79928eb5e6508cd33

Observation dae16513-ac8f-466d-a8b0-8c28aaec925c · outbound

This paper cites MegaScale: Scaling large language model training to more than 10,000 GPUs.

IC-Cache: Efficient Large Language Model Serving via In-context Caching MegaScale: Scaling large language model training to more than 10,000 GPUs

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.320664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.312290Z digest=sha256:68b4396c55c8ba4ed896e1f634abcd716a78e05f9bbf1e3915433e713f065701

Observation e7f5ad9c-8aad-4a38-ad0e-7897ee31dd45 · outbound

This paper cites Billion-scale similarity search with GPUs.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Billion-scale similarity search with GPUs

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.310486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.315989Z digest=sha256:816cbea8ba65f181b7153586adde3810170076aab5fd8e46d64354e0bcd2ebc3

Observation ead06160-c83b-40fe-912f-208322b47337 · outbound

This paper cites Tanh works better with asymmetry.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Tanh works better with asymmetry

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.300514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.319029Z digest=sha256:e86b48f44faf5c0f733995a1260415e3de7f9161661d8f8b15ce197ca485b74d

Observation 391c229e-cd53-430d-aed8-b10dcc8d28ef · outbound

This paper cites Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.322004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.322004Z digest=sha256:a68c008dd549f1c5356daa812b5d68856b46b6791038b069a0ba94dcf49f9142

Observation c08c3152-cb8f-4692-b3d7-6337991f75ba · outbound

This paper cites Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.325281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.325281Z digest=sha256:9e75113029e51bf931e0c624d58df9c0514afd378c233c3e0ed25c2fd0ced5d9

Observation 946dc1bd-1535-494a-ab91-c7ea8e33cf37 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Efficient memory management for large language model serving with pagedattention

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.328661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.328661Z digest=sha256:f5f976cf68552a1d6aabcf3b75069a3c878a8816b1eb4c6ff54fc447dc2f7be5

Observation 6b3e0f31-9917-4819-9713-88ad72c33937 · outbound

This paper cites Auto-GDA: Automatic Domain Adaptation for Efficient Grounding Verification in Retrieval-Augmented Generation.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Auto-GDA: Automatic Domain Adaptation for Efficient Grounding Verification in Retrieval-Augmented Generation

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-10T16:59:01.797977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.332595Z digest=sha256:1a11e8b16d346bc21f8c62c79f0fd897201331e39c038a8634229054020e699a

Observation e7d1cf4a-5f0f-42f2-a50a-6e3b78c0ed52 · outbound

This paper cites Retrieval-augmented generation for knowledge- intensive nlp tasks.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Retrieval-augmented generation for knowledge- intensive nlp tasks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.278868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.336033Z digest=sha256:dfaac16feb12295b6418e50cde759fdd247d1aa267f189e8bc4680c1e005f4ea

Observation c4163a19-a8e2-496d-b1cf-f4fd42b79230 · outbound

This paper cites Dpsynthe- sizer: differentially private data synthesizer for privacy preserving data sharing.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Dpsynthe- sizer: differentially private data synthesizer for privacy preserving data sharing

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.269666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.339123Z digest=sha256:9f03b664d534e35f9acc38e899518fc50e4fb85cf4b139a15694ae7802d08fb9

Observation c38910f0-5b83-440e-9b42-dd41ee577771 · outbound

This paper cites Schapire.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Schapire

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.260458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.342208Z digest=sha256:b267e34c53e302f81db6dfc77f34ffa1098b322fbda400c02448c9f5b59323c7

Observation 3de0911c-4795-4685-b6fe-85e16e4bce7a · outbound

This paper cites Gon- zalez, and Ion Stoica.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Gon- zalez, and Ion Stoica

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.251889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.345353Z digest=sha256:d21b44e38c6731cce2dd37860dca34eb9e4654f2b62ddf94a8c3878ade77778f

Observation 6f803b86-85e5-468a-aa9f-53d9479639e6 · outbound

This paper cites AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding.

IC-Cache: Efficient Large Language Model Serving via In-context Caching AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.348128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.348128Z digest=sha256:4f5742d70862f71bef1044252b0bef5569bf267393d39ecd4fa1ccc9975cfc81

Observation 3537dd21-d42d-472a-83eb-32702659733f · outbound

This paper cites Openorca: An open dataset of gpt augmented flan reasoning traces.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Openorca: An open dataset of gpt augmented flan reasoning traces

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.243226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.350881Z digest=sha256:d02d26df8006c2512b5928711fcb6a0b349535b37fb9da427b452892caffc547

Observation 1ca1e0b8-f16c-4aef-a0bb-dc53bfee0978 · outbound

This paper cites Parrot: Efficient serving of llm-based applications with semantic variable.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Parrot: Efficient serving of llm-based applications with semantic variable

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.234904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.353771Z digest=sha256:4f854c1d27b12652c390285e82dcc643db54f833749ef783cbb0a86e06bc2580

Observation e2b46443-4199-4a28-b683-bdf3080bd62e · outbound

This paper cites Andes: Defining and enhancing quality-of- experience in llm-based text streaming services.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Andes: Defining and enhancing quality-of- experience in llm-based text streaming services

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.225238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.356375Z digest=sha256:9d7f376696aa52f9e18fc274bc9f86eda2226ef0fd9bde6450b12299adcb201a

Observation 95a57241-3c22-4ef3-909f-094181b591e5 · outbound

This paper cites In-context Learning with Retrieved Demonstrations for Language Models: A Survey.

IC-Cache: Efficient Large Language Model Serving via In-context Caching In-context Learning with Retrieved Demonstrations for Language Models: A Survey

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.358903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.358903Z digest=sha256:bf646acf8c96cef35b5dd0f346536c1313ceef84cb7f3117ef7ff9dac9847815

Observation 522b6ef3-5e6b-411e-9d0f-a74d285a6b0c · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Simpo: Simple preference optimization with a reference-free reward

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.362089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.362089Z digest=sha256:c3d155877628f0bcc92e44dcfb161e7f77171d9741972157b5e21e4cb82c0d2e

Observation 33dffffc-c17f-42f0-ab03-fb5177ecbbf9 · outbound

This paper cites MS MARCO: A Human Generated MAchine Reading COmprehension Dataset.

IC-Cache: Efficient Large Language Model Serving via In-context Caching MS MARCO: A Human Generated MAchine Reading COmprehension Dataset

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.364722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.364722Z digest=sha256:f16ad84746207ce49416f74e815bd52857435e04481e003f8375e8a638b28398

Observation 035a3950-fb3e-470a-b310-124f02b63984 · outbound

This paper cites RouteLLM: Learning to Route LLMs with Preference Data.

IC-Cache: Efficient Large Language Model Serving via In-context Caching RouteLLM: Learning to Route LLMs with Preference Data

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.367502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.367502Z digest=sha256:264572978d630397e10145ca8144a7005c0fd5e3285f765a7f0a70b53aa7a615

Observation b70c6d1a-371a-4f0d-a4bc-6fad3bf9fb16 · outbound

This paper cites Training language models to follow instruc- tions with human feedback.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Training language models to follow instruc- tions with human feedback

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.209813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.370325Z digest=sha256:8c2c914e0fd012741b49659fb4f33fe26482111009b35f9087acaa0d119a156a

Observation 5d3ec71e-b22e-4de4-856e-e8c4bfd0f1ba · outbound

This paper cites Splitwise: Efficient gen- erative llm inference using phase splitting.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Splitwise: Efficient gen- erative llm inference using phase splitting

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.373036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.373036Z digest=sha256:e937fbf2ddf0787630a3d9a185a59ce9c835d96e9525bb207e9a8fac391cac7b

Observation 605cb783-6603-4532-b2bd-4b797d720249 · outbound

This paper cites ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving.

IC-Cache: Efficient Large Language Model Serving via In-context Caching ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.376047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.376047Z digest=sha256:7356fda87c80034b529dcb940fc3a3b4e480c7714e558524cf0176fc3a53c436

Observation d477c421-d779-42f1-919d-6a1ad668f1bb · outbound

This paper cites Modserve: Scal- able and resource-efficient large multimodal model serving.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Modserve: Scal- able and resource-efficient large multimodal model serving

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.379428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.379428Z digest=sha256:1615997892ce5e9e0a2035d74ed90c35629cbb7cdbab939936259a684c848227

Observation d1204164-7432-4757-a02d-d0e905e17bc4 · outbound

This paper cites an unresolved cited work.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:59:02.193782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.382463Z digest=sha256:88e143de72ac1166aefd5fdcd7693236eb2b87a9db5be37b64e818f979a84ad2

Observation 2a15ae50-03c3-4a63-8cee-834b174626cb · outbound

This paper cites The probabilistic relevance framework: Bm25 and beyond.

IC-Cache: Efficient Large Language Model Serving via In-context Caching The probabilistic relevance framework: Bm25 and beyond

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.184115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.385557Z digest=sha256:61bfbde224d68d771497d1813efc073a8efc6ce9ddfe8b5535542ddf830aa2ed

Observation 6b26c39d-d4b9-496e-a29f-ca20b9578756 · outbound

This paper cites Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.389424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.389424Z digest=sha256:ccf68975b9275dd2b588245366cdfbf26f8b9dd6ccdc0ba799721a01c825c9f3

Observation 540059ff-f982-4dc7-b9d3-a703af6d1982 · outbound

This paper cites Gonzalez, and Ion Stoica.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Gonzalez, and Ion Stoica

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.392854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.392854Z digest=sha256:ad61aa5da6c37853175996530f5a2d8d633a1bc28c08ce3d8e61b71b8066d6ef

Observation f03d290f-c6ca-42d2-9b55-dfa4dac278ac · outbound

This paper cites A statistical interpretation of term specificity and its application in retrieval.

IC-Cache: Efficient Large Language Model Serving via In-context Caching A statistical interpretation of term specificity and its application in retrieval

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.168557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.401232Z digest=sha256:8bd2b4251d83b8bde55b2ae580d9a9f1f45e132e05325074e7ee8358861cc68a

Observation e4714e4c-3b96-4f0f-835b-eb5358916a14 · outbound

This paper cites Hygen: Efficient llm serving via elastic online-offline request co-location.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Hygen: Efficient llm serving via elastic online-offline request co-location

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.404692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.404692Z digest=sha256:150a3af24bf749c07d59528e3be454d411682bac8a856dce465c3fc4ce705906

Observation 48874c01-f19b-4c00-985c-05ded19a3a08 · outbound

This paper cites Hashimoto.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Hashimoto

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.407925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.407925Z digest=sha256:73ec74f7dcc5cab9a48aecda7796d0c86f5c37a6c3584a41369bcb31c35f43d8

Observation 95cd782d-7f5b-481f-8aa0-a38814cc11e8 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Gemma 2: Improving Open Language Models at a Practical Size

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.411181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.411181Z digest=sha256:f180d16ddd168cedc6dc8568623e1408553d1c1ce983290cd05adec20a492608

Observation 567a4a46-69a7-4919-a334-1efc461420a6 · outbound

This paper cites an unresolved cited work.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:59:02.152835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.414581Z digest=sha256:c6bb454ccd640a35b4697b2c9f87760035e24a0762a66c9124ef6bb25333a91f

Observation e129664e-46f7-4ad3-8439-a31b468b8046 · outbound

This paper cites Emergent Abilities of Large Language Models.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Emergent Abilities of Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.417721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.417721Z digest=sha256:d71b8b48069236bee0c935ce80e71c3bdccf03a669ca15724feab6c008eb17ed

Observation 52f8979c-3250-407b-8261-f1ee720fae31 · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Fast Distributed Inference Serving for Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.420549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.420549Z digest=sha256:81d2c419124251a9c532910ce73096c1ddf6dc7ec62224fed074476949978077

Observation 5ec0c356-036f-414a-9093-2aff9256d5b1 · outbound

This paper cites dLoRA: Dynamically orchestrating requests and adapters for LoRA LLM serving.

IC-Cache: Efficient Large Language Model Serving via In-context Caching dLoRA: Dynamically orchestrating requests and adapters for LoRA LLM serving

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.143575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.423310Z digest=sha256:e624382e8c7739013254f7954e757954fc042b93ff9f5a0399a30a2e647daed1

Observation 4e61e9fa-c230-405c-b437-f0c5dc32a5e8 · outbound

This paper cites Why in-context learning models are good few-shot learners? In ICLR, 2025.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Why in-context learning models are good few-shot learners? In ICLR, 2025

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.133859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.426023Z digest=sha256:8692b0fe6e4d5533ed718d5789ba85a6413ca90b49da92dd510926c7ffb21863

Observation 7c569c36-6cf0-4a3b-8347-14970b0e2415 · outbound

This paper cites PowerInfer-2: Fast Large Language Model Inference on a Smartphone.

IC-Cache: Efficient Large Language Model Serving via In-context Caching PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.428571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.428571Z digest=sha256:f7190d03adbbcb0934f2a193382a8cb3b6143019c1f84c62727020c15b7fc26d

Observation a437b52c-ed8b-4a15-8e53-6e5611e3efb2 · outbound

This paper cites CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion.

IC-Cache: Efficient Large Language Model Serving via In-context Caching CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.431548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.431548Z digest=sha256:694ad8b37336d6ec56ae578d9b1a0179dd9abc92f32301f1d8071e0f825b52d1

Observation ac1ef0f7-dd5d-4b55-9f7f-c44df8a7be25 · outbound

This paper cites Generating Data for Symbolic Language with Large Language Models.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Generating Data for Symbolic Language with Large Language Models

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-08-10T16:59:01.539371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.434253Z digest=sha256:90e60da798f03fdd1a222a8a12e6c74fcc52ed4b51ef416757d696e17ec9afd7

Observation fd80b406-6514-4059-b32c-6213b6ee35f1 · outbound

This paper cites Compositional exemplars for in-context learning.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Compositional exemplars for in-context learning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.124312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.437310Z digest=sha256:8065a9f1f9c244a8e72b1aae4421bc4dd41690fa36cbff36aae16bf3c69fb1b9

Observation 5d0ae6f1-7cb1-472d-bc84-92625cdb4022 · outbound

This paper cites Orca: A distributed serving system for{Transformer- Based} generative models.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Orca: A distributed serving system for{Transformer- Based} generative models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.115345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.440027Z digest=sha256:9950cda22b15484d3095a23cf00495fc95166bdda08b15a12d0ab52cc4ef3d07

Observation 28518b5b-a337-4f2c-9f10-67ee85eb4b12 · outbound

This paper cites Longrag: A dual-perspective retrieval- augmented generation paradigm for long-context question answering.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Longrag: A dual-perspective retrieval- augmented generation paradigm for long-context question answering

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.106458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.442999Z digest=sha256:3ceaac7367b7fe7b31a4b47d85d3b980a5ee52f1a6d0f85afa86c113b935b5bd

Observation 299158f3-1257-4348-870e-49c0a38a585e · outbound

This paper cites LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset.

IC-Cache: Efficient Large Language Model Serving via In-context Caching LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.446099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.446099Z digest=sha256:e935a8b0e85760bbadd23bf4d10bbcddd5044e57405e28b2066da3df8c48e133

Observation 745952cf-51da-46fe-b11f-2cf26ae7ff9a · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.NeurIPS, 2023.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Judging llm-as-a-judge with mt-bench and chatbot arena.NeurIPS, 2023

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.097402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.449572Z digest=sha256:47747c3f014701155dbc7e568f632299453cbf14c297310d379216bd33c9db01

Observation 72e24730-f6f0-4451-ac9f-bf749807b0f8 · outbound

This paper cites Gonzalez, Clark Barrett, and Ying Sheng.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Gonzalez, Clark Barrett, and Ying Sheng

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.088761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.452502Z digest=sha256:8161f4289f069d5dd69805509d017f3942db5a47dd94ceff43779b2d5008124d

Observation 52914cab-38cd-435e-ba70-e6f7090abdba · outbound

This paper cites DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving.

IC-Cache: Efficient Large Language Model Serving via In-context Caching DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.455880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.455880Z digest=sha256:b0bda032d68f3e14cbddd5125fee6facea457a05a5cfb9c0510565777e97e095

Observation 5b4d1226-6145-435e-961e-ed29311ae7da · outbound

This paper cites Distillspec: Improving speculative decoding via knowledge distillation, 2024.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Distillspec: Improving speculative decoding via knowledge distillation, 2024

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.079945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.459544Z digest=sha256:dbefbd797bbc29ef11ea78b8e2780fb9941d72ee14254e451e5d3321b26ddf4e

Observation 0a8c8721-424c-4d62-9124-cb322f5dac56 · outbound

This paper cites We can bound this with the union bound: 𝑃(ˆ𝑖𝑇 ≠ 1)≤ 𝑁∑︁ 𝑖=2 𝑃(𝜇𝑖 >𝜇1) (2).

IC-Cache: Efficient Large Language Model Serving via In-context Caching We can bound this with the union bound: 𝑃(ˆ𝑖𝑇 ≠ 1)≤ 𝑁∑︁ 𝑖=2 𝑃(𝜇𝑖 >𝜇1) (2)

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.070905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.462700Z digest=sha256:166f0d8d207abf60e5747515fe587acc6e027428c0c0ad7c5da352108275a508

Observation fdef6cda-73fe-4b83-a309-51dda7a6e446 · outbound

This paper cites We can state this more formally for the number of comparisons,𝑚𝑖(𝑇), for a sufficiently large T: 𝑚𝑖(𝑇)≥ 𝐾 log(𝑇) Δ2 𝑖 (3) where𝐾 is a positive constant.

IC-Cache: Efficient Large Language Model Serving via In-context Caching We can state this more formally for the number of comparisons,𝑚𝑖(𝑇), for a sufficiently large T: 𝑚𝑖(𝑇)≥ 𝐾 log(𝑇) Δ2 𝑖 (3) where𝐾 is a positive constant

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.061618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.465836Z digest=sha256:2b4294441c8ec3a0944989b23dbe527c88d982399145b30e20e3f4843bc256b6

Observation 29f295ec-864f-4b79-b281-308edef2edff · outbound

This paper cites Let the em- pirical difference be ˆΔ𝑖(𝑚) = 𝜇1−𝜇𝑖 after𝑚 compar- isons, whose true mean is the utility gap Δ𝑖 =𝑈1−𝑈𝑖.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Let the em- pirical difference be ˆΔ𝑖(𝑚) = 𝜇1−𝜇𝑖 after𝑚 compar- isons, whose true mean is the utility gap Δ𝑖 =𝑈1−𝑈𝑖

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.051814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.468915Z digest=sha256:235a7a9ef6438c416b4ee347e755592d2bb10e203f0150b02f5d756e5768df5e

Observation f823556b-fdd4-4bbf-b6c4-68c904ce3010 · outbound

This paper cites an unresolved cited work.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:59:02.042112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.472164Z digest=sha256:ff1fe613adcb3a8573c3a49c86d4f8e99e21665bfc143635589e4f461c2650fe

Observation 8b09cd53-4a64-4bdd-978b-c201b058ab8d · outbound

This paper cites Substitut- ing this result back into the union bound from step 1 gives the final bound.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Substitut- ing this result back into the union bound from step 1 gives the final bound

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:59:02.032253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.475312Z digest=sha256:6c8bac8f02b812725ac5917e73b02a8a7c909936ffabe631e723ca0d248acefa

Observation 6df4f04f-d3fe-4924-9082-5c8d802e3434 · outbound

This paper cites an unresolved cited work.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:59:02.022550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.478487Z digest=sha256:6c8fa694567a49f7110cdc667cfea20d24a923e3a241e7dc47bc876bc01c2388

Observation 9152419b-e6e1-4fb5-af5b-d387d722330a · outbound

This paper cites an unresolved cited work.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:59:02.012709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.481708Z digest=sha256:4d4b5624564dfb7392f713d331571b43772ea5fb87f9240028b3095f44d36262

Observation aac84641-c28b-4673-86ee-cb524d096197 · outbound

This paper cites an unresolved cited work.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:59:02.003421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.484719Z digest=sha256:35dc0e19e6863ee9121e0bec661eeb0deae09a9159e3e3a2ae4881042fafa604

Observation bb69b9a2-e6f5-45fd-b009-dbb4adbd88f4 · outbound

This paper cites an unresolved cited work.

IC-Cache: Efficient Large Language Model Serving via In-context Caching Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:59:01.994092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T16:59:01.487312Z digest=sha256:70449d69684b5b72e22302d057e62d5c8fd3eb5a6f95442b62101ada807e21c6

Pith citing papers

Observation b0f988cc-6209-4367-8174-1ca3a132931f · inbound

CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems cites this paper.

CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems IC-Cache: Efficient Large Language Model Serving via In-context Caching

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:35.999530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:42:35.999530Z digest=sha256:507e4a257e0ff9f0a35fde1036ba91d5c2394869feb4ad62387ea5cb465f79d2