Pith. sign in

Paper Citation Record · LEDGER

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference

As of 12 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2607.20507.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.20507 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:58:08.181911Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 90b49b74-0e22-45b6-8ac5-94023f6cc8ee · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Advances in Neural Information Processing Systems , volume =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:04.635644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:04.635644Z digest=sha256:c774a6eb6257d4e018abebc5907ef7aea341d926f0337e5c435c1227aab7c406

Observation 16baf39c-1956-4da2-9250-d913920b9975 · outbound

This paper cites Transactions on Machine Learning Research , year =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Transactions on Machine Learning Research , year =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:04.714808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:04.714808Z digest=sha256:7c282b15ecd0733191c51a0c381a23ff1b8a7416e6279541457799194e923687

Observation 268a6c57-1b20-4e03-acd2-6185d20afc28 · outbound

This paper cites 2023 , volume =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference 2023 , volume =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:04.797819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:04.797819Z digest=sha256:8e2d9dfe67f30712050a11affaaca3ef5efffd69f21f556f3e35336faa6fe761

Observation 24610a3e-bf75-4c98-9182-2942538ff241 · outbound

This paper cites 2023 , url =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference 2023 , url =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:04.936357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:04.936357Z digest=sha256:f8dd1e4f634c907a289666a36604c987bf927f197759fc4f9853b4b57a66b3a9

Observation 516aff8b-85a5-4c06-ac77-23b9d7442bd7 · outbound

This paper cites Proceedings of the 40th International Conference on Machine Learning , pages =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Proceedings of the 40th International Conference on Machine Learning , pages =

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.030510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.030510Z digest=sha256:e682b5f1bb7215c5c75023a786a6a50bb0c9fc23665ef41c0a73b88a2a5123a7

Observation 675aff8d-7a30-4427-b801-c4a7f47ca0e1 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Accelerating Large Language Model Decoding with Speculative Sampling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.103454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.103454Z digest=sha256:f0165c41257739fce3ff4fa7f6fa76db35a2a43a8baeeb584cb710e5d1621567

Observation 67e44686-6a96-4b7e-9eaa-42f3d3d44e65 · outbound

This paper cites 2024 , publisher =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference 2024 , publisher =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.177707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.177707Z digest=sha256:07bb50666525928a7c3e1003ca0ef08a6965773e37854e4489d55def6f5f973b

Observation 8659404b-334c-4bd0-9e3a-650210a9c112 · outbound

This paper cites and Chen, Deming and Dao, Tri , booktitle =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference and Chen, Deming and Dao, Tri , booktitle =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.365116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.365116Z digest=sha256:2eee0d82b232f62371541dd84f1d865d02821b8f094be5f2d4a64275bc42ae28

Observation c987e5ba-0a3b-4b70-8003-9fd8ce8a9ffb · outbound

This paper cites 2024 , volume =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference 2024 , volume =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.394024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.394024Z digest=sha256:1564307e5afc77faa8b534dc8f649ef87a964ad5f27c958aac6dd1f858a22d97

Observation 26ce685d-ddb2-48cf-8eb1-4e963efc8dff · outbound

This paper cites Break the Sequential Dependency of.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Break the Sequential Dependency of

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.457930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.457930Z digest=sha256:6800bca9e8956c9cd9b596176a307b357dbd3c29fc268227afc3e3850e648227

Observation 942d0c99-998d-464c-87ca-35504046ffb9 · outbound

This paper cites an unresolved cited work.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.510043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.510043Z digest=sha256:0e1bb13f803057254ffc9848e8f8a94e9732de0774a272d871bb813ff720c18f

Observation 452c4d0e-da0a-4830-84e3-f69b0a3a79bf · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , pages =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Findings of the Association for Computational Linguistics: ACL 2024 , pages =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.669849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.669849Z digest=sha256:53df359bd034cd38dc4f2eb7e2973878ab56d350606bb2a7ab8aab57d12ec40e

Observation 09be17be-f8f3-44b2-b617-ee20682ec72e · outbound

This paper cites 2023 , address =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference 2023 , address =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.824091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.824091Z digest=sha256:d5538cc4bed0345a3c908a521c2bb11e4fa07a8b3c8fef0af718e11da4bc11c1

Observation 068aa273-416e-483d-8142-29866a5d3d24 · outbound

This paper cites Proceedings of the Seventh Annual Conference on Machine Learning and Systems , year =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Proceedings of the Seventh Annual Conference on Machine Learning and Systems , year =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:05.986193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:05.986193Z digest=sha256:e006b6c54edae08ea751a3cedf28dc273ff0273607ff1fb80391496cfb45edf7

Observation 3d07fdbd-a640-4bb8-975e-5ed2c1cd8115 · outbound

This paper cites MeanCache: User-Centric Semantic Caching for LLM Web Services.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference MeanCache: User-Centric Semantic Caching for LLM Web Services

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:06.169493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:06.169493Z digest=sha256:c4ad042a3729e0d6fafbe590778b8bc6989df3ad8dee054c80d53aecf5876de2

Observation 36b6fbfb-565b-4601-9a2c-2b3cae74170e · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Advances in Neural Information Processing Systems , volume =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:06.367143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:06.367143Z digest=sha256:f9d82c68d0db84f10e1aca1c3b1d7ee063519645f482a90183d21c46e3c19af4

Observation da5eb0de-d486-4885-be94-7280f4fa4f76 · outbound

This paper cites KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:06.513882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:06.513882Z digest=sha256:500a3b93c81906aae203fce15d4f0f7031d57b9d2ffb832d85cb226631bd9bb4

Observation b708ea27-3567-42ca-a72d-27a74d2b3016 · outbound

This paper cites Qwen3 Technical Report.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Qwen3 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:06.716260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:06.716260Z digest=sha256:d690b398a6132220dc8169c00f87f1639da4b036343df8829ee5c99bbd66c778

Observation b169a79f-5a0a-43b9-83c6-e2dbf5aa5980 · outbound

This paper cites Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing , pages =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing , pages =

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:06.854370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:06.854370Z digest=sha256:17d1f26de88ff06250f74c584cf4411ced23ae3915e209832c162bbe2539a954

Observation 9f652980-3369-4f4b-8804-b24b5a33ed3a · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Advances in Neural Information Processing Systems , volume =

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:07.038196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:07.038196Z digest=sha256:8826e1ef0dcbeca656e9616731855abab1ece4ad006df46e06e46c49520cf49a

Observation ac43dd8a-f21a-47ef-9732-0a01388b6ae8 · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Advances in Neural Information Processing Systems , volume =

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:07.142971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:07.142971Z digest=sha256:6d6c4400f5febb245b573dffc4f6f5bc6e23d1bfb750a9a4c2253408d68fb40e

Observation e81a133e-06aa-4a36-ad3c-ad4b9f5c0ae2 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:07.291602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:07.291602Z digest=sha256:eea13d3471b4aadec8036578c6e5849e506b03613f6cb990786f6c95a9051957

Observation f664783c-749e-496c-8d20-92e82c2a7108 · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Advances in Neural Information Processing Systems , volume =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:07.460354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:07.460354Z digest=sha256:653ce9387197a9cb47cf4886a82672b6b15f9701c679ca0f262a3652e2c9eed3

Observation fb5594d6-8ec9-416d-a30c-17b0611d0036 · outbound

This paper cites Scaling Laws for Neural Language Models.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Scaling Laws for Neural Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:07.611796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:07.611796Z digest=sha256:0ee58ab294186fac2cfdc65e119948b4c4cc52329ab04d22c6cf4032edf207ae

Observation ff9d53e7-7b09-4ae1-bb68-44a35ba17889 · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference Advances in Neural Information Processing Systems , volume =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:07.794419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:07.794419Z digest=sha256:790d7eecb138598a796f5dfb760a02cb39f19e655808faf382996ae0c29a511a

Observation bf4bba0c-d969-4a6a-97fb-4ea8a2a3d06d · outbound

This paper cites The Llama 3 Herd of Models.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference The Llama 3 Herd of Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:07.998164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:07.998164Z digest=sha256:f96abeeccd560b42d1d50737c05e771e17cd5ed489c433c313c032513bc406d8

Observation 5498f6bd-3387-4ca8-a9a4-e225290f5fea · outbound

This paper cites DeepSeek-V3 Technical Report.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference DeepSeek-V3 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:08.069399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:08.069399Z digest=sha256:13dcb039d15aa3fd327ebeefe6b7ca84c1c7c3cfa7f08b2fe0b907d5593f99f2

Observation 2a653bc0-de73-4415-a241-44863d1ff648 · outbound

This paper cites FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets.

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T08:58:08.181911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:58:08.181911Z digest=sha256:8ef11a823bd9148570ee30a30c235f6ebfbd201470b4be04cb97e057acac6758

Pith citing papers

No inbound Pith citation observations are available.