Pith. sign in

Paper Citation Record · LEDGER

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching

As of 20 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2607.15516.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15516 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T23:14:56.217903Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a0809f22-7d37-49db-bcae-391ce1b28604 · outbound

This paper cites ACM ICAIF ’25 Workshop on LLMs and Generative AI for Finance.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching ACM ICAIF ’25 Workshop on LLMs and Generative AI for Finance

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:54.266217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:54.266217Z digest=sha256:7ebe8e42c7ecd2609fffd534fd883a3839dc02797f5d5416fcba5790c990b6ae

Observation 19eeaa77-46ec-4283-8dda-3a9df4e977ba · outbound

This paper cites A Unified Approach to Routing and Cascading for LLMs.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching A Unified Approach to Routing and Cascading for LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:54.415492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:54.415492Z digest=sha256:ed12431a3e64e1fd6c35e20750a017eed865663bb3d130169b48587d9bb0d047

Observation e02edf72-1abe-4718-9950-6e60fbc097ef · outbound

This paper cites Auditing Prompt Caching in Language Model APIs.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching Auditing Prompt Caching in Language Model APIs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:54.505885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:54.505885Z digest=sha256:deafea359a0c66815673cf9b90f716d3f45db6d7ebb3371be17a5b575e666f0a

Observation fd748da0-4d6e-40ff-9f9d-b36300ca1cc9 · outbound

This paper cites LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:54.635607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:54.635607Z digest=sha256:207d9c9069e1878c8c77b7c6a0adf2e1f004916f17b46aff5e3f0824ffc2754a

Observation 46dea7f4-22b0-4d06-a52d-df2dee069eef · outbound

This paper cites LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:54.779144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:54.779144Z digest=sha256:f3b9c7cd64bfe9221a7faf6f426a7742b37ed76e471ac7a501e44a5feb0e9c11

Observation 0a4b7d72-aae2-46c4-93b9-407f4c3e6bdf · outbound

This paper cites 500xCompressor: Generalized Prompt Compression for Large Language Models.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching 500xCompressor: Generalized Prompt Compression for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:55.032941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:55.032941Z digest=sha256:1adbb6b0607561bfbda3df5c9577249864b6ba2ed2f8850757a00b21bec03e99

Observation 5dbfb937-72ed-4373-8c73-3dc72786679e · outbound

This paper cites an unresolved cited work.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:55.091512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:55.091512Z digest=sha256:6cd183942597d3991ddc9f5a9cf072927c7d4a34369cd8e4ff4cbd79676d877d

Observation 5c176990-f89c-46d2-a762-9301e750ae2d · outbound

This paper cites Google; 15 authors.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching Google; 15 authors

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:55.190278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:55.190278Z digest=sha256:70194fc8961d0a814b266aea73ac4151cea2815b7e374ecdcc11da07d237afb6

Observation 3a274ccf-32db-45c9-8427-627e218378d6 · outbound

This paper cites Evaluation across OpenAI, Anthropic, Google on DeepResearch Bench.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching Evaluation across OpenAI, Anthropic, Google on DeepResearch Bench

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:55.353824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:55.353824Z digest=sha256:b71bbc2cf55953518a81493cc962231353ca237fda9669ce876c006e80bbbb78

Observation 4d43ba17-ac2f-4da1-bcc0-d9bfa014d235 · outbound

This paper cites Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:55.449411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:55.449411Z digest=sha256:00eabe50e965cfab3cb2d92de9b4dc4f974d49a8d4c3d84429dd9a60d260f2a6

Observation e9066041-dd6e-4d9b-b383-804dfec0f00f · outbound

This paper cites RouteLLM: Learning to Route LLMs with Preference Data.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching RouteLLM: Learning to Route LLMs with Preference Data

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:55.597409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:55.597409Z digest=sha256:9526e1c7f7cdc15c5050c78c27c1299b15401616279453610d09e62aded3899c

Observation 0601da62-3a69-4f86-925e-e54011795ef7 · outbound

This paper cites Salesforce AI Research + UIUC.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching Salesforce AI Research + UIUC

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:55.772582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:55.772582Z digest=sha256:a208cf288602cd9c05427693fa4d4b65e257222f89d4185edbc7c9f8f2b6b1dc

Observation a733540b-2c0a-402e-9374-f26b32b22f98 · outbound

This paper cites Accepted ICLR.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching Accepted ICLR

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:55.841987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:55.841987Z digest=sha256:60270f951799c0dfd078f0dbd257bfcbe6ad461ae7859c2c5763869bae694496

Observation b7a8842c-8dd5-4631-8fcf-d1e7c6c76164 · outbound

This paper cites Converts code, documentation, PDFs, and images into a NetworkX knowledge graph with Leiden community detection and LLM-extracted concept edges.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching Converts code, documentation, PDFs, and images into a NetworkX knowledge graph with Leiden community detection and LLM-extracted concept edges

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:55.923546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:55.923546Z digest=sha256:9dc9b6b3382bcca71ebd892d82f6995f2d5090f897988da7cf9c6691eca3e643

Observation bf0e75c4-ceb1-4ce9-aa4e-8197b8d7d40c · outbound

This paper cites Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan.τ-bench: A benchmark for tool-agent-user interaction in real-world domains.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan.τ-bench: A benchmark for tool-agent-user interaction in real-world domains

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:55.980635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:55.980635Z digest=sha256:9ddc97c09a26fee80e2ecfe5a6f35dab61c6687280dc33ab61b108bafce78523

Observation 619f889d-f4dc-42ee-af05-cc6a2eb03f25 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:56.055295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:56.055295Z digest=sha256:6b4239f4c6a243530b744f562a8cb381cb1a9c891ccb5d8f9459c5824b53fc18

Observation bb75c710-b8c2-4b17-a3fb-8b8755367b24 · outbound

This paper cites Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:56.115834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:56.115834Z digest=sha256:d5f52717aa23534a45a4ed0f808f35b4adff17af897c41a745eb1c7c327263c7

Observation 2801332b-4c07-439b-bc95-e961a678a4d8 · outbound

This paper cites "" Feed one observed document version.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching "" Feed one observed document version

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:56.217903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:56.217903Z digest=sha256:19b6d4ab7c7ee809aeb6414d7ba437d6fdc55acd4d44f172d3069b6078afc87c

Observation 540d357a-8690-43e4-ac86-9461ff7113a7 · outbound

This paper cites FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:54.054486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:54.054486Z digest=sha256:0e0d1c68e927588da50939aa08daf59227af11cda118312ec2318cea418d100b

Observation a7f1a10d-9790-47e9-a803-8c988cda7288 · outbound

This paper cites LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:53.715023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:53.715023Z digest=sha256:03bacd71cb0adad16b8d283ac284fa680b32d0f341533d608d07d1936ce4a109

Observation ffe15007-a4c7-4182-82c4-3865fb982245 · outbound

This paper cites NeurIPS 2025 Poster; 83% hit rate on production prompts.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching NeurIPS 2025 Poster; 83% hit rate on production prompts

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:53.875090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:53.875090Z digest=sha256:1e6cf80679987e90e51a0805241f335d7fa94b76a06cd067f5e7a009760f54f8

Observation f79c36e8-5ca0-49ed-ae22-5db0b9c5880d · outbound

This paper cites Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:54.909730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:54.909730Z digest=sha256:13cf0c2edf5eacdbeb82c18cf02b704341440afb7022f247216e90a35c6e8a71

Pith citing papers

No inbound Pith citation observations are available.