Pith. sign in

Paper Citation Record · LEDGER

LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 59 inbound Pith citation observations for arXiv:2310.05736.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.05736 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 59 of 59 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:27:03.120485Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:36:44.322096Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f72dd5e3-3850-44c9-8af0-63c3c00b9f94 · inbound

A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions cites this paper.

A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:46:27.647448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T02:46:26.957539Z digest=sha256:fd7b3318f2755a218d5cbfc711ff8a142359e31da41a11ef806e217c6254ceb2

Observation bec1097a-429a-4dc0-bea3-2c1bb52f6112 · inbound

AdaComp: Extractive Context Compression with Adaptive Predictor for Retrieval-Augmented Large Language Models cites this paper.

AdaComp: Extractive Context Compression with Adaptive Predictor for Retrieval-Augmented Large Language Models LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:08:26.041444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-23T21:06:59.776841Z digest=sha256:a80de2883d346144e0b6111d757dca429cec6de6468627a88e0ef355568c644a

Observation 038d4429-156d-4e4e-89ce-8cdbf1905052 · inbound

E2LLM: Encoder Elongated Large Language Models for Long-Context Understanding and Reasoning cites this paper.

E2LLM: Encoder Elongated Large Language Models for Long-Context Understanding and Reasoning LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:38:24.967159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-23T20:36:55.159302Z digest=sha256:e4b0eaebad039806b4610649641822e35c8372d0e1b0c7537d141c4162ebc7b6

Observation 6afb25e3-bfd5-40c4-a724-80806a27129d · inbound

Compressed Chain of Thought: Efficient Reasoning Through Dense Representations cites this paper.

Compressed Chain of Thought: Efficient Reasoning Through Dense Representations LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:47:40.400393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T04:47:40.335477Z digest=sha256:363f001511659db7cb062e57ee5f1c1deb544c2ac4b2fdcbb0d54e1241d60e9d

Observation c0805d18-4818-4e42-8e52-cf6f52f50a36 · inbound

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation cites this paper.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.685933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.685933Z digest=sha256:33b132f555fe97e2b05e699c2ac15c0cc4413ab84f6b2bf4faf513695449d0a7

Observation e48c0a43-7e26-48da-9d2d-4c97cf3dd3fa · inbound

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention cites this paper.

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T23:46:30.078718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T23:46:29.975858Z digest=sha256:941a245b69c449b9c6d1d4679c1a90a32d35b16b58d223bd7d294f3115795b6f

Observation 6e2ceb50-3c8b-41dc-9d28-4039d3bc7f44 · inbound

Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation cites this paper.

Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T05:40:21.491572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:40:21.491572Z digest=sha256:d7103855085f67469fd0c452232b069b93f0b41b60d685eee558aa6ecb44ba81

Observation 445cf6ed-50af-4bb1-a32e-798ee654dd31 · inbound

QwenLong-CPRS: Towards $\infty$-LLMs with Dynamic Context Optimization cites this paper.

QwenLong-CPRS: Towards $\infty$-LLMs with Dynamic Context Optimization LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:52.446651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:52.446651Z digest=sha256:6bf90184d26037f2ceb61698da928803442306f5cb369561f6ebaffb3cd3c42a

Observation afafabaa-3390-46ed-81b2-6029bc351209 · inbound

A Survey of LLM $\times$ DATA cites this paper.

A Survey of LLM $\times$ DATA LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 189

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:13.213167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:13.213167Z digest=sha256:3daa52475ad8b5fd723574487df9be131e00f14519e3600bd129434c3612e2d5

Observation 7fb1ec33-d153-40f9-9093-18fd3a040d60 · inbound

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling cites this paper.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:05.729533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:05.729533Z digest=sha256:67994df6ebbd9a7d7952023896d6525f77fac2a4423fcedf082bb2f5ac335000

Observation 0427e401-f35a-4281-b7c8-c7fb90d3b9e6 · inbound

Lossless Token Sequence Compression via Meta-Tokens cites this paper.

Lossless Token Sequence Compression via Meta-Tokens LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:14:26.584176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:14:26.584176Z digest=sha256:a0604846cb9756548e89e53b6f89a551eea233357606aa3faecf95fa22bd7101

Observation 9fb66aae-c29f-4639-b5d7-2df33cc99457 · inbound

Cartridges: Lightweight and general-purpose long context representations via self-study cites this paper.

Cartridges: Lightweight and general-purpose long context representations via self-study LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:31.063414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:31.063414Z digest=sha256:42ade0a020f7f9f918dbcef201bdc85c2dd01c5b1d071371dc8e6f7b68724abf

Observation b22fdf2f-2736-4930-a9f5-d6438f68cbe7 · inbound

Brevity is the soul of sustainability: Characterizing LLM response lengths cites this paper.

Brevity is the soul of sustainability: Characterizing LLM response lengths LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:09.648283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:09.648283Z digest=sha256:201762766fe33b9b47328be6e1d8a5e314739eecf289a727c63064a7f762dc89

Observation 12d84b6d-e533-4cd1-aaac-86815c522721 · inbound

LoRA-Gen: Specializing Large Language Model via Online LoRA Generation cites this paper.

LoRA-Gen: Specializing Large Language Model via Online LoRA Generation LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:24.837387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:09:24.837387Z digest=sha256:07fd3a7e4707138a7dc2d2003c8692aa9e18d6863ff9449999707147856f6ef9

Observation ea12e6e0-b53f-4cc3-a8ce-11b12b42f102 · inbound

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent cites this paper.

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:17:24.592535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T11:17:24.406028Z digest=sha256:6737fc9d60365422d3d5b3744bf98420fd7c3a58467210c601624402eb8aa4c0

Observation d4cf843a-3838-4780-ac24-684bd888cb39 · inbound

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent cites this paper.

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:40:10.564649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:40:10.564649Z digest=sha256:6902afdd01146d01bcd7f5ab2af7509c2237ae51b639a4363f80a93bb5b30404

Observation 9c490a55-a641-41ed-979e-524003d7b96a · inbound

DaMoC: Efficiently Selecting the Optimal Large Language Model for Fine-tuning Domain Tasks Based on Data and Model Compression cites this paper.

DaMoC: Efficiently Selecting the Optimal Large Language Model for Fine-tuning Domain Tasks Based on Data and Model Compression LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:01.641529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:52:01.641529Z digest=sha256:1ab6620cb8b5371631392ac503647cf71e00196dd97e64727b87099a8b879a4c

Observation cc7a1180-8b1d-4f7c-8378-d88f4f88a128 · inbound

How can we assess human-agent interactions? Case studies in software agent design cites this paper.

How can we assess human-agent interactions? Case studies in software agent design LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T10:36:37.392346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:36:37.392346Z digest=sha256:cbc8dbb4e5723afb564d7c02a95c57746fd58c8d6337172421c721940de8f640

Observation 9fd21ffa-571e-4863-abbb-a76960a04134 · inbound

ARC-Encoder: learning compressed text representations for large language models cites this paper.

ARC-Encoder: learning compressed text representations for large language models LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T08:28:49.536406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:28:49.536406Z digest=sha256:723eb69a80ebebd9b8e1afc72635cc035c246eb86c26dc2857acb9cc1dd81c76

Observation 32c3b362-8b83-428a-a01d-689ab3ba003b · inbound

When Compression Becomes an Attack Surface: Black-Box Attacks on Prompt-Compressed LLM Agents cites this paper.

When Compression Becomes an Attack Surface: Black-Box Attacks on Prompt-Compressed LLM Agents LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T08:06:11.078142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:06:11.078142Z digest=sha256:84076d8248c808652776a37c73f036f3f5c9dedce0c56b72cdec4f68b3ce4c57

Observation 1840f4d0-04f9-4b48-9270-0fabf218d965 · inbound

Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression cites this paper.

Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:44:11.449562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T13:43:51.127429Z digest=sha256:d1afe2fe5f9769e4db993a8996102aca29ebc59b1dafefdbad74dc45adba56c8

Observation 33b6f1fa-13fa-4b47-9045-1a7d96d29872 · inbound

Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression cites this paper.

Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T03:25:23.076743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:25:23.076743Z digest=sha256:d26a696bf9f5fe9e1d2dfba4bf7b5f78b312d2822151b4719e30c72ea4857b20

Observation 24041952-da97-4dc5-9b99-fdc3f1637114 · inbound

Learning to Configure Agentic AI Systems cites this paper.

Learning to Configure Agentic AI Systems LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T13:10:10.359193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T13:06:56.207692Z digest=sha256:ac164741d60fc64cdcdd6c095595d0da5933a20e10747f870fc60e9f038c6230

Observation 10c5e7c7-b318-48a3-8ac3-99ae685b5df0 · inbound

Learning to Configure Agentic AI Systems cites this paper.

Learning to Configure Agentic AI Systems LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T11:31:29.533983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T11:29:45.134060Z digest=sha256:23ecb2219c54cecdbeaaf3381cf3efdf7245118b45824b547d4cc5536feb7eef

Observation b67fc6f0-ae67-4b8b-80ce-9d7915ae65f8 · inbound

LLM-assisted Agentic Edge Intelligence Framework cites this paper.

LLM-assisted Agentic Edge Intelligence Framework LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:00:03.440252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T13:56:03.688192Z digest=sha256:4fbbf6972feac4334bc76cd128144c7ad5d6253a71f03c0d55b1ab42d19d1754

Observation f698696a-d1b7-4447-b945-3c592ab986b2 · inbound

On the Effectiveness of Context Compression for Repository-Level Tasks: An Empirical Investigation cites this paper.

On the Effectiveness of Context Compression for Repository-Level Tasks: An Empirical Investigation LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:40:27.553422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:00:56.512868Z digest=sha256:45a44c169c9951fd3f3531eaf568ee88bbd51d37ba2669727e030201201f2d30

Observation 3dc12810-508d-4da3-ab4e-f63eee353003 · inbound

Compressed-Sensing-Guided, Inference-Aware Structured Reduction for Large Language Models cites this paper.

Compressed-Sensing-Guided, Inference-Aware Structured Reduction for Large Language Models LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:55:10.698328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T06:52:18.563146Z digest=sha256:6299099a6f8c2276df4063d10e07dcb8883fb0d1b5f4e5b690cb183f52e896b5

Observation d77242a0-5e09-4624-adee-921596b1bbc4 · inbound

ONTO: A Token-Efficient Columnar Notation for LLM Input Optimization cites this paper.

ONTO: A Token-Efficient Columnar Notation for LLM Input Optimization LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.788133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T05:21:56.826488Z digest=sha256:d8c655af6cfeaea8d6a3d9ece85d314a7c6b67ea0c8e72f697e2c8710d7d183f

Observation ff9c3f20-8131-48f6-a877-a1e59cafc4cf · inbound

Supplement Generation Training for Enhancing Agentic Task Performance cites this paper.

Supplement Generation Training for Enhancing Agentic Task Performance LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:59:49.528114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T00:58:27.655909Z digest=sha256:1c527b6a0a2f379b4d828d469f72199ed8f694794616b626e0e96ee4b61d9456

Observation 075c5251-e277-4d6d-aeec-c40261aeb228 · inbound

SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference cites this paper.

SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:07.779041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T14:09:30.821354Z digest=sha256:10b475860cf4df9432aa9e497ffe55a5d17cfe64315413cace8cc9c44d323f35

Observation a234ca39-dfe8-4eb4-b367-d6070dd1cf0f · inbound

OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory cites this paper.

OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:31:25.537117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-07T10:45:48.976501Z digest=sha256:c476f06da07a5d0f05a06519c42272b01f9f6f2b82446adf6b27cf007c3dbf92

Observation 3726b935-c200-4b81-b842-0f3460351fb0 · inbound

Budget-Aware Routing for Long Clinical Text cites this paper.

Budget-Aware Routing for Long Clinical Text LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:10.192361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T19:59:00.888734Z digest=sha256:5ca1a7aae6addd8a848af4d96d9ac3d4ea0c22b2aa9306e22c0589e3bd9a6b90

Observation adc10e6b-b90c-4c19-b7d5-189f57fa8d6b · inbound

LLM-Oriented Information Retrieval: A Denoising-First Perspective cites this paper.

LLM-Oriented Information Retrieval: A Denoising-First Perspective LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:01:19.749858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:54:06.144968Z digest=sha256:19c5c3629f1497e1a41b21729824eeadb6ef942dbaf5240adff65f9b704991f9

Observation c2fec4e9-8772-4d5f-9473-5a56d9782452 · inbound

LLM-Oriented Information Retrieval: A Denoising-First Perspective cites this paper.

LLM-Oriented Information Retrieval: A Denoising-First Perspective LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:19:16.570350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T00:18:32.423103Z digest=sha256:9079ce260742336d198d9bfc14142d79080cb6f61a492dbe6e3c01154cc68d78

Observation 570d6085-1061-43d9-ba01-bf90ae8330e2 · inbound

Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks cites this paper.

Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:24.196333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:18:11.836537Z digest=sha256:c4bb228669f24575d5d42833174a1706560513ed0e83d78f8f2c4ea49044907a

Observation 0ce26496-205a-4f76-9f79-39760693bd75 · inbound

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions cites this paper.

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:48.513036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T15:17:37.904831Z digest=sha256:355738dd3a091ea43ae00f7025e46e4706233ac26063d00df7277d54f72a82b1

Observation 7f1a60ae-bca1-40cc-b45f-89d97c3143b8 · inbound

Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets cites this paper.

Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:24:01.270362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T23:23:49.298204Z digest=sha256:9eb951ae3c3691a922bfc658843e6c666c07f89d43f2e9e53a03fa2b3afa5a07

Observation 94c48765-afd0-491c-8695-d7eb7ac9dc0e · inbound

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories cites this paper.

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:26.695917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:56:13.058872Z digest=sha256:60212167c17a6728ac2a9a389c11b0da7bc8574d76e5138f5aa04e8ba40b1459

Observation 53d80952-6f94-4501-b0fa-8058eeabff75 · inbound

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories cites this paper.

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-13T07:44:25.325808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:44:25.325808Z digest=sha256:f938e31403416b15e49a7d5f2d1df014f7478b3e187547c4dc07746f348c6114

Observation ddddc700-5c17-4948-8aa4-b4d4e80ec9e8 · inbound

QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving cites this paper.

QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:57.478212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T01:47:43.240850Z digest=sha256:a3ab1cbffac429aac31c74579d026d40bb46d0f7cb3b9c643550dc5b946122ac

Observation 2736a578-a94a-418a-9eeb-13cec35d42d4 · inbound

Decision-Aware Memory Cards: Counterfactual-Inspired Context Selection and Compression for Tool-Using LLM Agents cites this paper.

Decision-Aware Memory Cards: Counterfactual-Inspired Context Selection and Compression for Tool-Using LLM Agents LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:25.260446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T19:47:08.844690Z digest=sha256:e3671ab016a0146a819feea2c1eb903fc1328a8902d89e5d39f683c4622d3420

Observation ca86e023-c913-4143-8476-247483c077a5 · inbound

End-to-End Context Compression at Scale cites this paper.

End-to-End Context Compression at Scale LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:31.616571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T16:36:54.699174Z digest=sha256:5375f18d7eb726f19ae7b73df4d340be9acbac6bbb2809d4d19651aca429b9ee

Observation 38dda257-2d76-4c89-a68f-c7290b5a80a0 · inbound

Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents cites this paper.

Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.619433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T16:10:25.191306Z digest=sha256:d1063de88341d79e7260ae67b7fb7a8a4e89604a2ac812242a36ce42470fd202

Observation 4555f900-c28b-4ff1-b6fd-9311a2aa28a7 · inbound

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns cites this paper.

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:19:23.905129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T19:47:59.090249Z digest=sha256:d99efe46708604730dbfcf21429a3a6326b53e8f8d4b2228c9d74d1312169c74

Observation e1434e72-7500-43ac-98d2-1752019b4273 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 130

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:09:36.936832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T16:15:22.543601Z digest=sha256:055c920ff38f7ed571b149ee073e492a40f29cd17521d69e21c811d8e72018a9

Observation a3c17276-531a-447f-b190-2fb8246f4ece · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:13.256103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:13.256103Z digest=sha256:2df19a8ff9631c49a17eb3de84af59b883d118a00970c55995011bd688690684

Observation 06e93a9e-23a9-4284-a100-c4a53829084e · inbound

HMARS: A Hierarchical Multi-Agent Memory System for Long-Context Reasoning cites this paper.

HMARS: A Hierarchical Multi-Agent Memory System for Long-Context Reasoning LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T11:34:37.884491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T11:30:52.871762Z digest=sha256:af3ed72ff4df40bc296de55bdda5c54ec848a206cdf84a1cc4c0e1d1ba161688

Observation 4404e395-a8a2-428c-8e87-7a2eed6babba · inbound

Mapping Text to Multiplex Graph: Prompt Compression as L\'evy Walk-Guided Graph Pruning cites this paper.

Mapping Text to Multiplex Graph: Prompt Compression as L\'evy Walk-Guided Graph Pruning LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:49:21.550972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T01:45:36.335508Z digest=sha256:7fd89c488ad2fe3bb0ec620dca6f36c1a2f62ad54b23ff8b06e5ea7ea40f12c3

Observation b4d2f544-20d4-495b-ae9b-7a7cbd752945 · inbound

ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair cites this paper.

ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:18:22.569952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-03T14:09:32.980488Z digest=sha256:8aea6a1582d09a51291d029d485196e1a53f1de6a61eb6a87333901fc5bd6cd2

Observation d4c13e6d-a18b-4ed4-bcea-7c1ee9e55a06 · inbound

ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair cites this paper.

ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T08:32:14.868469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:32:14.868469Z digest=sha256:bcc3c0f7a2464aaa214064fa5b1f9e3aa845d9528227b8400d07ac295cd91d9c

Observation 051587d5-d924-4796-a7b1-950773355e3a · inbound

Compression, structure, and executor capability: a controlled real-cost decomposition of language-model agent skill optimisation cites this paper.

Compression, structure, and executor capability: a controlled real-cost decomposition of language-model agent skill optimisation LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T05:14:41.315225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:14:41.315225Z digest=sha256:f1a1961d3d668814f3b0a45b95c469a9d579c6aebc9a6b2d3bbad9716d1d695b

Observation c0bd2724-36b3-4170-98b8-ba9ce4697d40 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.323273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:6b724805231ea90ffd989a17d3df08024d2f708a3c081c4113b1c4965a5b6081

Observation 410403c7-29ae-4e64-a6a4-897774022f78 · inbound

Mach-Mind-4-Flash Technical Report cites this paper.

Mach-Mind-4-Flash Technical Report LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T03:29:34.486347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:29:34.486347Z digest=sha256:7e9e877394e8a6a2af53a67f9d0a3c3fd02ad0fe550e0dc0e48ab19cedcf223d

Observation 3a5f3d21-5cf8-4a9d-800c-d8a8c972387b · inbound

What Context Does a Coding Agent Actually Need to Act? cites this paper.

What Context Does a Coding Agent Actually Need to Act? LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T17:40:33.397259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:40:33.397259Z digest=sha256:af7f31a07d702115f046845140abd4dc5499cbc8752f90131d5c7c5cce09be10

Observation fd748da0-4d6e-40ff-9f9d-b36300ca1cc9 · inbound

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching cites this paper.

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:54.635607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:54.635607Z digest=sha256:f686104df83e207d19ad96b8013a62dedd02068660046c320f3a7ed7454bb91b

Observation 0a17e1cd-b03b-4492-a933-1c39358fb63b · inbound

Shapley Context Pruning: A Cooperative Game Perspective for Context Reranking and Pruning cites this paper.

Shapley Context Pruning: A Cooperative Game Perspective for Context Reranking and Pruning LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T14:32:14.315923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:32:14.315923Z digest=sha256:b57a7c21c050acf91304001accd497dde0821e8dbaec1f014b259485c50d5e4e

Observation e27cbc51-3fc1-4a48-905d-fe59b0e9ded6 · inbound

Hierarchical Domain Generalization cites this paper.

Hierarchical Domain Generalization LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T20:54:09.109669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T20:54:09.109669Z digest=sha256:26e3c9bfcd9f1c00a75b2baf437fcce0fe18d9ff4e1b35eef73d02c794ba51e9

Observation b0406970-d688-4c57-a030-8dbe36cf166a · inbound

Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing cites this paper.

Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T11:37:06.990703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:37:06.990703Z digest=sha256:68a3e2767f277f43e6476e77b64d7bc008bea790739147a8e59a0402fd66aee4

Observation f2c13c4a-5ffe-4121-8948-f0309e808f58 · inbound

MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents cites this paper.

MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T15:27:03.120485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:27:03.120485Z digest=sha256:90e3c70d99d2e482dbba7317c4affe5a23314fbfdd47b31829f8ee970b61950a