Pith. sign in

Paper Citation Record · LEDGER

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding

As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2607.22389.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.22389 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T04:58:48.072701Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4a5eb13f-1a44-446a-a3a8-4ad8c73ae0c6 · outbound

This paper cites Mistral 7B.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Mistral 7B

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:41.116950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:41.116950Z digest=sha256:0fe995504815cf1773043a8db2912f83ae2c0442ca76711e7764a99180a6b80c

Observation d393a459-3d3a-45d5-8569-c0bd97478b75 · outbound

This paper cites Qwen2 Technical Report.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Qwen2 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:41.221977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:41.221977Z digest=sha256:5e5d9e8a297094fd9ee42b51d35ab476008b00b426402bb447cc73a881f77935

Observation b9b2bd9c-f749-42eb-9cb7-980212cc8ceb · outbound

This paper cites Llama 3 model card,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Llama 3 model card,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:41.318115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:41.318115Z digest=sha256:1564eabf657b97a6e01a263205661f465d33c03ece8451d4ce97f02cb6d25c03

Observation da7909b2-9490-49a6-86de-26876e8dbc64 · outbound

This paper cites How long can context length of open-source llms truly promise?.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding How long can context length of open-source llms truly promise?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:41.469197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:41.469197Z digest=sha256:8085b642b207dd405b89ace9736f2a7cae58718b645457de572a077d682b477b

Observation aa6817ee-8830-46ac-a1cf-e0de2eed627d · outbound

This paper cites Language models are few-shot learners,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Language models are few-shot learners,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:41.553081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:41.553081Z digest=sha256:2977532506a2ce8158157f42f610a36189077814d0cb40a7c4fef0ecd75edb6d

Observation a9d9bc7c-6939-467d-bd4f-49f6c3eb6c83 · outbound

This paper cites Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:41.719217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:41.719217Z digest=sha256:069030b16706ac3069ecb3c6907981f4dd9c06a93f6122840f47dda43f2e81c4

Observation 265aae74-fe63-4ce2-bff8-76c421346c67 · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Longbench: A bilingual, multitask benchmark for long context understanding,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:41.837251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:41.837251Z digest=sha256:07add6b0b2a70b3b2a49249a5685e7886fcc3626baabbb6148690d446db9a795

Observation b66a1dd3-0fe1-46a0-9064-8ec084fbd602 · outbound

This paper cites P3-llm: An integrated npu-pim accelerator for llm inference using hybrid numerical formats,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding P3-llm: An integrated npu-pim accelerator for llm inference using hybrid numerical formats,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:41.915830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:41.915830Z digest=sha256:d0a932ddc6f58a4fb72bfd24598f206144604bb04b289b2cd16a8d991800af53

Observation 64458efa-d583-4833-b270-826b924a6558 · outbound

This paper cites A survey on large language model acceleration based on kv cache management,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding A survey on large language model acceleration based on kv cache management,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.060439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.060439Z digest=sha256:e33340849208175052322ec5e8e2980741e3be95a428d006e54c3dd131e0b394

Observation 1efbd9f7-260f-488c-b7d1-d0c26ae5ee0b · outbound

This paper cites Codec: Prefix-shared decoding kernel for llms,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Codec: Prefix-shared decoding kernel for llms,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.210027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.210027Z digest=sha256:ae2d4916dd568d105be8a1777d5df9686fcd07d806e0477f7deec82c1ad7e78c

Observation d7ddd9b6-e43d-45d6-aeea-f49e8bbffe2f · outbound

This paper cites Orca: A distributed serving system for transformer- based generative models,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Orca: A distributed serving system for transformer- based generative models,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.300437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.300437Z digest=sha256:7dd94add3f7d9c8d442037d0b90bf177f2e1584d15423b766a1ae6b458b1faf2

Observation f2bc9d2b-61b4-461b-8350-8052822ef91a · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Splitwise: Efficient generative llm inference using phase splitting,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.410203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.410203Z digest=sha256:9570545fe20cd6e7aeacdff69846c863d490330b2f9822ec95023a7fe5f98ad9

Observation 91226077-408b-4a62-8895-c59e6dead0e7 · outbound

This paper cites Kivi: A tuning-free asymmetric 2bit quantization for kv cache,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Kivi: A tuning-free asymmetric 2bit quantization for kv cache,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.511116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.511116Z digest=sha256:cda14e8e3e4699755eed189fac3379732b81702f509395351ccd1c51b9d858eb

Observation 9444d6e0-37ff-4da2-b119-2d5ac9ffbcad · outbound

This paper cites Efficient streaming language models with attention sinks,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Efficient streaming language models with attention sinks,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.619086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.619086Z digest=sha256:aa4e6fedbd60e4f0846a616a7afaa2506f0a5a0f74539efd449f48325c8cfab5

Observation 64b35529-fa0f-45f7-9700-bc509221e412 · outbound

This paper cites Duoattention: Efficient long-context llm inference with retrieval and streaming heads,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Duoattention: Efficient long-context llm inference with retrieval and streaming heads,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.712200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.712200Z digest=sha256:34366b45ea9243169f6e3100cceef4c43ed2fe4e648a4f3cd1918c882714c304

Observation 17e9a2ab-f619-4ccc-8439-672908bf54cf · outbound

This paper cites Snapkv: Llm knows what you are looking for before gener- ation,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Snapkv: Llm knows what you are looking for before gener- ation,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.821069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.821069Z digest=sha256:c901d9d6c1fc05355408e890dc512f4848834b343cfdfe105e709ba3bcbbcba7

Observation f01aef9d-f544-4f9e-81b1-22137343fa9d · outbound

This paper cites Sepllm: Accelerate large language models by com- pressing one segment into one separator,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Sepllm: Accelerate large language models by com- pressing one segment into one separator,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.925778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.925778Z digest=sha256:c4f4d873aa2e7782cbd602ea27786e26cb8e1bb6060af8c2cb596edbe08f1ea4

Observation 71c03b91-62dc-4df3-88b6-452aaa2d9d7b · outbound

This paper cites Lm-infinite: Zero-shot extreme length generalization for large language models,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Lm-infinite: Zero-shot extreme length generalization for large language models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.999165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.999165Z digest=sha256:bbee0d8ccce3ce97e946f41c113869df17e85bc029e16d45fae78bc9cb9d9d53

Observation 9f39ce39-6cc0-495b-966b-d541fcf165f6 · outbound

This paper cites Longformer: The Long-Document Transformer.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Longformer: The Long-Document Transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:43.070342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:43.070342Z digest=sha256:c74aac908a7ddcb2e6d6f9871bb319054385ecaa343a0fe9160d933632a53c8d

Observation 23a69917-216f-40a3-b323-06d044da6719 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding H2o: Heavy-hitter oracle for efficient generative inference of large language models,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:43.182267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:43.182267Z digest=sha256:c9e6764e9d8df9875a838301d22a52b80c7b87c1a1093299fac41232cc379945

Observation 54487552-cbc4-45c4-a1d4-2f33166ce04a · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:43.305148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:43.305148Z digest=sha256:20adf9abdc6d51936e9a3e070b00b15537afd9a27c3e469187811c7fc45adfed

Observation 318a4ef6-4f36-4e23-939a-687ecb7d8edd · outbound

This paper cites Kvo-llm: Boosting long-context generation throughput for batched llm inference,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Kvo-llm: Boosting long-context generation throughput for batched llm inference,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:43.394240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:43.394240Z digest=sha256:f12ce67315ad3b4adbfd52998e1445269240a0ef8983d8aed080e88837ee3f07

Observation 8e2cabfc-b120-4daf-86f8-207a0a322d8b · outbound

This paper cites Keyformer: Kv cache reduction through key tokens selection for efficient generative inference,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Keyformer: Kv cache reduction through key tokens selection for efficient generative inference,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:43.506772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:43.506772Z digest=sha256:92a661b970c1a67793615cd31bacdda65dcb53f2f74809be87d62aa747476733

Observation 313f3b18-dce6-4810-b5ce-f1a393fa2957 · outbound

This paper cites Alisa: Accelerating large language model inference via sparsity-aware kv caching,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Alisa: Accelerating large language model inference via sparsity-aware kv caching,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:43.606157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:43.606157Z digest=sha256:54c38457027f8167b53f7022a1e3d2b11d0d4b5ef8e86ac7c221356dae28604e

Observation 2d78b5e8-4f14-4873-93fa-b52b7a00c4d3 · outbound

This paper cites Mata: A memory-efficient attention accelerator for llms exploiting look-back kv cache pruning,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Mata: A memory-efficient attention accelerator for llms exploiting look-back kv cache pruning,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:43.696933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:43.696933Z digest=sha256:d6b40307046323640ed2c0bf57ad05b0716f727630f428f7f21b73e9bfc42156

Observation 88bcf029-6fc7-4492-8f9d-7280cd5933e4 · outbound

This paper cites Unicaim: A unified cam/cim architecture with static- dynamic kv cache pruning for efficient long-context llm inference,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Unicaim: A unified cam/cim architecture with static- dynamic kv cache pruning for efficient long-context llm inference,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:43.829479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:43.829479Z digest=sha256:28a099f69d93a768a25cb20d9fc45424e8ececb415bf7c5fa5b3f0f881e6ed43

Observation d746782c-237a-4ea6-850b-e105d2e2bfc5 · outbound

This paper cites Token-picker: Accelerating attention in text generation with minimized memory transfer via probability estimation,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Token-picker: Accelerating attention in text generation with minimized memory transfer via probability estimation,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:43.934750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:43.934750Z digest=sha256:ac06560b2fd2b69e29680fec144f5bc0761520a0f3e3731aaacd0eb3317d251a

Observation 79c477fa-51e5-4b0a-8746-14a8ec912734 · outbound

This paper cites Dias: Distance-based attention sparsity for ultra-long- sequence transformer with tree-like processing-in-memory architecture,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Dias: Distance-based attention sparsity for ultra-long- sequence transformer with tree-like processing-in-memory architecture,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:44.011211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:44.011211Z digest=sha256:92cfa3709d437771d0c0dc16f13aedcf667ee90df3a4d7554b8738b895c97e36

Observation fd5d4989-9e38-45d6-93e6-795bbd388645 · outbound

This paper cites Veda: Efficient llm generation through voting-based kv cache eviction and dataflow-flexible accelerator,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Veda: Efficient llm generation through voting-based kv cache eviction and dataflow-flexible accelerator,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:44.075196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:44.075196Z digest=sha256:871f67364d417eb1b5baab8d1c1008ad1bdc58395a3cda35860eba32875bcba4

Observation 20e531de-bec9-4f9a-b671-a9e308b1d187 · outbound

This paper cites Kv-cache oriented query-aware sparse attention accelerator with cross-stage precision-configurable digital cim,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Kv-cache oriented query-aware sparse attention accelerator with cross-stage precision-configurable digital cim,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:44.170208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:44.170208Z digest=sha256:761d4c33715d8c7590bbe0ac2492e9a685e038472a6608904d0cb59a7aa96bf8

Observation fc0f6cc6-865f-4b67-af00-7818900e3b05 · outbound

This paper cites End-to-end acceleration of generative models with runtime regularized kv cache management,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding End-to-end acceleration of generative models with runtime regularized kv cache management,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:44.321904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:44.321904Z digest=sha256:a225317614cf90aff58f779550f7020af0172db82263f708e9d9a684f1684585

Observation deeb2018-df6f-49d8-af1b-701c3b66cca0 · outbound

This paper cites Edgellm: A highly efficient cpu-fpga heterogeneous edge accelerator for large language models,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Edgellm: A highly efficient cpu-fpga heterogeneous edge accelerator for large language models,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:44.430466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:44.430466Z digest=sha256:364318b64a238cf64a55a6104d9dc2e0e1faa494d0b78e3625fa6dddb2dab5b7

Observation 082819a2-cf6e-4748-b2e9-7663597bb079 · outbound

This paper cites Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:44.529874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:44.529874Z digest=sha256:3344eb6d00e124eec747211346182e5a431150c761c94331dde241ebbbd3eeb9

Observation 63bd56cf-985d-4172-ba8b-b72d0e54149a · outbound

This paper cites Flightllm: Efficient large language model inference with a complete mapping flow on fpgas,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Flightllm: Efficient large language model inference with a complete mapping flow on fpgas,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:44.760684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:44.760684Z digest=sha256:1329a84e24c99ddf3565fc6b9711f88ae529b71616eda5d5ec7001c7f2b4abed

Observation 7f0df5c5-732d-48a8-902d-ee7f519581f4 · outbound

This paper cites Ofq-llm: Outlier-flexing quantization for efficient low- bit large language model acceleration,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Ofq-llm: Outlier-flexing quantization for efficient low- bit large language model acceleration,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:44.943518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:44.943518Z digest=sha256:b078a3bb73e9929a4d233e0a667979b037f65579e1f450337c09217e5b443e3f

Observation ca253343-8b28-448c-901b-7b5abf8a5582 · outbound

This paper cites Kv cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Kv cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:45.144166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:45.144166Z digest=sha256:173a6f0217efc50c18f30458835096a1d05e6d10c7b412598731f5704bf6b6c7

Observation e20f6a49-0c49-43c5-94ad-f2b9c415ab98 · outbound

This paper cites Apt-llm: Exploiting arbitrary-precision tensor core comput- ing for llm acceleration,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Apt-llm: Exploiting arbitrary-precision tensor core comput- ing for llm acceleration,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:45.303812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:45.303812Z digest=sha256:af7545c9f849b2abce4197a37597b817672cf8f65141a2c26058d4ed60e87301

Observation abc3d7e8-46f9-48a6-b05e-c1b19e0be6fb · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding A Survey on Efficient Inference for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:45.474492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:45.474492Z digest=sha256:90a85033391f6e6935f697ccbf59b9a2867b02788cfbb4f26f1bcd7a2dfb6483

Observation b3a0208c-2ad0-46de-b0b8-7c0713d15844 · outbound

This paper cites When to stop? towards efficient code generation in llms with excess token prevention,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding When to stop? towards efficient code generation in llms with excess token prevention,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:45.605674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:45.605674Z digest=sha256:4d89cf21bb97de1c9108c6ec048eda2773ee46eab28763a5b527ae02f2a3c659

Observation d6ec201e-38ec-4d68-a31d-44fe07b3fea5 · outbound

This paper cites Llmcompass: Enabling efficient hardware design for large language model inference,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Llmcompass: Enabling efficient hardware design for large language model inference,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:45.795483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:45.795483Z digest=sha256:e4bdc8fadd76171ea07528a9478cae77eec9c14159ff9d134ea795479bef6e4b

Observation 604e267d-01bf-49e1-805d-3ab2733c5df8 · outbound

This paper cites Energy cost modelling for optimizing large language model inference on hardware accelerators,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Energy cost modelling for optimizing large language model inference on hardware accelerators,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:46.038243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:46.038243Z digest=sha256:047cf8d66c347ee73fac3ff20636a8f15350f4e613d058619f97db3192e3e450

Observation 9168e695-d063-4ad4-9634-4fa85f8a9393 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:46.181845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:46.181845Z digest=sha256:1b4b743e5f524218d38282cc34e2d9c64ea1de6d4d6b8d891a788b834f213190

Observation 89491f21-4613-47d2-8be2-e7c1ee4649b6 · outbound

This paper cites Skipkv: Selective skipping of kv generation and storage for efficient inference with large reasoning models,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Skipkv: Selective skipping of kv generation and storage for efficient inference with large reasoning models,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:46.374845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:46.374845Z digest=sha256:242fa51e6f2f8cad3347deb290f4d5f90e9c3d875c84b1afb7937d9f2abfd844

Observation 9a0e496a-1570-4ff2-900b-8d6699310dce · outbound

This paper cites Titanus: Enabling kv cache pruning and quantization on-the-fly for llm acceleration,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Titanus: Enabling kv cache pruning and quantization on-the-fly for llm acceleration,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:46.548695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:46.548695Z digest=sha256:6a7c6ef0027704157ca32c5b69b4132950cc7e88ac377270727f07a7e2675193

Observation 605bc04d-cf93-41d7-a727-1a6322b69f05 · outbound

This paper cites Infinigen: Efficient generative inference of large language models with dynamic kv cache management,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Infinigen: Efficient generative inference of large language models with dynamic kv cache management,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:46.763518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:46.763518Z digest=sha256:59591e4d90aea870c02a93fb098df4102d55f80bd8d64843f2b2f3c25e60efe7

Observation 3fe095db-cf41-4bfe-b822-d71c0a6c13f6 · outbound

This paper cites Sparq attention: Bandwidth-efficient llm inference,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Sparq attention: Bandwidth-efficient llm inference,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:47.016141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:47.016141Z digest=sha256:6c43fd90100ea59c180a3d6a5052f2e8662097cf91b4ca3e53485ee2284e648d

Observation 610a51a3-5140-4cf8-873c-e171f7a54f2f · outbound

This paper cites Algorithm 232: Heapsort,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Algorithm 232: Heapsort,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:47.163798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:47.163798Z digest=sha256:ac46a5f21887b44712a1efbbe3cb5c0f035d969baf68f85782fa8161d89f362b

Observation f3832966-788f-4656-a365-4e0ff33f09a0 · outbound

This paper cites Sorting networks and their applications,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Sorting networks and their applications,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:47.269194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:47.269194Z digest=sha256:4bb020c7045cc3d78a4d5fd8f4e8015ff3074e85d63c157832998be6c5523367

Observation 1da2e8d6-b1c9-4d46-ab3b-8d346b62474f · outbound

This paper cites Efficient memory management for large language model serving with PagedAttention,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Efficient memory management for large language model serving with PagedAttention,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:47.421112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:47.421112Z digest=sha256:ef8deb48c951999e05e466264fd45d18207a7419e23458417279924fffabeccb

Observation 5bbae7ab-3070-43ee-8ecf-b4e80e042fe3 · outbound

This paper cites Anda: Unlocking efficient llm inference with a variable- length grouped activation data format,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Anda: Unlocking efficient llm inference with a variable- length grouped activation data format,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:47.516573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:47.516573Z digest=sha256:c8b9925e4293738942ba157ac15b3723d5bd400f28af23655282b85a707eac6a

Observation ef6216b1-3f70-4659-8ed8-230610832737 · outbound

This paper cites Ten lessons from three generations shaped google’s tpuv4i: Industrial product,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Ten lessons from three generations shaped google’s tpuv4i: Industrial product,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:47.631238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:47.631238Z digest=sha256:f26fffff17f7d2596d05f57245a646ba83d4bc4ac5bd84364f5a5d6d57c7f77f

Observation ab6c5676-5725-4df3-8714-ccfd8c9d5462 · outbound

This paper cites Dramsim3: A cycle-accurate, thermal-capable dram sim- ulator,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Dramsim3: A cycle-accurate, thermal-capable dram sim- ulator,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:47.771700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:47.771700Z digest=sha256:b391fce2c56f0cdfc7e684f9c36b059dec2df24330a8c875aebe2d056aff14e7

Observation 6d45b6d6-2d5b-49e7-83f1-77b9cf4008f3 · outbound

This paper cites Ada-kv: Optimizing kv cache eviction by adaptive budget allocation for efficient llm inference,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Ada-kv: Optimizing kv cache eviction by adaptive budget allocation for efficient llm inference,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:47.864332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:47.864332Z digest=sha256:1b7605a980701dce308536ae359998bdfd92b46d352e678fca2f9b1fcbc0cc03

Observation 469204c9-e339-46d2-a2b8-a54f42bcba56 · outbound

This paper cites Dynamickv: Task-aware adaptive kv cache compression for long context llms,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Dynamickv: Task-aware adaptive kv cache compression for long context llms,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:47.969320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:47.969320Z digest=sha256:beb640362bd8a69a505366bac4e19d1814dbd6d3d6bc47401495f58caf595bd0

Observation 9345b92f-f78d-4469-8e47-d447e96ef55d · outbound

This paper cites DeepScaleTool: A tool for the accurate estimation of technology scaling in the deep-submicron era,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding DeepScaleTool: A tool for the accurate estimation of technology scaling in the deep-submicron era,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:48.072701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:48.072701Z digest=sha256:faa2c515df7cddcc5e9d126e7861e98dc61d51c0d198429029eccfcbefe00d59

Pith citing papers

No inbound Pith citation observations are available.