Pith. sign in

Paper Citation Record · LEDGER

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding

As of 10 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2607.22389.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.22389 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T04:58:48.072701Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4a5eb13f-1a44-446a-a3a8-4ad8c73ae0c6 · outbound

This paper cites Mistral 7B.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Mistral 7B

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:41.116950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:41.116950Z digest=sha256:7effe1a4eab1446db70942d6647dcd6bb5ba534e00d44ed717f6529d24ee27fe

Observation d393a459-3d3a-45d5-8569-c0bd97478b75 · outbound

This paper cites Qwen2 Technical Report.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Qwen2 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:41.221977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:41.221977Z digest=sha256:89c3485172cedc202c0715fc550d781e7cae80f99f2c42c03d1e7e03c6f37ab9

Observation b9b2bd9c-f749-42eb-9cb7-980212cc8ceb · outbound

This paper cites Llama 3 model card,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Llama 3 model card,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:41.318115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:41.318115Z digest=sha256:9010334f3e3d4d5ad1cdc8070fbd4dbf815471e937004c47eaa5f283a0ebfcb1

Observation da7909b2-9490-49a6-86de-26876e8dbc64 · outbound

This paper cites How long can context length of open-source llms truly promise?.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding How long can context length of open-source llms truly promise?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:41.469197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:41.469197Z digest=sha256:87133f8725bdf20b3b57bfff016cc89f63ac6744d6c978f67c409d2277cbed9c

Observation aa6817ee-8830-46ac-a1cf-e0de2eed627d · outbound

This paper cites Language models are few-shot learners,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Language models are few-shot learners,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:41.553081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:41.553081Z digest=sha256:f5cefedb2fcd133974f6ca9b8123608c13b2351cdce5b185116a3e5975ce0d56

Observation a9d9bc7c-6939-467d-bd4f-49f6c3eb6c83 · outbound

This paper cites Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:41.719217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:41.719217Z digest=sha256:d21ce8b3f3864dc993ed401ba0dcefa2ed8106f54c2119fe2bd0150b9acea67f

Observation 265aae74-fe63-4ce2-bff8-76c421346c67 · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Longbench: A bilingual, multitask benchmark for long context understanding,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:41.837251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:41.837251Z digest=sha256:42a597280dd478edf964ac2d0aa5ad3abb815c498fa9049f3ad4142d6c66a067

Observation b66a1dd3-0fe1-46a0-9064-8ec084fbd602 · outbound

This paper cites P3-llm: An integrated npu-pim accelerator for llm inference using hybrid numerical formats,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding P3-llm: An integrated npu-pim accelerator for llm inference using hybrid numerical formats,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:41.915830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:41.915830Z digest=sha256:e32a85891c94b1d45f9201a4cdee8294ff1f5b022909d06d6a357509a29bb308

Observation 64458efa-d583-4833-b270-826b924a6558 · outbound

This paper cites A survey on large language model acceleration based on kv cache management,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding A survey on large language model acceleration based on kv cache management,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.060439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.060439Z digest=sha256:fda6c5058358843f8aaee88ed7e34e01cf7724697d85f5e665144f862e91762a

Observation 1efbd9f7-260f-488c-b7d1-d0c26ae5ee0b · outbound

This paper cites Codec: Prefix-shared decoding kernel for llms,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Codec: Prefix-shared decoding kernel for llms,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.210027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.210027Z digest=sha256:fdc584eadab4adaa0f0c6511ee8457fee1eb41c0151845e77e4555e525ee24c7

Observation d7ddd9b6-e43d-45d6-aeea-f49e8bbffe2f · outbound

This paper cites Orca: A distributed serving system for transformer- based generative models,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Orca: A distributed serving system for transformer- based generative models,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.300437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.300437Z digest=sha256:ef71d866e353334d2ef843c903fad9c8cb8178052a9665e8db273a9099d004d2

Observation f2bc9d2b-61b4-461b-8350-8052822ef91a · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Splitwise: Efficient generative llm inference using phase splitting,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.410203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.410203Z digest=sha256:66fa8d5310d144469db942e9260c3fe9a728da69589ca933481c41d28d95850d

Observation 91226077-408b-4a62-8895-c59e6dead0e7 · outbound

This paper cites Kivi: A tuning-free asymmetric 2bit quantization for kv cache,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Kivi: A tuning-free asymmetric 2bit quantization for kv cache,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.511116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.511116Z digest=sha256:57bb4f5270a96b9ac2d163554d2880fdd402461114867eb8b856bb8efab86a80

Observation 9444d6e0-37ff-4da2-b119-2d5ac9ffbcad · outbound

This paper cites Efficient streaming language models with attention sinks,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Efficient streaming language models with attention sinks,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.619086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.619086Z digest=sha256:6c41f3c1c6c4f868258a9a9925efe32a09fd4471702cba4b2d670f95bb0c1a14

Observation 64b35529-fa0f-45f7-9700-bc509221e412 · outbound

This paper cites Duoattention: Efficient long-context llm inference with retrieval and streaming heads,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Duoattention: Efficient long-context llm inference with retrieval and streaming heads,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.712200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.712200Z digest=sha256:7839c1eabdebd3fbb9fe551da92700153bcab87c558bc51ea5ef09366c5f8891

Observation 17e9a2ab-f619-4ccc-8439-672908bf54cf · outbound

This paper cites Snapkv: Llm knows what you are looking for before gener- ation,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Snapkv: Llm knows what you are looking for before gener- ation,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.821069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.821069Z digest=sha256:02ba103b90ef7f582ce114f26bfe1d2907f25a4c617a7b0b72cb3916167e8233

Observation f01aef9d-f544-4f9e-81b1-22137343fa9d · outbound

This paper cites Sepllm: Accelerate large language models by com- pressing one segment into one separator,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Sepllm: Accelerate large language models by com- pressing one segment into one separator,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.925778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.925778Z digest=sha256:9c98f7bd4a24f4feb2be19879b7fa67427b2979c5aad67dbe4b50c623518ddcb

Observation 71c03b91-62dc-4df3-88b6-452aaa2d9d7b · outbound

This paper cites Lm-infinite: Zero-shot extreme length generalization for large language models,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Lm-infinite: Zero-shot extreme length generalization for large language models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:42.999165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:42.999165Z digest=sha256:6b174f06687fd8b72c589cc2c29b6b6aff9bed2bbc2cb10f333e82fee39e91cb

Observation 9f39ce39-6cc0-495b-966b-d541fcf165f6 · outbound

This paper cites Longformer: The Long-Document Transformer.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Longformer: The Long-Document Transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:43.070342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:43.070342Z digest=sha256:a41c0485b349c238045dd5f342c3a7088d1dd91ea318770c68d5a167e4bee5a9

Observation 23a69917-216f-40a3-b323-06d044da6719 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding H2o: Heavy-hitter oracle for efficient generative inference of large language models,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:43.182267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:43.182267Z digest=sha256:5151134c24ce90dfdf01e2c3f3af87d5c8c8e43553a4d4c74ac71bf89ed99635

Observation 54487552-cbc4-45c4-a1d4-2f33166ce04a · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:43.305148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:43.305148Z digest=sha256:df0432110875580bfec0623b80d23ad9b50487e5fd1e9eaae719402fce7c8a03

Observation 318a4ef6-4f36-4e23-939a-687ecb7d8edd · outbound

This paper cites Kvo-llm: Boosting long-context generation throughput for batched llm inference,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Kvo-llm: Boosting long-context generation throughput for batched llm inference,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:43.394240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:43.394240Z digest=sha256:fe9556db8e651a0dca051708c895bcc4bc66f01150d2ceebd0036d59b26da4cb

Observation 8e2cabfc-b120-4daf-86f8-207a0a322d8b · outbound

This paper cites Keyformer: Kv cache reduction through key tokens selection for efficient generative inference,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Keyformer: Kv cache reduction through key tokens selection for efficient generative inference,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:43.506772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:43.506772Z digest=sha256:3f055425c946361cc09b71a2051255e0e2c7252556d7ff8d64e6ec985816ed3d

Observation 313f3b18-dce6-4810-b5ce-f1a393fa2957 · outbound

This paper cites Alisa: Accelerating large language model inference via sparsity-aware kv caching,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Alisa: Accelerating large language model inference via sparsity-aware kv caching,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:43.606157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:43.606157Z digest=sha256:1072b2c59bb00ee5e8099f55d468534fcbc4e95a720a21dd7b07a7ca440f0b5a

Observation 2d78b5e8-4f14-4873-93fa-b52b7a00c4d3 · outbound

This paper cites Mata: A memory-efficient attention accelerator for llms exploiting look-back kv cache pruning,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Mata: A memory-efficient attention accelerator for llms exploiting look-back kv cache pruning,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:43.696933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:43.696933Z digest=sha256:16e49425b07773c012765eeedcbc9057b2de649bf23be3dc64413de5ce203786

Observation 88bcf029-6fc7-4492-8f9d-7280cd5933e4 · outbound

This paper cites Unicaim: A unified cam/cim architecture with static- dynamic kv cache pruning for efficient long-context llm inference,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Unicaim: A unified cam/cim architecture with static- dynamic kv cache pruning for efficient long-context llm inference,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:43.829479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:43.829479Z digest=sha256:24f351be6712e6b016c89d3bf11e6c5427c4d55b51031cc5615173b5014c7ecc

Observation d746782c-237a-4ea6-850b-e105d2e2bfc5 · outbound

This paper cites Token-picker: Accelerating attention in text generation with minimized memory transfer via probability estimation,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Token-picker: Accelerating attention in text generation with minimized memory transfer via probability estimation,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:43.934750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:43.934750Z digest=sha256:8fe371b60c651e4884cca9cc11dbfaed6ac8fd9a627d860c8cdf9fee00e0fa65

Observation 79c477fa-51e5-4b0a-8746-14a8ec912734 · outbound

This paper cites Dias: Distance-based attention sparsity for ultra-long- sequence transformer with tree-like processing-in-memory architecture,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Dias: Distance-based attention sparsity for ultra-long- sequence transformer with tree-like processing-in-memory architecture,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:44.011211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:44.011211Z digest=sha256:86c654422f6d59b20b45d5a4b233c40a04c3097793b1a7bb1af752c578b85a07

Observation fd5d4989-9e38-45d6-93e6-795bbd388645 · outbound

This paper cites Veda: Efficient llm generation through voting-based kv cache eviction and dataflow-flexible accelerator,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Veda: Efficient llm generation through voting-based kv cache eviction and dataflow-flexible accelerator,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:44.075196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:44.075196Z digest=sha256:cea0dc46c193d32adbb6ac9ed5033d36fde2fba5a6193bc4ade395c27d434dea

Observation 20e531de-bec9-4f9a-b671-a9e308b1d187 · outbound

This paper cites Kv-cache oriented query-aware sparse attention accelerator with cross-stage precision-configurable digital cim,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Kv-cache oriented query-aware sparse attention accelerator with cross-stage precision-configurable digital cim,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:44.170208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:44.170208Z digest=sha256:cf15d8eb074830c7d8317f0c28c170e5984da9d53581d901499eb0d40cf3e83d

Observation fc0f6cc6-865f-4b67-af00-7818900e3b05 · outbound

This paper cites End-to-end acceleration of generative models with runtime regularized kv cache management,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding End-to-end acceleration of generative models with runtime regularized kv cache management,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:44.321904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:44.321904Z digest=sha256:cccef5eb0999d9923e6e8cd596010f4b196564371f4bd10329aa91dd32b55481

Observation deeb2018-df6f-49d8-af1b-701c3b66cca0 · outbound

This paper cites Edgellm: A highly efficient cpu-fpga heterogeneous edge accelerator for large language models,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Edgellm: A highly efficient cpu-fpga heterogeneous edge accelerator for large language models,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:44.430466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:44.430466Z digest=sha256:a657e82dc2f02d9abe858e513d7cee94242977665483793d9a12e53b35d20c60

Observation 082819a2-cf6e-4748-b2e9-7663597bb079 · outbound

This paper cites Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:44.529874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:44.529874Z digest=sha256:c0a73418180761d65bf61d611659f5759bca7930f1b9d1e8eacf7da36992c962

Observation 63bd56cf-985d-4172-ba8b-b72d0e54149a · outbound

This paper cites Flightllm: Efficient large language model inference with a complete mapping flow on fpgas,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Flightllm: Efficient large language model inference with a complete mapping flow on fpgas,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:44.760684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:44.760684Z digest=sha256:69d343d51053833a3f4dd800d86a509626af49bb78e58173e4180df729c43a04

Observation 7f0df5c5-732d-48a8-902d-ee7f519581f4 · outbound

This paper cites Ofq-llm: Outlier-flexing quantization for efficient low- bit large language model acceleration,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Ofq-llm: Outlier-flexing quantization for efficient low- bit large language model acceleration,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:44.943518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:44.943518Z digest=sha256:1dbdb690cbff3f5be55f80ff71f8abcb6b1c9084666bc80cfd10f9f0f41cfa4c

Observation ca253343-8b28-448c-901b-7b5abf8a5582 · outbound

This paper cites Kv cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Kv cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:45.144166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:45.144166Z digest=sha256:63632077d0f01581d3a3d452b1e0b45ad565bba007b07751581720467c57079f

Observation e20f6a49-0c49-43c5-94ad-f2b9c415ab98 · outbound

This paper cites Apt-llm: Exploiting arbitrary-precision tensor core comput- ing for llm acceleration,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Apt-llm: Exploiting arbitrary-precision tensor core comput- ing for llm acceleration,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:45.303812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:45.303812Z digest=sha256:23f5a29fb6c4d967871e576abe9c7523439f9a96cff4a201093630cb58f71a7a

Observation abc3d7e8-46f9-48a6-b05e-c1b19e0be6fb · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding A Survey on Efficient Inference for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:45.474492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:45.474492Z digest=sha256:ee976e548dacd15673d56cce870ce7a8ba470e08537727392ca216ed5789068f

Observation b3a0208c-2ad0-46de-b0b8-7c0713d15844 · outbound

This paper cites When to stop? towards efficient code generation in llms with excess token prevention,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding When to stop? towards efficient code generation in llms with excess token prevention,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:45.605674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:45.605674Z digest=sha256:329eb7659848556a8574707c2a11f1224f1bcba87819ae65bf6f3903d8edfe4c

Observation d6ec201e-38ec-4d68-a31d-44fe07b3fea5 · outbound

This paper cites Llmcompass: Enabling efficient hardware design for large language model inference,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Llmcompass: Enabling efficient hardware design for large language model inference,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:45.795483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:45.795483Z digest=sha256:77e5eaa85a7e3f285d9ce156397471007689661ae63e293fa3317ef57b5bc83c

Observation 604e267d-01bf-49e1-805d-3ab2733c5df8 · outbound

This paper cites Energy cost modelling for optimizing large language model inference on hardware accelerators,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Energy cost modelling for optimizing large language model inference on hardware accelerators,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:46.038243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:46.038243Z digest=sha256:df44be2d9db793be57fb8ba4b01afe34b31cf061db7d5337bfeff3a3d5a76ce6

Observation 9168e695-d063-4ad4-9634-4fa85f8a9393 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:46.181845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:46.181845Z digest=sha256:9e98dfe006272f54865c773964a8d06a4ebb3db50503f3bae2340e4023357a52

Observation 89491f21-4613-47d2-8be2-e7c1ee4649b6 · outbound

This paper cites Skipkv: Selective skipping of kv generation and storage for efficient inference with large reasoning models,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Skipkv: Selective skipping of kv generation and storage for efficient inference with large reasoning models,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:46.374845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:46.374845Z digest=sha256:3f167cb6a15de9c27ada759373e2aec7b5e82fc065bcf702efc9f5e2a9ab2be3

Observation 9a0e496a-1570-4ff2-900b-8d6699310dce · outbound

This paper cites Titanus: Enabling kv cache pruning and quantization on-the-fly for llm acceleration,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Titanus: Enabling kv cache pruning and quantization on-the-fly for llm acceleration,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:46.548695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:46.548695Z digest=sha256:3d34160378610f6930212cc1f31f146e416a998a6b0ac2b3aaed4878e5366519

Observation 605bc04d-cf93-41d7-a727-1a6322b69f05 · outbound

This paper cites Infinigen: Efficient generative inference of large language models with dynamic kv cache management,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Infinigen: Efficient generative inference of large language models with dynamic kv cache management,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:46.763518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:46.763518Z digest=sha256:5fc0ac3e514ab3de795119b1904b95b24b5189828e140a1e2d02716eef0c45c7

Observation 3fe095db-cf41-4bfe-b822-d71c0a6c13f6 · outbound

This paper cites Sparq attention: Bandwidth-efficient llm inference,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Sparq attention: Bandwidth-efficient llm inference,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:47.016141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:47.016141Z digest=sha256:541f686afaf4d8d6a3c03171091a08e4f181ef5203d984c40d44673b45c98fff

Observation 610a51a3-5140-4cf8-873c-e171f7a54f2f · outbound

This paper cites Algorithm 232: Heapsort,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Algorithm 232: Heapsort,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:47.163798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:47.163798Z digest=sha256:b9a65064e427eb03e60f1a68c4091d5329ca211bb16130bfacac2b8554d387d6

Observation f3832966-788f-4656-a365-4e0ff33f09a0 · outbound

This paper cites Sorting networks and their applications,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Sorting networks and their applications,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:47.269194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:47.269194Z digest=sha256:c13a2125e318a1457b241765be3a07fe75f4757a4e545d24ef9224afca1a6644

Observation 1da2e8d6-b1c9-4d46-ab3b-8d346b62474f · outbound

This paper cites Efficient memory management for large language model serving with PagedAttention,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Efficient memory management for large language model serving with PagedAttention,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:47.421112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:47.421112Z digest=sha256:8eefb61976636dffd5f861e9c7e391455bba16f376bbf0958567b85bd38afffc

Observation 5bbae7ab-3070-43ee-8ecf-b4e80e042fe3 · outbound

This paper cites Anda: Unlocking efficient llm inference with a variable- length grouped activation data format,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Anda: Unlocking efficient llm inference with a variable- length grouped activation data format,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:47.516573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:47.516573Z digest=sha256:5ca2e567bb06a402b6546b1ce2ecc75258098e78edeea17c3074a0f415cc9e3b

Observation ef6216b1-3f70-4659-8ed8-230610832737 · outbound

This paper cites Ten lessons from three generations shaped google’s tpuv4i: Industrial product,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Ten lessons from three generations shaped google’s tpuv4i: Industrial product,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:47.631238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:47.631238Z digest=sha256:4b49d1c57a03d8e96b6eb46d70d186b3e778c3449d40c36e19c02ed83966d371

Observation ab6c5676-5725-4df3-8714-ccfd8c9d5462 · outbound

This paper cites Dramsim3: A cycle-accurate, thermal-capable dram sim- ulator,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Dramsim3: A cycle-accurate, thermal-capable dram sim- ulator,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:47.771700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:47.771700Z digest=sha256:616aa98da17d2a13f84c91b8d9f0f2e8814b5505bd736a8134b883a71cbfe566

Observation 6d45b6d6-2d5b-49e7-83f1-77b9cf4008f3 · outbound

This paper cites Ada-kv: Optimizing kv cache eviction by adaptive budget allocation for efficient llm inference,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Ada-kv: Optimizing kv cache eviction by adaptive budget allocation for efficient llm inference,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:47.864332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:47.864332Z digest=sha256:2862d41271b1d0865badd2daa6640a8d5d616ee6ae32edaab1254cd24081ab81

Observation 469204c9-e339-46d2-a2b8-a54f42bcba56 · outbound

This paper cites Dynamickv: Task-aware adaptive kv cache compression for long context llms,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Dynamickv: Task-aware adaptive kv cache compression for long context llms,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:47.969320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:47.969320Z digest=sha256:f8eafbfc7c4ecd2b051d8e1d192a4903bc2ffc0a3f8cb3d7482a054fe498b70b

Observation 9345b92f-f78d-4469-8e47-d447e96ef55d · outbound

This paper cites DeepScaleTool: A tool for the accurate estimation of technology scaling in the deep-submicron era,.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding DeepScaleTool: A tool for the accurate estimation of technology scaling in the deep-submicron era,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:48.072701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:48.072701Z digest=sha256:60921dda9584aebf11f40090cff9d3b2f78d8bf1a730b4013363e782e287c657

Pith citing papers

No inbound Pith citation observations are available.