Pith. sign in

Paper Citation Record · LEDGER

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference

As of 20 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 4 inbound Pith citation observations for arXiv:2507.03153.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03153 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:23:11.429005Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T07:06:53.318182Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T08:55:35.129021Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 283541c0-7eae-45bd-94b7-ec057d567d2a · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:15.328653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:06.495867Z digest=sha256:a50a5095adeeff9e11d238bd1ce9396e256c42b23af73b9441ee8dcd0e55a77c

Observation b596e2d0-6399-4cf9-a0de-0c5b61b62bd6 · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:06.616998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:06.616998Z digest=sha256:0153c98ef538791c28a69bbc8440d84736f63ba76e8371b8e587389011f71f1d

Observation 3d9f2aa4-a207-4c12-9005-4d935434b456 · outbound

This paper cites Longformer: The Long-Document Transformer.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Longformer: The Long-Document Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:06.773480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:06.773480Z digest=sha256:7ad3b4433c65645d37f7e08dc9ddadc2c52dc5bc708dd7abf927715ba08d50b0

Observation 4dc6f1ec-f1bd-48fc-909b-5a8d15a18bba · outbound

This paper cites GPT-NeoX-20B: An Open-Source Autoregressive Language Model.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:06.929626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:06.929626Z digest=sha256:ae5718c439f12224201b15982272fe18ed636dcaa0b30f51b1e3473045a2a324

Observation 4c16654e-8ea8-45e1-93e7-fe4bfd9029f7 · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:15.172036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:07.055162Z digest=sha256:24f4c4b41e95d49c6d51030c5d7da24a4d2df8f962bc0ea5caebb0ccd27410fd

Observation b58bdb54-51f7-452a-9701-6a74709b43ac · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:07.305255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:07.305255Z digest=sha256:040c47af66d104e41ac17cb7af7d9c0b947fd3afe93ccebb7eb98dba0b814102

Observation 4d79de82-e945-491d-b9ae-c7fa6d4d4158 · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:15.020570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:07.554340Z digest=sha256:ce34bb29167eef04888bd8c800a322a8e77360881ec64a3b3d196da1a0bf7edb

Observation baaaabff-b972-4e5e-b815-5aec8ec9e06b · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:14.887903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:07.611925Z digest=sha256:07728f9024b34d77d99a62a891429b015617133cb98660afa083f4941dd37b88

Observation a3bcea62-8d5c-458c-942f-7e698426e0d6 · outbound

This paper cites 2024.{Cost- Efficient} large language model serving for multi-turn conversations with{CachedAttention}.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference 2024.{Cost- Efficient} large language model serving for multi-turn conversations with{CachedAttention}

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:23:14.740866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:07.676775Z digest=sha256:6e40071e50ec07c9a2d4b01aa0718d806594f1d8d1e9638bc3d67e8be4d829ba

Observation a4f8178d-b21c-4288-bb66-550d6a64f86f · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:07.738948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:07.738948Z digest=sha256:e17f7a77036e390c47a3c2f64150dbcf5260d962fde35cd57e8587066bc69b94

Observation 51290079-dd5d-49e2-becb-197b6672b7bb · outbound

This paper cites Reformer: The Efficient Transformer.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Reformer: The Efficient Transformer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:07.805883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:07.805883Z digest=sha256:a498dcff77e6bf1453c04c8bf99da40ac375b866e3c95f30f1053d9acef0c9dd

Observation f005e9bb-00fd-43ba-ae2e-564e9bfdc5a8 · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:07.892996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:07.892996Z digest=sha256:af7c245f87014ac684a983a6fd9c0b9b23f1d7a12ec87b82a564225c66dcd94c

Observation 673f8366-0a7c-4916-8c3e-704d82699165 · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:14.551089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:08.044730Z digest=sha256:73b33ed258ef0e5cd45a3a35c900042cea5640e461d0ffb6b8f59ac425feea2f

Observation 8e17e11e-569d-4aa1-8c22-918e25924b4a · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:14.388386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:08.111084Z digest=sha256:9a4a6bcfaf376ada28ec4d517cd3ebc926972b3e199e7b3a34657e2e5d79e371

Observation 821624bc-6404-49cd-beb8-71168644d3a4 · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:08.199783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:08.199783Z digest=sha256:83f103ee6f73267f11bf0e92714d475044bf86493215cb88aa0aa30c1ea8749d

Observation a88891cc-7e72-4ad8-b2bc-9b95e0e3169c · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:08.356733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:08.356733Z digest=sha256:bc691af6fc5c17a7c635b099c38601b60c6bb4751061cd12bec6585bf56dc71a

Observation 9f95e3c7-396e-45e7-a18c-415c57366bd0 · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:14.151458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:08.450785Z digest=sha256:4d49e07bc6331014f8d7f9cb9fb79551ee240ca8a90f7ab12394080cb5034223

Observation 0ddab3eb-4b04-4d4b-86dd-934d3b0a6c6a · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:08.546761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:08.546761Z digest=sha256:c180328232f16e4c2ea4ba997a9f5a5e9e96d2855d3b10d8269b6ca71e050f34

Observation 1bde7af9-e12f-464e-940f-a70578758c78 · outbound

This paper cites Advances in Neural Information Processing Systems 37 (2024), 22947– 22970.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Advances in Neural Information Processing Systems 37 (2024), 22947– 22970

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:23:14.260718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:08.289775Z digest=sha256:d54416313ec1043f65c259508bb82ba7df0ceebc3cfdb35ed378ea6a2440e8a9

Observation 4cfc5cd7-ebf8-4bf4-91fb-20b43ebd3780 · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:08.724360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:08.724360Z digest=sha256:13e23f95f156122dd997020717dd63f320c8a0c9553ef0dc9d04724d42228ccb

Observation dd570bcb-6f0c-45d5-9dd8-8bc3a5fab274 · outbound

This paper cites s1: Simple test-time scaling.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference s1: Simple test-time scaling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:08.922591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:08.922591Z digest=sha256:1349741b36f4a7cf2f9d13075b223f036d1d6ca313488fe726387f71127b5fc9

Observation a6a83586-5834-405a-a6b6-f5e81933883a · outbound

This paper cites InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:08.996865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:08.996865Z digest=sha256:34cc17d25d23d5fe4ad380b387fdf7fb13fda9e69a66fa281f75b3c976b2cc20

Observation 73b00083-ebe0-4ad6-b566-9c8e9e4d6279 · outbound

This paper cites HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:08.655628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:08.655628Z digest=sha256:20ca8021520f3c2e20bd3fb8f9a71ad66739c1caf01c3bd445f20be413b57286

Observation 5cdbeddd-2d21-4162-aa5f-5982cab441b1 · outbound

This paper cites vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:09.183968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:09.183968Z digest=sha256:a9842fefaf9ef989ca82193a2890883e6121136d9edaf5f2d0f3588c82373a7d

Observation fbb6a93f-78c8-48a2-af27-994cb7722caf · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:13.924966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:09.287184Z digest=sha256:a5a9c541a70d286d1ca583f1ad9eb105776c8d9adeb7eef4e55eb4fe2d1e32aa

Observation cdcfa3c6-f8c5-4f3a-b2cf-7f7d0504f755 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Code Llama: Open Foundation Models for Code

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:09.372823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:09.372823Z digest=sha256:9db38393555efa90e73c5cf7e996182a04f729d9905687b915ac0986c2b092c9

Observation 37b0486f-032e-4737-905f-db21fa971db7 · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:13.789727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:09.449443Z digest=sha256:9f0f320b8e5d9e8235afa058a77491c6bed52b3197ae2516458410c61b3a554e

Observation 8d5ada3d-7956-4b8f-a0e9-bab06d28c0a3 · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:14.023156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:09.083238Z digest=sha256:a31f49a6c912f3273f6853f89b60bbc05b3139497607f8adb74ec88417e93289

Observation fa13dd56-c45f-45b3-9124-c34322e448b1 · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:09.634676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:09.634676Z digest=sha256:0c33562aed562a6a9595ad8ea4b051f9112a4670b391a5ed4342dcb299fbbec3

Observation b87760f9-9b4d-4089-8ba3-851c6ef17fc6 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:09.715033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:09.715033Z digest=sha256:0887ad2e728afa27a9b6aca0faf0691fdf61fd6f94aa36f77d38216722f0a2e9

Observation 07f4e39e-5f34-4aae-a0e3-33f962a8c334 · outbound

This paper cites Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:09.786594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:09.786594Z digest=sha256:ef46efab673d0e24954ed371895fb94731f1bb1c15cba04bf4af03a91c335282

Observation b02aaafc-504c-4ac7-a1b7-fdee5b54e97c · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:09.870644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:09.870644Z digest=sha256:1163cab276bed28703033132c0e19ff9abb33939991e94d7e0aa3ce13f473223

Observation 7d142924-a365-45e9-9f08-d8cdae1f92f4 · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:13.611617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:09.526833Z digest=sha256:c59a12ae9c5130f5e387bff67dd0967b3d9d05e172d3d960704636824f6a3118

Observation 30d3ef68-2644-49ff-86db-2a3f9113d671 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Efficient Streaming Language Models with Attention Sinks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:10.042409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:10.042409Z digest=sha256:a9c6ec02dfe11019613c74d1afb150191d625d34270867cc1cec3497d03a50a1

Observation ca89c058-e067-484c-a479-72a3b1eb8b75 · outbound

This paper cites XAttention: Block Sparse Attention with Antidiagonal Scoring.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference XAttention: Block Sparse Attention with Antidiagonal Scoring

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:10.145162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:10.145162Z digest=sha256:62fda2d6c00a0f76d1a572a5e8876a16ddf19b5dfe0b7a59cf6aceb4d7b29051

Observation ce532b92-7fce-410b-84e2-38f491b37009 · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:13.433429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:10.248132Z digest=sha256:10eec84b3d837c2a232f52b74d9307c176953f72a2eeaf877988840f6bd3242a

Observation 1b89ed75-3b16-4829-bdb5-889e414fcf96 · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:13.221464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:10.351457Z digest=sha256:ea8aa8f64546020ed5e7b384f8bfc4165e171537d9dcd261d08443383e5d61ec

Observation 30b42c94-a547-4810-941f-147bf261f7ce · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:09.940130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:09.940130Z digest=sha256:9f3f7090f56a25fe3bd8c1c77455c5fc9a30c1948cd5fc4450773da13d4fa168

Observation 23595a6e-6812-41cf-bb8e-053175fa418d · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:12.833036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:10.538778Z digest=sha256:59c991f3c4d5caa83ef9dea10837327a76da04282dd85876a397d3ac649454b4

Observation 379f6667-7c1b-4af3-a2aa-8d206fe92ad7 · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:12.594045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:10.679923Z digest=sha256:4d6e951f1a4913493f9b065ffe30799dbb51830cfe3a66165b50bc18c857f378

Observation bd63e885-b224-42e4-9093-0fa42a612a38 · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:10.761186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:10.761186Z digest=sha256:fabd58a27849ce4dfc5b8a5c94f42d79a1b32e498ad711e451a70807ff5b3b7f

Observation 82a4b9e8-a88a-4978-85f7-292e48558915 · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:12.432400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:10.854045Z digest=sha256:04ebfa64dcba58df3a0af9e0e39c8b8ced7d13515a76adfba9d16338ddea1d96

Observation 81360729-f072-4c54-b3ad-1f54066d5c5c · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:13.045392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:10.459568Z digest=sha256:6dd8ed3fef595df3f3b5b82cc339ed5bc79b28d5c078d4e63f61f75402b0817f

Observation 312b2387-692d-428f-b045-4841fb73ad65 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference OPT: Open Pre-trained Transformer Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:11.007396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:11.007396Z digest=sha256:5f80fb161bd515332ef851346c08bf6dcf2a1820a0f44c77043d395192739acd

Observation 7cbd7307-683a-477e-b16c-06b410c55bbe · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:12.292772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:11.147498Z digest=sha256:0c6305ca225714319995bf696ac763c4f9211bfaf416ddc070572d2ca9f20ef5

Observation 2ca5ed73-8e11-4321-8050-b223b8889972 · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:12.068144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:11.236416Z digest=sha256:351983f9a366f38c115cffaefdf661a0c304e4752212aa2c4c01a2a910a3b92f

Observation 8aa3366e-4b03-4a9e-97e7-c694e98953dd · outbound

This paper cites HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:11.321915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:11.321915Z digest=sha256:3356ce447fab03609fd49c093923f0c2256b4601974f882d892cb9844462f662

Observation 72cd1495-ea5b-4234-864d-b2433786a270 · outbound

This paper cites an unresolved cited work.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:10.932828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:10.932828Z digest=sha256:4835bd6b74ffc8df5ce060dcd6d89896cb672a9108998fde827c689986721100

Observation 4d750438-4284-444a-9063-5f37783b2aa0 · outbound

This paper cites 2024.{DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference 2024.{DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:23:11.898805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:23:11.429005Z digest=sha256:f0a9a3d8e7f55164600a6a905512bc952bc53199c49ad08f0693f01be0b33d45

Observation 845365b8-9b6d-474c-b622-dd86c4a0ac78 · outbound

This paper cites Pointer Sentinel Mixture Models.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Pointer Sentinel Mixture Models

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:08.819309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:08.819309Z digest=sha256:4f807a9187fe43342b4b90a3248fd4748c2b0fe8046b7c0bb2f1796e5ce9b748

Observation 1ad134e0-f174-484e-9c75-853194db01f2 · outbound

This paper cites Advances in Neural Information Processing Systems 35 (2022), 16344–16359.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Advances in Neural Information Processing Systems 35 (2022), 16344–16359

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:07.433412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:07.433412Z digest=sha256:80253a106b6e021a901756d97d7acf116c163bb413c6a4cf7e4535b835154980

Observation 2a831765-b1c5-4afe-8fc2-00a250eba7be · outbound

This paper cites In Proceedings of the 29th Symposium on Operating Systems Principles.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference In Proceedings of the 29th Symposium on Operating Systems Principles

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:07.959084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:07.959084Z digest=sha256:b8f67d7ee191ef0ace02761f628dd3beee63249c8d1b65184782cc0aafd522e3

Observation 30a9b1a4-667b-4e1b-be27-98305084f558 · outbound

This paper cites A Complete Survey on LLM-based AI Chatbots.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference A Complete Survey on LLM-based AI Chatbots

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:07.166141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:07.166141Z digest=sha256:ca174e9f5d6c2f1c00c2bcd578bdd0913463b83cfae4861ad9b0dea4341da427

Pith citing papers

Observation 40b19ae0-e370-40bd-b243-539118cd83c6 · inbound

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference cites this paper.

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:05:54.167753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T02:56:28.828593Z digest=sha256:8959eabcecdd50bf9e7af2f9a38d70339b3fe765ebfcf37ce90f9e5fa36eb640

Observation c2597739-d9ed-4c33-b905-30feae5dc31f · inbound

ChunkFlow: Communication-Aware Chunked Prefetching for Layerwise Offloading in Distributed Diffusion Transformer Inference cites this paper.

ChunkFlow: Communication-Aware Chunked Prefetching for Layerwise Offloading in Distributed Diffusion Transformer Inference HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:27:02.584535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T01:23:46.931949Z digest=sha256:c92bac28f4dfc04ebea7e538eb5e08ef7357ac2500f78cc4c3b88d168d9e7e9e

Observation f74a384a-6f6a-4d7f-82a1-563ce4dd2dd2 · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:04:45.301914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T05:37:13.211613Z digest=sha256:6d2938f926e8068f3edf90ee56fa01a093d7575680f5ebf727e53d949d16f687

Observation 306398cd-4251-4152-b2c4-6f9056b11667 · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.130368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:06:53.318182Z digest=sha256:316a987f2a781cf4190b4f7d5147c17deab2a05129a18eef0a095208accb3b19