Pith. sign in

Paper Citation Record · LEDGER

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories

As of 15 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2508.08457.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.08457 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:36:25.895941Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f47553a0-a5f2-4537-8721-603c27270229 · outbound

This paper cites FLAT: An optimized dataflow for mitigating attention bottlenecks,.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories FLAT: An optimized dataflow for mitigating attention bottlenecks,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:24.793088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:24.793088Z digest=sha256:965396e819c506c336dcd7581760da1d30776a9c3efcc11df6a148b96930ce40

Observation 0f131db9-5af6-470f-8c10-a70d592c3874 · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:24.860506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:24.860506Z digest=sha256:f2fc398539affc784709ec96a98e9c90db740037556a88b28bf11d67e27c3712

Observation 5a701258-f509-428e-8b57-655d2572f0a7 · outbound

This paper cites PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:24.928403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:24.928403Z digest=sha256:d819f256f8d1d08c69fb02a0ee4f14da2ca8bff6aca8f4e5e89dce966c049fb2

Observation 089691a3-5ddb-47da-9315-a00db09fd2eb · outbound

This paper cites CMOS+X: Stacking Persistent Embedded Memories based on Oxide Transistors upon GPGPU Platforms.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories CMOS+X: Stacking Persistent Embedded Memories based on Oxide Transistors upon GPGPU Platforms

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:36:26.158866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T21:36:24.993295Z digest=sha256:9e18a8baebe408fe03c1a63f0027c8dd4b312f873ff72c5b3762bc6fed50624f

Observation d91127b3-f289-493d-af8e-e55852b8fb28 · outbound

This paper cites Timeloop: A systematic approach to DNN accelerator evaluation,.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories Timeloop: A systematic approach to DNN accelerator evaluation,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:25.127584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:25.127584Z digest=sha256:6f05a3f983139f878753ecdf657a150585f9f40e528ca3a6e80670fbec47ed21

Observation c7ce50ea-adca-467b-bfdf-d3b001847b5e · outbound

This paper cites A-IGZO FETs with High Current and Remarkable Stability for Vertical Channel Transistor(VCT) / 3D DRAM Applications,.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories A-IGZO FETs with High Current and Remarkable Stability for Vertical Channel Transistor(VCT) / 3D DRAM Applications,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:25.239314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:25.239314Z digest=sha256:3bede02cc64ccceafd6bfd45a48eb61be95e412aa6353343d2a489cc02400416

Observation 99e2900d-b5fd-47d9-8030-c4632dcf7d21 · outbound

This paper cites Integration of 0.75V VDD Oxide-Semiconductor 1T1C Memory with Advanced Logic for An Ultra-Low-Power Low-Latency Cache Solution,.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories Integration of 0.75V VDD Oxide-Semiconductor 1T1C Memory with Advanced Logic for An Ultra-Low-Power Low-Latency Cache Solution,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:36:27.102419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T21:36:25.334829Z digest=sha256:a8fb2e72577778b2e6a4fc0a861a939f5aaab19e70cd218ac90b429b67bc6631

Observation c1ae9546-5e4d-4e5b-8d9b-d6b9f42c9ea0 · outbound

This paper cites AMD Next-Generation “Zen 4.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories AMD Next-Generation “Zen 4

Reference 8

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T21:36:26.658354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T21:36:25.407922Z digest=sha256:b0ba0b90f93c73b1f9e57687e3b656cba3dfff83c4ff67ecfcb4a15febf6c841

Observation dabc142d-8cc0-4cbc-9102-09e6060fb135 · outbound

This paper cites Taming throughput-latency tradeoff in LLM inference with Sarathi‑Serve,.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories Taming throughput-latency tradeoff in LLM inference with Sarathi‑Serve,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:25.462008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:25.462008Z digest=sha256:0ec369436437f0849d19581f9149821d49da8a91b37dfce33b4b1b04358b0d35

Observation 53b1e519-fb93-4279-bff8-0706bb83a058 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:25.580836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:25.580836Z digest=sha256:63749dbc45145976d9eee297424d18ae58cf3a2e063d0799e408ddd890519c9a

Observation c5f0cfb0-89f4-4d86-ae1b-1a65cf08a766 · outbound

This paper cites The Llama 3 Herd of Models.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:25.698784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:25.698784Z digest=sha256:b3ed6cdf04bbe2cf66e095b4e3a605192878a95a70a809e471aa36926587b944

Observation 4610b7e7-14a7-40ff-a59a-8da2b5758fe4 · outbound

This paper cites GPT-4 Technical Report.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories GPT-4 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:25.814551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:25.814551Z digest=sha256:a03acff150215cc73f03a645427e04f7225c6d562721f32afe23a006a951f05a

Observation b34dcaa5-f101-49a5-8375-06d237f6b7f2 · outbound

This paper cites an unresolved cited work.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:25.895941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:25.895941Z digest=sha256:2eba3c0e28792c1cfb8ac6e596805ed71655ce41edc3215bc0fc400a046462e1

Pith citing papers

No inbound Pith citation observations are available.