Pith. sign in

Paper Citation Record · LEDGER

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference

As of 14 August 2026, this Paper Citation Record lists 100 of 121 outbound references and 0 inbound Pith citation observations for arXiv:2502.07578.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07578 v3

Coverage vector

measured 100 of 121 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:19:43.247215Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 121 outbound references displayed

  • verified exact1
  • verified fuzzy37
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 07bac523-f739-4068-87a5-c1bcda05eaac · outbound

This paper cites URL: https://azure.microsoft.com/en-us/ pricing/calculator/.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference URL: https://azure.microsoft.com/en-us/ pricing/calculator/

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.861786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.861786Z digest=sha256:f693a6ec9ec13c341171650ee2949d65ffe0ad8d12ab549468f27d7987357025

Observation f50db636-43f6-481c-b096-45b2a2d8b8c8 · outbound

This paper cites Enabling cxl memory expansion for in-memory database management systems.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Enabling cxl memory expansion for in-memory database management systems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.866044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.866044Z digest=sha256:85619aab1535a2cba18bdb418fe4d5a3d1dc1c0b5cb6493101f9e4475847e136

Observation 5089bf48-f3a9-4e3d-91cb-1453f5c22089 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.869539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.869539Z digest=sha256:e0a5297ce16f00b94b651d8965f36db1f6f996652cb0399c95240418c163a14d

Observation a3d5d758-2180-4b31-b598-d9b0e46d418d · outbound

This paper cites Introducing the next generation of claude.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Introducing the next generation of claude

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.873549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.873549Z digest=sha256:6e2f3366afd82d9a35131fa7b040e24226a03a3d2b58fa5a198418ccd3af388d

Observation ea5c3df5-c2e5-413d-ab17-fb19b56d3d23 · outbound

This paper cites Exploiting CXL-based memory for distributed deep learn- ing.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Exploiting CXL-based memory for distributed deep learn- ing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.877646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.877646Z digest=sha256:431898222650af136c188a0c2322047021f0696910f49b4f63d2bccc3b3927bb

Observation 5a0cd7bb-5661-46a5-b1b2-fb361d8d7820 · outbound

This paper cites Llama 2 70b: An mlperf inference benchmark for large language models.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Llama 2 70b: An mlperf inference benchmark for large language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.881241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.881241Z digest=sha256:84ed246e063e51c67a351c5a10b3ad7a5d430b92ed9a35aaafec8ac2d0826799

Observation a3965f40-faf1-48a8-8170-6704fd7edf22 · outbound

This paper cites Improving image generation with better captions.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Improving image generation with better captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.884793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.884793Z digest=sha256:a672324dca2634f0eb92064424a75ec8bbb47d326ed6236feccdca833f1742a2

Observation 24294e30-8d96-4e91-bc5c-60cdcf8f0080 · outbound

This paper cites 144-lane, 72-port, pci express gen 5.0 pex89144 express- fabric platform.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference 144-lane, 72-port, pci express gen 5.0 pex89144 express- fabric platform

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.888201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.888201Z digest=sha256:06c91ecf822635d279ce2d77c4ecb2e0a43be58005d70f071c36b6047160a06b

Observation 7d94f444-8173-4e1f-9ad4-c70dc261920b · outbound

This paper cites The berkeley out-of-order machine (boom): An open-source industry- competitive, synthesizable, parameterized risc-v processor.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference The berkeley out-of-order machine (boom): An open-source industry- competitive, synthesizable, parameterized risc-v processor

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.891765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.891765Z digest=sha256:9ef176a2351ac6a30dfcbd7c5a74865d6d6b05f36e694006fd683c14dd06b981

Observation 5bbabc2f-30e6-4979-aaf5-7321ed902eef · outbound

This paper cites Patterson, and Krste Asanović.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Patterson, and Krste Asanović

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.896429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.896429Z digest=sha256:6f3b7f843b5bd692ed9551c88b0920d96562144afb6e3b130e3b77ad893e4309

Observation bbaeb890-7147-48e2-9b19-32d4dddaccfe · outbound

This paper cites Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.899955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.899955Z digest=sha256:d4c748ca5f79d23d074801ca1ae1d74c9d904cd908b1d9aafd2b2aa3021634ce

Observation 5a3e52fd-7d4b-4141-b10d-809ffeee38ff · outbound

This paper cites LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.903454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.903454Z digest=sha256:895fe105d3549070167e7ee0237f5cf25966f7681384c06662ffc92b9253a208

Observation 2469b9de-ad81-44cc-9436-8113f685dad5 · outbound

This paper cites The true processing in memory accelerator.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference The true processing in memory accelerator

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.907410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.907410Z digest=sha256:612b9b7072d96af951a3d17a4b9035b58598055124c6da387b0dbe78a3ab8c1a

Observation 004a9ef1-3dcb-4734-b9bd-d136be30c219 · outbound

This paper cites To pim or not for emerging general purpose processing in ddr memory systems.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference To pim or not for emerging general purpose processing in ddr memory systems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.910828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.910828Z digest=sha256:53e68792e21d9c5575862a928414e608ead922aa16024278db0d71add85b0ce8

Observation 338a6725-26d1-48f7-b93b-c2cf52df2949 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.917902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.917902Z digest=sha256:02042cdd951773c53fc65360be5b125acd9133401c27b08e92bb105f9b8e2dea

Observation cd1e20ff-666e-4342-a52f-208a3b487efe · outbound

This paper cites Dram spot price.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Dram spot price

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.921715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.921715Z digest=sha256:980c57984fa0e0cea86a0326d349c39ddb2cc68a2fcdbd7702564089b24d2706

Observation 49a53810-bc51-4a1c-9f6b-e73e391f1e6e · outbound

This paper cites The Llama 3 Herd of Models.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.926165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.926165Z digest=sha256:67f44c4f9b252a5a2bfa4e654e1b1cd1c8ee6d2cb1779bd0e7dc0af27b714df5

Observation 67955ca7-fbde-4b9f-8eb0-223e80525b45 · outbound

This paper cites The inference cost of search disruption – large language model cost analysis.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference The inference cost of search disruption – large language model cost analysis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.930080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.930080Z digest=sha256:929ca2543a9460726da565f19acebf2a41c56d87c400114d1d9f87f83298bca5

Observation 12ba10d4-ae69-4261-a872-0cc372b2d438 · outbound

This paper cites Nvidia tesla a100 80gb gpu sxm4 deep learning com- puting graphics card oem.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Nvidia tesla a100 80gb gpu sxm4 deep learning com- puting graphics card oem

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.934251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.934251Z digest=sha256:2b4e5df3d9ffb5ce10e84bb5cf6c763dc697635632fb07a9e708e870e716c412

Observation c3f77fda-f121-4c79-a0de-2993503f6b9d · outbound

This paper cites Pci interface ic.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Pci interface ic

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.938373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.938373Z digest=sha256:63ea7869d0dcab77cbb71310b6dda74a719db4e69bd1017076763ac414b0b93d

Observation 2f0894fd-ced2-40d0-91ef-bbda10643519 · outbound

This paper cites Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.942471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.942471Z digest=sha256:46e7d8139ef7a1fae4c0afb4ee143ff67d693667184f86d3463535c1654ef471

Observation fde56906-0e9f-4ce9-a19b-78ad37b29a88 · outbound

This paper cites Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-design.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-design

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.946555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.946555Z digest=sha256:a70a0468031949bd6f00a29b5a527f40d179a1aa3d3cc9898e6cd94dc816d4a4

Observation 40162986-cd5d-489c-9cb5-40d60ace121d · outbound

This paper cites NDA: Near-DRAM acceleration architecture lever- aging commodity DRAM devices and standard memory modules.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference NDA: Near-DRAM acceleration architecture lever- aging commodity DRAM devices and standard memory modules

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.950787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.950787Z digest=sha256:1ff1d7a86259148fdaaa8a805ea9f14ff2b04ed4981bbcfae7207aa2882b6292

Observation 5c1f2059-3400-449b-b31b-ff9222b4fa4d · outbound

This paper cites Reinhardt, Adrian M.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Reinhardt, Adrian M

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.954152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.954152Z digest=sha256:ca400853739003563e11920714cfcb4d9ccfb0dd0784f5e6977cc2cf23cfd454

Observation 6456e43a-4f47-40ce-ba0a-2873b1e4bee8 · outbound

This paper cites Sparsep: Towards ef- ficient sparse matrix vector multiplication on real processing-in- memory architectures.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Sparsep: Towards ef- ficient sparse matrix vector multiplication on real processing-in- memory architectures

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.961275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.961275Z digest=sha256:b368a07d107a150277a12a90ebf7263ebeaeb434bc6b3ee4df2dfaf8bed11f5a

Observation 77b8ec34-0861-4634-9898-0ae60995e6eb · outbound

This paper cites Benchmarking a new paradigm: Experimental analysis and characterization of a real processing-in- memory system.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Benchmarking a new paradigm: Experimental analysis and characterization of a real processing-in- memory system

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.964664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.964664Z digest=sha256:bd55b1274093a2d4f3ce759e42ff76e05312fc0b2866d58f3e7d4d2b3440bc45

Observation 931f5695-aaf7-4f1c-9352-0a50f9f63405 · outbound

This paper cites Evaluating machine learningworkloads on memory-centric comput- ing systems.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Evaluating machine learningworkloads on memory-centric comput- ing systems

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.968712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.968712Z digest=sha256:9fa84c0fb9f6c436db08077f59170ac5364814df01aa5b9cf61d64dbbc23abb5

Observation c6cc619e-5bb2-46e2-a359-a533ca66c343 · outbound

This paper cites Our next-generation model: Gemini 1.5.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Our next-generation model: Gemini 1.5

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.972030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.972030Z digest=sha256:285a4401e36cc5edf1a3a44dc6bfc8dd520ff2ca2e3dbbdc52e5637c29ae2b51

Observation 834a3a4f-a801-4228-9c53-1d452bad5f9f · outbound

This paper cites Memory pooling with cxl.IEEE Micro, 43(2):48– 57, 2023.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Memory pooling with cxl.IEEE Micro, 43(2):48– 57, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.975332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.975332Z digest=sha256:c80f6f402155273cdbf4be8a588f664aa8acbc1c26af2ae96da2752ebbf9ed8b

Observation 49574b5b-b973-451a-8eed-66df7d31407a · outbound

This paper cites Direct access, High-Performance memory disaggregation with DirectCXL.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Direct access, High-Performance memory disaggregation with DirectCXL

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.978839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.978839Z digest=sha256:08f18f967cfb4325262992a46f63cc534bc0267cc4b98f4472fd234f81eca429

Observation f5295c49-69b5-4e67-8994-0e33a8f01a05 · outbound

This paper cites OliVe: Acceler- ating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference OliVe: Acceler- ating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.982214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.982214Z digest=sha256:91c32c0639865ca2840f958d2741a861c96598a5bdf4e09a29752bd8fdaeac7f

Observation 61eaf1de-bd24-4108-903c-c4c1a42c4044 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.985722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.985722Z digest=sha256:7584373cc4c4664c8f5b366fcdff58cc1a72bf4f1ecc938ba97ee5b44901eb6c

Observation 9b69936e-63a4-47c0-955d-64219e5c4490 · outbound

This paper cites ELSA: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference ELSA: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.989518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.989518Z digest=sha256:cf73e5e7e71432a39ad0901e119a39c9c02595ab7b2ef65a0db2edd5e2d08df0

Observation 5e296e87-2f2c-4d9a-b59c-ac1b3e695651 · outbound

This paper cites Deep residual learning for image recognition.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Deep residual learning for image recognition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.992996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.992996Z digest=sha256:3f52682f31df097d4fa83badf369d9f776210b0a9260340c0f2a2fb865bf8391

Observation 55134ec4-639a-426d-8520-975c003e047c · outbound

This paper cites Gaussian Error Linear Units (GELUs).

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Gaussian Error Linear Units (GELUs)

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.996285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.996285Z digest=sha256:33da3736ffcdd53733535a15908c12c94c9a1e959afd46db48ac23f2111e87d3

Observation cb395ee5-4e1b-416a-ae1e-4824d597c81c · outbound

This paper cites Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:42.999831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:42.999831Z digest=sha256:a84af0fb4e76e1b0d4fee70eafeb118d2cf6cd71a8efdc5526c11d59774cb6f8

Observation 63c96630-b1d1-448b-b1fc-ed12aa84ec4d · outbound

This paper cites Le, Yonghui Wu, and Zhifeng Chen.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Le, Yonghui Wu, and Zhifeng Chen

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.007401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.007401Z digest=sha256:4c67de74b238b4f57bdec8d146614764322e8ff3e28ce9b12e181dc4489c95e6

Observation 6ffc5613-87bf-401f-8af7-621eb91e81f2 · outbound

This paper cites BEACON: Scalable Near-Data-Processing Accelerators for Genome Analysis near Memory Pool with the CXL Support.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference BEACON: Scalable Near-Data-Processing Accelerators for Genome Analysis near Memory Pool with the CXL Support

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.011573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.011573Z digest=sha256:d30117a73e5a3a4e589f6148e2224f7369b618b978f543eb382e3aebd5edda65

Observation ba3a06f7-092e-4df1-a37c-510a3da0fa64 · outbound

This paper cites Floatpim: In-memory acceleration of deep neural network training with high precision.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Floatpim: In-memory acceleration of deep neural network training with high precision

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.015085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.015085Z digest=sha256:d8b0f3cafc96656282c325acccf0517d4e235505f5c7cdae91332e2a7fb19f18

Observation 9ba074ee-9d04-4044-a038-a29e28cb3de0 · outbound

This paper cites Ultra-efficient processing in-memory for data intensive applications.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Ultra-efficient processing in-memory for data intensive applications

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.018239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.018239Z digest=sha256:2ac3042027aab4404050cfdc66457ed81dd7dfbfdfa47c87a808cd718a921449

Observation fed0dce8-a0f9-46db-bfbe-c59958960241 · outbound

This paper cites Intel xeon gold 6430 processor, 60m cache, 2.10 ghz.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Intel xeon gold 6430 processor, 60m cache, 2.10 ghz

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.021491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.021491Z digest=sha256:0c0c5de60e33dc336ff5f366f2ffbc8832d9ea34a97edccc33659691919b1da4

Observation bfa1399a-05e0-4d73-a5a6-e7ba13adfdf0 · outbound

This paper cites Intel ® xeon® gold 6430 processor.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Intel ® xeon® gold 6430 processor

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.025589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.025589Z digest=sha256:5868709ba0af47b53c63bdbad0cb02f2e4b8d48aa0e8c96e3f2a8b02edad1849

Observation 3d974faa-ce32-4c60-9c7f-19365e6705cd · outbound

This paper cites CXL-ANNS: Software- Hardware Collaborative Memory Disaggregation and Computation for Billion-Scale Approximate Nearest Neighbor Search.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference CXL-ANNS: Software- Hardware Collaborative Memory Disaggregation and Computation for Billion-Scale Approximate Nearest Neighbor Search

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.029008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.029008Z digest=sha256:182d5d95d12e902aa184354747b7b0eb305fd1eaecd8e03c738ed22a0fb2b115

Observation bf1c202c-5700-43bd-b906-d8d3953b23e6 · outbound

This paper cites Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.032314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.032314Z digest=sha256:6eb03436bf4ac67ea4bf369dfa9864a793584a13e300bc1957332c089b0711d2

Observation f3183397-8d8e-4962-b7fd-85ba7ed408bb · outbound

This paper cites Noise-resilient DNN: Tolerating noise in PCM- based AI accelerators via noise-aware training.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Noise-resilient DNN: Tolerating noise in PCM- based AI accelerators via noise-aware training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.035525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.035525Z digest=sha256:ce36ee9e735579dce414934d01ba9740019d23e7f9010b2dc08792a02c4e0f23

Observation 551d1872-b606-4835-927f-e7255e0853e1 · outbound

This paper cites ChatGPT for good? On opportunities and challenges of large language models for education.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference ChatGPT for good? On opportunities and challenges of large language models for education

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.039355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.039355Z digest=sha256:1706c60b18692a2982e7008740d102dd414415cf09be33b234256939ac0048a8

Observation fc0b1791-a46a-4a6c-bef5-518b86cd102e · outbound

This paper cites Rec- nmp: Accelerating personalized recommendation with near-memory processing.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Rec- nmp: Accelerating personalized recommendation with near-memory processing

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.043715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.043715Z digest=sha256:c3acca1e2b4644aa8a876c4c8089672a58775f4eb33decd407ae38f938456c7b

Observation 40984173-3fef-493d-8435-905a989496c9 · outbound

This paper cites Moonwalk: Nre optimization in asic clouds.ACM SIGARCH Computer Architecture News, 45(1):511–526, 2017.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Moonwalk: Nre optimization in asic clouds.ACM SIGARCH Computer Architecture News, 45(1):511–526, 2017

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.047402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.047402Z digest=sha256:57027c31a4194b4c9bdb192a045cd44fc0bc16969638d136ffaa48227d0fcfcf

Observation 21964e30-ba5f-42c4-93fe-95844210fa51 · outbound

This paper cites Aquabolt-XL HBM2-PIM, LPDDR5-PIM with in-memory processing, and AXDIMM with accel- eration buffer.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Aquabolt-XL HBM2-PIM, LPDDR5-PIM with in-memory processing, and AXDIMM with accel- eration buffer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.051215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.051215Z digest=sha256:90015156070933adbb4d3f45cc8375bf7de604dd20b2da3ed72509c167367607

Observation 815bf70f-471f-4ee4-9d71-5072dcb57c42 · outbound

This paper cites Samsung PIM/PNM for Transfmer Based AI: Energy Efficiency on PIM/PNM Cluster.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Samsung PIM/PNM for Transfmer Based AI: Energy Efficiency on PIM/PNM Cluster

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.424550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.055256Z digest=sha256:d834e7f72ef5f498e9b7abc664084c7a13ef24703721ea0418bfa04e9a877784

Observation 24dbc041-3b61-4467-884a-3dfc5ade2e11 · outbound

This paper cites A 1ynm 1.25v 8gb 16gb/s/pin gddr6-based accelerator-in-memory supporting 1tflops mac operation and various activation functions for deep learning application.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference A 1ynm 1.25v 8gb 16gb/s/pin gddr6-based accelerator-in-memory supporting 1tflops mac operation and various activation functions for deep learning application

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.059497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.059497Z digest=sha256:02dea6e6022a43acbac16ff90a844fc2735b207d1934e36b29da892a4f2bbc70

Observation b6cef06e-7a03-471a-b491-2f78bf7b2609 · outbound

This paper cites Maeri: Enabling flexible dataflow mapping over dnn accelerators via recon- figurable interconnects.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Maeri: Enabling flexible dataflow mapping over dnn accelerators via recon- figurable interconnects

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.411377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.063790Z digest=sha256:01f5beefcab20f8532d232617087c52d7183b4fb0ca93018cc767d3d972ff550

Observation 6588b9e2-114f-418f-a03a-f1a17dc44c75 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Efficient memory management for large language model serving with pagedattention

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.067968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.067968Z digest=sha256:90f51fb204f451ec1deff26789dadaf9039cfbb414d5a443e41f2d413ded1330

Observation 0be80827-8f78-46d0-aab6-14544a8d6650 · outbound

This paper cites Memory-centric computing with sk hynix’s domain-specific memory.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Memory-centric computing with sk hynix’s domain-specific memory

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.072246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.072246Z digest=sha256:cee02188214522583b9bf39d664f3b2ca2e1d81bf811ec140ac007890b8a60a6

Observation c7b25c34-da8e-45b6-a33f-722a70341f7f · outbound

This paper cites System architecture and software stack for GDDR6-AiM.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference System architecture and software stack for GDDR6-AiM

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.391031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.076737Z digest=sha256:f6ff24d52ffd452069b8e6b1684b23738aff6c4b63aec4853afcfccc726d999b

Observation 7b24437a-9736-4834-b18c-7bb0c0f435ba · outbound

This paper cites 25.4 a 20nm 6gb function-in-memory DRAM, based on HBM2 with a 1.2 tflops pro- grammable computing unit using bank-level parallelism, for machine learning applications.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference 25.4 a 20nm 6gb function-in-memory DRAM, based on HBM2 with a 1.2 tflops pro- grammable computing unit using bank-level parallelism, for machine learning applications

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.380426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.080673Z digest=sha256:a60c966fc266c70f139c04ebb3b138f8943a523e1bc1af8651492c75f8fb92d8

Observation 66bdc8f1-1e91-4704-bee7-d804f371433f · outbound

This paper cites Improving in-memory database operations with acceleration DIMM (AxDIMM).

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Improving in-memory database operations with acceleration DIMM (AxDIMM)

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.369207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.084956Z digest=sha256:92393159f309f94e6982c5beeb804e1f8e694e654d78cd9c7267b93e5d60e5fa

Observation bc0b94e1-27db-470f-a9be-167adc88e9e6 · outbound

This paper cites Using machine learning to increase yield and lower packaging costs.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Using machine learning to increase yield and lower packaging costs

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.357587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.088957Z digest=sha256:eddf253656a5418a76aa860846877e762b6ddce3725a04435d95ec84277c0d9b

Observation 5d1065ab-8c94-415c-9c2c-93facde77733 · outbound

This paper cites A 1ynm 1.25 v 8gb, 16gb/s/pin gddr6-based accelerator-in-memory supporting 1tflops mac operation and various activation functions for deep-learning applications.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference A 1ynm 1.25 v 8gb, 16gb/s/pin gddr6-based accelerator-in-memory supporting 1tflops mac operation and various activation functions for deep-learning applications

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.347503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.093213Z digest=sha256:b982f638249a650521986871de6beba823079cc6c3de034e1c8fd3d20ec89961

Observation 06c69c93-84af-4188-b7d6-b4ebf002e139 · outbound

This paper cites Berger, Lisa Hsu, Daniel Ernst, Pantea Zardoshti, Stanko Novakovic, Monish Shah, Samir Ra- jadnya, Scott Lee, Ishwar Agarwal, Mark D.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Berger, Lisa Hsu, Daniel Ernst, Pantea Zardoshti, Stanko Novakovic, Monish Shah, Samir Ra- jadnya, Scott Lee, Ishwar Agarwal, Mark D

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.336520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.097580Z digest=sha256:cdc944c254bdb8ee4644aa3e641e101bea2d428f568b4ee3d16aaedb0b756a48

Observation 4cdfb794-01a7-46ab-a526-cb5ae6f9a145 · outbound

This paper cites Accelerating distributed reinforcement learning with in-switch computing.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Accelerating distributed reinforcement learning with in-switch computing

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.326047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.101725Z digest=sha256:151c09c2ead3d63faaa9aa10adca25da7fd3dc70482271c13056d23ee0810d5a

Observation 8ed1efdb-7d91-4d21-8099-e68bb427a939 · outbound

This paper cites Specification.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Specification

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.315839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.105664Z digest=sha256:2c302b61c523f6a766272739c9c62ef61f8ab43e67c8654652bd4d3d243dbf68

Observation 7f709682-02b2-49d6-829d-af86e7374b5a · outbound

This paper cites DeepSeek-V3 Technical Report.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference DeepSeek-V3 Technical Report

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.109950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.109950Z digest=sha256:600274888de67b2bcd5421f5b3556b77bddf8e5073138bd493f8482979a988aa

Observation 510e802c-84e8-47b2-af67-c7f69a80ffe1 · outbound

This paper cites Enmc: Extreme near-memory classification via approximate screening.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Enmc: Extreme near-memory classification via approximate screening

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.305602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.113850Z digest=sha256:2c374b1355d24624c9ab4c83df2e8b6c348937461e954541bf79a3e8079f20cf

Observation 9ccfcc69-ef8b-44b3-a38d-05032c7029c6 · outbound

This paper cites Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.294964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.118085Z digest=sha256:e918582f9e7b14f297efbb4a0cfcf288f6fc42083058774da838e81a2b73c516

Observation 0ae30fee-06ff-434c-8ca7-5fb36673ebfc · outbound

This paper cites Ramulator 2.0: A Modern, Modular, and Extensible DRAM Simulator.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Ramulator 2.0: A Modern, Modular, and Extensible DRAM Simulator

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.122136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.122136Z digest=sha256:ffe21daeadb91ff1522d991d4aa3bb42dcd2abc28cf6efc7e5fdbad6eb23ceb4

Observation bac7644a-2243-4259-a206-b048edc8780a · outbound

This paper cites A Binary-activation, Multi-level Weight RNN and Training Algorithm for ADC-/DAC- free and Noise-resilient Processing-in-memory Inference with eNVM.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference A Binary-activation, Multi-level Weight RNN and Training Algorithm for ADC-/DAC- free and Noise-resilient Processing-in-memory Inference with eNVM

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.283985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.126343Z digest=sha256:5c6a5b1f458fe98a9e8102dd3b730923c32fd3ed421d2e07086b67a19dee19ea

Observation 91daf12f-993f-4c47-a980-e9d882c50483 · outbound

This paper cites Dram power calculator.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Dram power calculator

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.273545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.129668Z digest=sha256:1da030f6bc8040cba5bfecdae0d8714374e641bcbdb0ce32a4465df64ddf8d47

Observation b421ff6f-4c59-477b-b23b-21d96fadea73 · outbound

This paper cites Nvidia shipped 3.76m data center gpus in 2023.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Nvidia shipped 3.76m data center gpus in 2023

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.263356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.133028Z digest=sha256:0a60c66953a0f17d5ae8d5373c13950366d37ce2b20e6d14f0977135467a9629

Observation ab8cb036-44dc-4216-915c-da2638c00132 · outbound

This paper cites Supply chain aware computer architecture.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Supply chain aware computer architecture

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.253293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.136219Z digest=sha256:9b9edd86d00d61cc56e4dc4a594b873165caaefcc688399729280a20a76779ed

Observation 7d85537c-4844-4d37-9329-0f38af56f576 · outbound

This paper cites 184QPS/W 64Mb/mm 2 3D logic-to-DRAM hybrid bonding with process-near-memory engine for recommendation system.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference 184QPS/W 64Mb/mm 2 3D logic-to-DRAM hybrid bonding with process-near-memory engine for recommendation system

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.243312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.139776Z digest=sha256:23c0613f5c3fc753b22134dde4249d3723bbdb321b020cde482fd29bd4f7af41

Observation ff55d8ee-85ff-4168-86b3-cdfc2e145b2e · outbound

This paper cites Introduction to the nvidia dgx a100 system.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Introduction to the nvidia dgx a100 system

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.232936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.143328Z digest=sha256:f8e1d753a43bebe7123309f3e2878d3aebfcae745b29c0d4472080cdd1fc77fe

Observation 87ffbc70-e3a9-4f62-a1c7-08606df975ee · outbound

This paper cites Nvidia a100 tensor core gpu.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Nvidia a100 tensor core gpu

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.222631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.146694Z digest=sha256:ae69bd10a16af703232e6bcf665c759fd4ddeaa658d3b9c96d9236b7737c888f

Observation b9db1fba-f9ff-41f1-839c-6a6dfc0ed566 · outbound

This paper cites Nvidia hgx a100, the most powerful end-to-end ai supercom- puting platform.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Nvidia hgx a100, the most powerful end-to-end ai supercom- puting platform

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.212585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.150077Z digest=sha256:4b1b0136f3db1e8d0da908cd9aef920265a733d76501600afa3657d708bf86c0

Observation 4f6571af-9e27-4194-9641-a9036d6ecbba · outbound

This paper cites Nvlink and nvlink switch.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Nvlink and nvlink switch

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.202214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.153989Z digest=sha256:9289957eca5ad1341c9fda13d3a2c370e9f344563520d55190b0206ab62bd3b3

Observation 3cb0cb57-9174-4a9d-8d7c-0057f438580d · outbound

This paper cites an unresolved cited work.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:19:44.191897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.157526Z digest=sha256:e42197aa72124fdf2694c23db28c5829d6c820350dc31dc47549b303bca39abe

Observation 48f7f33b-48ef-4d04-9e6f-567b0376ba51 · outbound

This paper cites Fine- grained dram: Energy-efficient dram for extreme bandwidth systems.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Fine- grained dram: Energy-efficient dram for extreme bandwidth systems

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.181932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.160924Z digest=sha256:068ff61b58c1fb9b37051656c37711f9ef02724edc34bc7df064336cab241ca5

Observation 001a95ce-8a7c-4781-a9a6-d1e1842e31df · outbound

This paper cites Bureau of Labor Statistics.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Bureau of Labor Statistics

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.171164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.164318Z digest=sha256:1ddba0a440fdc47d49bcd6985743a075fdc17a42c82dc83a3f89d6c72f4c858d

Observation 93dc2701-7455-4c73-a6e6-33f54ddaa771 · outbound

This paper cites Accelerating neural network inference with processing-in-dram: From the edge to the cloud.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Accelerating neural network inference with processing-in-dram: From the edge to the cloud

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.161590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.168553Z digest=sha256:3c11126ac806eba9eff076bf3c227ece7e69b926e2e5f07941cd7f20870678e4

Observation 53c05752-51aa-45d7-a70a-58eac2fb0983 · outbound

This paper cites Gpt-4 turbo and gpt-4.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Gpt-4 turbo and gpt-4

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.151325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.172611Z digest=sha256:774a05251d588ae491d110c46161e8434ed2d6eafc0128b821849cf798e00f1c

Observation 89e01c35-16d1-4ced-bc9e-a20aeb8f2d32 · outbound

This paper cites Learning to reason with llms.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Learning to reason with llms

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.140848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.176883Z digest=sha256:c485dcb595026c6a4f3f3a577b1f373c4bbcf0f8e633704eb6aa0088a7d91df6

Observation 72d9eac6-99da-4454-9711-51cc0a1679c7 · outbound

This paper cites Video generation models as world simulators.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Video generation models as world simulators

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.130360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.180173Z digest=sha256:aed27e5392d8db2fc48ebdf9f2295ef61e3af227d22807dadfa8390d66c0394b

Observation 05f5e22d-b5ed-44e4-91e8-f04aa921dd0d · outbound

This paper cites GPT-4 Technical Report.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference GPT-4 Technical Report

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.183524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.183524Z digest=sha256:55b72c29e4da387a442e30dab0fb2039382b24cebbfba3604ada6c232bd1d1b6

Observation b83c9968-c212-40fa-9f43-80d845986aac · outbound

This paper cites Cost and yield analysis of multi-die packaging using 2.5 d technology compared to fan-out wafer level packaging.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Cost and yield analysis of multi-die packaging using 2.5 d technology compared to fan-out wafer level packaging

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.116094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.186916Z digest=sha256:ff1493976365dfb769ed9592e9e216f757099e410915faec25b4ce74ae4312ec

Observation 1d76f3a5-652f-412f-a3fa-fa8c10785fb6 · outbound

This paper cites Attacc! unleashing the power of pim for batched transformer-based generative model inference.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Attacc! unleashing the power of pim for batched transformer-based generative model inference

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.104257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.190248Z digest=sha256:dbf0e06f0b1fd340bf85a11f1e6ecf4cdb8a432505b21b8b95207d587a0d586d

Observation 49de6f65-d730-407d-8c97-2d2b9afe42ae · outbound

This paper cites Trim: Enhancing processor-memory interfaces with scalable tensor reduction in memory.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Trim: Enhancing processor-memory interfaces with scalable tensor reduction in memory

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.092592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.197770Z digest=sha256:60472841acdfabc12e49b88357acb326491e977033ce6a0cb4ac0e04417fa2b1

Observation a1d117e2-4a53-43a7-8f1d-f5aca6135e4e · outbound

This paper cites An LPDDR-based CXL-PNM Platform for TCO-efficient Inference of Transformer-based Large Language Models.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference An LPDDR-based CXL-PNM Platform for TCO-efficient Inference of Transformer-based Large Language Models

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.081069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.201368Z digest=sha256:5316e5c5a0669769d83a2ab0798e6f11f98a6149adcfd406a25eae7ba7af1d06

Observation c2b6d119-1ac8-4b0c-8f60-3eb64c6beb38 · outbound

This paper cites an unresolved cited work.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Unresolved cited work

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.193719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.193719Z digest=sha256:38f25aa69b11b7197872df14ceafe6890695f5466de16e86423fd8dddba59423

Observation 5b2912e3-6b3b-4e5b-ba80-6981b4c294a7 · outbound

This paper cites Splitwise: Efficient gen- erative llm inference using phase splitting.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Splitwise: Efficient gen- erative llm inference using phase splitting

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.059017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.208695Z digest=sha256:76e10bfc95e006e7e138b98e3bd94ea4b553f68a9f80b75a685f1467e12e5a8f

Observation b30cd212-45ff-4832-aeff-ec2c75b2361f · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Movie Gen: A Cast of Media Foundation Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.212984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.212984Z digest=sha256:9a53079ede78fecb5ab231d8f3f1397c8c9d7ef740a61169627cf4f1f72bf63e

Observation cb0aafde-3aa6-426a-ac18-5890ceeaef73 · outbound

This paper cites Nvidia ada lovelace leaked specifications, die sizes, architecture, cost, and performance analysis.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Nvidia ada lovelace leaked specifications, die sizes, architecture, cost, and performance analysis

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.070811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.204959Z digest=sha256:ac741f7a2c0d55bc4bce49a6cdf9d564aad411d8ffc5621a9642b0e188ecd413

Observation 50869f67-984a-48bc-a6f2-53098521ce11 · outbound

This paper cites FACT: FFN- Attention Co-optimized Transformer Architecture with Eager Cor- relation Prediction.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference FACT: FFN- Attention Co-optimized Transformer Architecture with Eager Cor- relation Prediction

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.046764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.221118Z digest=sha256:16f9c93b6c8a4c9377ec11706dd9f9d1eeac6f6c503f05660575c2b03932cb91

Observation 7770edb4-b003-40ce-8c5b-e8e21091736f · outbound

This paper cites Dota: detect and omit weak attentions for scalable transformer acceleration.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Dota: detect and omit weak attentions for scalable transformer acceleration

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.036444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.225224Z digest=sha256:2c01487ee46bbebd6284aa480ac26a958a13e5e58458d1a65986350a8d6a27c3

Observation 31adadca-0d46-4a13-9f5d-f7a7364b8f78 · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.216975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.216975Z digest=sha256:9ad05097ea0959dc4be82f0a8b7e0165fbda6ab871e8c83039bf5f4ac479026b

Observation a8ad555e-ed1a-4fa2-9fc5-6b722b428f21 · outbound

This paper cites Impala: Algorithm/architecture co-design for in- memory multi-stride pattern matching.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Impala: Algorithm/architecture co-design for in- memory multi-stride pattern matching

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.025911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.233212Z digest=sha256:9ef8bd6bf1d587e46ad8b7a536d7b1eda19fcb37f76552804102f5ba892987f6

Observation 52e7404c-5666-46c4-bf2d-99a4b3108973 · outbound

This paper cites 8gb gddr6 sgram c-die.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference 8gb gddr6 sgram c-die

Reference 97

Resolution
verified exact
raw_fallback, observed 2026-08-08T12:19:43.535628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.236583Z digest=sha256:b4389bafe231ab5505f2236b2710ad03e903e076c98353d6f92b4c811a6afdc5

Observation 347e29e6-146d-4f10-8c3c-f9e18db03e1d · outbound

This paper cites Searching for Activation Functions.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Searching for Activation Functions

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.229683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.229683Z digest=sha256:efd17d738be8b7e117a8df29938d9ac43510dab12165ba38f1ba2be3acff5bad

Observation 36b826bd-1753-4dde-8de6-b8fb57a89e45 · outbound

This paper cites GLU Variants Improve Transformer.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference GLU Variants Improve Transformer

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.243256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.243256Z digest=sha256:aaa55b3da2f3899af5544ec00c23a2e043408d5db175cb68284a7273d78d67ca

Observation 1106256b-df8e-4e41-a921-c63c4186a17a · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.247215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.247215Z digest=sha256:057f71e00bb0b6e43d0a5340b47dd698d4c2de19bf072b738ce0e03632678946

Observation 7992e68b-9baf-43c4-87c3-91a2b39d4c87 · outbound

This paper cites Generative ai winds in memory semiconduc- tors, total demand for server drams is declining.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Generative ai winds in memory semiconduc- tors, total demand for server drams is declining

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:19:44.015080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T12:19:43.239755Z digest=sha256:419081244eb4ce18cdb78871165eb6cd1d9fa3f5b6a663eba0fd9ddfa9dc5aa8

Pith citing papers

No inbound Pith citation observations are available.