Pith. sign in

Paper Citation Record · LEDGER

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats

As of 14 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 2 inbound Pith citation observations for arXiv:2511.06838.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.06838 v4

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T00:16:27.497363Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T09:52:39.415504Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

79 of 79 outbound references displayed

  • verified exact10
  • verified fuzzy66
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 91c0cbda-410f-4f64-9607-3dcedc025466 · outbound

This paper cites AMD INSTINCT™ MI350X GPU.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats AMD INSTINCT™ MI350X GPU

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.381545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:29610b5bce74384b4fa9ab00f6a18bf230b6b9afcd84ca48690faf102df85790

Observation 0466a4a0-b073-41c4-90d2-aa0ad4611b80 · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.385570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:958354fa9276bfccacac666c1c443c52693d0f0477d86807c033520ff3f87e14

Observation 5ffe8591-3f1b-422f-be20-b97d8420901e · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:20:32.193780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:ea0b66d5a83f35ae661b70f124b176015a3af1f23e535e51acfd417b8150c3f7

Observation c9058645-61af-400b-9d8b-72c2ae33df5f · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.388138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:2dd89a53d9ef4bb4c2e0bb46ad5f1ade9d0385b174947d42daa2839fd0111464

Observation 8ca165b9-7a3c-442c-ad7f-7a252ee8e115 · outbound

This paper cites CACTI 7: New tools for interconnect exploration in innovative off-chip memories.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats CACTI 7: New tools for interconnect exploration in innovative off-chip memories

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.473674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:b8beaed737b85449e07c5dc8818074b3934f402f41686919d3ca7995083f83da

Observation 86fd75e1-3aa8-4d7a-be30-4c2a9b166c8f · outbound

This paper cites Language Models are Few-Shot Learners.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Language Models are Few-Shot Learners

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.358350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:22416fc7b1ec6caf32e8dbff560b403338dcb8227aa667f2bc79ed52b0d72081

Observation 94742906-3c9b-4726-92fd-5aafa83d2ef5 · outbound

This paper cites BitMoD: Bit-serial Mixture-of- Datatype LLM Acceleration.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats BitMoD: Bit-serial Mixture-of- Datatype LLM Acceleration

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.501615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:55e80060df9d944148210506d9a5383c7af359cdb473141113f5106be2871d24

Observation 3bfbabd8-5e52-45fe-956a-2cf4d8fd5442 · outbound

This paper cites Ecco: Improving Memory Band- width and Capacity for LLMs via Entropy-Aware Cache Compression.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Ecco: Improving Memory Band- width and Capacity for LLMs via Entropy-Aware Cache Compression

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.476680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:51b839fe84e1ed4c26dc1cd566493b65ced203cf7d6974728b80b0c4ac5d59f0

Observation 5b80588a-b746-4f55-9cb6-9fa6eb4302f7 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:20:32.226910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:16f6fdc597d24a31af87a31d6552d326b1298be2ff54d5b63c8c794c1085ed66

Observation 0a46fa4d-337c-4a22-a73d-0f95d19bc213 · outbound

This paper cites DeepSeek R1.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats DeepSeek R1

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.351770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:f1b9e8821b11a40435d2a6822e7069ae7325871b14e63e436b2a95fa79ec761a

Observation 7ce4bcba-230c-4174-a0e2-6326a867e548 · outbound

This paper cites LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:20:32.221843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:ff438eef14863baa8f9527ee97326d507debfb926a470b2ed403d35c435d6c66

Observation d89f3c01-9891-409b-ae6c-ab27d4be13f7 · outbound

This paper cites The true Processing In Memory accelerator.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats The true Processing In Memory accelerator

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.344619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:45d05b7ebd022e9dda0d9db14828a13c5965e11f12945a298c4c92595cb63202

Observation 0aa2b922-6b06-47ce-a15c-99b451485171 · outbound

This paper cites Documenting large webtext corpora: A case study on the colossal clean crawled corpus.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Documenting large webtext corpora: A case study on the colossal clean crawled corpus

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.349072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:f3d53613686f3a8e480f2b0b8e4b32e47535731f6440c24140bfe33bfe2a91d4

Observation 4d617caa-5ed4-4368-a3c5-f0e4732d6285 · outbound

This paper cites Learning from Students: Applying t-Distributions to Explore Accurate and Efficient Formats for LLMs.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Learning from Students: Applying t-Distributions to Explore Accurate and Efficient Formats for LLMs

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.342113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:62610f4fe718741ca8e123624ca908c77633f8c540eb276042993f9a8be07a64

Observation 92c72bcc-f4ed-4cf7-82c9-048c642bef3b · outbound

This paper cites Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.433912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:11c9324077f5ae5194341e18bf78f5115826bc26c786c6fda2c500a8378e5e8d

Observation fd0a485e-3ccd-4422-9514-ede7d9e78e35 · outbound

This paper cites GPTQ: Accurate Post-training Compression for Generative Pretrained Transformers.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats GPTQ: Accurate Post-training Compression for Generative Pretrained Transformers

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.338936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:ce487ae89c9f408d88d8b473f9b13b14d099be4c7416ec37d51450b0240070e0

Observation fe1e1bfc-e012-415b-a232-62872a374461 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:20:32.232926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:daccb0075c395b5842a9ab33075d8a9cb4ec804ebef130a9e5e9b7ba2ed605ba

Observation a7d3720c-d2c1-4338-80dc-2e0477ee3806 · outbound

This paper cites He, B., Yin, L., Zhen, H.-L., Liu, S., Wu, H., Zhang, X., Yuan, M., and Ma, C.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats He, B., Yin, L., Zhen, H.-L., Liu, S., Wu, H., Zhang, X., Yuan, M., and Ma, C

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:20:32.211612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:aa6812130d570695a4bf4a2844b3babf3081370d5f0df49f6c0138b114d8d519

Observation 730c8e44-014a-4199-bf83-7497416de775 · outbound

This paper cites Energy Cost Modelling for Optimizing Large Language Model Inference on Hardware Accelerators.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Energy Cost Modelling for Optimizing Large Language Model Inference on Hardware Accelerators

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.391034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:59410f1336f8e6198f861511a86c75dfad192063ac611719c41dfbd32f1bf594

Observation b3e9ca14-f00a-4855-871a-65d16eeecdcd · outbound

This paper cites OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.403649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:7b065b9af71ce0269c89974d26a610f19de74f336681554d7bd08989d9b1cc92

Observation a301124c-f97e-48c7-bce5-84946469ebf0 · outbound

This paper cites ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network Quantization.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network Quantization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.555058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:91c937b87f6a1746835934cf8db9bac675018c6d2e805ac0f0b72f0f0c135ca0

Observation f4ae97eb-9a0a-448a-a701-865a7cfd32ab · outbound

This paper cites Newton: A DRAM-maker’s Accelerator-in- Memory (AiM) Architecture for Machine Learning.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Newton: A DRAM-maker’s Accelerator-in- Memory (AiM) Architecture for Machine Learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.557822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:e1b1494a9b5335846a68aa77450206097c3d68148565e62f59881b80d3a47272

Observation 4258aa8e-5cc0-4a2e-89a0-78261b683fac · outbound

This paper cites LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture- Dataflow Co-Optimization.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture- Dataflow Co-Optimization

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.562943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:70012779b07896bce3ab777966297cae55d72080df2b04656320441df9f0fb8f

Observation 7636e3b1-8978-4041-a7ac-96c90721699b · outbound

This paper cites How Would the Viewer Feel? Estimating Wellbeing from Video Scenarios.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats How Would the Viewer Feel? Estimating Wellbeing from Video Scenarios

Reference 24

Resolution
metadata mismatch
raw_fallback, observed 2026-05-18T00:20:33.547004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:b84048e955eaf3e7b01b5f08f405333d88a8eb9fab19536255b8cc3dc62fdd90

Observation 223b3dd1-693a-4b74-bb2a-7b2810876f34 · outbound

This paper cites NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.549665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:a1dc8ff4a1f39be409ab9c9bb95016222dacc93e974a6970558367db033beec6

Observation ec386510-1f74-4678-8581-679a8b1da606 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.565565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:f205921249467403199d731f3a3972295ffa7fece427ef72d9787b41f095691b

Observation b18f5743-db24-4491-acf6-16b565ecb9f0 · outbound

This paper cites M-ANT: Efficient Low-bit Group Quantization 13 for LLMs via Mathematically Adaptive Numerical Type.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats M-ANT: Efficient Low-bit Group Quantization 13 for LLMs via Mathematically Adaptive Numerical Type

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.537830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:d02a52f74542c53ae1f57e182cc8f53bfe635be26accf7b026850293f5916034

Observation 80654548-de0e-471d-8be6-49fe29b618bc · outbound

This paper cites PLAIN: Leveraging High Internal Bandwidth in PIM for Accelerating Large Language Model Inference via Mixed-Precision Quantization.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats PLAIN: Leveraging High Internal Bandwidth in PIM for Accelerating Large Language Model Inference via Mixed-Precision Quantization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.543646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:2eecfe8a91832b9098fb953d9782889b86647b9f3882631ea549d8cf547ab8dc

Observation 51813f2e-16d3-4639-ac90-41c9a433f04b · outbound

This paper cites FIGNA: Integer unit-based accelerator design for fp-int gemm preserving numerical accuracy.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats FIGNA: Integer unit-based accelerator design for fp-int gemm preserving numerical accuracy

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.531034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:5217cf85dd29501887bbf6732178d6a62c17428ed06d11a66a17c5f852bef3ac

Observation 4d924579-6091-423b-b186-3eb49ad1d2e8 · outbound

This paper cites BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.533641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:1a6f4ebdf395639d18f6afa41d32c463d6e5aee5023b53d489dfb741d5ca3c19

Observation da94f300-dd5a-4261-a94a-4a3b8d313b0b · outbound

This paper cites High Bandwidth Memory DRAM.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats High Bandwidth Memory DRAM

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.525525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:196114f9348c32dd084b21dc588b44143d8c43da518dd86863fb50241f279163

Observation 2529d643-ad30-41e9-8c70-a908eea353ea · outbound

This paper cites High Bandwidth Memory (HBM3) DRAM.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats High Bandwidth Memory (HBM3) DRAM

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.516901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:0f89cbe572b26a8a17a25195cea4a363642742a65f9d88f4023e8ea5eb99fd8c

Observation b29bf160-1749-4b53-921e-7d66abece790 · outbound

This paper cites High Bandwidth Memory (HBM4) DRAM.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats High Bandwidth Memory (HBM4) DRAM

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.519297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:7ec8d5a2a01202292eb6ac1556b2f8e4208a4bf88c5dc4b9d8c3dfb5d5bcbf88

Observation 3200a273-e128-414a-9922-878b8eddc9a0 · outbound

This paper cites Mistral 7B.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Mistral 7B

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:20:32.203892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:1292f7bfff541d7e6f006237de685c89f8fbc61568eefd51f65726259d15f790

Observation 57c2c5a6-3de8-43ed-80b8-51209280a230 · outbound

This paper cites Ten Lessons From Three Generations Shaped Google’s TPUv4i: Industrial Product.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Ten Lessons From Three Generations Shaped Google’s TPUv4i: Industrial Product

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.510404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:c51fe0dc71f6bed1677882928ed8a1c7b96e02379d72e0e9bf3a177c6310056f

Observation 18815ced-57ba-4c65-9637-f4cad0e07032 · outbound

This paper cites SK Hynix AI-Specific Computing Memory Solution: From AiM Device to Heterogeneous AiMX-xPU System for Comprehensive LLM Inference.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats SK Hynix AI-Specific Computing Memory Solution: From AiM Device to Heterogeneous AiMX-xPU System for Comprehensive LLM Inference

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.512917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:6064221eeb705b8973110e29633d67a650f469b2dbbbe173bce669c6d9b96033

Observation 98a5828b-bcab-47f5-b38c-8744fa088213 · outbound

This paper cites Samsung PIM/PNM for Transfmer Based AI : Energy Efficiency on PIM/PNM Cluster.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Samsung PIM/PNM for Transfmer Based AI : Energy Efficiency on PIM/PNM Cluster

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.521864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:669db77392a214241be7475db3b55e37f0ce589ae9182f829cc27a326258066e

Observation 4119634a-03db-4c84-8e7c-bf3219ed4ba0 · outbound

This paper cites Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.540602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:df00d3445f3904dcbc8d993610ab91567b89b3dc07fb66d7f5a0b42680ddcde1

Observation 13268a75-31ce-41c5-afb3-3b5db08816e7 · outbound

This paper cites Pimba: A Processing- in-Memory Acceleration for Post-Transformer Large Language Model Serving.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Pimba: A Processing- in-Memory Acceleration for Post-Transformer Large Language Model Serving

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.560520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:9f42824e077292686adab6edc9126792f84169f95c1ad9347e32dd210cbe5729

Observation 4bed79c1-8e12-4a1a-9097-fb40f4493735 · outbound

This paper cites Tender: Accelerating Large Language Mod- els via Tensor Decomposition and Runtime Requantization.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Tender: Accelerating Large Language Mod- els via Tensor Decomposition and Runtime Requantization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.495865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:5cdba090d74b81ab5026582af1d6d2e90f71a27bcdd2a5e281b388d6507554ae

Observation 2d268b78-fc49-4542-b8da-6d5404da4f2d · outbound

This paper cites MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.505237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:a536f1e7244659506148d8192ed956d9481a4013e58a310043e99da93ef5ddc4

Observation cfe3a37d-6b99-4a4b-9a51-624b39404277 · outbound

This paper cites A 1ynm 1.25V 8Gb 16Gb/s/Pin GDDR6- Based Accelerator-in-Memory Supporting 1TFLOPS MAC Operation and Various Activation Functions for Deep Learning Application.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats A 1ynm 1.25V 8Gb 16Gb/s/Pin GDDR6- Based Accelerator-in-Memory Supporting 1TFLOPS MAC Operation and Various Activation Functions for Deep Learning Application

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.491616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:2f6335a631339ec0b1a36f22503b630e4df8c2be94f4b2f539393d21b6b6d72b

Observation 86221d5b-e162-42e8-ab32-738c4119ca63 · outbound

This paper cites Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology: Industrial Product.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology: Industrial Product

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.486287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:43c82f00593d71ca5b8b48fcb7a6ca65bb5a530927c73c184d3c4535e66d2530

Observation 910caeb7-a98b-4182-a0ef-9451d363122b · outbound

This paper cites H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM Inference.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM Inference

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.488924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:492175172fd3b26de7b3b4e9ae78306410d5c09a08e39109b0eac61c37e67f23

Observation f029eb9a-5252-4338-bc54-5c292db6e0fe · outbound

This paper cites ORCHES: Orchestrated Test-Time- Compute-based LLM Reasoning on Collaborative GPU-PIM HEteroge- neous System.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats ORCHES: Orchestrated Test-Time- Compute-based LLM Reasoning on Collaborative GPU-PIM HEteroge- neous System

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.507758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:ce628731b95577caeb36db4f9d6f36a51878a875218095c82ff5649836271d89

Observation a7c7dd20-0c88-4026-adb0-4a3e17011c33 · outbound

This paper cites AWQ: Activation-aware Weight Quan- tization for LLM Compression and Acceleration.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats AWQ: Activation-aware Weight Quan- tization for LLM Compression and Acceleration

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.483421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:3f3a52ccf7ef91aaf6b36073ff9da5155b8e213c84e66777d8ef47dada11ed4a

Observation 7eb3f5a8-fed2-4360-b683-6b1fbe435314 · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.463754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:12bc82d8a3b8da6ddc3e77d464a2a5fae012e15148e2ac6f553569cede75708f

Observation d00ef8dd-72a6-4f6e-aa24-477d713a4d0a · outbound

This paper cites SPARK: Scalable and Precision-Aware Acceleration of Neural Networks via Ef- ficient Encoding.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats SPARK: Scalable and Precision-Aware Acceleration of Neural Networks via Ef- ficient Encoding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.467168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:e13759323db193b28e5702b0f4811bf672c8411dadcf451f8b90a9b41b484865

Observation 4b9b8b7c-91c6-44dc-a0c2-0e6b37189d8c · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.470279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:6c7d63f88a843cb33cf059ab61a8c5ae22980657b67fdf608142f9b7fd486f55

Observation 370a2073-acda-4e41-a9a1-f4bbd43360f2 · outbound

This paper cites Ramulator 2.0: A Modern, Modular, and Extensible DRAM Simulator.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Ramulator 2.0: A Modern, Modular, and Extensible DRAM Simulator

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.479541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:c4ffda7e1f5292f03a08224f562ac2c508711ab088626d2e2bd394cd6405d83a

Observation f560dd81-1e11-4663-be9e-d10ed09d5862 · outbound

This paper cites Pointer sentinel mixture models.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Pointer sentinel mixture models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.498845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:5eca2a5306f4b1891f4e597a428ceebd18a6472934a3a99ee2ea9405f10b34cb

Observation 04b1a893-0387-4d43-9ed3-4cbe65cf78bd · outbound

This paper cites Introducing Llama 3.1: Our most capable models to date.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Introducing Llama 3.1: Our most capable models to date

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.452529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:24b148bf0c4c2554d31f6016a4908f8cb55798310b9b9ff46c4d09ed51ece119

Observation 5b87c798-2369-44d1-8b9a-d1c7a15dd751 · outbound

This paper cites Llama-3.2-90B-Vision-Instruct.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Llama-3.2-90B-Vision-Instruct

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.458550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:736efa00713262dd152d0fabe717f8c8f93ca21522cae20cbbd071fc18b0f721

Observation b2b80751-7dd5-49d2-8f89-c2600d8ab179 · outbound

This paper cites Llama 3.2: Revolutionizing edge AI and vision with open, customizable models.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Llama 3.2: Revolutionizing edge AI and vision with open, customizable models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.448130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:b2193b328d8b93322e46993eda5baf9eda90d0512fc38efab5e6974382c54439

Observation e16d8a53-e957-4f1f-8ae8-234a62a67c34 · outbound

This paper cites Meta Llama 2.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Meta Llama 2

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.455706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:5561349da0768c8620565d6ba6fc4643e3c805c2476e8a5e7d2fbcda9b220ab9

Observation 474e1aa6-4ec8-40a1-b4fe-a129476d8109 · outbound

This paper cites The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.528443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:c0535af8bc0cdfb3db4541f9bd2cd2e6e06252305dffc01069b5bc98096706a4

Observation bb3509a2-068d-4a9b-b2d5-f5c029e1ec85 · outbound

This paper cites FP8 Formats for Deep Learning.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats FP8 Formats for Deep Learning

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:20:32.216570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:f68c3e19b13f6f3e502108e41c19d961e656dafa5ccb9768803bd599acd71a08

Observation 5093e4e9-adfc-4cd1-828d-101a3284c8da · outbound

This paper cites Introducing NVFP4 for Efficient and Accurate Low-Precision Inference.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Introducing NVFP4 for Efficient and Accurate Low-Precision Inference

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.437286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:ac5e0f88f8ee1827b21cfa14bb5d53328261fc7600072139d0bac7f7f40cad24

Observation c1b5315e-0c17-41a2-9933-d09ccb181804 · outbound

This paper cites NVIDIA Blackwell GPU Architecture.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats NVIDIA Blackwell GPU Architecture

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.441718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:22deb7d9f827183f862ebb5b5ab825f23a56502d82e466e178d71fd9810b8c16

Observation ac45eb04-401d-49ec-99e3-160476b0c899 · outbound

This paper cites Openai o3-mini.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Openai o3-mini

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.429274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:a264d8786a28b146e86cac4256645d44476ec9049c263c2dadb6495173d4b0b8

Observation 93764bbe-b56a-4fb8-91c1-507055548e46 · outbound

This paper cites Gsm8k dataset.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Gsm8k dataset

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.426603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:c3c1d889faa8a6f19d62bcedd888a6a9230636a0e171e99438ab49b28ae4c068

Observation ec768789-ed4b-46d9-b381-57899850a37f · outbound

This paper cites FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.393649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:75a1cb385f63f8fef1c76af8e6581e0573373c497eb671077bb3a6fada25eb2e

Observation 93f03422-2c83-4024-b65d-1f3bfbff073f · outbound

This paper cites AttAcc! Unleashing the Power of PIM for Batched Transformer- based Generative Model Inference.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats AttAcc! Unleashing the Power of PIM for Batched Transformer- based Generative Model Inference

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.444850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:77daddec826174eb4c43eaac7f1142d2fe53c193e93e960fdbacad9b34517e3a

Observation 913f3582-5df7-48b6-83e0-a79860ba73bc · outbound

This paper cites MicroScopiQ: Ac- celerating Foundational Models through Outlier-Aware Microscaling Quantization.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats MicroScopiQ: Ac- celerating Foundational Models through Outlier-Aware Microscaling Quantization

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.354666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:965c2c6d3733f8a00852f98105695c840ff890cd3261ba29bd5a8c970ba78af5

Observation 82a7c380-c6ea-4179-8f22-dd0fa2eceba7 · outbound

This paper cites With Shared Microexponents, A Little Shifting Goes a Long Way.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats With Shared Microexponents, A Little Shifting Goes a Long Way

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.421658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:75fc844b7c1c4502db6180684a9933b531351d6b6dfe6b1ff9d4e3e4861be013

Observation bf252da6-c9db-4ab1-a5d5-5bf9851bd74e · outbound

This paper cites IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.424042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:797732b5f9387342efa5f9d7fbe5d94220577dc85fda3e213f00b34e033260e8

Observation bbb88c1a-9395-482d-8aad-0a3394ea7e9c · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.418893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:057090c9d571eba2c340e1114a100e44d27078c47aed9d3ec6c5fafd1b9824c5

Observation dfc59e7d-6eae-497f-9578-b4905e295fd1 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.361607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:baefc03cce8ca0218255d5b511a6fa5440876edf2ed2f941253cbc0ac99ebdd3

Observation 461f9047-e068-43af-98fe-113a9a93a2bf · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats LLaMA: Open and Efficient Foundation Language Models

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:20:32.170271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:ed808433ac3712d0b90bbbed223e67e5261e508a26a66446a425ab795d49a6d0

Observation a0879e73-781c-4dcd-a536-8001cd9782b6 · outbound

This paper cites FP8 versus INT8 for efficient deep learning inference.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats FP8 versus INT8 for efficient deep learning inference

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:20:32.176875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:509dc515576dca5d99858fcbc802b7529b7919edd69215c975f286ceaad50ccb

Observation 3c4815e1-f3c6-4828-9b6f-ad776c5e7580 · outbound

This paper cites ZeroQuant-FP: A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats ZeroQuant-FP: A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:20:32.183663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:5d3f87bdf0a5ef8e07c27a4ed13554f9d039c60787514a80e720ec7a6de203ea

Observation 49cc89dc-9e41-440a-8596-e2db780c35d4 · outbound

This paper cites SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.413716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:a116c986a141790d528f9c6f32fe8f58b3d6825f398038ba4cd9134f4db4070a

Observation 2a03912f-72ba-44c6-9c7c-0db39722b8e6 · outbound

This paper cites Amove: Accelerating LLMs through Mitigating Outliers and Salient Points via Fine-Grained Grouped Vectorized Data Type.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Amove: Accelerating LLMs through Mitigating Outliers and Salient Points via Fine-Grained Grouped Vectorized Data Type

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.416502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:ae3cfa51e7d8d39de4b1c860bbc222af62800e0752bc35c506cd991e82dd6b64

Observation 5eb34fdd-7deb-4ff6-aee9-ffd31afc069a · outbound

This paper cites Qwen2.5 Technical Report.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Qwen2.5 Technical Report

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:20:32.188473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:660b0767aa63950a4dd2f852b00747e251b46cffa76ca2b71f193df617efc9b2

Observation 6aca5194-9473-4521-8077-568ed4212ed6 · outbound

This paper cites FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.411176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:5ebfe9960c8e001c3be0661d27b924b4d40cb86e1c17caf8a66e03bdd9a443f5

Observation a23971a4-014b-412e-b08c-da7d65ce5ceb · outbound

This paper cites Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batch- ing.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batch- ing

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.400872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:e8dcb85fa123c038a1723fac1f3ef98587650783f46d1a51035d843c31de2ec6

Observation 7cce09f4-7b68-494d-bc38-8972158ae4e1 · outbound

This paper cites SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.396613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:e0f75ae4776ec1ec734417f6708b7cd1f6fd6b8376d92f3ebb1db0bc87a295b0

Observation 0e81c5fb-20f9-4e49-a0a4-3ed5e59f177f · outbound

This paper cites DistServe: Disaggregating Prefill and Decoding for Goodput- optimized Large Language Model Serving.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats DistServe: Disaggregating Prefill and Decoding for Goodput- optimized Large Language Model Serving

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:20:33.406530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:96d948acf9b1ad6edf72a88fadf5f45c047a1f0f0a54abf1e50c93de104b1cb2

Observation f067f8df-4765-48fc-aab2-61a31f82e5ad · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats A Survey on Efficient Inference for Large Language Models

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:20:32.198469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:16:27.497363Z digest=sha256:88b33c07f838fb94ea486806cc8a9d3bd39b7db32c76e87ca79862e3f994f527

Pith citing papers

Observation 6e5d9942-adc4-496a-9eb2-3d3afd6144d7 · inbound

CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device cites this paper.

CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T09:52:39.415504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:52:39.415504Z digest=sha256:729fbc95b5acb9e677bb9253e3bf4bcffabfabd07242c169e65869c31462b884

Observation 3c97f8a5-465f-478d-86b0-258194387a0e · inbound

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference cites this paper.

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T22:36:29.218635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T22:36:29.218635Z digest=sha256:76d064d476ffe71e90e3121f663571e48a598198852796388d43d96263dda4ef