Pith. sign in

Paper Citation Record · LEDGER

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM

As of 21 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2608.06989.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06989 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:07:38.016961Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact2
  • verified fuzzy35
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 080f7c07-e0e8-47ca-b1b9-04d6b9377184 · outbound

This paper cites cublas docs,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM cublas docs,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:39.062564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.695553Z digest=sha256:22a9e3721e19a5498d7584f74e8838cd25874e38b33d89b5d2983e0152462225

Observation 3b74aec1-f5df-4baa-8530-498f67f321dd · outbound

This paper cites High bandwidth memory dram (hbm1, hbm2) jesd235d,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM High bandwidth memory dram (hbm1, hbm2) jesd235d,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:39.047523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.700788Z digest=sha256:ccf79ee3ef3a967f6ef5a4b40f3945f801ba41582b112b4391f603722e1ac69b

Observation 7468b991-2e9c-4c07-bbf3-33ac2f79386e · outbound

This paper cites Analyzing cuda workloads using a detailed gpu simulator,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Analyzing cuda workloads using a detailed gpu simulator,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.705355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.705355Z digest=sha256:a171d59a1a71ebfc51d2b945b12663e2329182c910c95d16f1e22ff5512f22ba

Observation 378b0234-3b45-48ff-a9de-f7567c195570 · outbound

This paper cites Moe-lightning: High-throughput moe inference on memory-constrained gpus,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Moe-lightning: High-throughput moe inference on memory-constrained gpus,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.711044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.711044Z digest=sha256:22f3ef4b2e3819fe8d052e04607be302bd7fb1a3141f8f2c25fa0eb04caa7444

Observation 444dc580-08d2-4c80-a5a2-a38ff775be57 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.715953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.715953Z digest=sha256:4be81eabecb3a7abf70133b36ef7c3b262874cec448726d2af51e3084f52d994

Observation e90bd290-db89-4fdc-9e77-870ed6ebfafc · outbound

This paper cites Palm: Scal- ing language modeling with pathways,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Palm: Scal- ing language modeling with pathways,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.720747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.720747Z digest=sha256:bdd25fb35d08574b04e71b5a461b75f7d91343144102f5a07fc40fc3f7682733

Observation 4e7cea13-109a-4cc0-89e2-fe66b3d6da83 · outbound

This paper cites Asap7: A 7-nm finfet predictive process design kit,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Asap7: A 7-nm finfet predictive process design kit,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.726009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.726009Z digest=sha256:a53f049a59cd0dfaa1db144c0d3c801896547664033ed30817af8a64e1fc72fd

Observation 20fef4f4-2bdd-4f8d-9c9b-1bbd8eb785b5 · outbound

This paper cites Projects – COIN-OR: Computational infrastruc- ture for operations research,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Projects – COIN-OR: Computational infrastruc- ture for operations research,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.983040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.730812Z digest=sha256:fae7c9b327dcfab1f315e7b40d34aef726a87c092b295b916268e13c4d652a5d

Observation 57bb3874-8eea-4b32-99e2-08dbf80f3092 · outbound

This paper cites Deepseekmoe: Towards ultimate expert special- ization in mixture-of-experts language models,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Deepseekmoe: Towards ultimate expert special- ization in mixture-of-experts language models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.965110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.735097Z digest=sha256:165a1d0a8e17433e06ecc3b8caa2aa392f12c2d418da0f26647f8fb9f549c275

Observation 7d215e54-60af-4e43-ac25-35cb46e632be · outbound

This paper cites The true processing in memory accelerator,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM The true processing in memory accelerator,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.739364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.739364Z digest=sha256:a57cac06b6e7a2057e1de3174cd4dda7bc52c839c42a83b612e2170b36b83424

Observation b71ffb43-2818-4c36-af5b-cf38e62eec34 · outbound

This paper cites Accelerating llm inference throughput via asynchronous kv cache prefetching,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Accelerating llm inference throughput via asynchronous kv cache prefetching,

Reference 11

Resolution
verified exact
raw_fallback, observed 2026-08-10T17:07:38.266524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.743717Z digest=sha256:cfecb62e324b57a3e3ecfc13cc7c46155e0cf5a9c6b36e1b73889fccf4316e45

Observation 50e9a289-18c4-43d8-9783-6937f46a02a9 · outbound

This paper cites The Llama 3 Herd of Models.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.748311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.748311Z digest=sha256:5f5f42346dc961d5cc0602eb873d91db98dc28c4a9993e433628e63d03f48ff1

Observation 2b136343-57ca-44f0-93c7-7c43136ff078 · outbound

This paper cites Klotski: Efficient mixture-of-expert inference via expert- aware multi-batch pipeline,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Klotski: Efficient mixture-of-expert inference via expert- aware multi-batch pipeline,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.936738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.753254Z digest=sha256:b62182df474cc50f47f166a560a557a2a4b5db5544360b4eefd529aa36f4e5c0

Observation 235f5503-dd76-40d3-b23f-ff66dc40534a · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.758985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.758985Z digest=sha256:a0e46f4a9dd3b045f3fac57de9ff47bd9d498c6a65b77ef6dd1494f57ea18e7a

Observation d274dc01-2f97-413f-ab86-9d08951e2c2c · outbound

This paper cites The future of low-latency memory: Why near memory requires a new interface,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM The future of low-latency memory: Why near memory requires a new interface,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.908916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.763452Z digest=sha256:d96420673d2362e20ec455cd0c26f5e1899ded33622ada0701ead384f5968023

Observation 3ff899c6-eaac-45f7-8bf5-314a966c18ac · outbound

This paper cites Papi: Exploiting dynamic parallelism in large language model decoding with a processing-in-memory-enabled computing system,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Papi: Exploiting dynamic parallelism in large language model decoding with a processing-in-memory-enabled computing system,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.892099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.767957Z digest=sha256:4b6d4d07e8cd4d1775d9d8249d98b6594532df921d32baa6046cf5abc4eaeb0c

Observation b22b43ca-6b2c-4ab5-bd37-ff753a166092 · outbound

This paper cites Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.876912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.772704Z digest=sha256:56d53046003f0c358c8c87b5c4085ba58adaa9de1039b676c7e3a6b45d1fdf09

Observation 4d57ed89-8e91-4723-b661-49ead71d38c0 · outbound

This paper cites Lightllm: A versatile large language model for predictive light sensing,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Lightllm: A versatile large language model for predictive light sensing,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.862011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.778011Z digest=sha256:d3b7618a547d1b31902fb103610692db1b6d126b25bf5f78d11550bcc0c5d1db

Observation d50537ad-5740-4fbb-94c1-ee347a01d742 · outbound

This paper cites Decoding llm performance—prefill phase is compute bound on npu,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Decoding llm performance—prefill phase is compute bound on npu,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.846499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.782728Z digest=sha256:23847cb3377ab579b2c51185d5b97d58ad700a3f52891860396bb7e56de8ebbd

Observation 705e2e2a-8662-45f5-9cdc-ad0240330214 · outbound

This paper cites Mixtral of Experts.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Mixtral of Experts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.787253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.787253Z digest=sha256:bced3dece683662f5cb4e404bc1a5cc3639cc165bb591f2468afd4dade48bf33

Observation ae003cce-039d-4a6a-aa1f-6631b5566e1d · outbound

This paper cites The battle of chatbot giants: an ex- perimental comparison of chatgpt and bard,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM The battle of chatbot giants: an ex- perimental comparison of chatgpt and bard,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.831974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.792301Z digest=sha256:65da4fac6845fd874f6fa32e00b9c930e700da96d6309cf5b740218d4edcf5b5

Observation c3281b42-77ec-45ea-bdbe-1c02dcc7cb58 · outbound

This paper cites Accel-sim: An extensible simulation framework for validated gpu modeling,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Accel-sim: An extensible simulation framework for validated gpu modeling,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.796835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.796835Z digest=sha256:7f9da49e8aeab33b48c9702bad9ea1ad7428b3f6f4a53f1dcebbfcf298c840b9

Observation 7933b684-d142-41de-a549-86ffa7766c08 · outbound

This paper cites Sk hynix ai-specific computing memory solution: From aim device to heterogeneous aimx-xpu system for comprehensive llm inference,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Sk hynix ai-specific computing memory solution: From aim device to heterogeneous aimx-xpu system for comprehensive llm inference,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.801418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.801418Z digest=sha256:65219ccc2b5cac257c2a70d41412ed7b0ab251fe112154bc759b68f96afead7b

Observation afacd144-eb5c-4f26-9bca-40082e3e05fc · outbound

This paper cites Monde: Mixture of near-data experts for large-scale sparse models,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Monde: Mixture of near-data experts for large-scale sparse models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.797868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.806290Z digest=sha256:7766d03f0f3c3966b87bf86b73d1cf1673dccba9012fb68ce5b3805a9e7e50be

Observation afe5b5f9-c85c-48ec-ba8e-8fde1e61f319 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Efficient memory management for large language model serving with pagedattention,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.810685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.810685Z digest=sha256:e68b8fc37b007b6af01c7abb73d9e13d53e0d1f6ccbea457eaa13430d3e1f5f5

Observation 6b992dcc-1456-4ffe-9245-e8a383c00909 · outbound

This paper cites vllm: Easy, fast, and cheap llm serving for everyone.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM vllm: Easy, fast, and cheap llm serving for everyone

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.771604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.815171Z digest=sha256:88db4bf117cab6ecee05683d075d83766a65df420f710b75fddb9d21ca16deee

Observation 1b98ec61-3a59-4a96-bfbc-b047e080cf63 · outbound

This paper cites A 1ynm 1.25 v 8gb, 16gb/s/pin gddr6-based accelerator-in-memory supporting 1tflops mac operation and various activation functions for deep-learning applications,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM A 1ynm 1.25 v 8gb, 16gb/s/pin gddr6-based accelerator-in-memory supporting 1tflops mac operation and various activation functions for deep-learning applications,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.819794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.819794Z digest=sha256:a7ed3da5f4438df3dbe9cfb95b8532b0baca946404ad7e4864a257a0b7e0f44e

Observation 42c4c93d-eb61-4471-be5a-5da3eb13c67a · outbound

This paper cites Hardware architecture and software stack for pim based on commercial dram technology: Industrial product,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Hardware architecture and software stack for pim based on commercial dram technology: Industrial product,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.825298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.825298Z digest=sha256:82ce330f12030f9dccbc7e049e8ad9f298589e78a0b7e57dd536170aefa01cc0

Observation e8f5c68a-4265-4551-9b81-93b293924ff5 · outbound

This paper cites H2-llm: Hardware-dataflow co-exploration for heterogeneous hybrid-bonding-based low-batch llm inference,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM H2-llm: Hardware-dataflow co-exploration for heterogeneous hybrid-bonding-based low-batch llm inference,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.736234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.829870Z digest=sha256:43bac3d18b6d17a20e469c7c7885232a3730d945265b9ecc27721c146abc41c8

Observation bf29786f-762f-4bd4-adeb-e0dace7d7ac5 · outbound

This paper cites Ascend: a scalable and unified architecture for ubiquitous deep neural network computing: Industry track paper,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Ascend: a scalable and unified architecture for ubiquitous deep neural network computing: Industry track paper,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.834220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.834220Z digest=sha256:6367438d0b2727972dc5977c068dae56d08e738aee31c8309acc136768392c40

Observation d54ef2b4-5686-4447-8c92-cb43a559cf07 · outbound

This paper cites Davinci: A scalable architecture for neural network computing,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Davinci: A scalable architecture for neural network computing,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.838777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.838777Z digest=sha256:c1a838ebe748ed1224f45e2cca143b872ed284dddf3719d30dfbd246ca56e8b7

Observation 883731a9-3936-4704-9b81-bae3e3063d8e · outbound

This paper cites DeepSeek-V3 Technical Report.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM DeepSeek-V3 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.843266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.843266Z digest=sha256:b8f73b53a13d3516dfabceacd56ca852f35eaf03718d673c3e33950c60dd65a9

Observation 04065920-91ef-4545-b53e-51ced19f61f5 · outbound

This paper cites Ramulator 2.0: A modern, modular, and extensible dram simulator,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Ramulator 2.0: A modern, modular, and extensible dram simulator,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.848227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.848227Z digest=sha256:4f515fbc4805e30b64751b3a0fde8daa1908eeb27feb52dd188ad20d310f48f5

Observation 3eeb1b74-e7fd-49a2-88cd-ca2bc0b4e623 · outbound

This paper cites The design process for google’s training chips: Tpuv2 and tpuv3,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM The design process for google’s training chips: Tpuv2 and tpuv3,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.852810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.852810Z digest=sha256:12d887abf0edbf59f6db76edda0d4d8afcf6dc96b8effa8a992692f52e38f87e

Observation 168726cf-77ee-4eea-8682-0b9035b2d1a5 · outbound

This paper cites Nvidia a100 tensor core gpu architecc ture,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Nvidia a100 tensor core gpu architecc ture,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.678996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.857378Z digest=sha256:0af1b2b09aa0bab2ecee40f91bc094d4e32aacf41dc74d2eb8ce86048531023c

Observation 2cbc5a3e-9369-4398-9bdb-33fc15fed075 · outbound

This paper cites Introduction to the nvidia dgx a100 system,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Introduction to the nvidia dgx a100 system,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.663185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.862209Z digest=sha256:ef6ae6d48ed68831a9986a5d4485ed31e4a83c5ae9f5db2b5fc45c83feb05af8

Observation 617a579e-4348-4adb-9dd0-e10c31e98954 · outbound

This paper cites Nvidia h100 tensor core gpu architecture,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Nvidia h100 tensor core gpu architecture,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.647557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.867070Z digest=sha256:0b04bb910d0aa1eeb9171dc4a2f3f3053c6265a8226e331f877277b83ae5b491

Observation acc7508a-f626-4fb4-bf34-0cabf49df959 · outbound

This paper cites Chatgpt,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Chatgpt,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.631915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.871919Z digest=sha256:e15aaace46c32451b2c5eb0de22967f3c95d356c7f899b14dd463a7a4db223a1

Observation f63776a2-4a47-46c1-8ed1-6407f5661d28 · outbound

This paper cites Gpt-oss-120b,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Gpt-oss-120b,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.616441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.877248Z digest=sha256:11398dde5ff5f431a42bd78a8bda2d67b2d7ad7d227efb6d03a86c911c070d36

Observation 498eb86c-0d7a-4e71-95d6-6352ef7bec5f · outbound

This paper cites VESPA: VIPT Enhancements for Superpage Accesses.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM VESPA: VIPT Enhancements for Superpage Accesses

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-10T17:07:38.126461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.882236Z digest=sha256:fc5fa9a8d852af1d934724a89db8a979edf9a603cb5874b137880cf367ca035a

Observation c8ed4d31-0f97-4ba3-ac9b-152c07ca6a22 · outbound

This paper cites Attacc! unleashing the power of pim for batched transformer- based generative model inference,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Attacc! unleashing the power of pim for batched transformer- based generative model inference,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.887482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.887482Z digest=sha256:f742ea4a25edd9b9df470b44694caee1f1a0d5a2a6a4bc439a7a6b441f432841

Observation c7f2f9ad-fd01-4eeb-b859-ee9cb25b790f · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Pytorch: An imperative style, high-performance deep learning library,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.892413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.892413Z digest=sha256:605431cb78e178ed16e2774cc44cb151cdd4e8b8b49f84c557b176409c7070c3

Observation adcbebf4-df7a-45c7-aab6-28ec5758aefd · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Splitwise: Efficient generative llm inference using phase splitting,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.896797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.896797Z digest=sha256:b8e240ce810c3dd290da4cc2e0a968b7dc4b85bd2204895b00925484050a9b5a

Observation 05ba398e-7adb-45ab-ba9c-aa9d5b2a9f79 · outbound

This paper cites {DRAMA}: Exploiting{DRAM}addressing for{Cross-CPU}at- tacks,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM {DRAMA}: Exploiting{DRAM}addressing for{Cross-CPU}at- tacks,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.901598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.901598Z digest=sha256:f5fa43996ff245b0b60819817e169faab048297e7df7c22132eddfccd18869db

Observation 99cb3109-344e-4acf-91cc-a742c7736005 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Code Llama: Open Foundation Models for Code

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.906475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.906475Z digest=sha256:619f6b738375dbb29e4e383529165744a8f57b9090d86822b38e6d872e01f390

Observation 1ca98075-9a99-47dd-82d6-f1495de28c4a · outbound

This paper cites Ramulator 2.0,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Ramulator 2.0,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.562417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.911242Z digest=sha256:3126f31bee19a9415b42dacec6379163c51c0f30f87db30937ef2a0b3f9a9664

Observation b8f15562-cb00-4ea5-b30e-e32804b6858b · outbound

This paper cites Khaa44801b-mc16: 8gb hbm2e flashbolt,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Khaa44801b-mc16: 8gb hbm2e flashbolt,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.546657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.915869Z digest=sha256:4635ebac58f39617557bb7f372b88d5a585991aa10a0fc2bf8e267b70fa0f96e

Observation 3cd728a0-7f39-48ec-a0e5-4e529ffa946f · outbound

This paper cites Ianus: Integrated accelerator based on npu-pim unified memory system,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Ianus: Integrated accelerator based on npu-pim unified memory system,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.515151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.925266Z digest=sha256:b6c503d985457669a4cec18430b2d9e3147a5e5e5ea092a06c4b7e26f268a41f

Observation 81964d28-58fe-4e09-9d97-aef75e69b67e · outbound

This paper cites Facil: Flexible dram address mapping for soc-pim cooperative on-device llm inference,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Facil: Flexible dram address mapping for soc-pim cooperative on-device llm inference,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.929827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.929827Z digest=sha256:bfaae789a76008c86ad4f4b3b9cf875c76b6fa0eb57edf6dd5e2ef7c3ce43f8e

Observation b49f68be-cefb-4d46-b4dd-8d348a120514 · outbound

This paper cites an unresolved cited work.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-10T17:07:38.489538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.934297Z digest=sha256:b1c310eae9b4a1929066f4178451f50f7f295063da174d58c4fd38c14e9a3391

Observation b5b64ee8-ab0d-4c2a-9d06-970ec14438ff · outbound

This paper cites Design compiler,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Design compiler,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.475000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.939466Z digest=sha256:de72f4638710dd6ef16ed7be0f0769fa6e4211982260eae18258fba765e8f5b5

Observation edce804e-4cf3-45d1-93c0-13f3a8e855f0 · outbound

This paper cites Qwen1.5-moe-a2.7b,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Qwen1.5-moe-a2.7b,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.460323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.944163Z digest=sha256:fdc8b55288208432fcdf6676d753c14621439f41ec51b87b975ad970e1c564bd

Observation 96078f9c-7b54-41d4-b7ab-f80ab14f83f5 · outbound

This paper cites Sglang: Efficient execution of structured language model programs,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Sglang: Efficient execution of structured language model programs,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.443975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.948535Z digest=sha256:c9be480a492350bd9889b5dc026f5514cbe0e0cd1c8f12778a1eb6805d60a654

Observation 910f8908-83a9-411e-8b17-1125c43caea1 · outbound

This paper cites Introducing the next generation of claude.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Introducing the next generation of claude

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.426565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.953117Z digest=sha256:761309e232b96e5fcc11f20f6749ca1abc19f5cb747d70405d2e698d1bd30fac

Observation 5a9f9032-2d6b-4cd1-a3fd-95e6b61f826e · outbound

This paper cites Upmem software development kit (sdk),.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Upmem software development kit (sdk),

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.410712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.957755Z digest=sha256:1ec5a7ae30a466e9a0972c6e0ff07fc50118d627799d8e59a4d76dba79a6fef1

Observation ac10e1cc-307d-4650-b64f-292aa15c45a8 · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.962298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.962298Z digest=sha256:b9c8fc9350cd0b52db60910bb4ed34737f8f86555763d21175754f6ecbb3583b

Observation 16927598-8583-4e6b-9871-b9556d42bcbb · outbound

This paper cites Pim gpt a hybrid process in memory accelerator for autoregressive transformers,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Pim gpt a hybrid process in memory accelerator for autoregressive transformers,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.395249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.967462Z digest=sha256:32fcec03a330c8db4700c78044d86510e219be008142b550fb3db00fe570f6ca

Observation bb2bf72f-8ddb-4c6a-8ea8-7a31bf9af2e5 · outbound

This paper cites Waitgpt: Monitoring and steering conversational llm agent in data analysis with on-the-fly code visualization,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Waitgpt: Monitoring and steering conversational llm agent in data analysis with on-the-fly code visualization,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.379994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.972628Z digest=sha256:6d787f20e03966e2e6eb53c1a99dc154e0d06685f70f8d61fd16211f1d80f9a9

Observation 99b76cda-51fc-4b0d-99ae-85fb04ce4241 · outbound

This paper cites Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.977874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.977874Z digest=sha256:24f034673d95d48022e8d85dbf394209bc5618c79642402e8180a64bb05bbf81

Observation 5f496768-78ad-464e-b06d-01e9d5510997 · outbound

This paper cites Duplex: A device for large language models with mixture of experts, grouped query attention, and continuous batching,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Duplex: A device for large language models with mixture of experts, grouped query attention, and continuous batching,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.364518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.982655Z digest=sha256:179d409ef4d89198d8ed43b419cfdccaadb86d6cde14dc489c3f26a359fcdbe3

Observation 1a943f7c-7f6a-461c-a0d9-e69183ee3b60 · outbound

This paper cites Llmcompass: Enabling efficient hardware design for large language model inference,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Llmcompass: Enabling efficient hardware design for large language model inference,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:37.987438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:37.987438Z digest=sha256:cc799bedaacb3abd1e332dc0a4defa9a1d4c260f1c996f67f54a4443f25e7bf2

Observation 561cccd9-80e0-41b3-b6fd-2de7b77a3ceb · outbound

This paper cites Um-pim: Dram-based pim with uniform & shared memory space,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Um-pim: Dram-based pim with uniform & shared memory space,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.339889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.992508Z digest=sha256:74bc7a5d4b522c939b32795ca8affbc6282303c358c0c649051b4715a5157d3a

Observation 89f41714-2435-494d-a2d4-fbf1fe817cec · outbound

This paper cites Lmsys-chat-1m: A large-scale real-world llm conversation dataset,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Lmsys-chat-1m: A large-scale real-world llm conversation dataset,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.324063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.997008Z digest=sha256:7ef2376d4241d75b55fc5618016855b6357e80ba658291b5d87cca81dcb07e85

Observation 71c53d55-488d-4640-9f8b-b233e79affa2 · outbound

This paper cites HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:38.001641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:38.001641Z digest=sha256:83b70cd539b244794ff529f5bc0283c00165d1c94516ba4855988fec3fcde2c4

Observation e4237319-436f-4a50-9577-d0bae12978e3 · outbound

This paper cites {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T17:07:38.006403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:07:38.006403Z digest=sha256:66ead60f727fd26f3a47a63ae31bc3c74d5aaf593097ca3b9be783085843cfc1

Observation 62ce0ff7-c17c-4158-8e2c-e5372e6186df · outbound

This paper cites Mixture-of-experts with expert choice routing,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Mixture-of-experts with expert choice routing,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.297303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:38.011039Z digest=sha256:62d72bc1e8f233ea09e9f6a9c8e66b315404dab8c9c36a9c9bd8f72002656098

Observation f3e9a925-87d8-4d46-b866-a74baf7971b1 · outbound

This paper cites A comprehensive analysis of superpage management mechanisms and policies,.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM A comprehensive analysis of superpage management mechanisms and policies,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.281956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:38.016961Z digest=sha256:fece8ef712eeb76dedf68cf8a03bd020c3b1876a39917da20c69d64633ae8782

Observation 93da23a9-244c-4f81-b22e-cd7f9d2d58e7 · outbound

This paper cites Available: https://semiconductor.samsung.com/dram/ hbm/hbm2e-flashbolt/khaa44801b-mc16/.

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM Available: https://semiconductor.samsung.com/dram/ hbm/hbm2e-flashbolt/khaa44801b-mc16/

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:07:38.531333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:07:37.920527Z digest=sha256:7b5e1f405f9cfbb724fe79c0a2820d8dbf71d895955a368ca057854ac6c7bc17

Pith citing papers

No inbound Pith citation observations are available.