Pith. sign in

Paper Citation Record · LEDGER

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge

As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2608.09291.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09291 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:11:14.200548Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact5
  • verified fuzzy30
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e32df82e-0b82-4bde-b96e-a6fc77a698a9 · outbound

This paper cites Attention is all you need,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Attention is all you need,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:13.880373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:13.880373Z digest=sha256:20acfb276af4fca9d9a1713d0a5383ed1dea673c4ede7ff40d0c4268248dd2ba

Observation 36b613a3-f314-42a5-aeaf-e9b2e1b0b291 · outbound

This paper cites GPT-4 Technical Report.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:13.893367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:13.893367Z digest=sha256:c4f6c8d872445e01879ac2295b77ea840665103645a645c8ec7b18c1b5294045

Observation a455d192-1b8b-4c51-a974-5444b2ff48f5 · outbound

This paper cites Edgellm: Fast on-device llm inference with speculative decoding,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Edgellm: Fast on-device llm inference with speculative decoding,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.322895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:13.898827Z digest=sha256:3cee6aa7a5d8447a7cbf63a0dbb36d6f0e8352fab78b186ab1b01345965f43d1

Observation 0a85edc2-af46-4ffc-8b5d-7f707be0105f · outbound

This paper cites Tenet: An efficient sparsity-aware lut-centric architec- ture for ternary llm inference on edge,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Tenet: An efficient sparsity-aware lut-centric architec- ture for ternary llm inference on edge,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:13.904026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:13.904026Z digest=sha256:a9c9583b3f9079be649c39ee0826398da8d3be3f8f7c32f3e7a54446f7c61b5c

Observation 1fddd498-6867-41f5-8f01-7f2b81bebb31 · outbound

This paper cites Efficient inference for edge large language models,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Efficient inference for edge large language models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.282071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:13.909424Z digest=sha256:ac533b831de1e53c013160e8481d488c3db2bfc607967ebdd5dc8993b33de972

Observation 3573c1ed-d213-43fe-a8a0-7e6dd3caa118 · outbound

This paper cites Hqp: Sensitivity-aware hybrid quantization and pruning for ultra-low-latency edge ai inference,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Hqp: Sensitivity-aware hybrid quantization and pruning for ultra-low-latency edge ai inference,

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-11T20:11:15.277640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:13.917125Z digest=sha256:0bd26c0cf9ec7a6c2243fecffae4929a45cc7d5ff2df9d0e9c7c5d944fa9ab17

Observation 7256c3b9-6579-4967-b4c3-72bcde7efa2e · outbound

This paper cites Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:13.922408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:13.922408Z digest=sha256:001f55540f7ae36d816efd6a7afcfdb85bb7b5411827dcb7b0121c3099430101

Observation b6dab895-3ee1-4ebe-bd9e-cebfc9fc7376 · outbound

This paper cites A unified and resource-aware framework for adaptive inference optimization on edge devices,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge A unified and resource-aware framework for adaptive inference optimization on edge devices,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.250404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:13.928976Z digest=sha256:300876d145d6d3142aa98bb8fe4521acc2ca78183107720661cd5c409964709c

Observation 57d5f87a-7660-4883-bf2d-aff7ae0f1811 · outbound

This paper cites Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:11:15.103327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:13.934956Z digest=sha256:9eb12b8dae8325321c7e4bc3929e06b45b7e8b16d951ee163f3ec588bb0d35ed

Observation 48e8ba41-e9f3-4d7f-ad30-6138c2f9806e · outbound

This paper cites H2eal: Hybrid-bonding architecture with hybrid sparse attention for efficient long-context llm inference,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge H2eal: Hybrid-bonding architecture with hybrid sparse attention for efficient long-context llm inference,

Reference 11

Resolution
verified exact
raw_fallback, observed 2026-08-11T20:11:15.063507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:13.939638Z digest=sha256:91ae3929c40e227bc697d75d6c41a3debccf883be0e59b1ef73f3744504a74e0

Observation d735169a-fb09-4d58-b1c0-f61e8b9e3d60 · outbound

This paper cites Gptq: Accurate post-training quantization for generative pre-trained transformers,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Gptq: Accurate post-training quantization for generative pre-trained transformers,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.218896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:13.945458Z digest=sha256:5433e3187be3f4a286d812054237483b8b9650f183cf51c3ff92f0dc4ad620ed

Observation 778cd9c9-cea3-4f58-906c-e3619c3399f9 · outbound

This paper cites Awq: Activation-aware weight quanti- zation for on-device LLM compression and acceleration,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Awq: Activation-aware weight quanti- zation for on-device LLM compression and acceleration,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.188325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:13.950811Z digest=sha256:9969c4546c6e34a0038dcaa2c6a141fae9c8bc0acda9f5a4ef0c5a9a8a57a685

Observation 53f80e26-2ee7-465b-b94e-33f9963f40e2 · outbound

This paper cites Qserve: W4a8kv4 quantization and system co-design for efficient llm serving,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Qserve: W4a8kv4 quantization and system co-design for efficient llm serving,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:13.955411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:13.955411Z digest=sha256:0ea62fa3c9d41b59983af4406b9863c927000ef9aec08c93d563fdaab63392be

Observation d368b82d-868b-453d-8ba0-de167cbda201 · outbound

This paper cites FlatQuant: Flatness Matters for LLM Quantization.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge FlatQuant: Flatness Matters for LLM Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:13.960472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:13.960472Z digest=sha256:82c54b2df85f8a6334fb4ca616680fa9b10d3001f4dfd00e3085f21d469467c0

Observation f1c7cdd9-0473-491b-aedb-ad8728347b6c · outbound

This paper cites LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:13.966910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:13.966910Z digest=sha256:42712984ef2bb09b72da05bd4dc432c50ef2db61ffdbb971ad577682bad0c0f3

Observation 6cdfe297-8f34-48e2-950f-3814763d6db7 · outbound

This paper cites SmoothQuant: Accurate and efficient post-training quantization for large language models,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge SmoothQuant: Accurate and efficient post-training quantization for large language models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.144878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:13.973125Z digest=sha256:497c50296d5bfb2f882c02cba09724f8901c353a5d49e696564f7d117b092893

Observation 046f9a38-d0fb-48d5-ae75-62d1045143ce · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Sparsegpt: Massive language models can be accurately pruned in one-shot,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.122255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:13.978271Z digest=sha256:89e9681383d44e975baccaab17f47857fbf488eeea51ff134b2e1a8833934bfa

Observation bb98d5e5-50cf-4d10-a4c9-069f951d2b3c · outbound

This paper cites Unleashing network/accelerator co-exploration potential on fpgas: A deeper joint search,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Unleashing network/accelerator co-exploration potential on fpgas: A deeper joint search,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.098421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:13.982786Z digest=sha256:0944f02238a2947b31b4fc85957e8c3121b2e62dd27a654a8b16d9416702fdb4

Observation cf7b5df2-95e6-4630-af16-49bcb67304b7 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge PaLM-E: An Embodied Multimodal Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:13.988363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:13.988363Z digest=sha256:d7e6ade7cc02f7c4fc3e0b951da913849be57906bbc73e20ea992322e7ffe40e

Observation 8bb0eb05-3851-4281-a247-d7730da7223b · outbound

This paper cites Foundation models in autonomous driving: A survey on scenario generation and scenario analysis,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Foundation models in autonomous driving: A survey on scenario generation and scenario analysis,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:13.993348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:13.993348Z digest=sha256:a5b07e2708f887d08d55661e3f9a1c82b18e9a4584c1cb75eb2543a9715b992b

Observation 44a86edf-1e17-4d32-9cb3-94e35d5ba7b7 · outbound

This paper cites Intelligent Assistant Language Understanding On Device.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Intelligent Assistant Language Understanding On Device

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:11:14.819211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:13.998090Z digest=sha256:57e250c9477fc4eb60dad9981bf3a1f90ae6a831e11ff8d9b807938fad021d56

Observation 4f410ef4-782c-4eb8-9ae9-3e3918c0dd92 · outbound

This paper cites Private llm inference on consumer black- well gpus: A practical guide for cost-effective local deployment in smes,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Private llm inference on consumer black- well gpus: A practical guide for cost-effective local deployment in smes,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:14.004286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:14.004286Z digest=sha256:5d5a3f2a46c5785a24127588fc4d356e5d5dc7299de6b633a7780f69b5e6bcc4

Observation 5df39e1f-77ec-492d-913f-d281794a4cc8 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Efficient memory management for large language model serving with pagedattention,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:14.009653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:14.009653Z digest=sha256:049aade6ab71f94e4353fc9f0bc04f6f32c10c279519dcd7f9ef50e930347645

Observation 096f2b0e-e92e-48ec-9a70-4aef88b8f3e1 · outbound

This paper cites Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.036825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.014423Z digest=sha256:31c11b66037e4bde40f1b48289ba54fffa1fea83e5ffee706e44ca636d799963

Observation 64b24222-b831-4819-b475-270e1a2c6a8a · outbound

This paper cites Wrp: Weight recover prune for structured sparsity,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Wrp: Weight recover prune for structured sparsity,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:16.008394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.019429Z digest=sha256:34190670a0980c9ff700b7f8926a39c5a578e30640510a614c82ff47b1cc56e5

Observation 73ef787b-cd4f-4dcb-a3d3-5df2d51b054a · outbound

This paper cites Structured pruning for large language models using coupled components elimination and minor fine-tuning,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Structured pruning for large language models using coupled components elimination and minor fine-tuning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.978334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.025339Z digest=sha256:92c99b1e61e14941ebdb1236e973d60d08657c51fb01f4515982c9713cc5af87

Observation c213d2d9-454f-4641-86e3-b607f2ef8cec · outbound

This paper cites Unisparta: A unified sparse tensor program tuning framework,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Unisparta: A unified sparse tensor program tuning framework,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.957569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.031314Z digest=sha256:e9a9c070c4bfceb93fa274fa96f1eeadb8264aeb745fe030d4f68609a0fdc39c

Observation 9f400c61-4807-41f8-8509-6baea235b01b · outbound

This paper cites Closertome: A unified framework for accurate and transferable latency prediction across heterogeneous devices,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Closertome: A unified framework for accurate and transferable latency prediction across heterogeneous devices,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.934916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.037833Z digest=sha256:26f83fa5d22f3f186c60c69c7ce76e4620f2d307863d2d7866ffc241bff57e50

Observation 87a07782-4901-43b4-b6c3-1f6bdbf0b776 · outbound

This paper cites Sparse gpu kernels for deep learning,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Sparse gpu kernels for deep learning,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:14.042758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:14.042758Z digest=sha256:8faf881f1656171c93c4a863f8de0a9333d4c08bcd587b5b6781b938a4a67545

Observation 476a3036-6e4d-4ff3-8694-2888c44c5441 · outbound

This paper cites Dtc-spmm: Bridging the gap in accelerating general sparse matrix multiplication with tensor cores,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Dtc-spmm: Bridging the gap in accelerating general sparse matrix multiplication with tensor cores,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.899565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.049399Z digest=sha256:354d4777bde963d82cc7edb24ba369098b4c8d79489246c0274a3431730848ef

Observation 9dc8fdac-502c-4518-a87a-36225c41839f · outbound

This paper cites High Performance Unstructured SpMM Computation Using Tensor Cores.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge High Performance Unstructured SpMM Computation Using Tensor Cores

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:11:14.386302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.057480Z digest=sha256:6d8399837cae5e839dfc02dd5123e6df787e1eabe11c78a7e6f3d17d0da980e4

Observation d625940a-3aa5-4f92-93e4-2177901cead1 · outbound

This paper cites Acc-spmm: Accelerating general-purpose sparse matrix-matrix multiplication with gpu tensor cores,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Acc-spmm: Accelerating general-purpose sparse matrix-matrix multiplication with gpu tensor cores,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.879741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.063332Z digest=sha256:d596a94224432b591e158e4cf8c46627bc12d87248c6f74af8e5fdca64b45c5c

Observation 9a51e21c-3b40-411b-b54c-891c67399322 · outbound

This paper cites V oltrix: Sparse {Matrix-Matrix}multiplication on tensor cores with asynchronous and balanced kernel optimization,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge V oltrix: Sparse {Matrix-Matrix}multiplication on tensor cores with asynchronous and balanced kernel optimization,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.849359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.069653Z digest=sha256:3f997e38e6d3e763e7c4e735be4434e14bdbc13ddff0a7e1b996e17e265c8f19

Observation ca3a53d4-3eb1-4f08-b9c0-1b8529d384b0 · outbound

This paper cites {SparTA}:{Deep-Learning}model sparsity via{Tensor- with-Sparsity-Attribute},.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge {SparTA}:{Deep-Learning}model sparsity via{Tensor- with-Sparsity-Attribute},

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.819336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.075875Z digest=sha256:dc3eb84a6da4a0475c85f5f84d9ebf564d17fb4de31c6449cd10cd036d455ad7

Observation e644f87f-7eff-4d7d-83f3-95f372ccef9c · outbound

This paper cites cusparse library user guide,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge cusparse library user guide,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.796539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.081652Z digest=sha256:6856cbb9c273c2129fca9038b386ee8505398e022ca126e92dd9b3e122d2f108

Observation 685a4805-e82b-42ba-8f36-c622fb9910e5 · outbound

This paper cites cuSPARSELt,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge cuSPARSELt,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.768465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.086991Z digest=sha256:37acae0655c3c0dc4d74c61cb2351994d322698a07560e8f1d1db92e30d72b6e

Observation adbbfad2-d07e-4028-99f4-d68fe5a3c4eb · outbound

This paper cites Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:14.093933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:14.093933Z digest=sha256:82e6b640be06b54759a6e8c539981e08eb50047815cae63df3d2e839864b32b2

Observation 040defa0-ce08-4624-9015-0898cbafeab1 · outbound

This paper cites Spinfer: Leveraging low-level sparsity for efficient large language model inference on gpus,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Spinfer: Leveraging low-level sparsity for efficient large language model inference on gpus,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.739592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.101175Z digest=sha256:199777d85dc0f4d4ba9f834066fc4ed3ae066561e15dc81fba9b498cd3fd7328

Observation b9119f88-8de1-4c92-bbe8-22b3a84e8f60 · outbound

This paper cites Sputnik: Sparse matrix multiplication on gpus,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Sputnik: Sparse matrix multiplication on gpus,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.714739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.110607Z digest=sha256:331bb3f9c35176a2bb567ffeaded8f79b75e02cc2d56ff7047f5f4d7607d29a3

Observation 8fcc6f41-34f2-4be4-9336-81e88b79fe16 · outbound

This paper cites Nvidia jetson agx orin developer kit,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Nvidia jetson agx orin developer kit,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.684324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.116153Z digest=sha256:c2a0161322ea58f0fa1c717ae2128b5e5d8acd6c5d4fa7f4d92249f5f7f3960f

Observation 8156f3f9-4bd9-431b-86ac-7285f34dd375 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:14.123256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:14.123256Z digest=sha256:2be0175caa5fa7b461d7711a52506a69861516301176baab992f37a6c521c064

Observation 86034bed-f3e5-4c73-8d50-8c8053f20aab · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge OPT: Open Pre-trained Transformer Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:14.129460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:14.129460Z digest=sha256:96e08a1795ac53caaed7759a7d58e4a718f710ff3b6a04c5be65bf087e844833

Observation b6761266-fdb3-45b4-9124-e88596bb7f7f · outbound

This paper cites Qwen2 Technical Report.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Qwen2 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T20:11:14.135742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:11:14.135742Z digest=sha256:057c21903c820200b1881aa45f9c1f00601b741c541a4b2c53f4db701b2ba98f

Observation a1e28cf9-2e28-41fb-99b9-4fa3ed02c662 · outbound

This paper cites Llama 3: Open foundation and instruction-tuned models,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Llama 3: Open foundation and instruction-tuned models,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.654539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.141376Z digest=sha256:7390cf314ca5cdf57f5e1a1bfa9226c2cfbdb5228c5619c6f3b23fae3584d81c

Observation d504adc6-2e79-4e11-bd1f-0cd281c722d6 · outbound

This paper cites Cutlass: Cuda templates for linear algebra subroutines and solvers,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Cutlass: Cuda templates for linear algebra subroutines and solvers,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.629743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.148396Z digest=sha256:728846744f0069743274a6e9dc37ebb5c371f9e671ed0bfa9790188e25e638fe

Observation d85dbe4e-a52c-4760-8aff-b0bb09672f72 · outbound

This paper cites cublas library user guide,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge cublas library user guide,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.604215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.159695Z digest=sha256:d498fed369d300829d8ef7694afd90606f5708ee74fc5cd1b12159f3c598a994

Observation d43fd8cd-d093-4e7e-819c-db8ce32c10cb · outbound

This paper cites Fastertransformer,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Fastertransformer,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.576894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.169034Z digest=sha256:92800c91460fcc10905c48e70a6d68aca1b365ce2c993565e74882eb9bb92b74

Observation ffc38f0b-64fa-4337-8039-4a66f96141f2 · outbound

This paper cites A simple and effective pruning approach for large language models,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge A simple and effective pruning approach for large language models,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.554020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.177358Z digest=sha256:f4caa9775f33d966c730ee26daff5315b49f44fc79f0dc5f681c87057c558181

Observation 2e7f0333-732b-4d8f-8bc7-9e1ba16b237d · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate llm serving,.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge Atom: Low-bit quantization for efficient and accurate llm serving,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.532365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.184734Z digest=sha256:cf9fa7119c209387798cdbc555486bef75bca4d997b8d6dc955b76b7e25c038d

Observation bd7c9003-2240-4437-a743-b7f1649931cb · outbound

This paper cites He examines various aspects of embedded systems, with a focus on performance, availability, flexibility, and energy efficiency.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge He examines various aspects of embedded systems, with a focus on performance, availability, flexibility, and energy efficiency

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.488007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.200548Z digest=sha256:7d962c7ec210e4e4e4280ea6cf9e61ea9dbe887f631d7ec3d57b9fb86782d107

Observation 2f1cdba3-8dda-42f6-86bb-783c038b9f1a · outbound

This paper cites He is currently pursuing the Eng.D.

UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge He is currently pursuing the Eng.D

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:11:15.510650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:11:14.192355Z digest=sha256:08a9915707fb8f73ce7f426295ca00f747f6578dd5f350454c77d74284391933

Pith citing papers

No inbound Pith citation observations are available.