Pith. sign in

Paper Citation Record · LEDGER

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration

As of 8 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2508.19087.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.19087 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:03:38.453321Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact1
  • verified fuzzy49
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2c0249ee-648e-420a-b3af-852c7066a201 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:29.940257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:29.940257Z digest=sha256:bf09376a744832b5cbbdfa472cd45f370aae8602d5f2212282cd77fd9bba98a6

Observation 0699aa5c-7ee9-4474-839f-22467db1afb1 · outbound

This paper cites The Llama 3 Herd of Models.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration The Llama 3 Herd of Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:30.103343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:30.103343Z digest=sha256:7196fee73ceee3728f7aae230258946232274cff3afacf3d40fca57feb7a6031

Observation f1046ed7-e9ae-4b3d-8fc0-ad7eac3940fa · outbound

This paper cites DeepSeek-V3 Technical Report.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration DeepSeek-V3 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:30.247494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:30.247494Z digest=sha256:09978a7866b24d27ac3c5918fa169b78cde1a285091e3433afb49be54d980455

Observation 20b2debd-4947-433b-b75d-e7d68e19e71d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:30.430398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:30.430398Z digest=sha256:11215151a779ce6185488c7dcc4630e07dad123995c7f07a14cb0cb6428f84a9

Observation de9a7b54-2bdf-4117-ad18-fa77863d036e · outbound

This paper cites Jumping nlp curves: A review of natural language processing research,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Jumping nlp curves: A review of natural language processing research,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:52.699089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:30.620469Z digest=sha256:74c092e0496ddd18027dbd6bb0f89c48e8d39ece88d536b8a27b89b49c0707a4

Observation 21a7223a-8da9-4462-b25f-84ba27760508 · outbound

This paper cites Scaling Laws for Neural Language Models.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Scaling Laws for Neural Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:30.727386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:30.727386Z digest=sha256:ec08276ed9164a785d75889d03bcdd9e673b780bb1b925c375c4b5e8ceaf1d43

Observation 3a9c9fab-7b3c-42e3-a932-c722a16c76a3 · outbound

This paper cites Chateda: A large language model powered autonomous agent for eda,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Chateda: A large language model powered autonomous agent for eda,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:52.411263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:30.864703Z digest=sha256:5405a77236d83b565b98fb2d65fe87c9b62c6e3a0e1b15059678c007ff90d91d

Observation 97ec05cb-66d7-42cd-be3c-4883c458397a · outbound

This paper cites Dtatrans: Leveraging dynamic token-based quantization with accuracy compensation mechanism for efficient transformer archi- tecture,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Dtatrans: Leveraging dynamic token-based quantization with accuracy compensation mechanism for efficient transformer archi- tecture,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:52.188830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:30.971935Z digest=sha256:ae2878bb44e4db836c8b98eddc7f47aed624cfd9b25cb526e0dd81ec64882b7d

Observation 285e7a34-e16c-40b1-9d04-e41ec4daf6e7 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Qlora: Efficient finetuning of quantized llms,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:51.958128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:31.093906Z digest=sha256:281f52fecb43b55fd92205ed9df4b0fb52f41eec83c3e1011db785a3c15b70f3

Observation 2a09aa32-f0e9-43b3-892a-b2d5ffa42574 · outbound

This paper cites Optq: Accurate quantization for generative pre-trained transformers,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Optq: Accurate quantization for generative pre-trained transformers,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:51.665625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:31.247082Z digest=sha256:77055c6450befd17e23a90d0b09ccfeb1c3a5609768da41626f0c976e7cb71aa

Observation 53251243-f0e0-4ebd-ac71-deab2d1b995a · outbound

This paper cites A precision-scalable risc-v dnn processor with on- device learning capability at the extreme edge,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration A precision-scalable risc-v dnn processor with on- device learning capability at the extreme edge,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:51.518375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:31.362494Z digest=sha256:4577db77628c092512bc26ce1990c7c87ca18ec0e9ef564bf5a3e2df19ff3e1f

Observation e942cdf7-30a4-4ea2-8417-6779523368fe · outbound

This paper cites Token-scaled logit distillation for ternary weight gen- erative language models,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Token-scaled logit distillation for ternary weight gen- erative language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:51.335615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:31.478915Z digest=sha256:40df7a719f1389bbde4fb9840efeae0608a212e2d6da87c13ade7f3887ff58a6

Observation 7715ab1b-072b-4683-af15-09bc88d460a3 · outbound

This paper cites OneBit: Towards Extremely Low-bit Large Language Models.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration OneBit: Towards Extremely Low-bit Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:31.642699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:31.642699Z digest=sha256:0b635d25af6c8da6c304ce80375cafa2198c60d384b90fa5bad35cd8620c3ad9

Observation 34a9228e-f7c6-4f26-9ea9-ffa5b255dbc8 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Smoothquant: Accurate and efficient post-training quantization for large language models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:51.062987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:31.776902Z digest=sha256:7a0eee94806b409bd83ca2b726dc0322a915095f5b26cfbe154d6d479e68ed7c

Observation 2dbe31b4-d068-44b5-8d19-b14b892fff01 · outbound

This paper cites Omniquant: Omnidirectionally calibrated quantization for large language models,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Omniquant: Omnidirectionally calibrated quantization for large language models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:50.801835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:31.837782Z digest=sha256:627be1a0215063fcda7e2f1e8511607158dde213343cae49bb56b4a0fa9191b2

Observation e29de1d8-be67-4cc6-a866-00e42673bd77 · outbound

This paper cites Holes: Boosting large language models efficiency with hardware-friendly lossless encoding,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Holes: Boosting large language models efficiency with hardware-friendly lossless encoding,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:50.513827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:31.893683Z digest=sha256:f117dbd6719ed02852acaa19fb269af720c86bd78224829a917354b78dc7bf7e

Observation 81e59714-90dc-4fd5-b33f-69d73af0889c · outbound

This paper cites Quantization via distillation and contrastive learning,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Quantization via distillation and contrastive learning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:50.218760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:32.017450Z digest=sha256:9ba89fcb84555da3f5ec63b1b30d6cb373178e77d7861b1e07b03ea05efb00d8

Observation a347db84-0d13-4c1c-bfb1-733796f494e4 · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate llm serving,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Atom: Low-bit quantization for efficient and accurate llm serving,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:49.940044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:32.133751Z digest=sha256:93be741d2d70c3fd19850c2542872d7659c459191b5e34679d1948f5ec4cf03b

Observation 78e738c1-ebaa-4270-bd8e-f0f32a7b3542 · outbound

This paper cites Nvidia a100 tensor core gpu: Performance and innovation,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Nvidia a100 tensor core gpu: Performance and innovation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:49.644008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:32.231960Z digest=sha256:ac4891ac48823a8837f0075aaaf958488e402839845f7d13ffeeaf1817cc05c6

Observation 8b2eb790-87f0-4345-95e9-5dad8cc06a8c · outbound

This paper cites Rtx on—the nvidia turing gpu,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Rtx on—the nvidia turing gpu,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:49.365835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:32.311442Z digest=sha256:25a224131bf41b3f6c16098a2e42da4c2122e5eb8d3c0f0ecf83abc9f8a28703

Observation 61402a7a-7864-4591-806f-256035d61ed8 · outbound

This paper cites Dissecting tensor cores via microbenchmarks: Latency, throughput and numeric behaviors,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Dissecting tensor cores via microbenchmarks: Latency, throughput and numeric behaviors,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:49.070155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:32.488863Z digest=sha256:4878c92c19bfca5f02fd1b24c76ac5102779b8719a0281d70d6c76b3d8233053

Observation f6e2c64e-11de-4867-be05-59d2e3d3b060 · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:32.624554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:32.624554Z digest=sha256:434307fdb114ff32ac6ee30817b645b32991cb498edc05b3e2d5734d643be35d

Observation 02687371-529f-495d-bdff-b1680de5b07c · outbound

This paper cites G-blastn: accelerating nucleotide alignment by graphics processors,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration G-blastn: accelerating nucleotide alignment by graphics processors,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:48.703589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:32.713355Z digest=sha256:12f5c6f4471691a511293501b213e4e202c03880ca914e0aab46e0b614a1179c

Observation 972eb02e-45bf-4945-8ff3-c7891e9db510 · outbound

This paper cites Accelerating performance of gpu-based workloads using cxl,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Accelerating performance of gpu-based workloads using cxl,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:48.438897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:32.816745Z digest=sha256:52ccaa06528ad11046c31cf2acaf70d9c991f7e1658b22f6e3130d40b06db9e2

Observation 0d767571-b9e6-4997-b49f-ecf36bb09708 · outbound

This paper cites Superneurons: Dynamic gpu memory management for training deep neural networks,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Superneurons: Dynamic gpu memory management for training deep neural networks,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:48.156464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:32.903184Z digest=sha256:8733602e329b326dc27df5526a53507bfc1cc911cfe83d0e438544c375c47107

Observation 862a8792-8b7c-43a0-9836-96f2dbbd69e9 · outbound

This paper cites Gpt3. int8 (): 8-bit matrix multiplication for trans- formers at scale,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Gpt3. int8 (): 8-bit matrix multiplication for trans- formers at scale,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:47.803596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:33.026595Z digest=sha256:bf143c1322b67b8fdc22f2635e10593afacfde3541c24472399ec8d7b84f7e87

Observation a7457989-dce0-4e7f-9a58-7ff4d12c6b58 · outbound

This paper cites Quant-llm: Accelerating the serving of large language models via fp6-centric algorithm-system co-design on modern gpus,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Quant-llm: Accelerating the serving of large language models via fp6-centric algorithm-system co-design on modern gpus,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:47.454561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:33.175416Z digest=sha256:497e95b628a738307bf2cd2e854a0b692d84d901dca7ac6fd6cd07ac77474bb4

Observation caa6ff66-940d-461e-992b-d09fa11ae585 · outbound

This paper cites Tsm2x: High-performance tall-and-skinny matrix– matrix multiplication on gpus,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Tsm2x: High-performance tall-and-skinny matrix– matrix multiplication on gpus,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:47.084547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:33.298313Z digest=sha256:e91fc04c940f03549c74cdb8bf9dcfb7f1afe875cab0eabb967ae5d0bc8d7a8f

Observation 73a8173f-ae2d-4e27-a9a0-f835c17507ed · outbound

This paper cites Stream-k: Work-centric parallel decomposition for dense matrix-matrix multiplication on the gpu,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Stream-k: Work-centric parallel decomposition for dense matrix-matrix multiplication on the gpu,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:46.669641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:33.396338Z digest=sha256:abe17401fcbfa71e7c18f630ba787ac381de3d776aefe4ecef8d4efbdff2011e

Observation efedc03e-0a91-4e01-9a05-fab1cfa26ab3 · outbound

This paper cites Warp-aware adaptive energy efficiency calibration for multi-gpu systems,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Warp-aware adaptive energy efficiency calibration for multi-gpu systems,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:46.295684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:33.509043Z digest=sha256:ebdc620376c2189d0bef66eb30679e5a2a13acf3fb914236bda44fb08bc13cdc

Observation c55015c0-a37a-4ba3-991e-9e6c536392ab · outbound

This paper cites Dg-replace: A dataflow-driven gpu-accelerated an- alytical global placement framework for machine learning accelerators,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Dg-replace: A dataflow-driven gpu-accelerated an- alytical global placement framework for machine learning accelerators,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:45.883712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:33.711616Z digest=sha256:20e25417681706d9ef7cf5c26c1178c15464b484aca927e5f89669ffff40bbba

Observation 6a408321-96c9-44fb-8927-41f1b0f5f099 · outbound

This paper cites Enabling efficient sparse multiplications on gpus with heuristic adaptability,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Enabling efficient sparse multiplications on gpus with heuristic adaptability,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:45.519740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:33.821701Z digest=sha256:4ee735e1be2b817c3efa6d7e46a57c6dd069103ea0f0f117438b2d14e7050859

Observation f204c8cb-b910-4361-9f64-5ac40004a4c5 · outbound

This paper cites Bstc: A novel binarized-soft-tensor-core design for acceler- ating bit-based approximated neural nets,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Bstc: A novel binarized-soft-tensor-core design for acceler- ating bit-based approximated neural nets,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:45.169637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:33.927448Z digest=sha256:ed7272fb1b518c2f6ff514579ad078398e6772d31bd3b68760e669ade6d73aa8

Observation 1f0e1a3d-b951-4eec-9265-5029184e20fe · outbound

This paper cites Accelerating binarized neural networks via bit-tensor-cores in turing gpus,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Accelerating binarized neural networks via bit-tensor-cores in turing gpus,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:44.917241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:34.072475Z digest=sha256:8442658588baa0cbb76a83595d07101ddcfc96515b65012d561c267c160e86de

Observation a7657730-372d-442d-9efe-85a35fbaa7d5 · outbound

This paper cites Demystifying the nvidia ampere architecture through microbenchmarking and instruction-level analysis,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Demystifying the nvidia ampere architecture through microbenchmarking and instruction-level analysis,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:44.637710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:34.214076Z digest=sha256:5e565ed0e4ef9c9b9dc8688969665dd366b2db022a8018759cff52eba3cd73f0

Observation ddd27145-4aa4-4b4e-bd39-8a1262afe34b · outbound

This paper cites Dissecting the NVidia Turing T4 GPU via Microbenchmarking.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Dissecting the NVidia Turing T4 GPU via Microbenchmarking

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:34.362212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:34.362212Z digest=sha256:072fb4d3c96bf5bf480741ef16cdaca01bbb7fdd4b77da90f5d19bb777778a5e

Observation c1cfc1ed-d322-48d7-9a5a-b4352db32e67 · outbound

This paper cites Gtco: Graph and tensor co-design for transformer-based image recognition on tensor cores,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Gtco: Graph and tensor co-design for transformer-based image recognition on tensor cores,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:44.121829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:34.545982Z digest=sha256:fa08c6d20a8b197395e6ca54ac7a78dcc13f48d7ecea0bfe0bfec28b62ba6261

Observation 10a85451-0f7f-44a6-b720-0855114b9c2a · outbound

This paper cites Reducing shared memory footprint to leverage high throughput on tensor cores and its flexible api extension library,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Reducing shared memory footprint to leverage high throughput on tensor cores and its flexible api extension library,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:43.817709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:34.626516Z digest=sha256:4ea29759ef46258ce77fd77c5c51ecbd6208b82304f8d98d6a434e449d455fd1

Observation d3ba3b1d-64d2-4aac-80b0-1c17207ef002 · outbound

This paper cites Tc-gnn: Bridging sparse gnn computation and dense tensor cores on gpus,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Tc-gnn: Bridging sparse gnn computation and dense tensor cores on gpus,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:43.542363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:34.781504Z digest=sha256:3bf0784b8129b2fda806473b46f5268a12fc1c6927d6733708784bf0ee22dd45

Observation bc41d50c-520f-4568-b483-fee0d23a90d6 · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration A Survey on Efficient Inference for Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:34.877201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:34.877201Z digest=sha256:db106e9630d9ba279cafa772e2cba1c140355c9c3c9bf503f2aea1ffb73a9349

Observation 936e10d2-b8f9-432d-b9d3-406727c0d6b3 · outbound

This paper cites Transformer tricks: Precomputing the first layer.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Transformer tricks: Precomputing the first layer

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:35.015302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:35.015302Z digest=sha256:7fd7183ce1fb41d6882b11d31b700518b760645a8373dcee0cf3ea9633782c39

Observation 20028b8c-8e9b-4be8-9aad-7f086a53eca0 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:35.155547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:35.155547Z digest=sha256:cf9d58b9b60e4971a360706e5ea27b2e2f1d8d5256d6f1414aac4635077a154d

Observation 0384a01f-6700-410c-b846-6077e1f05f29 · outbound

This paper cites Efficiently scaling transformer inference,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Efficiently scaling transformer inference,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:43.271413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:35.276363Z digest=sha256:a526290f9f0619c96afc8d8ee74bdbcfaab2b78a89d88f98e5bdf76a812a8b65

Observation 50dbda44-73ad-4c8c-b464-a047058de65d · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:35.382312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:35.382312Z digest=sha256:dc556a236f8c0ababf22751f8bcc4135af7986c88d05149634586d34a7008f06

Observation 23f16808-6923-48eb-9bd2-dbfd81312a1f · outbound

This paper cites Squeezellm: Dense-and-sparse quantization,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Squeezellm: Dense-and-sparse quantization,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:42.979238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:35.525034Z digest=sha256:6b7e84911a7651b770727afd730158e4fff5a97874d1476fa448fd3823f0eeea

Observation 602f8122-5394-423d-9363-362753a53a0f · outbound

This paper cites Quantsr: accurate low-bit quantization for efficient image super-resolution,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Quantsr: accurate low-bit quantization for efficient image super-resolution,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:42.735145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:35.654288Z digest=sha256:74326c74857c5bca54dcdaec21ad111530cf646e97f7b9381c9a7dc93c28cb7d

Observation 8c9c4e4d-e22e-4f6a-a269-72f65297aff9 · outbound

This paper cites Accurate lora-finetuning quantization of llms via infor- mation retention,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Accurate lora-finetuning quantization of llms via infor- mation retention,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:42.456579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:35.740233Z digest=sha256:051eda60453de299e2c3dd720876bb6dc02f0c69697af048724851b61177941c

Observation aa6adcfe-69e2-4d00-bdd2-aa522c72ac1a · outbound

This paper cites Bimatting: Efficient video matting via binarization,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Bimatting: Efficient video matting via binarization,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:42.173545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:35.795656Z digest=sha256:fbfd27436aaea1cde397c86e29d3fe43ebfabfd717c847a000b3ab7b559222ed

Observation dc0e5c26-be37-45c6-a541-59cffa191625 · outbound

This paper cites Bibert: Accurate fully binarized bert,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Bibert: Accurate fully binarized bert,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:41.854183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:35.926778Z digest=sha256:bf94f98bc34a81b36ce7e6b2cff9040284c6771d84c6ce25df892e66b781be68

Observation 6b77b327-62a3-4095-bb3d-bf53d3e02ad0 · outbound

This paper cites Bebert: Efficient and robust binary ensemble bert,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Bebert: Efficient and robust binary ensemble bert,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:41.556430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:36.063989Z digest=sha256:20e324a5606a77acaddc588343c6ee4380fc88d63b75a43b35ac7695a7b0dc52

Observation febe2237-c4f7-4ba0-bdb7-eb4f062c9e7f · outbound

This paper cites Anda: Unlocking efficient llm inference with a variable- length grouped activation data format,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Anda: Unlocking efficient llm inference with a variable- length grouped activation data format,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:41.285241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:36.192211Z digest=sha256:cd22cacb43d7f084161642fa0124cdf54892edd76ed166e13cd1c7fec9e6e65d

Observation e934cee2-b33f-4156-9766-e77fc31f79e5 · outbound

This paper cites Binaryconnect: Training deep neural networks with binary weights during propagations,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Binaryconnect: Training deep neural networks with binary weights during propagations,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:41.038528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:36.289139Z digest=sha256:09fa9e7af3601d5d9168a175e9306d2ba28a47f10d20568eec97890be0757f87

Observation 32badd89-5357-4eb1-bf96-f3abad82b614 · outbound

This paper cites Apnn-tc: Accelerating arbitrary precision neural net- works on ampere gpu tensor cores,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Apnn-tc: Accelerating arbitrary precision neural net- works on ampere gpu tensor cores,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:40.746968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:36.427395Z digest=sha256:bdb5d97d5c1155d13e0e2bb133c2d07af7a39ccc2687ae06fad157868f863899

Observation 85c7506c-6611-476b-bd56-2eafefebde65 · outbound

This paper cites O3bnn: An out-of-order architecture for high- performance binarized neural network inference with fine-grained prun- ing,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration O3bnn: An out-of-order architecture for high- performance binarized neural network inference with fine-grained prun- ing,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:40.550324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:36.587831Z digest=sha256:4c703c0189d959e901db2439ef2c1dba5b6fb40010de8b99a2e73a2d271241ac

Observation 2e159bb0-57f2-4c08-8b4a-a89501d8ca9f · outbound

This paper cites O3bnn-r: An out-of-order architecture for high- performance and regularized bnn inference,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration O3bnn-r: An out-of-order architecture for high- performance and regularized bnn inference,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:40.290730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:36.716924Z digest=sha256:00b121e40fc627136d51e82b7374955604cb02569e7f8f3b7f26dd110a5d9e29

Observation a82f4c78-4e48-4753-9e7d-fca5a7664637 · outbound

This paper cites Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:36.820892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:36.820892Z digest=sha256:aef53186c48ecc9c88001732e3829344e69746c707897d427ff960953d8d56f0

Observation d6f7676a-04bf-4c3b-9020-9fc997bb10ac · outbound

This paper cites Energy-efficient neural network accelerator based on outlier-aware low-precision computation,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Energy-efficient neural network accelerator based on outlier-aware low-precision computation,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:40.054971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:36.973470Z digest=sha256:e3c1003cfdcd62696abdeb1be0ed9aa9165e52ee1e441db3bc4373d39b7bbffc

Observation beb297e3-6573-4407-8cf0-a383adc8e25b · outbound

This paper cites Haq: Hardware-aware automated quantization with mixed precision,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Haq: Hardware-aware automated quantization with mixed precision,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:39.822948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:37.059456Z digest=sha256:d96eee243c679efc89e9b17646b4d4283218108d913c146d266a9b40cd4daccb

Observation 2f35f174-06ed-4ec5-90b4-044f816a4ede · outbound

This paper cites Lq-nets: Learned quantization for highly accurate and compact deep neural networks,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Lq-nets: Learned quantization for highly accurate and compact deep neural networks,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:39.574486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:37.163159Z digest=sha256:74863c51527899dad150ece0314ca1ee0711bcc856f191b23eb982562399e222

Observation ed54c9ca-a242-4b44-ad33-9e83c77fd0f8 · outbound

This paper cites DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:37.317001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:37.317001Z digest=sha256:454941f769a0be3ee090cd9a0982426e7eae6f0655f73f07b5b66cb24c0ccb3a

Observation 64b701fa-203f-4860-86af-cb433ba3e196 · outbound

This paper cites BitNet: Scaling 1-bit Transformers for Large Language Models.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration BitNet: Scaling 1-bit Transformers for Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:37.480444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:37.480444Z digest=sha256:7ce2788b22420cf4e222bb194a9f6ed16d65653d2240afbb726714c2aadc568a

Observation c4a7e274-a1e0-4630-8f06-594cc0e859fa · outbound

This paper cites CUTLASS,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration CUTLASS,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:39.296976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:37.615424Z digest=sha256:f7c271cadc1a36b4ac95b2f2acfe471c0582a16d20060d1bb47ed88fd48988e4

Observation 8b8e1fab-3214-4f98-ab05-99f3338c3ecd · outbound

This paper cites Understanding GEMM Performance and Energy on NVIDIA Ada Lovelace: A Machine Learning-Based Analytical Approach.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Understanding GEMM Performance and Energy on NVIDIA Ada Lovelace: A Machine Learning-Based Analytical Approach

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-08-05T16:03:38.709223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:37.761446Z digest=sha256:fce5d8dadd27baf459cafecd01a9fac5506930815e627edac751525273252bba

Observation 647c40c9-853c-4d34-83ae-be0dbcf3c981 · outbound

This paper cites Benchmarking and dissecting the nvidia hopper gpu architecture,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Benchmarking and dissecting the nvidia hopper gpu architecture,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:44.427189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:37.863736Z digest=sha256:6ed3d24f5a38fa1b00702cb4ee359e465ddf10d7fa84329efcf2e27b990ff885

Observation 218ea8a5-470c-47c8-9151-ce529186440c · outbound

This paper cites Qwen2.5 Technical Report.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Qwen2.5 Technical Report

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:37.961793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:37.961793Z digest=sha256:a1c41d134243134fc7d50800bac3df0ed89717d635f877e6cb8f11d87c82fbe5

Observation 23525f4e-8c54-465c-9ef4-0e17cad9dccd · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration OPT: Open Pre-trained Transformer Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:38.057130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:38.057130Z digest=sha256:f5635a578159f62b898cb1a2f233c8fbeaa5d54591452fe8f01b503f113e5b04

Observation 372fa1e2-d486-4352-ad8f-e8c613a2b8e1 · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:38.218549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:38.218549Z digest=sha256:25b227a79313e661bf05c70a76cc1d4f156332fb93a4d3085ece01756ff2b76a

Observation 831a1e11-38ba-443b-8d03-a72fe0122d8c · outbound

This paper cites Quarot: Outlier-free 4-bit inference in rotated llms,.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Quarot: Outlier-free 4-bit inference in rotated llms,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:03:39.050426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:03:38.331692Z digest=sha256:3683c26373b467122d5ba9d606a4016a99be1ce7265f690ec30ded44fe644eb6

Observation 90c16b91-f1f3-4421-860f-f5104b84fcdd · outbound

This paper cites Pointer Sentinel Mixture Models.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration Pointer Sentinel Mixture Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:38.453321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:38.453321Z digest=sha256:73f6035e61b83203104c373691028d4c9748c47deb52bf9a7708174d95e65e9a

Pith citing papers

No inbound Pith citation observations are available.