Pith. sign in

Paper Citation Record · LEDGER

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format

As of 16 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2411.15982.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15982 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:46:01.321581Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

89 of 89 outbound references displayed

  • verified exact1
  • verified fuzzy65
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eb908304-59b4-4082-8b4e-8ca13af78c73 · outbound

This paper cites Resq: Residual quantization for video perception,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Resq: Residual quantization for video perception,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.916127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.916127Z digest=sha256:5f65c216774b8099e5c336e92b5a1383fe57f529df3f25a026caebbccab24cc8

Observation 07b6e85e-10eb-4742-ad69-eb44c17f2a83 · outbound

This paper cites Bit-pragmatic deep neural network computing,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bit-pragmatic deep neural network computing,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.924633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.924633Z digest=sha256:86044a8c676d94bfd0793f24cffdd17d48c9c83f29a58b2c6cffe3b879fad3bc

Observation 109156cc-ec88-4c3e-b740-788a6f30fc11 · outbound

This paper cites Explaining neural scaling laws,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Explaining neural scaling laws,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.928808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.928808Z digest=sha256:f0bb1bf2da89b1e001693b6b4463fa0dd2380f0b2dee81e5e3b6fba0678e24cc

Observation f22e992c-6df8-4344-b35e-d24c47398f35 · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Longbench: A bilingual, multitask benchmark for long context understanding,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.933063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.933063Z digest=sha256:b697cbc468926a910a46f3acf0bde52374a88969a90ed0c6cf898db24e086e8f

Observation c2fbf310-6db7-48f4-b945-3546b592e458 · outbound

This paper cites Demystifying chatgpt: An in-depth survey of openai’s robust large language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Demystifying chatgpt: An in-depth survey of openai’s robust large language models,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.937312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.937312Z digest=sha256:2c9a33ad3978f88a6b684bc00654f93a09b06f72ec403ba92a79493e33599922

Observation 512cd457-8d26-4023-b16d-7a04113f5a87 · outbound

This paper cites Genus synthesis solution,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Genus synthesis solution,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.941946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.941946Z digest=sha256:ef43344244feae3caa7103770eb8b7c874d0b3188ae573fff4b9f0d196cebab8

Observation 5c53dec2-f8da-4aba-a7c7-4e8f421615c9 · outbound

This paper cites General purpose deep learning accelerator based on bit interleaving,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format General purpose deep learning accelerator based on bit interleaving,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.946229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.946229Z digest=sha256:bbcf15bfa3813295791d616b5f0b729619651ac349ef19ba9dd6b29ba001f5b6

Observation f864ed24-45c9-401f-ae54-5c8143377e09 · outbound

This paper cites Quip: 2-bit quanti- zation of large language models with guarantees,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Quip: 2-bit quanti- zation of large language models with guarantees,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.950195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.950195Z digest=sha256:f1aca7a1c62fd625beb723a608be22a255e9ef64c07c1c3bc4fb8dbd08ca7031

Observation db8e462a-25cb-46fb-b2b5-1b5a23573c3e · outbound

This paper cites EfficientQAT: Efficient Quantization-Aware Training for Large Language Models.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format EfficientQAT: Efficient Quantization-Aware Training for Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.954375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.954375Z digest=sha256:3dc398679bc6549fba9cbf02bcbb1d86765a359a9af9cf763d3ca8adec10e95a

Observation c08c524f-0c14-4d61-b5d6-f0409de0d6e1 · outbound

This paper cites Nacl: A general and effective kv cache eviction framework for llm at inference time,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Nacl: A general and effective kv cache eviction framework for llm at inference time,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.959299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.959299Z digest=sha256:3c4ba2d1ebbdca4457ef8db19d8ffacd0a329f919f0a9f943a19fd876f64301b

Observation 3e140eb0-9c01-4fba-b3db-04b06150bd15 · outbound

This paper cites Palm: Scaling language modeling with pathways,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Palm: Scaling language modeling with pathways,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.688552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:00.963367Z digest=sha256:dbc076cedc1b62f673794c6dfcc11f67c206115a885d3ae08c5eb7a1b09e167c

Observation 0e391695-d4ee-48c0-a897-a30f0f232543 · outbound

This paper cites Vs-quant: Per-vector scaled quantization for accurate low-precision neural network inference,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Vs-quant: Per-vector scaled quantization for accurate low-precision neural network inference,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.673582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:00.967701Z digest=sha256:74cd03fc7e026f9f15bc323dedf85af7e130561c84fa7b4f8a28193daab91a57

Observation cbfde174-f710-43b3-bc1e-c0588e604377 · outbound

This paper cites Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.660384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:00.971544Z digest=sha256:e481f0e29eb91e474b27fb1d107ad257d2bfcd04f1877e221139bb3612f18b6c

Observation 987c2874-d2d8-4184-a3d2-f795bbdf96e0 · outbound

This paper cites With shared microexponents, a little shifting goes a long way,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format With shared microexponents, a little shifting goes a long way,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.647418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:00.976155Z digest=sha256:6e96aa2bf1be39e859ecd75e0001146f99e4cde9682324f8c1abf3fbaa9f0962

Observation d0088e03-3677-4f8a-8050-e4d1ccb8efc8 · outbound

This paper cites A timing-driven approach to synthesize fast barrel shifters,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format A timing-driven approach to synthesize fast barrel shifters,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.635399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:00.980180Z digest=sha256:4446159e4c4fdcb010206815f5c2a67a24425abff0c314c88548bacc850a84ee

Observation ec788826-bdd6-4922-b1b5-c725344a1534 · outbound

This paper cites Llm.int8(): 8- bit matrix multiplication for transformers at scale,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Llm.int8(): 8- bit matrix multiplication for transformers at scale,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.622334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:00.984391Z digest=sha256:3ce0cc9e46cf672434f91a523c35f56dddc3c3c53ae55193b5eca352b7040db5

Observation 8480b9a4-4201-4813-a0d8-68b85334e80b · outbound

This paper cites The case for 4-bit precision: k- bit inference scaling laws,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format The case for 4-bit precision: k- bit inference scaling laws,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.606453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:00.988162Z digest=sha256:6f4a88555a1ede99ebd63651e7be10fd0d75d9cb4410e345aacc3f5b4c655ba6

Observation 00ecbf6f-8afa-4325-9fb0-89afa02558c9 · outbound

This paper cites Hawq: Hessian aware quantization of neural networks with mixed-precision,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Hawq: Hessian aware quantization of neural networks with mixed-precision,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.588365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:00.992180Z digest=sha256:c92992f347b52c9b0d669f5d25e29daadf93fe45ae3be560d8b069d42efbee5f

Observation e6510551-5932-40cc-b624-1173024cb859 · outbound

This paper cites Training dnns with hybrid block floating point,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Training dnns with hybrid block floating point,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.571178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:00.996132Z digest=sha256:f0bcf8532afc04c7ccea372b76bba64356e70630d7fa79d8926632d17043c510

Observation dce06533-ae2f-40b8-af4b-eab6c87e6b06 · outbound

This paper cites Skvq: Sliding-window key and value cache quantization for large language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Skvq: Sliding-window key and value cache quantization for large language models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.557598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.000362Z digest=sha256:75d2c6109c9d046ff5cb6f8ac935667720e810d6d7d555b491eec0f860e9e886

Observation 178aaf15-301b-40f6-8b13-bc2fbe940d64 · outbound

This paper cites Extreme compression of large language models via additive quantization,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Extreme compression of large language models via additive quantization,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.545100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.004503Z digest=sha256:210d7ab2edb35bd7f3ed1a56aa57d84579b5078a799d3dc1350327b35bd66834

Observation f389ac12-1c55-47f5-80ad-99753d11967f · outbound

This paper cites Reconfig- urable acceleration of 3d-cnns for human action recognition with block floating-point representation,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Reconfig- urable acceleration of 3d-cnns for human action recognition with block floating-point representation,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.532602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.008439Z digest=sha256:eb7acf1638ac4116c1fa117c31e3292a767cafd1ac864a606b9885edce0098a2

Observation 01002d37-24f7-4610-8d4a-c2d9c229c290 · outbound

This paper cites Static block floating-point quantization for convolutional neural networks on fpga,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Static block floating-point quantization for convolutional neural networks on fpga,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.520801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.012376Z digest=sha256:1bf1665d2e8ec09533ba5f6943a6487316e227676b0e05d15c950d77a6f2a4c1

Observation 553aa229-a25a-4769-a67b-c312da294cc8 · outbound

This paper cites Optq: Accurate quantization for generative pre-trained transformers,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Optq: Accurate quantization for generative pre-trained transformers,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.506326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.016683Z digest=sha256:8c8eb18c67021e47685afd6f8366b5268026e82c55341bd30e62068fd89db054

Observation be801d61-b0fc-4bb3-a27d-e66b430b3b5c · outbound

This paper cites LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.020686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.020686Z digest=sha256:4e31b3f2c3296394508fb2c1a38f27ea95de5b91322aafaaca077b20772ab494

Observation 6fcd9706-878e-4c1c-a67b-2c44d6dc4687 · outbound

This paper cites Boost: block minifloat-based on-device cnn training accelerator with transfer learning,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Boost: block minifloat-based on-device cnn training accelerator with transfer learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.488289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.025336Z digest=sha256:6196614f064580dd4f8467239cac4e195f2c842169d6acbb32fe8c079e1d9f72

Observation 2efccf6d-a519-4127-b9b4-af72090c2eac · outbound

This paper cites Olive: Accelerating large language models via hardware- friendly outlier-victim pair quantization,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Olive: Accelerating large language models via hardware- friendly outlier-victim pair quantization,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.472758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.029221Z digest=sha256:911ee00f0c38be39abdadffcb76cbb6b880400ebdc3224ab65f6f9f8a7b78eb2

Observation 734655d4-95bf-4ea4-b953-1923a5d8a9fe · outbound

This paper cites Ese: Efficient speech recognition engine with sparse lstm on fpga,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Ese: Efficient speech recognition engine with sparse lstm on fpga,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.459389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.033280Z digest=sha256:a3203e64a619cc08452ec55e5290ef46127cf154f831f61be1728f29ae174382

Observation e31e43a2-f16c-464b-a096-57ffd31fb592 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.037498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.037498Z digest=sha256:c08479929cff05c34c881fc5c8dcebf848b0e0c5a5a1804299306482631bc7fc

Observation acb4a88b-71fd-425a-a4d4-7557df0a3b61 · outbound

This paper cites A precision-scalable risc-v dnn processor with on-device learning capability at the extreme edge,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format A precision-scalable risc-v dnn processor with on-device learning capability at the extreme edge,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.445583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.042306Z digest=sha256:c5abd6688ec410b954bfd386b0042ce4f0623ff11d5d993bc5fb00f159ece07a

Observation 3c4a177a-2705-4842-99c8-b7d2c8a2e8a7 · outbound

This paper cites Mind the gap: Attainable data movement and operational intensity bounds for tensor algorithms,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Mind the gap: Attainable data movement and operational intensity bounds for tensor algorithms,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.429805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.046677Z digest=sha256:e888d7bfdb707b1bb7d8c0e502d43561e610e7bba1debee8177f201bad5ae1f9

Observation 9feedc6f-68ee-45ee-b614-0cd0da64d635 · outbound

This paper cites Figna: Integer unit-based accel- erator design for fp-int gemm preserving numerical accuracy,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Figna: Integer unit-based accel- erator design for fp-int gemm preserving numerical accuracy,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.410522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.050530Z digest=sha256:0c6cce2f24c181de3a68d9d03b5ecf7cce9865499e3dab56dbed264f1d89c6e7

Observation b971c7e3-66b4-44ab-9994-26903bdfdec2 · outbound

This paper cites Perplexity—a measure of the difficulty of speech recognition tasks,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Perplexity—a measure of the difficulty of speech recognition tasks,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.392252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.054333Z digest=sha256:aeeae73a561883fcc46d236359fa5c5e89a636fb61e0a7df2f38400e64fc7a5f

Observation cca7ac48-4bed-454b-91a7-e725e24eecce · outbound

This paper cites Mr. biq: Post-training non- uniform quantization based on minimizing the reconstruction error,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Mr. biq: Post-training non- uniform quantization based on minimizing the reconstruction error,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.371118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.060235Z digest=sha256:1a0cb7e75b73de29c6a7f995f93257a6c92c8730fe83ca4a9512e1dc705051f2

Observation 5a868e5d-84ae-43b1-a764-91759d50d4b7 · outbound

This paper cites Biqgemm: matrix multiplication with lookup table for binary-coding-based quan- tized dnns,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Biqgemm: matrix multiplication with lookup table for binary-coding-based quan- tized dnns,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.356272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.064416Z digest=sha256:233124a272900fa6fb798c66657d05920283c9eb60f7525a7fbb038d58fe690f

Observation 28571475-c8a3-430d-97f8-fb952857a7fa · outbound

This paper cites Ten lessons from three generations shaped google’s tpuv4i: Industrial product,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Ten lessons from three generations shaped google’s tpuv4i: Industrial product,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.341385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.069776Z digest=sha256:8f12adedf201964a35063eb169b3ac07d0e936bb4c48ef957aa3443b926fbc7e

Observation fe860abe-bc83-47e5-b316-3a520278d183 · outbound

This paper cites Stripes: Bit-serial deep neural network computing,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Stripes: Bit-serial deep neural network computing,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.326657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.073787Z digest=sha256:dda0579c506b765dec739a8a53a6ae2397c248370de4e6c1f118a76b3a691e1d

Observation 94942890-0859-4f55-b5a3-4a08a3024924 · outbound

This paper cites A survey of gpt-3 family large language models including chatgpt and gpt-4,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format A survey of gpt-3 family large language models including chatgpt and gpt-4,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.078272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.078272Z digest=sha256:141fa133a771344ce374177dc9008c382c1f2642a003ec45603d0190951f80a8

Observation 02267acb-faf9-41a0-8f47-18c580f1c56c · outbound

This paper cites A 95.6-tops/w deep learning inference accelerator with per-vector scaled 4-bit quantization in 5 nm,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format A 95.6-tops/w deep learning inference accelerator with per-vector scaled 4-bit quantization in 5 nm,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.304176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.082459Z digest=sha256:34de1ad3d98da272ddbd5de2fafa2ff14a34b90075314fb61e1d3794b6d90b63

Observation e01546cb-3e38-4a57-a7db-f2803705ecdf · outbound

This paper cites Compressed context mem- ory for online language model interaction,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Compressed context mem- ory for online language model interaction,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.288380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.086446Z digest=sha256:315a647229abf0551b845f4bedc268ce66ffee5363ca07a1288e3d7e930f1181

Observation ba442fe0-bdc1-4a37-be6d-7f271a5796e7 · outbound

This paper cites Dacapo: Accelerating continuous learning in autonomous systems for video analytics,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Dacapo: Accelerating continuous learning in autonomous systems for video analytics,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.273986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.090475Z digest=sha256:f948127a03ef4721807a7c7599dba5d25563ed76f5ab1ddbd80b09b9adc1f8af

Observation 048c2a80-8910-4127-b228-ffba59500712 · outbound

This paper cites Winning both the accuracy of floating point activation and the simplicity of integer arithmetic,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Winning both the accuracy of floating point activation and the simplicity of integer arithmetic,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.260399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.096261Z digest=sha256:4ba0c7120448ad589a8e5b0b75f5f6283d0bb81e630ef269a123e165eee1188d

Observation 05f82623-88e2-48b1-9229-de66136b466e · outbound

This paper cites One-shot model for mixed-precision quantization,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format One-shot model for mixed-precision quantization,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.244419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.101397Z digest=sha256:7201eea8e2bcac8fe6e10dbfea22b754f8287e2aa315b093d79c74e3f542dd7f

Observation 6de4f875-f11c-4ed3-9f6a-a87b72e63cee · outbound

This paper cites Flexpoint: An adaptive numerical format for efficient training of deep neural networks,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Flexpoint: An adaptive numerical format for efficient training of deep neural networks,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.231230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.106244Z digest=sha256:089455fafb333ce91caea89304c73cc9f8ca1bf5f29be110174de2c1fe52c801

Observation 6ff985f6-a23d-486b-8f4d-8d73abf68337 · outbound

This paper cites Tender: Accelerating large language models via tensor decomposition and runtime requantization,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Tender: Accelerating large language models via tensor decomposition and runtime requantization,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.218324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.110394Z digest=sha256:70aecfe609323b188fa648b9192f0c6d88333718a47a8d7aeab83aab327fe7c5

Observation 9c0e6188-8cdf-4159-85a6-5069701f9205 · outbound

This paper cites Bitcluster: Fine-grained weight quantization for load-balanced bit-serial neural network accelerators,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bitcluster: Fine-grained weight quantization for load-balanced bit-serial neural network accelerators,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.206101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.114610Z digest=sha256:0b46a9c21a5507964bddce540aa14f348576ce3fdec2b5d4d96e63ae16b5b1df

Observation 19bd6321-1c54-4f04-8dc9-116e8de5dff2 · outbound

This paper cites Norm tweaking: High-performance low-bit quantization of large language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Norm tweaking: High-performance low-bit quantization of large language models,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.190206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.118186Z digest=sha256:cc3ecd1062f4c2aa0cad96057a7565f95437da0b4bb3902910fc12414babc77f

Observation e3b6e19c-f006-4094-a4f2-de465d4cb624 · outbound

This paper cites Geo: Generation and execution optimized stochastic computing accelerator for neural networks,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Geo: Generation and execution optimized stochastic computing accelerator for neural networks,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.177003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.122265Z digest=sha256:b8b2938ec6b02b8e104082c4ee762d27ffa2fb523d2244973eec5607299293a5

Observation fd869e77-bcc5-47ed-ae71-1f27620566bb · outbound

This paper cites Quasar-vit: Hardware-oriented quantization-aware architecture search for vision transformers,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Quasar-vit: Hardware-oriented quantization-aware architecture search for vision transformers,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.160011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.126214Z digest=sha256:d38da65d4597ddc0ae2c6bac0a39e4ac4479af98824be2c7fa33a78b4ca1499f

Observation 995b9f46-e3c0-4229-aeed-91e1b6697a4a · outbound

This paper cites High-performance fpga-based cnn accelerator with block-floating-point arithmetic,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format High-performance fpga-based cnn accelerator with block-floating-point arithmetic,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.144687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.130338Z digest=sha256:f51a135efbf91f2278c5b19b018ad9c7184a53252f008a1fe26704d0cb37aedc

Observation ac37e816-1469-47bf-b6f8-e6b369beb443 · outbound

This paper cites Awq: Activation-aware weight quan- tization for llm compression and acceleration,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Awq: Activation-aware weight quan- tization for llm compression and acceleration,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.127154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.134742Z digest=sha256:da404237652727b023ca4738082a4be6d1e64b8d88183876c9b6a18d74614858

Observation 0bd651d8-2194-4958-bb11-36374a362fd2 · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.139346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.139346Z digest=sha256:e01fdb2e353680266323c35e50f4ea610b3b25ad8fb927f8b742aaec9d6b04e9

Observation 96581e0b-fb9d-4487-9eeb-24b973f4971b · outbound

This paper cites LLM-QAT: Data-Free Quantization Aware Training for Large Language Models.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LLM-QAT: Data-Free Quantization Aware Training for Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.145166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.145166Z digest=sha256:c478a5f3f592d208a297e12214b757fb05bc24cd35bfe8d0d915547993ae3328

Observation 938449d2-5a6d-4ae6-93af-c87b1417f81a · outbound

This paper cites Kivi: A tuning-free asymmetric 2bit quantization for kv cache,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Kivi: A tuning-free asymmetric 2bit quantization for kv cache,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.111294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.149988Z digest=sha256:78e0568f76c760489eca3c28f25b381c2941209cf3ab00593a6566913b38f49e

Observation 9fac23e3-43b4-470b-9fe0-a571e64a12c3 · outbound

This paper cites Dis- tilling bit-level sparsity parallelism for general purpose deep learning acceleration,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Dis- tilling bit-level sparsity parallelism for general purpose deep learning acceleration,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.094997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.154781Z digest=sha256:affa81bc81ae1277d6ab5ed49f38c4b60ee5417cc743f5700837f2483fa94592

Observation 9f43329e-2dac-445d-80dc-259f8ef78496 · outbound

This paper cites Keep the cost down: A review on methods to optimize llm’s kv-cache consumption,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Keep the cost down: A review on methods to optimize llm’s kv-cache consumption,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.071054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.159281Z digest=sha256:3381d17c83417a64068e1fccaf1d106983b11b2daf23e3b02b06b52746cac0d4

Observation fda3f6a9-e9e4-4d15-b6ee-fed49585f9ee · outbound

This paper cites Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-12T13:46:01.510626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.164716Z digest=sha256:09b5b38347105877cc80d4565afcc78e2b70d01d5f25c7da57f4dbaac629b28c

Observation 92cb1726-fe3b-4af0-9975-a35027aa97d6 · outbound

This paper cites Fpnew: An open-source multiformat floating-point unit architecture for energy-proportional transprecision computing,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Fpnew: An open-source multiformat floating-point unit architecture for energy-proportional transprecision computing,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.047111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.169746Z digest=sha256:dc5e1cddcf506a7be714860b2d280cf393a5b22acb9a2f81274940118e4f06ba

Observation ab2800b6-3577-4014-99cc-55e9691b7a41 · outbound

This paper cites The penn treebank: Anno- tating predicate argument structure,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format The penn treebank: Anno- tating predicate argument structure,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.029380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.175099Z digest=sha256:8049f39c3977c581b1db5b1ef4a5bd28f15f0d64bf7d5835f979b4198c7e90e3

Observation 53655596-273c-437f-8332-4a6e8b23254c · outbound

This paper cites Pointer sentinel mix- ture models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Pointer sentinel mix- ture models,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.011649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.180120Z digest=sha256:05ae69be7feb44fc555deec8bc3781008186b1f83651f401c2f7b9537014bf55

Observation c70b0fc3-7afe-4fca-9688-f206d80ec28d · outbound

This paper cites Flexblock: A flexible dnn training accelerator with multi-mode block floating point support,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Flexblock: A flexible dnn training accelerator with multi-mode block floating point support,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.994256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.184647Z digest=sha256:9ce5def4c7cf0770d0719aa323b0fb28862b42837a255795e3360f00682332ee

Observation 42f71867-f06c-4e26-8f0f-df6241d81ce7 · outbound

This paper cites Cutlass,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Cutlass,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.977580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.189062Z digest=sha256:fa631fc3b98aabc66e6e285af46e31b16bd793391ec7869a4f68141bf519bfcc

Observation aaa2dfc7-d314-412d-b7be-4e114a599e64 · outbound

This paper cites GPT-4 Technical Report.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format GPT-4 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.193821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.193821Z digest=sha256:281857e6851c3afd39d0f87da9d112b3753184f0b22166e1ddca50710d917a03

Observation 38f48fab-2824-488e-99e6-22a062854016 · outbound

This paper cites LUT-GEMM: Quantized matrix multipli- cation based on LUTs for efficient inference in large-scale generative language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LUT-GEMM: Quantized matrix multipli- cation based on LUTs for efficient inference in large-scale generative language models,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.956504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.198965Z digest=sha256:e56a23ca1c1868ebbbb4f4ed5b045c5e505eb44aaf82cbb8056b6a1164f89066

Observation 842a34fe-7a8a-445c-af99-adb9fc1a11ab · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.937328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.203051Z digest=sha256:f05ab0f6ffbe7827757688c603ce7a847da8a515c3315739c8db342b24aa213c

Observation a2e29963-82d3-4df9-bd66-09724eabba27 · outbound

This paper cites Omniquant: Omnidirectionally calibrated quantization for large language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Omniquant: Omnidirectionally calibrated quantization for large language models,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.918778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.207048Z digest=sha256:1070dc65d431c698c6709b228026184b6e7c29df5d991f733dcd41f507eb38ad

Observation 0c831a98-4474-4151-be7d-780a15bb36ad · outbound

This paper cites Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural network,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural network,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.902106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.211499Z digest=sha256:ec07890f5553392f1b7dae1b4c7d0d59e67a875b0748461dc45765f6079891d6

Observation 6fc9dfc1-34c9-4723-bbb3-a58c01237086 · outbound

This paper cites Bitwave: Exploiting column-based bit-level sparsity for deep learning accelera- tion,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bitwave: Exploiting column-based bit-level sparsity for deep learning accelera- tion,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.883768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.215760Z digest=sha256:f03ba71df84d07c48f551f3196fa7847a7d095c4e6d4709a28a9783916edc6b4

Observation 29c5eea0-0148-490a-bccc-b7a1fa42f4ca · outbound

This paper cites Dissecting tensor cores via microbenchmarks: Latency, throughput and numeric behaviors,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Dissecting tensor cores via microbenchmarks: Latency, throughput and numeric behaviors,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.866350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.219759Z digest=sha256:e607c3d529322c0d27cf5dd6c9d2cf748c7e88f6939d141bc9b30faf00f14c46

Observation 9f38003d-2c2e-46f5-a22d-73d7d17f5816 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Gemma: Open Models Based on Gemini Research and Technology

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.224872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.224872Z digest=sha256:b4a210f0012cfca0cab697f29736bbc112055edd69b755d3939d8269399df9f9

Observation 141d78a5-df15-482d-b957-e9c47fd67d6a · outbound

This paper cites Bebert: Efficient and robust binary ensemble bert,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bebert: Efficient and robust binary ensemble bert,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.847626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.230172Z digest=sha256:24342c87aec5fcdc9b7fc357ab66be5e0f720dfd79fd86de22845341a62de640

Observation 11d12725-bcca-489b-a2a9-cfa99513dfc3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LLaMA: Open and Efficient Foundation Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.235170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.235170Z digest=sha256:23ea68a926c6ffbba2c01a73589e3e5a54050db984bd6aa2eed24f55f0286a29

Observation 5c8e640e-b043-4082-8752-c673a4b40678 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.241211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.241211Z digest=sha256:884a63f36a99c00fd2839ca32817cb35b09b5ab68f1cf8b2884d981496ca256b

Observation 46eab912-d74a-47a8-b1b4-90ff15b5743c · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.246476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.246476Z digest=sha256:e002ef0b2d018af486eed1d0489913e201da06f80063a9c2770384b6bbfd788b

Observation 8714c3f3-58f9-4e62-b876-960241b8ce18 · outbound

This paper cites Bsvit: A bit-serial vision transformer accelerator exploiting dynamic patch and weight bit-group quantization,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bsvit: A bit-serial vision transformer accelerator exploiting dynamic patch and weight bit-group quantization,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.830986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.252804Z digest=sha256:e3ddb77c2b7aa5d85e7b37ba04eabee2a6c12e0d115629e69d1b7042c3238f55

Observation 7117e91a-1dc4-4735-8b60-08834df4a5c0 · outbound

This paper cites Haq: Hardware-aware automated quantization with mixed precision,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Haq: Hardware-aware automated quantization with mixed precision,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.813986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.257602Z digest=sha256:9a4bed363821d7cfba898ab321ead237d573bf909ffd1073973dc11c9b4ea0ba

Observation 7357fcfa-01dc-47ce-984c-24f3c5b8c6d7 · outbound

This paper cites Outlier suppression: Pushing the limit of low-bit transformer language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Outlier suppression: Pushing the limit of low-bit transformer language models,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.796440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.263418Z digest=sha256:cdc8c2ba7caf8ac18a9a502a42b3125fe7adad441f183bb45e497e0216659171

Observation e45a072e-fb67-41d6-b90b-9c748cf3f417 · outbound

This paper cites Quant-llm: Accelerating the serving of large language models via fp6- centric algorithm-system co-design on modern gpus,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Quant-llm: Accelerating the serving of large language models via fp6- centric algorithm-system co-design on modern gpus,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.780173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.268642Z digest=sha256:4c7840bdf4b645b3acc43c688ad47d74f51eef5939a4737457a1eace7da3ea5c

Observation a71eaf87-6323-4f73-9d05-19aca2ce0d6b · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Smoothquant: Accurate and efficient post-training quantization for large language models,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.762681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.274153Z digest=sha256:1c2174e0879a21e448b7352573bd1441e51aea994d7d864ed9b139edd14a2b2a

Observation bcbc3551-c787-4ab8-a8e0-5a7133c1a315 · outbound

This paper cites Efficient streaming language models with attention sinks,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Efficient streaming language models with attention sinks,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.745057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.279979Z digest=sha256:4b072739115cd9e731897048ece31473e1d0b3061bea8ab0715e2f0d838a9938

Observation 40732215-a09e-41ab-9da8-fce2c45ab70d · outbound

This paper cites OneBit: Towards Extremely Low-bit Large Language Models.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format OneBit: Towards Extremely Low-bit Large Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.284612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.284612Z digest=sha256:dfcfa50addf95cb19834c46b7aa985c66d175273e671cdfcb010fa790461eca5

Observation 486b959b-41fa-45b4-854b-7615f86f023f · outbound

This paper cites Kv cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Kv cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.726050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.289754Z digest=sha256:735150e94e0f075be98d1f2090b551e8c4372d302bb7b03284bad367e8c5ff17

Observation ff4d496e-4129-4775-9c74-ebbb851dd57d · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.294427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.294427Z digest=sha256:43ad59970007e45967eb6f02087dd7043307b143ef443290c940ae431215cb39

Observation 85aeb8e5-3c60-483a-87ce-7b2378c2c551 · outbound

This paper cites Mokey: Enabling narrow fixed-point inference for out-of-the-box floating-point transformer models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Mokey: Enabling narrow fixed-point inference for out-of-the-box floating-point transformer models,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.705368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.299569Z digest=sha256:2ddf38fd48a7bf783fb8b3d11eec6e894a25c7c5f68248811e810f933462a27a

Observation aca341d1-8958-4999-ade4-1e25ceec5f95 · outbound

This paper cites Fast: Dnn training under variable precision block floating point with stochastic rounding,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Fast: Dnn training under variable precision block floating point with stochastic rounding,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.682402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.303655Z digest=sha256:863555efbd2cf30e25c986a114760b5d810f542bcc9deffda1a82e83b891f0c3

Observation 50f869ac-ab0d-42bf-85f6-51a7c069feca · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format OPT: Open Pre-trained Transformer Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.308601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.308601Z digest=sha256:9ba67f03ebfc59b95b2a98c2b1de291b5019023a1777d33c7f34e10671b4a4d4

Observation 60706211-3401-4a13-b602-09912f9bbf66 · outbound

This paper cites Cam: Cache merging for memory-efficient llms inference,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Cam: Cache merging for memory-efficient llms inference,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.659492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.313207Z digest=sha256:ccf40ef318237dae903b58da7519ce037ea55224aa497dc9af9dc008fc7e2590

Observation a419e5d6-a114-4446-9c43-703baa852076 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format H2o: Heavy-hitter oracle for efficient generative inference of large language models,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.639814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.317190Z digest=sha256:79600064875dd45ce1b7e03e4eee8d914aaef0b84f13e467c7c65fa14968d565

Observation 5e21a48c-cfbe-473b-a856-68a0bb2292e9 · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate llm serving,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Atom: Low-bit quantization for efficient and accurate llm serving,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.621468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:46:01.321581Z digest=sha256:22e8e4833a69c1fee86d6fd0e585d8780eaf22faaab5b6c72ad85e7174c0e5cc

Pith citing papers

No inbound Pith citation observations are available.