Pith. sign in

Paper Citation Record · LEDGER

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format

As of 16 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2411.15982.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15982 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:46:01.321581Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

89 of 89 outbound references displayed

  • verified exact1
  • verified fuzzy65
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eb908304-59b4-4082-8b4e-8ca13af78c73 · outbound

This paper cites Resq: Residual quantization for video perception,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Resq: Residual quantization for video perception,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.916127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.916127Z digest=sha256:0bd67e8ef098eb1e393cb258a104def6711bc5dc187a19c1cc5b49e5b1eb8f79

Observation 07b6e85e-10eb-4742-ad69-eb44c17f2a83 · outbound

This paper cites Bit-pragmatic deep neural network computing,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bit-pragmatic deep neural network computing,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.924633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.924633Z digest=sha256:18c2fa9d657ab7e1b0e3542ec7b44a9fb9e763b439bc133c43fbb4e7ea6b77af

Observation 109156cc-ec88-4c3e-b740-788a6f30fc11 · outbound

This paper cites Explaining neural scaling laws,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Explaining neural scaling laws,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.928808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.928808Z digest=sha256:43d96805e38cbc2ec20dffc061e95b293c577a915cadad80856371ebb9d1ba98

Observation f22e992c-6df8-4344-b35e-d24c47398f35 · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Longbench: A bilingual, multitask benchmark for long context understanding,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.933063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.933063Z digest=sha256:623218d2755b80ca29bb661b7340adedab0494332d041276ed6ffadb9eab3889

Observation c2fbf310-6db7-48f4-b945-3546b592e458 · outbound

This paper cites Demystifying chatgpt: An in-depth survey of openai’s robust large language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Demystifying chatgpt: An in-depth survey of openai’s robust large language models,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.937312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.937312Z digest=sha256:a16c5175750ecee07b0c9b27a07215c65defd88bff8156e08378d7c300d6c6f4

Observation 512cd457-8d26-4023-b16d-7a04113f5a87 · outbound

This paper cites Genus synthesis solution,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Genus synthesis solution,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.941946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.941946Z digest=sha256:2d46f839660002ee5766208ec14c56f92c7687a24bd644f3b42773f1d055b555

Observation 5c53dec2-f8da-4aba-a7c7-4e8f421615c9 · outbound

This paper cites General purpose deep learning accelerator based on bit interleaving,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format General purpose deep learning accelerator based on bit interleaving,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.946229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.946229Z digest=sha256:9960d0af2f74512b7963c6fe5a437c320b6de75b6cfe4c51d62400e2194fa0df

Observation f864ed24-45c9-401f-ae54-5c8143377e09 · outbound

This paper cites Quip: 2-bit quanti- zation of large language models with guarantees,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Quip: 2-bit quanti- zation of large language models with guarantees,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.950195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.950195Z digest=sha256:8227b1fa5b1d852b83932e0a7d32f8cc200f3de4d0926b85cce1912c42116476

Observation db8e462a-25cb-46fb-b2b5-1b5a23573c3e · outbound

This paper cites EfficientQAT: Efficient Quantization-Aware Training for Large Language Models.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format EfficientQAT: Efficient Quantization-Aware Training for Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.954375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.954375Z digest=sha256:5872936f1e149aff067cde0dda2085cb985d1b9b5be4bbeaca3aff98f258b319

Observation c08c524f-0c14-4d61-b5d6-f0409de0d6e1 · outbound

This paper cites Nacl: A general and effective kv cache eviction framework for llm at inference time,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Nacl: A general and effective kv cache eviction framework for llm at inference time,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:00.959299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:00.959299Z digest=sha256:9202a5e3ddc37753b9599855e25b51e3e6be3d8ff48bad46befaa595c00622a7

Observation 3e140eb0-9c01-4fba-b3db-04b06150bd15 · outbound

This paper cites Palm: Scaling language modeling with pathways,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Palm: Scaling language modeling with pathways,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.688552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:00.963367Z digest=sha256:21c9d95558ddef1639a85bf41590e1ade54b74c853143f520f7c44301355f98d

Observation 0e391695-d4ee-48c0-a897-a30f0f232543 · outbound

This paper cites Vs-quant: Per-vector scaled quantization for accurate low-precision neural network inference,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Vs-quant: Per-vector scaled quantization for accurate low-precision neural network inference,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.673582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:00.967701Z digest=sha256:9f273b4b4559bf727a3743eef15e89d6cb002d4a9822b3d2ab54f70ff0dda76b

Observation cbfde174-f710-43b3-bc1e-c0588e604377 · outbound

This paper cites Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.660384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:00.971544Z digest=sha256:c1aa2f090b0a0255658a7b936873715a34ca4a35433210d61e06029616b1a905

Observation 987c2874-d2d8-4184-a3d2-f795bbdf96e0 · outbound

This paper cites With shared microexponents, a little shifting goes a long way,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format With shared microexponents, a little shifting goes a long way,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.647418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:00.976155Z digest=sha256:1c76a53376ed55e7af41e6c7fe3bd6bc9dc51ddf488902443d22bda38767bcc4

Observation d0088e03-3677-4f8a-8050-e4d1ccb8efc8 · outbound

This paper cites A timing-driven approach to synthesize fast barrel shifters,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format A timing-driven approach to synthesize fast barrel shifters,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.635399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:00.980180Z digest=sha256:a65699ef1b58a4a0f2aef76dcb7a7b1f1807a69ae338a10051381c5bc893fb57

Observation ec788826-bdd6-4922-b1b5-c725344a1534 · outbound

This paper cites Llm.int8(): 8- bit matrix multiplication for transformers at scale,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Llm.int8(): 8- bit matrix multiplication for transformers at scale,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.622334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:00.984391Z digest=sha256:7b428c1179801875ba1d58b6fb9ce0698de2e3322a4c9521b1f7eee7c25ce320

Observation 8480b9a4-4201-4813-a0d8-68b85334e80b · outbound

This paper cites The case for 4-bit precision: k- bit inference scaling laws,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format The case for 4-bit precision: k- bit inference scaling laws,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.606453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:00.988162Z digest=sha256:f253f8b64ea77e10d76550e1b43f30c045bc8cd5c6350b785a69d341eebf7457

Observation 00ecbf6f-8afa-4325-9fb0-89afa02558c9 · outbound

This paper cites Hawq: Hessian aware quantization of neural networks with mixed-precision,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Hawq: Hessian aware quantization of neural networks with mixed-precision,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.588365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:00.992180Z digest=sha256:b93713d73e9199228e40fb83ae3bbf01f8ba99404d12affe1ee5a6d808bb5e1b

Observation e6510551-5932-40cc-b624-1173024cb859 · outbound

This paper cites Training dnns with hybrid block floating point,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Training dnns with hybrid block floating point,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.571178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:00.996132Z digest=sha256:6924586fa6d254267fc4f59e037037721bae31efb65b441906d741404527ecfb

Observation dce06533-ae2f-40b8-af4b-eab6c87e6b06 · outbound

This paper cites Skvq: Sliding-window key and value cache quantization for large language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Skvq: Sliding-window key and value cache quantization for large language models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.557598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.000362Z digest=sha256:a5c59d157fd49b6f404ff3e50174ca844615b2bc7af709ee2fec5e9b0413e0f1

Observation 178aaf15-301b-40f6-8b13-bc2fbe940d64 · outbound

This paper cites Extreme compression of large language models via additive quantization,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Extreme compression of large language models via additive quantization,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.545100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.004503Z digest=sha256:89a5069d907c688d6b08ee07c4f66d56e1a879637e4f56d68260452b22fb0cc9

Observation f389ac12-1c55-47f5-80ad-99753d11967f · outbound

This paper cites Reconfig- urable acceleration of 3d-cnns for human action recognition with block floating-point representation,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Reconfig- urable acceleration of 3d-cnns for human action recognition with block floating-point representation,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.532602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.008439Z digest=sha256:dfe656226f00ac7dddb4f4c26b55dfe408052e7fd711cd589bc5956cc1f91190

Observation 01002d37-24f7-4610-8d4a-c2d9c229c290 · outbound

This paper cites Static block floating-point quantization for convolutional neural networks on fpga,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Static block floating-point quantization for convolutional neural networks on fpga,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.520801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.012376Z digest=sha256:583070639f033ef025f96cf6e47afddb14868f59f7274fc4b2c268f310e5c667

Observation 553aa229-a25a-4769-a67b-c312da294cc8 · outbound

This paper cites Optq: Accurate quantization for generative pre-trained transformers,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Optq: Accurate quantization for generative pre-trained transformers,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.506326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.016683Z digest=sha256:b32e521329de11e870c265e83cb95f21f569fdc8607dccba468e6447e465834f

Observation be801d61-b0fc-4bb3-a27d-e66b430b3b5c · outbound

This paper cites LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.020686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.020686Z digest=sha256:51510eab31169c1bb008e6ea6b74f3634803d254b7406414be3ea2208ef196ac

Observation 6fcd9706-878e-4c1c-a67b-2c44d6dc4687 · outbound

This paper cites Boost: block minifloat-based on-device cnn training accelerator with transfer learning,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Boost: block minifloat-based on-device cnn training accelerator with transfer learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.488289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.025336Z digest=sha256:cc15c43b32d5084432b33279b60c8421a0e3a241f559e9455f3eaa7a1871f62c

Observation 2efccf6d-a519-4127-b9b4-af72090c2eac · outbound

This paper cites Olive: Accelerating large language models via hardware- friendly outlier-victim pair quantization,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Olive: Accelerating large language models via hardware- friendly outlier-victim pair quantization,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.472758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.029221Z digest=sha256:ba7b797fc0188b19f7cd03653a7679febaa506385bb0d2638bf9b07e99ab0cc3

Observation 734655d4-95bf-4ea4-b953-1923a5d8a9fe · outbound

This paper cites Ese: Efficient speech recognition engine with sparse lstm on fpga,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Ese: Efficient speech recognition engine with sparse lstm on fpga,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.459389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.033280Z digest=sha256:a4bd9e24c112f6f42577d989a3fac8c3943b049df03484775b895daac594d055

Observation e31e43a2-f16c-464b-a096-57ffd31fb592 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.037498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.037498Z digest=sha256:534a8b48db9b47dd61b37a8a64003477f860679679d3f6077adbb1d5c325b126

Observation acb4a88b-71fd-425a-a4d4-7557df0a3b61 · outbound

This paper cites A precision-scalable risc-v dnn processor with on-device learning capability at the extreme edge,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format A precision-scalable risc-v dnn processor with on-device learning capability at the extreme edge,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.445583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.042306Z digest=sha256:1248ad69ff40fc8446450f4e85a9c2898958244d61fd1f80f3567ea7d7adf45e

Observation 3c4a177a-2705-4842-99c8-b7d2c8a2e8a7 · outbound

This paper cites Mind the gap: Attainable data movement and operational intensity bounds for tensor algorithms,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Mind the gap: Attainable data movement and operational intensity bounds for tensor algorithms,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.429805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.046677Z digest=sha256:35cad199731ba71c5fd047a7fc8c6f191e5d4a860300b5356da1afbfd2bd2642

Observation 9feedc6f-68ee-45ee-b614-0cd0da64d635 · outbound

This paper cites Figna: Integer unit-based accel- erator design for fp-int gemm preserving numerical accuracy,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Figna: Integer unit-based accel- erator design for fp-int gemm preserving numerical accuracy,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.410522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.050530Z digest=sha256:4fbc0a3b8e6b6aa90184296bdfad99fcf358de0019c94f55ed62c4857adcf63a

Observation b971c7e3-66b4-44ab-9994-26903bdfdec2 · outbound

This paper cites Perplexity—a measure of the difficulty of speech recognition tasks,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Perplexity—a measure of the difficulty of speech recognition tasks,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.392252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.054333Z digest=sha256:483bdae84e7cc66232cf51ff462889c2ead9cecf54504cb06a3f1a3e4bf625ac

Observation cca7ac48-4bed-454b-91a7-e725e24eecce · outbound

This paper cites Mr. biq: Post-training non- uniform quantization based on minimizing the reconstruction error,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Mr. biq: Post-training non- uniform quantization based on minimizing the reconstruction error,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.371118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.060235Z digest=sha256:c8c6ecccf0038de7a2dfa36863bf1f2fe827700b35d274722c637546eecf2f86

Observation 5a868e5d-84ae-43b1-a764-91759d50d4b7 · outbound

This paper cites Biqgemm: matrix multiplication with lookup table for binary-coding-based quan- tized dnns,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Biqgemm: matrix multiplication with lookup table for binary-coding-based quan- tized dnns,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.356272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.064416Z digest=sha256:3412fb6f9a285faa55f0d8ddf32ae3001a5d220b7a81a230bbdf53f021f518b6

Observation 28571475-c8a3-430d-97f8-fb952857a7fa · outbound

This paper cites Ten lessons from three generations shaped google’s tpuv4i: Industrial product,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Ten lessons from three generations shaped google’s tpuv4i: Industrial product,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.341385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.069776Z digest=sha256:003b47b777dbad41d6d3ffd2ffa01bfac8c5c36fe435003240c21a603d14a798

Observation fe860abe-bc83-47e5-b316-3a520278d183 · outbound

This paper cites Stripes: Bit-serial deep neural network computing,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Stripes: Bit-serial deep neural network computing,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.326657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.073787Z digest=sha256:951fcf93cda864984d44388711d0189c471d7a79b6b57c3a8615966898d915cc

Observation 94942890-0859-4f55-b5a3-4a08a3024924 · outbound

This paper cites A survey of gpt-3 family large language models including chatgpt and gpt-4,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format A survey of gpt-3 family large language models including chatgpt and gpt-4,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.078272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.078272Z digest=sha256:058f1d3a2e981f45c6bb565b5e93e66dcfef9f5a63751436d76aa7af954d30a0

Observation 02267acb-faf9-41a0-8f47-18c580f1c56c · outbound

This paper cites A 95.6-tops/w deep learning inference accelerator with per-vector scaled 4-bit quantization in 5 nm,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format A 95.6-tops/w deep learning inference accelerator with per-vector scaled 4-bit quantization in 5 nm,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.304176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.082459Z digest=sha256:771c9781c5c44d6f55b77c992ddd31b668f9c22cfe4714d874d055fa0d899f2c

Observation e01546cb-3e38-4a57-a7db-f2803705ecdf · outbound

This paper cites Compressed context mem- ory for online language model interaction,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Compressed context mem- ory for online language model interaction,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.288380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.086446Z digest=sha256:01a09a38897d9a0491dc3655993dcd7eba2ef070768bb0e7570b5ccacb390123

Observation ba442fe0-bdc1-4a37-be6d-7f271a5796e7 · outbound

This paper cites Dacapo: Accelerating continuous learning in autonomous systems for video analytics,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Dacapo: Accelerating continuous learning in autonomous systems for video analytics,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.273986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.090475Z digest=sha256:37203a6d2747d622f330807eed0158ed9782ff03fa9038dac8bb07deb141a24c

Observation 048c2a80-8910-4127-b228-ffba59500712 · outbound

This paper cites Winning both the accuracy of floating point activation and the simplicity of integer arithmetic,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Winning both the accuracy of floating point activation and the simplicity of integer arithmetic,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.260399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.096261Z digest=sha256:2d0cb456bcbaea45767f0e7280671ac914d4b0ea5b2592172745268c170015bd

Observation 05f82623-88e2-48b1-9229-de66136b466e · outbound

This paper cites One-shot model for mixed-precision quantization,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format One-shot model for mixed-precision quantization,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.244419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.101397Z digest=sha256:d3d642c948d4a7b4553aeb5e9d3629fdb6262a11d147f78d08968456999b1edb

Observation 6de4f875-f11c-4ed3-9f6a-a87b72e63cee · outbound

This paper cites Flexpoint: An adaptive numerical format for efficient training of deep neural networks,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Flexpoint: An adaptive numerical format for efficient training of deep neural networks,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.231230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.106244Z digest=sha256:8d769c9ae6161477514c4a74426a762bda49ee34345ac6c20328d213b0a7a208

Observation 6ff985f6-a23d-486b-8f4d-8d73abf68337 · outbound

This paper cites Tender: Accelerating large language models via tensor decomposition and runtime requantization,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Tender: Accelerating large language models via tensor decomposition and runtime requantization,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.218324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.110394Z digest=sha256:dc895cac510e3c058ca09e05e004a57c552cc17413c08cda497c23b356124367

Observation 9c0e6188-8cdf-4159-85a6-5069701f9205 · outbound

This paper cites Bitcluster: Fine-grained weight quantization for load-balanced bit-serial neural network accelerators,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bitcluster: Fine-grained weight quantization for load-balanced bit-serial neural network accelerators,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.206101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.114610Z digest=sha256:bdcceef54bf2afee4839f14760f0c689e961c0defade1a5bde7d31478c25a687

Observation 19bd6321-1c54-4f04-8dc9-116e8de5dff2 · outbound

This paper cites Norm tweaking: High-performance low-bit quantization of large language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Norm tweaking: High-performance low-bit quantization of large language models,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.190206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.118186Z digest=sha256:301008aa6ee9035b701438c2e231ae99e32e08e7fb08f6e487ea6d4eb1f8c59d

Observation e3b6e19c-f006-4094-a4f2-de465d4cb624 · outbound

This paper cites Geo: Generation and execution optimized stochastic computing accelerator for neural networks,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Geo: Generation and execution optimized stochastic computing accelerator for neural networks,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.177003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.122265Z digest=sha256:ead91e3289fd2b248c9adb202dbf017eb5c1f4841f7c56bb77b94752657d9dfd

Observation fd869e77-bcc5-47ed-ae71-1f27620566bb · outbound

This paper cites Quasar-vit: Hardware-oriented quantization-aware architecture search for vision transformers,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Quasar-vit: Hardware-oriented quantization-aware architecture search for vision transformers,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.160011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.126214Z digest=sha256:8e573fc85f523eabe7706e5bd2dcaa24857a5ceae21b1042f47280fd203286d5

Observation 995b9f46-e3c0-4229-aeed-91e1b6697a4a · outbound

This paper cites High-performance fpga-based cnn accelerator with block-floating-point arithmetic,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format High-performance fpga-based cnn accelerator with block-floating-point arithmetic,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.144687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.130338Z digest=sha256:8e3f71e135b4ae8579a29f19ead5931cb9495707f83a33becf327069014bd7c2

Observation ac37e816-1469-47bf-b6f8-e6b369beb443 · outbound

This paper cites Awq: Activation-aware weight quan- tization for llm compression and acceleration,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Awq: Activation-aware weight quan- tization for llm compression and acceleration,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.127154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.134742Z digest=sha256:56a13e7330c624ff120b7164905e9a6ed73debb240a4c0b46ccdd70bac9c4a21

Observation 0bd651d8-2194-4958-bb11-36374a362fd2 · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.139346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.139346Z digest=sha256:cd8a62a9fc66e05a51171a23f5e161040cd8c45e37aeb8478426ffabb7aaf759

Observation 96581e0b-fb9d-4487-9eeb-24b973f4971b · outbound

This paper cites LLM-QAT: Data-Free Quantization Aware Training for Large Language Models.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LLM-QAT: Data-Free Quantization Aware Training for Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.145166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.145166Z digest=sha256:566b560be9ecd013d4191f17bdcb9f6cf903e12a082f39854c9e046414a84725

Observation 938449d2-5a6d-4ae6-93af-c87b1417f81a · outbound

This paper cites Kivi: A tuning-free asymmetric 2bit quantization for kv cache,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Kivi: A tuning-free asymmetric 2bit quantization for kv cache,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.111294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.149988Z digest=sha256:a88ce515358337efdbe8006b85b6eac8c0f67fe5de9e94ed5b61132b4325a92e

Observation 9fac23e3-43b4-470b-9fe0-a571e64a12c3 · outbound

This paper cites Dis- tilling bit-level sparsity parallelism for general purpose deep learning acceleration,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Dis- tilling bit-level sparsity parallelism for general purpose deep learning acceleration,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.094997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.154781Z digest=sha256:e5c9106cb87e84229fad015150c302efb198d6bbaef980bea4eefd580f14e844

Observation 9f43329e-2dac-445d-80dc-259f8ef78496 · outbound

This paper cites Keep the cost down: A review on methods to optimize llm’s kv-cache consumption,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Keep the cost down: A review on methods to optimize llm’s kv-cache consumption,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.071054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.159281Z digest=sha256:52ec9d053a1c5caee1dff31e9a7140c85330da41f695fe997c034eebcb00cc8e

Observation fda3f6a9-e9e4-4d15-b6ee-fed49585f9ee · outbound

This paper cites Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-12T13:46:01.510626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.164716Z digest=sha256:696879e0e16bbfeca1f3954015cccce49e59b1bc45c0d77ca8e782717e4d7c3d

Observation 92cb1726-fe3b-4af0-9975-a35027aa97d6 · outbound

This paper cites Fpnew: An open-source multiformat floating-point unit architecture for energy-proportional transprecision computing,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Fpnew: An open-source multiformat floating-point unit architecture for energy-proportional transprecision computing,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.047111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.169746Z digest=sha256:1f6a64180571b8c61abf80fb5239f131a26d2ab5a41b9069beee0eb9c8a2a467

Observation ab2800b6-3577-4014-99cc-55e9691b7a41 · outbound

This paper cites The penn treebank: Anno- tating predicate argument structure,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format The penn treebank: Anno- tating predicate argument structure,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.029380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.175099Z digest=sha256:f3a64910fd5033947c510e4368364b5add232cbe18c6304c99ea1a58c8104262

Observation 53655596-273c-437f-8332-4a6e8b23254c · outbound

This paper cites Pointer sentinel mix- ture models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Pointer sentinel mix- ture models,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:02.011649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.180120Z digest=sha256:30215a425f3b0081ea67d5f33ba2273a57cd7dba9e4eed8d44d9250dc3b0ba7c

Observation c70b0fc3-7afe-4fca-9688-f206d80ec28d · outbound

This paper cites Flexblock: A flexible dnn training accelerator with multi-mode block floating point support,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Flexblock: A flexible dnn training accelerator with multi-mode block floating point support,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.994256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.184647Z digest=sha256:6e370142b2c622b1656478ab0f7452da3d032e87fed320ccb93bfd173163a8d3

Observation 42f71867-f06c-4e26-8f0f-df6241d81ce7 · outbound

This paper cites Cutlass,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Cutlass,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.977580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.189062Z digest=sha256:76271d490f691af4fa0758ff8f7c5f74575e43514d21d10bb7adfcfa4d9bc2ca

Observation aaa2dfc7-d314-412d-b7be-4e114a599e64 · outbound

This paper cites GPT-4 Technical Report.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format GPT-4 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.193821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.193821Z digest=sha256:5c4740b540eb676587edd0cb9497ceb7602cddf1ecf1386b2c40eb9dcfa7cc8c

Observation 38f48fab-2824-488e-99e6-22a062854016 · outbound

This paper cites LUT-GEMM: Quantized matrix multipli- cation based on LUTs for efficient inference in large-scale generative language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LUT-GEMM: Quantized matrix multipli- cation based on LUTs for efficient inference in large-scale generative language models,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.956504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.198965Z digest=sha256:1884d24a5c88e75e2874d30725b9e96da661c9b07ba42a20a00e2f958628f0fb

Observation 842a34fe-7a8a-445c-af99-adb9fc1a11ab · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.937328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.203051Z digest=sha256:adea567744cbb1fc49c7610ad1237db876b7b6868205c6c1f7328813e1c62c4c

Observation a2e29963-82d3-4df9-bd66-09724eabba27 · outbound

This paper cites Omniquant: Omnidirectionally calibrated quantization for large language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Omniquant: Omnidirectionally calibrated quantization for large language models,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.918778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.207048Z digest=sha256:348673bf942ce65976dd1bc3a178fcf52e5363fb61dac8a5a77f7126b4e76283

Observation 0c831a98-4474-4151-be7d-780a15bb36ad · outbound

This paper cites Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural network,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural network,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.902106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.211499Z digest=sha256:249543b28facf38c77cc67f1ccd27b602382c76dc15c73c5288faaf44ca7c9ac

Observation 6fc9dfc1-34c9-4723-bbb3-a58c01237086 · outbound

This paper cites Bitwave: Exploiting column-based bit-level sparsity for deep learning accelera- tion,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bitwave: Exploiting column-based bit-level sparsity for deep learning accelera- tion,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.883768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.215760Z digest=sha256:3939dc4f82dc8ccdb286c17df06664f3eaf833e5dea9a763ff8d46e9bf19293a

Observation 29c5eea0-0148-490a-bccc-b7a1fa42f4ca · outbound

This paper cites Dissecting tensor cores via microbenchmarks: Latency, throughput and numeric behaviors,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Dissecting tensor cores via microbenchmarks: Latency, throughput and numeric behaviors,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.866350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.219759Z digest=sha256:dc56ab7717ed816173c964a12c507345300e912f491e19234a79942021ddc692

Observation 9f38003d-2c2e-46f5-a22d-73d7d17f5816 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Gemma: Open Models Based on Gemini Research and Technology

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.224872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.224872Z digest=sha256:56fbd4e1000a9c0276438b73ff520e579954011f0bd1f352456f203adec645cd

Observation 141d78a5-df15-482d-b957-e9c47fd67d6a · outbound

This paper cites Bebert: Efficient and robust binary ensemble bert,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bebert: Efficient and robust binary ensemble bert,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.847626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.230172Z digest=sha256:23e1445dfa2b8cbdf09768c10096792c68be6a652c01e90fbd815ec8dea0ea46

Observation 11d12725-bcca-489b-a2a9-cfa99513dfc3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LLaMA: Open and Efficient Foundation Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.235170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.235170Z digest=sha256:0a3a74e206bd7d0272b6120238db679d3771f75ab5b4727a262feac37b019890

Observation 5c8e640e-b043-4082-8752-c673a4b40678 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.241211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.241211Z digest=sha256:15d74d7fda24da88e2d859c74fe0cd5b358a8cdd97bd9ecec704c5e372d24da1

Observation 46eab912-d74a-47a8-b1b4-90ff15b5743c · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.246476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.246476Z digest=sha256:cc2f7329f63debc833000b0c049ab8f2e9f54ac20b909518b07fe18e867f1be0

Observation 8714c3f3-58f9-4e62-b876-960241b8ce18 · outbound

This paper cites Bsvit: A bit-serial vision transformer accelerator exploiting dynamic patch and weight bit-group quantization,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Bsvit: A bit-serial vision transformer accelerator exploiting dynamic patch and weight bit-group quantization,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.830986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.252804Z digest=sha256:3278f990b7841a926c6ee814f32a6e0f96599e878692328fe9b7a79c5b2f9f72

Observation 7117e91a-1dc4-4735-8b60-08834df4a5c0 · outbound

This paper cites Haq: Hardware-aware automated quantization with mixed precision,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Haq: Hardware-aware automated quantization with mixed precision,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.813986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.257602Z digest=sha256:0cac980a0e31c3e0d06ba220ebd51e283ee948989f1295e4b2a0438db3214761

Observation 7357fcfa-01dc-47ce-984c-24f3c5b8c6d7 · outbound

This paper cites Outlier suppression: Pushing the limit of low-bit transformer language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Outlier suppression: Pushing the limit of low-bit transformer language models,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.796440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.263418Z digest=sha256:38f8eb1a42625454f078fc22ea3c1be77b6ec342fed501cd8b37683b3731bf27

Observation e45a072e-fb67-41d6-b90b-9c748cf3f417 · outbound

This paper cites Quant-llm: Accelerating the serving of large language models via fp6- centric algorithm-system co-design on modern gpus,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Quant-llm: Accelerating the serving of large language models via fp6- centric algorithm-system co-design on modern gpus,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.780173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.268642Z digest=sha256:c09edb2c0cdd3857881ff7d6bbb8663338541c45a1b1143b4a85dac289cf0076

Observation a71eaf87-6323-4f73-9d05-19aca2ce0d6b · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Smoothquant: Accurate and efficient post-training quantization for large language models,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.762681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.274153Z digest=sha256:f67aef51b1c6cfedd82a9cad5c6ef6d4b4cc64d7c1f5cf392608ee64c4d6c1af

Observation bcbc3551-c787-4ab8-a8e0-5a7133c1a315 · outbound

This paper cites Efficient streaming language models with attention sinks,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Efficient streaming language models with attention sinks,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.745057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.279979Z digest=sha256:87e789962e45d7770a25af360742425921d55f567fdaf49fb9f0395fc2a27426

Observation 40732215-a09e-41ab-9da8-fce2c45ab70d · outbound

This paper cites OneBit: Towards Extremely Low-bit Large Language Models.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format OneBit: Towards Extremely Low-bit Large Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.284612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.284612Z digest=sha256:b61df1d6d2fdc08ae7fb621bc83c2ca133bf8275cf5bc767b71a6dcb3774d417

Observation 486b959b-41fa-45b4-854b-7615f86f023f · outbound

This paper cites Kv cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Kv cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.726050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.289754Z digest=sha256:614bddf10ebe552a539510aabf50f14503c0447f5a703aa1e01f74ff0bcf03bf

Observation ff4d496e-4129-4775-9c74-ebbb851dd57d · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.294427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.294427Z digest=sha256:4ae17df42d5169a83e1064ce86e1c069b62fd64c930797ce7bb41a8b117f4132

Observation 85aeb8e5-3c60-483a-87ce-7b2378c2c551 · outbound

This paper cites Mokey: Enabling narrow fixed-point inference for out-of-the-box floating-point transformer models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Mokey: Enabling narrow fixed-point inference for out-of-the-box floating-point transformer models,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.705368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.299569Z digest=sha256:c31485efc785763b51acb843c0ffb83fdbeb7ccf11275927cee3e1ca398d0192

Observation aca341d1-8958-4999-ade4-1e25ceec5f95 · outbound

This paper cites Fast: Dnn training under variable precision block floating point with stochastic rounding,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Fast: Dnn training under variable precision block floating point with stochastic rounding,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.682402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.303655Z digest=sha256:3cd887a6d1577198388722757189cf59136cb90d290083e9c945ebf09dca83cf

Observation 50f869ac-ab0d-42bf-85f6-51a7c069feca · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format OPT: Open Pre-trained Transformer Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.308601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.308601Z digest=sha256:e19c6507d3aeb63c377606e54600ddea6f743e1aad34be7ed031782acde7c470

Observation 60706211-3401-4a13-b602-09912f9bbf66 · outbound

This paper cites Cam: Cache merging for memory-efficient llms inference,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Cam: Cache merging for memory-efficient llms inference,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.659492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.313207Z digest=sha256:548a9df403337a5fc2b1ac3d79cb23b0d44ea95e96ebd077dad71789964ab8b7

Observation a419e5d6-a114-4446-9c43-703baa852076 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format H2o: Heavy-hitter oracle for efficient generative inference of large language models,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.639814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.317190Z digest=sha256:9b7c29313c93ecf89d9fc80add308de215033bad62e7eb4a67296bed7f8ac25c

Observation 5e21a48c-cfbe-473b-a856-68a0bb2292e9 · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate llm serving,.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format Atom: Low-bit quantization for efficient and accurate llm serving,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:46:01.621468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:46:01.321581Z digest=sha256:e6f4ac55fa4ee0b64a48134568d89047fc044793e8cc2c26d177100fad63f586

Pith citing papers

No inbound Pith citation observations are available.