Pith. sign in

Paper Citation Record · LEDGER

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference

As of 10 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 0 inbound Pith citation observations for arXiv:2502.02040.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02040 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:40:56.623496Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact1
  • verified fuzzy41
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4125d7d7-4910-4038-a307-f73154be051b · outbound

This paper cites Phi-3 technical report: A highly capable language model locally on your phone, 2024.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Phi-3 technical report: A highly capable language model locally on your phone, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.331392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.383105Z digest=sha256:ec7de914b05c0a1781ce1d3ece5af2044fd2d257f9ec054f63de6a7ce2ed08ae

Observation d0f03468-926a-41b4-8a48-0838424117a1 · outbound

This paper cites Llm inference performance engineering: Best practices., 2023.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Llm inference performance engineering: Best practices., 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.322083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.386863Z digest=sha256:99a0519269538409f137552318c4bdece01cab1a8040f36b5e99f5aea484be4e

Observation 71a53602-126b-415c-b675-abe2ed4fdebc · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.390827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.390827Z digest=sha256:113d6bfdb5a7c41a09f5f5e32af3dc191f94df7813e5945414ce0d2da1f498ab

Observation 0fd5c2b9-a9c3-4750-aed2-84c18df135ad · outbound

This paper cites Colt5: Faster long-range transformers with conditional computation.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Colt5: Faster long-range transformers with conditional computation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.311780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.394133Z digest=sha256:1668ed909623f41c69bc76cea59aee53f1d6da85161b09e8ffd7669451c70d9e

Observation ba66cd9e-843e-44c9-bdbe-c1c10b38aae4 · outbound

This paper cites Hydra: Sequentially-dependent draft heads for medusa de- coding, 2024.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Hydra: Sequentially-dependent draft heads for medusa de- coding, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.302208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.397230Z digest=sha256:8ebdc27645667c0478ee8bccbb2b6e9fec67bc51534cd44cbfdb121ef5110c1a

Observation 6d74ad29-afe5-4099-95e7-84ca590ae6c9 · outbound

This paper cites Massively multilingual sentence embeddings for zero- shot cross-lingual transfer and beyond.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Massively multilingual sentence embeddings for zero- shot cross-lingual transfer and beyond

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.292785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.400215Z digest=sha256:74adf153e48690e49112ef32a732caeeba2e4662a54a2bed039e830af4545f03

Observation 41fab238-730c-4342-91a8-007410e1ba14 · outbound

This paper cites an unresolved cited work.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-09T13:40:57.283289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.403333Z digest=sha256:7df7116d6327b062f82fd5f124bfb27fca9278e1daa7e12d575588c81f4370ed

Observation b7497dd7-1fc9-440e-a26d-dc0aaff5af52 · outbound

This paper cites Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.273661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.406063Z digest=sha256:f65673a4b8040d7c56a7b43771bc168b67b70ef3a6ea24e6a649410d863db09f

Observation 0d7e73a3-b19b-4020-8b21-4c906ee9a86a · outbound

This paper cites Speculative Streaming: Fast LLM Inference without Auxiliary Models.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Speculative Streaming: Fast LLM Inference without Auxiliary Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.409002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.409002Z digest=sha256:19fd203d088b8738829dc6b3f9a3b88039d1eb4c9174cd5651011ca23cdf24c8

Observation 0fd1e4e1-0fe1-44e3-ab53-3f16e10cdc7f · outbound

This paper cites Language models are few-shot learners.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Language models are few-shot learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.412320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.412320Z digest=sha256:29cf8744cd31a9f53adff88537b1721ca38a739976a8ff4913ec132b457ab523

Observation 4303cf81-200f-4f73-b870-0c06615b8231 · outbound

This paper cites Medusa: Simple framework for accelerating llm generation with multiple decoding heads.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Medusa: Simple framework for accelerating llm generation with multiple decoding heads

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.258058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.415675Z digest=sha256:2e8534aaad2e2894c434a093c08ef5e1ea70a73d5ea8d46848ee3a7221ddbed2

Observation aac75849-a2df-4b96-a798-56abaf1bfb20 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Accelerating Large Language Model Decoding with Speculative Sampling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.418582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.418582Z digest=sha256:c7df65470e45dca9aca7497fe1c8bbb434ad3e49b2c885ecf2611d820c829d2a

Observation 130c8be1-1ea0-485f-89b6-1dc16780f1fd · outbound

This paper cites EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.421746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.421746Z digest=sha256:135eafe4ac4881fb7d10615f7dc37e525924312477a85cdebc9b04e65cbfa94d

Observation 8ac1876a-3b63-4d21-a25c-97cf08a7e166 · outbound

This paper cites DialogSum: A real-life scenario dialogue summarization dataset.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference DialogSum: A real-life scenario dialogue summarization dataset

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.248545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.425216Z digest=sha256:b8ba4d7d8ca3e9f2fe5bdeb31c525024a9ab034fe6bfc2bd14aaa009a8075d63

Observation 21915d66-c976-49e7-91c7-a3ebf8c8865f · outbound

This paper cites Koala instruction set documentation, 2023.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Koala instruction set documentation, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.239004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.428185Z digest=sha256:ece222e8419436162366e970a14bfe822a6ee7002eae6f93a8b7a57fd9562a70

Observation 227accd1-0895-45b8-bca0-66d1ef3531a3 · outbound

This paper cites Dai and C.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Dai and C

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.229623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.431458Z digest=sha256:74ce66a382de564e1aa77780d282bc73a1733dad42dc83fa82a23d926944f131

Observation e26b19c1-73ca-4a02-901d-1d9b158496c8 · outbound

This paper cites SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.434433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.434433Z digest=sha256:dbad74f72489def64552b0e8e08a20e5a30fb94de03bd83a1a095d9a3934e676

Observation e41a4539-c671-4783-bc6a-9b7406b46111 · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.437585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.437585Z digest=sha256:09a841150e51cef8d6972902b4a612158b2cef24a7f06ae5383798a51b2247f6

Observation f6e487e9-a792-4216-8855-6e30d3dee94a · outbound

This paper cites GLaM: Efficient Scaling of Language Models with Mixture-of-Experts.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.440954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.440954Z digest=sha256:374ea653bb85c7c2533983ea81b979e3b752409bc45200ab6177034aa63822cd

Observation 160c048d-2e2c-45f4-b846-0c76eae1e813 · outbound

This paper cites Evaluating the State-of-the-Art of End-to-End Natural Language Generation: The E2E NLG Challenge.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Evaluating the State-of-the-Art of End-to-End Natural Language Generation: The E2E NLG Challenge

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.219584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.443911Z digest=sha256:f4dfe75f6e46e88894f566c67c593fcdafa376994a984c5b23e687f169748c38

Observation 609d15ac-e549-4076-9b5a-8e8fab22c3dc · outbound

This paper cites Depth-adaptive transformer.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Depth-adaptive transformer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.209682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.446767Z digest=sha256:e199de8ad4a8fcca291b79ebf2127fb5530af30953369f67f428ad8cdac95a8d

Observation f0c36f0c-7c4e-4754-bc62-a18a23bbc08f · outbound

This paper cites Predictive Exit: Prediction of Fine-Grained Early Exits for Computation- and Energy-Efficient Inference.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Predictive Exit: Prediction of Fine-Grained Early Exits for Computation- and Energy-Efficient Inference

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-09T13:40:56.806396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.450038Z digest=sha256:def4ca51469a089d73e01ff55c5c30c0b8980e8cce7586e0a3dc51811913b845

Observation f9c0b924-8153-49ff-ba7b-e93db0e1bb26 · outbound

This paper cites Decoupled early time series classification using varied-length feature augmentation and gradient projection technique.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Decoupled early time series classification using varied-length feature augmentation and gradient projection technique

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.200580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.453137Z digest=sha256:32a33bc1311810cceb2f55c66f6941bb8069992da4aef5478680f8088de46c5e

Observation 7084d49a-8e22-478c-a066-6a00445acd05 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.191175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.456254Z digest=sha256:ad6ab0ebd965c030a80731e29d8661b1e06704f235fe1d1f2f125989ebf22b09

Observation 4431a023-fc9e-4a15-b5d5-2dcf933839a3 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.182197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.459296Z digest=sha256:1df94175ce90dd1e14e00deff62609f58af344e82f8565825c14d4b35d392967

Observation 1b7f3946-3ba4-45cb-9ffe-e1ee086430ec · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Sparsegpt: Massive language models can be accurately pruned in one-shot

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.462238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.462238Z digest=sha256:dc6ef85e76fc3660de3e789534519563ee5bf76d72280f6ee977818dda7830c7

Observation 6bb98c18-d580-4034-8917-807fbbb8ad18 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.465244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.465244Z digest=sha256:4c17500ddd678d2fb02512ee22ea9a878585b8e759225588432c3c5dcba39204

Observation 5d4d0187-f8ff-4d4b-b8aa-da7a922e888a · outbound

This paper cites Breaking the sequential dependency of llm inference using lookahead decoding, November 2023.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Breaking the sequential dependency of llm inference using lookahead decoding, November 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.166407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.468436Z digest=sha256:84a14f40f6f669c13f302e0567c3e04c6219d509a8f49f4c9ad9d8d01dc22ec6

Observation 4aa6fb13-2c4e-41b8-aacc-e8cb3ca8c476 · outbound

This paper cites Garncarek and J.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Garncarek and J

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.156345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.471577Z digest=sha256:f1fbbe910e2d74397526c8d62d19b76c315937357cd2e2f06b11a2217fc885e4

Observation ecfd5849-273a-42c5-98b4-890529b78c3a · outbound

This paper cites Koala: Dialogue-based fine-tuning improves factuality and safety of llms.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Koala: Dialogue-based fine-tuning improves factuality and safety of llms

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.146814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.474505Z digest=sha256:fce5ca1e918fa1fb8d271d999d997f86469fe45b5f55cf491e4d3bb9bc23f989

Observation 85ed8cbd-fb4c-4221-b2c5-350088de0f74 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference MiniLLM: On-Policy Distillation of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.477410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.477410Z digest=sha256:0e9303607be41f6ea2a56ed6797e2ec23b0b8e2894fb350a3d50b0b3c110d82c

Observation bca9d024-aab7-43b6-9906-3b91c0ed9afd · outbound

This paper cites Identity mappings in deep residual networks.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Identity mappings in deep residual networks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.136782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.480814Z digest=sha256:b12559131ca9135869c82e81862a58f7a974bd7553c1e43fe6d30e5e308da66f

Observation 26d78d7a-9e40-4440-85c6-0dca8b7a41ed · outbound

This paper cites Dynabert: dynamic bert with adaptive width and depth.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Dynabert: dynamic bert with adaptive width and depth

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.127174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.484066Z digest=sha256:ff923a4922fff54755cb953664ad170e26cd70df4ceefc14d5bbd29173e41913

Observation faac1eb5-92f3-46ec-9546-12b43bf53dec · outbound

This paper cites Adaptive mixtures of local experts.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Adaptive mixtures of local experts

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.117441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.487126Z digest=sha256:a3918d0a7d47066ed49a7d36330d8f1a6de3121c10ba39611686021ee1148f60

Observation 9f701b72-db33-40fa-87e4-9b0568cc5a4c · outbound

This paper cites Mixtral of Experts.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Mixtral of Experts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.490622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.490622Z digest=sha256:007a8d4bbc98c7356877c2b5688160f199a3bdb6a6fcaff438a52de25faef987

Observation b1bca0b5-6a53-4789-89b7-97a58e7df821 · outbound

This paper cites Hierarchical mixtures of experts and the em algorithm.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Hierarchical mixtures of experts and the em algorithm

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.107682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.494312Z digest=sha256:b2ef21645f6210f675a1449a04e06c600e93c14db47b8964752e1bd8e96fab79

Observation 8fdaa0e2-c1c9-4eea-b6c6-0c14f2a38fa0 · outbound

This paper cites Ten lessons from three generations shaped google’s tpuv4i.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Ten lessons from three generations shaped google’s tpuv4i

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.097080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.497877Z digest=sha256:d7c3eefd5b9f28a25ae005793fd85605976d96427b162208d4279c4db56e0133

Observation 5e1ca558-39b8-43f0-b235-0172eaaed0ea · outbound

This paper cites In-datacenter performance analysis of a tensor processing unit.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference In-datacenter performance analysis of a tensor processing unit

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.087012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.501446Z digest=sha256:b906dc7ce9cd41d2f87879fe9635cd9f1bcb68d71759013e33da3a7ad6f881ed

Observation 1e7dd89a-5e9b-4a86-9cba-04cb45604405 · outbound

This paper cites Gpus and the future of parallel computing.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Gpus and the future of parallel computing

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.076898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.504762Z digest=sha256:9d271a1d3aa612ce094d7da5fdafed7d2e4700bc0175ac6c21472e684fc33d2c

Observation 1fe90847-784b-4f93-afd3-b61115c1f186 · outbound

This paper cites Ai and machine learning acceleration in mobile devices: A survey of architectures, hardware, and algorithms.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Ai and machine learning acceleration in mobile devices: A survey of architectures, hardware, and algorithms

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.067134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.508416Z digest=sha256:d861d1ef7c4130ba80968ff62cf851fb3b6a88c6e4148d160805223dbfd82050

Observation 178775d1-af33-4df4-9fab-7f32a86e552a · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.512100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.512100Z digest=sha256:77807d914260a5795634c284f7b5a2994c05a5010ba2aa4dcae67af4ce27bfdb

Observation c83a3b46-ca25-44a7-89b4-fdadeadc8250 · outbound

This paper cites GShard: Scaling Giant Models with Condi- tional Computation and Automatic Sharding.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference GShard: Scaling Giant Models with Condi- tional Computation and Automatic Sharding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.056698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.515875Z digest=sha256:443f0d4940ddd0616a8ec3935c7bf05d86a5cb316a21d77aa147077a52a1ef28

Observation bf0974c2-d4f8-4464-b951-dc8ea5df6241 · outbound

This paper cites Fast inference from transformers via speculative decoding.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Fast inference from transformers via speculative decoding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.519530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.519530Z digest=sha256:348674aa897340bf0fd1e78101d65a8bb2b3f706b96bd186b5afd28d56b7d623

Observation a989bc8e-992a-4d64-9e55-4178ae7d21e2 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Gemma: Open Models Based on Gemini Research and Technology

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.523044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.523044Z digest=sha256:22d2df6d3d7d35e2885c5084a5a181c4c6720352794d84b61127faf9ac62611c

Observation a7e6870a-1e77-4fe2-9c32-194f4a17984c · outbound

This paper cites SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.526776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.526776Z digest=sha256:d4dc18ecd0f7ea5de1b96ca8b22800ac6414b5ae833aa260d8e7f8717960b070

Observation a878b6c1-c1d3-4107-8066-86a96b964406 · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference OLMoE: Open Mixture-of-Experts Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.530396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.530396Z digest=sha256:0abb562935b323c65916a84191a9a95ec9db847b8283e4f30ea3e12833bc302a

Observation 7c104600-4e33-4238-8012-774f324cb833 · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Efficient large-scale language model training on gpu clusters using megatron-lm

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.041818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.533930Z digest=sha256:1d0bf1b5b7b7659104a9bce21004bdb1ad8f92e5a55082d027f4ed21a0d26fec

Observation f9aa0612-9a4f-46d0-9529-450ac437ab24 · outbound

This paper cites Cuda c++ programming guide, 2021.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Cuda c++ programming guide, 2021

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.032278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.536413Z digest=sha256:d59a05ffadd86a1e78f0cb3deb05df47eff5d9aa80318f6fb65630ee6a8193c7

Observation 6e7f2e82-0775-4321-b5de-c29ebf5ee8af · outbound

This paper cites GPT-4 Technical Report, 2023.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference GPT-4 Technical Report, 2023

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.538683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.538683Z digest=sha256:06f6085dccb04edec002dd13a19372748eafe50fb616ace27994b4e4405bbb37

Observation eaa65c5c-e415-45b0-a151-d99761e0f97d · outbound

This paper cites Language models are unsupervised multitask learners, 2019.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Language models are unsupervised multitask learners, 2019

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.016826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.541014Z digest=sha256:c6c4748df5ced3f1e3d109cd2bf46d335b1e042c14149f4968d1fd32c995cf2c

Observation 4b394c27-284a-410a-9876-ede0ab73677d · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.543541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.543541Z digest=sha256:63122b950099c7e0f2dae694a1721c2be4f9973e1798db9237c80a6cdf4f41ce

Observation 9ba1dab6-d238-47d8-b6e7-397f3f7d009d · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.546001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.546001Z digest=sha256:1807ca14fb3326779b559c2dea74c7f24763115e24a9e28e13bbd24dda8376da

Observation 99140b0c-0b9b-45c6-bccf-dd6f79f30650 · outbound

This paper cites Tran, Yi Tay, and Donald Metzler.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Tran, Yi Tay, and Donald Metzler

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.006694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.548559Z digest=sha256:f7b8d244e74995f1805520dec50c3a6e929b5cbee4fa4a4c39a47365b303bf0e

Observation fc8612f8-4408-4a35-ba5e-20b685fb18db · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.551350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.551350Z digest=sha256:8e2a43b2324598b951b4999f9a8e729f30042255e2d12637e7dcea59325b0f6d

Observation 82d54c11-66cc-49ee-b739-1d0bc93339bc · outbound

This paper cites an unresolved cited work.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-09T13:40:56.995345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.555030Z digest=sha256:7538de0cb03b85f881f54d9b14af1c33f64b15610fb7bada47e538bef466fbe0

Observation 88d5c12a-7214-4143-a264-5035e6bc7063 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.558179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.558179Z digest=sha256:a19a6a3d90c83b8cb6b086d29aa8fab0697617761b36fa70774e5bc5cef0155b

Observation f77c6337-9423-4108-8396-e2ff63b2e4f7 · outbound

This paper cites Accelerating LLM Inference with Staged Speculative Decoding.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Accelerating LLM Inference with Staged Speculative Decoding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.561768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.561768Z digest=sha256:3be30cbdc6117f05ebb93537ed7712f57fa170267f63304ac38550950c9a1144

Observation 8082d3ff-2a22-46ec-b727-441fb821c775 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference A Simple and Effective Pruning Approach for Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.565238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.565238Z digest=sha256:28c8e205a055da3a3baf53ab304861c82557edb76c066d1bc33d92b0160af24c

Observation 204e7fcc-a2aa-4a6b-a5f9-999fb832453c · outbound

This paper cites Manmatha.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Manmatha

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.985328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.569007Z digest=sha256:bdaa23008c8e3e6385ca31be82ad7a905717eb7dced6eb6f7db0fe3d7d430592

Observation aab5ecd6-5ddd-4812-9a99-a9d0e01cbdea · outbound

This paper cites Tuan Pham, S.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Tuan Pham, S

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.975515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.572295Z digest=sha256:1dbcfcc18f3dbbde011517abefa70a1edb19cf673f50bb9f46593341bdae0b80

Observation 3340cb68-c9a3-433b-ac70-c5a9fc519aa3 · outbound

This paper cites Model cascading: Towards jointly improving efficiency and accuracy of nlp systems.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Model cascading: Towards jointly improving efficiency and accuracy of nlp systems

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.965723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.575573Z digest=sha256:265f7dc7c30516bdf0f0cc85236d1329f6bf90d15fd754141c647f1a10444850

Observation a84b56e0-c4d7-43ac-9b3e-6fa0effefe23 · outbound

This paper cites Investigating acceleration of LLaMA inference by enabling intermediate layer decoding via instruction tuning with ‘lite’.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Investigating acceleration of LLaMA inference by enabling intermediate layer decoding via instruction tuning with ‘lite’

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.955798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.578809Z digest=sha256:e349490068d17a8b1fc2dc8e0b5b5f4c0a6e3833f9d4334078b6c7aed928ff62

Observation eaccb3ce-68cd-4c3f-86a9-c4aa429352f3 · outbound

This paper cites Attention is all you need.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Attention is all you need

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.581952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.581952Z digest=sha256:00d44729abb2f5574066cf6f7e22a816fe796beb292a0b68cd52b1a9750b14a5

Observation 941f8570-0355-4682-9b53-c23bb9ffce91 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.585245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.585245Z digest=sha256:a8503d254d949381101f02ff009663ecfb252a87d7ad8cd73ff802cd8d508487

Observation eb6434e6-a44b-44f0-8041-5c12645a8b5f · outbound

This paper cites Transformers: State-of-the-art natural language processing, 2020.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Transformers: State-of-the-art natural language processing, 2020

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.940331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.588725Z digest=sha256:d41d09a8afcbcdadcff96b65e5488c16aa791fe4079d45facafe6c820aa0bb58

Observation f6407729-53b6-45dd-acbe-16cf625e36cf · outbound

This paper cites Speculative decoding: Lossless speedup of autoregressive translation, 2023.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Speculative decoding: Lossless speedup of autoregressive translation, 2023

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.930511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.591822Z digest=sha256:9eca1c70305cec3f460af8c613984ac0376e66481a591db2def2c6dfa332053c

Observation 09a028c5-939f-4d78-ba5c-e2473c964c4d · outbound

This paper cites Deebert: Dynamic early exiting for accelerating bert inference.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Deebert: Dynamic early exiting for accelerating bert inference

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.920409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.595373Z digest=sha256:3eeab2ae011a6eefc9501dbbf7a71605a3862c19fdc8472ba3ae78e1f1a334d5

Observation 772efeba-ed59-4ef5-a62b-88fe4ebbe676 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.598812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.598812Z digest=sha256:fab7dd6fbb1b95b96d5a1192d06a021277c97576746689820f44ad521c46081a

Observation a3788be9-4e4b-4053-ad3c-d74088f4b257 · outbound

This paper cites Zeroquant: Efficient and affordable post-training quantization for large-scale transformers.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Zeroquant: Efficient and affordable post-training quantization for large-scale transformers

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.909812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.602300Z digest=sha256:b644a7e5f7ded9675aa9e7541f5306fd99948b4e22aaa6312c788151f1616b16

Observation 350e3e94-af68-4af6-9d1a-37eba339cc6f · outbound

This paper cites Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.605732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.605732Z digest=sha256:b7cfe9a84286d25cddbb3d2abe2df9e0e8ae5b7a0d6a0cde2d9dda7ad2e709f4

Observation e22e5582-7e5d-45ca-aeec-b77e78ccb127 · outbound

This paper cites P Xing, Hao Zhang, Joseph E.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference P Xing, Hao Zhang, Joseph E

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.609457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.609457Z digest=sha256:1c1c84964ba8dadb320dc42b934b493b561f86e98a3a188b94ef1aab52808b83

Observation 334aef28-f2e3-4be0-b144-280af8aee97c · outbound

This paper cites Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.612810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.612810Z digest=sha256:656fb8c48e02a8243bd099eb3680f8f02914e6eb24f7526f5348682c349bcfa5

Observation da1245c6-485f-4294-88a2-171311cd026f · outbound

This paper cites Zhou and et al.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Zhou and et al

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.893641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.616348Z digest=sha256:53b7d69a21be7a11bb393b9e05a96b968656000eb11389c42cc98b39d38c2cc1

Observation e687d854-34be-4b62-9d06-1176b7fdb6f1 · outbound

This paper cites Designing efficient sparse expert models.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Designing efficient sparse expert models

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.883584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.619675Z digest=sha256:2b48dc2ef18f35772561805717c0415523f48e9b9670beaf3a0e1462defb17ce

Observation f72b9031-d6d4-4fbb-9136-bdaa7c543a1b · outbound

This paper cites an unresolved cited work.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-09T13:40:56.874916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T13:40:56.623496Z digest=sha256:5e0322e78f87fa37313978dc92e3454af5c6634bcc75d4506ac63d27fc4932ad

Pith citing papers

No inbound Pith citation observations are available.