Pith. sign in

Paper Citation Record · LEDGER

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters

As of 23 August 2026, this Paper Citation Record lists 98 of 98 outbound references and 0 inbound Pith citation observations for arXiv:2502.07832.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07832 v1

Coverage vector

measured 98 of 98 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:44:00.720633Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

98 of 98 outbound references displayed

  • verified exact3
  • verified fuzzy13
  • unresolved82
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 609a0388-61f5-4d4f-8bab-334ba9f061f7 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.394550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.394550Z digest=sha256:4b139bb6f8412c32f23da5986cc19befbf8d83e6d51d5da820ba57b5e24a4c70

Observation f78eda28-a31b-4362-9066-1676b007fe72 · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.399688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.399688Z digest=sha256:3a2885c7f688f1abd33f65a526300ccaf7faa1131a8f79a1217a5b409b2fa3a0

Observation ad89bf29-a54e-4554-993e-c2a38d02cae8 · outbound

This paper cites Mathqa: Towards interpretable math word problem solving with operation-based formalisms, 2019.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Mathqa: Towards interpretable math word problem solving with operation-based formalisms, 2019

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.403441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.403441Z digest=sha256:6452551865f5634e5c98c85a758a724f99037ac40bba9b16d568f7c8aec1d39f

Observation 4c65fb43-6d32-422b-b5f5-f7e98e5c1f8b · outbound

This paper cites Qwen Technical Report.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.407280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.407280Z digest=sha256:8e0b2f23654f2766a05174d8acaef703a31804b102e2713722c5ccfffdaefd32

Observation 5219e199-9660-4a5e-ba13-6eda36712582 · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Pythia: A suite for analyzing large language models across training and scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.411395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.411395Z digest=sha256:81cf75640ae1e9fc38007877146c0cd2a9c87bc77db30c74798d4d834cdb71b1

Observation adf91926-3d4f-4a60-a854-f1f9a93f7d7b · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Piqa: Reasoning about physical commonsense in natural language

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.415307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.415307Z digest=sha256:7594e3ecc8a2de016edbb2994b9b73b737369d813e980ec5fa1b59d4429e2bc3

Observation cc9c7778-8471-48d6-ac99-e2e7b0a8d294 · outbound

This paper cites GPT-NeoX-20B: An Open-Source Autoregressive Language Model.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.419177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.419177Z digest=sha256:7f4d34c349daa990b51a7a22163d7afd58a07763d0c5db6f694200a76e895011

Observation 2fa7b9b5-d346-44e6-a89f-9ac1bcf1e561 · outbound

This paper cites Language Models are Few-Shot Learners.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Language Models are Few-Shot Learners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.422737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.422737Z digest=sha256:6a7011b86656d21c6e7a03a545397781ffd251923d4d2b7801ee2ad745f4d422

Observation 33217203-abf2-4e99-9754-d8f3a24df587 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.426468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.426468Z digest=sha256:3c0ef8984152540b0bd9eabd1207027b072a1b47c69a641752619bf4108f8553

Observation 99af52a5-86ed-477c-97da-3b0ec35d1cdc · outbound

This paper cites Code alpaca: An instruction-following llama model for code generation.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Code alpaca: An instruction-following llama model for code generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.429889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.429889Z digest=sha256:07fb0d31a5f907ca8c7df1f8e1239e68ab3dee98472fe28c9f4a8bb06018da5d

Observation ac8a45f6-28b3-4949-b2bd-0700b1623bde · outbound

This paper cites Learning to maximize mutual information for chain-of-thought distillation.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Learning to maximize mutual information for chain-of-thought distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.433149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.433149Z digest=sha256:7a4e52dd9d2e59f6a91aa634c696c9552db38e5c2777a81ba2ac51eded3bed39

Observation e4d976e9-10fa-47c7-b95f-0791502939b2 · outbound

This paper cites LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.436326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.436326Z digest=sha256:c5c739ed93b59028cf836a3886c76bf4cbd0dd0b01035a91d9da45be296bf70f

Observation 0dda9464-5803-46e1-828d-e2c8405ddf5c · outbound

This paper cites D ialog S um: A real-life scenario dialogue summarization dataset.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters D ialog S um: A real-life scenario dialogue summarization dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.439774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.439774Z digest=sha256:66829b28ce404d030c7e5fd4e02a14a81a3d086bb642801a260d19433b735525

Observation dca0e07c-86e1-4425-bf77-755de8711dbe · outbound

This paper cites Palm: Scaling language modeling with pathways.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Palm: Scaling language modeling with pathways

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.442950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.442950Z digest=sha256:9f4cf7c971254d0f5a170a8ea12b4e32482a22eea476739a33e0afbe4ae2ed95

Observation af5fd9e9-6e76-4b30-9820-56d76ac01d54 · outbound

This paper cites Boolq: Exploring the surprising difficulty of natural yes/no questions.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Boolq: Exploring the surprising difficulty of natural yes/no questions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.445935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.445935Z digest=sha256:b522e7d4f81b2d6b0cc9dfb4b574d4ca795e412855a613baf944a6d6b1ac5b48

Observation 76976001-ca3e-4cc4-8e76-993290329eff · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.448983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.448983Z digest=sha256:c5938f623d18ad78169388ea73b6f34247f2a31a3ffefe18a53647b0f46b66f1

Observation 384bbdd8-c598-4cc6-8732-558b9d69691f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Training Verifiers to Solve Math Word Problems

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.452482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.452482Z digest=sha256:de746b4e74bdac1887c35a76f78540d2fee6baf82b793354edae52c85157ae0c

Observation 55dcea1e-fcb7-4a65-85bc-ba36521e10a9 · outbound

This paper cites Free dolly: Introducing the world's first truly open instruction-tuned llm, 2023.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Free dolly: Introducing the world's first truly open instruction-tuned llm, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.456036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.456036Z digest=sha256:534bed946ba63de1cb175d5aa6d44a7ca0bbe13fdf71f39c7edd35b98e43de9d

Observation 5144d1d7-b9ea-43d0-b3b9-ac95b4ea587f · outbound

This paper cites Mutual: A dataset for multi-turn dialogue reasoning.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Mutual: A dataset for multi-turn dialogue reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.459132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.459132Z digest=sha256:4c9e03761288c7ead3ee957202fdead6f6fbb3c968d3a848147ce53c14460882

Observation f4f976b7-223a-4a05-98af-bec7aced9cd8 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.462157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.462157Z digest=sha256:40d490573a0e0fe12693e1e8d4d70fbc9adbe165433e851dad78eaceff16f0c0

Observation cd351f96-3ef2-4134-8c23-5cb350db3938 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.465407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.465407Z digest=sha256:f9d8d62eeba2128a31d47b7c97aa845ddaaa11a272310c3f35ec14676bf8ed8d

Observation d2947059-ec14-420b-b640-69bc0a070c9d · outbound

This paper cites an unresolved cited work.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.468388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.468388Z digest=sha256:16190968a8673f280e83cb396012eb297b45b512e5a92295ca9cadf32dff8b37

Observation f04ece8a-781a-4cf6-981a-a8cf6b514c72 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Qlora: Efficient finetuning of quantized llms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.471354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.471354Z digest=sha256:7c3a616fc45a3e8b40d1c57a4d14581f86a18656305c4ce6d4afda3c3d9ab31f

Observation 8d69a650-a8bd-4bfe-bada-f1598e01d261 · outbound

This paper cites Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.474327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.474327Z digest=sha256:766591b8aaf0ba6a801a8e36a6782c81ea28e287e2632b131369047aa7bb871e

Observation 9bb8c23d-e8c0-4e2b-bae4-dbe1d46a9d9a · outbound

This paper cites Blockwise compression of transformer-based models without retraining.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Blockwise compression of transformer-based models without retraining

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.477757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.477757Z digest=sha256:d046f3cdd087c8bb542e38ebfcd9488cffd9e64cfb7bbeff3e2ec2bc45bee178

Observation 6d9b0a07-1127-4b4f-8b2b-b5fab7e721f3 · outbound

This paper cites The Llama 3 Herd of Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters The Llama 3 Herd of Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.480714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.480714Z digest=sha256:ea2995db165037a2589a3ef55a1255ce56e7fd00adaaa41a049f5fd77a3ff1ef

Observation d1547bb2-eb02-4cf9-ad61-c92ecfa99937 · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Sparsegpt: Massive language models can be accurately pruned in one-shot

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.483877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.483877Z digest=sha256:3df5d63e89f18eb2a314458ec1659be4d2fb91c122aefb124af9a67fa4d849c8

Observation 9373bec1-19a2-4e7b-855b-716f523cf8a0 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.486906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.486906Z digest=sha256:333c3693889d7bde768fd7022af84d1674f7040dba67893858cb91e78d8a69df

Observation d554e6de-e130-4979-a77b-ffaab0f0b5f7 · outbound

This paper cites Learn-to-share: A hardware-friendly transfer learning framework exploiting computation and parameter sharing.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Learn-to-share: A hardware-friendly transfer learning framework exploiting computation and parameter sharing

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.489949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.489949Z digest=sha256:f133e57a0e3640fa7fca33ffd727e0d7849432f95da810678e8c61dd873b40b9

Observation f77d0b6a-ef8e-45b7-aa1e-c0768c072392 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters A framework for few-shot language model evaluation, 07 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.492841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.492841Z digest=sha256:c657e7f1eb7fc5c129e044fa0c67297d6a35315d35e92033df6f93c008d065c2

Observation a2a0077f-9906-4d40-a7c2-332cf89e403b · outbound

This paper cites The Unreasonable Ineffectiveness of the Deeper Layers.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters The Unreasonable Ineffectiveness of the Deeper Layers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.495709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.495709Z digest=sha256:e77f37c11971d191064e22078dc614e8707da865f2270862688c07a45b213082

Observation 990270c8-cd96-4223-a541-89660cfc90ba · outbound

This paper cites Textbooks Are All You Need.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Textbooks Are All You Need

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.498897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.498897Z digest=sha256:b6ed0ba2c851f50137eafaf6dc156c0a6842fcc4e96e5587870992351df9fc81

Observation 126127a4-5cba-4546-9584-de4734c56e1e · outbound

This paper cites Compressing pre-trained language models using progressive low rank decomposition.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Compressing pre-trained language models using progressive low rank decomposition

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.502449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.502449Z digest=sha256:7558f5e6236f73a035c9c1f86eecc4cd8f647a1c8cf6f8d13a34726719e8737c

Observation 1bdf172f-4dc9-4564-81cb-d56f9a209242 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Distilling the Knowledge in a Neural Network

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.505557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.505557Z digest=sha256:9a7fbc4461e37a75939bc2186e0bcbda2e91bea7f3edfdc6f00e2e2eec171261

Observation b740735b-53ed-46ac-8c77-911256352e1a · outbound

This paper cites Training Compute-Optimal Large Language Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Training Compute-Optimal Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.508785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.508785Z digest=sha256:524183b493fd8a6c5ff742ae099c29c5983acd90c8f59f2acbdfc91be0822479

Observation 24872e0f-5eed-430f-a35c-01ad60f58d30 · outbound

This paper cites Language model compression with weighted low-rank factorization.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Language model compression with weighted low-rank factorization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.512142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.512142Z digest=sha256:470cbc15b74b497a0452ab8218a0e1b0ae6d08fafd0b09af726ac6e59e5cd65a

Observation 2eae92fb-3812-48f7-b5b7-4b9ef31c70f8 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters LoRA: Low-Rank Adaptation of Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.515423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.515423Z digest=sha256:ccc7bdde9b8a00b472a02cf6dbfc37b8a997193c13b9bb65dfa8783b1822fe84

Observation ec55e8c8-09c2-491b-bdac-9477daf7f5fe · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.518499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.518499Z digest=sha256:8f27ca9dba4c2d574a74833181ea55f131a097b76e08b81ea28dd58395ed2b00

Observation d397db10-374b-451a-a7ef-ff47182143d8 · outbound

This paper cites safetensors.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters safetensors

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.521757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.521757Z digest=sha256:2627249fb0e513ca0aa7af32fc2bca1fab53acba8f237d2adb76bab1c82dc972

Observation 9bebc7d0-9c82-473a-8ad4-c3ac6329add6 · outbound

This paper cites Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.524776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.524776Z digest=sha256:3c048bd30697ba2245f1dd3155ad7b62db314d184930f60edea7605ef6f8caa5

Observation dbea67cc-d6dc-4143-8b6c-4fdf9f4d9c7c · outbound

This paper cites Mistral 7B.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Mistral 7B

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.528254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.528254Z digest=sha256:762c0f0e93df2b534b4193c03c0bfbb2cea159a6207434d8294266f2fef1885f

Observation 1146dde1-570b-456e-b301-574fa10fb48e · outbound

This paper cites Multi-Domain Neural Machine Translation with Word-Level Adaptive Layer-wise Domain Mixing.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Multi-Domain Neural Machine Translation with Word-Level Adaptive Layer-wise Domain Mixing

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-08T13:44:01.008160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.531798Z digest=sha256:3941c0fcb64220617922267ef161e0c5f64d58e2d27a5cd3c1f60ea44efe370b

Observation d05d8b5b-34a8-4366-8d4e-c58a586a8720 · outbound

This paper cites Weld, and Luke Zettlemoyer.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Weld, and Luke Zettlemoyer

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.535267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.535267Z digest=sha256:eae92b30fcdcc9dc2cc7686cd0b3a8b2e58179ffdc2ddf1d4288390128c447b2

Observation 0872bdf5-f05e-4ce2-b24d-f984a37244e5 · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.538291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.538291Z digest=sha256:3d5d05247330f62c8161c5b53a7c9102a83faef738ed16b316341db51e3a6cbf

Observation df2d3abd-ca64-4747-977d-7f5f91b1e126 · outbound

This paper cites arxiv-math-instruct-50.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters arxiv-math-instruct-50

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.569680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.541523Z digest=sha256:91f264fbc81c3572a4c0e955f10263fd079a8ea2b3d55df2182dfca0ba2e82a1

Observation 82fa24fe-d6f5-4e80-a3b1-bf4e25e3196a · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters SqueezeLLM: Dense-and-Sparse Quantization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.544467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.544467Z digest=sha256:bfb896d10beaeb867f661d5a359226aa7490f0b1ea4b88e6fb2a4ea381065963

Observation 1dce0f8e-e1f0-4d76-84ba-eaf82b5e4aa3 · outbound

This paper cites Full Stack Optimization of Transformer Inference: a Survey.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Full Stack Optimization of Transformer Inference: a Survey

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.547913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.547913Z digest=sha256:fb4345219b1e1bfda18c9de3f9dbc741e62c3100719aa2412db6ee0fa5c7bdbe

Observation bbe9f86d-16ac-4127-8751-b4fb6bcc56f5 · outbound

This paper cites Speculative decoding with big little decoder.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Speculative decoding with big little decoder

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.559923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.551373Z digest=sha256:43a15678fa792567e4a1d6d05fa49254014f8acf62c923f7a162879855862721

Observation 99319401-96ec-478c-95e5-3543a552728b · outbound

This paper cites Reformer: The Efficient Transformer.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Reformer: The Efficient Transformer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.554454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.554454Z digest=sha256:55b33c84ccd5b3e98c0adba4a654d58e52e0c3a6afeab52f1359ad4da1b3a695

Observation 5605fd54-6203-4d08-841f-2d13198d77bf · outbound

This paper cites o pf, Yannic Kilcher, Dimitri von R \.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters o pf, Yannic Kilcher, Dimitri von R \

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.550621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.557832Z digest=sha256:c0793bcea5fdda9f72ddc8a0193b382ff077ea6f17e13f946752b4fbb9601084

Observation 448f8bc4-5c02-4654-a90c-a552d35cd2f8 · outbound

This paper cites LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.560727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.560727Z digest=sha256:da435909e19f836d9a5bdd1ad084dabaf85fbe060cc670886eb9ab7cabe04f0c

Observation 1ce2381e-10f6-46c4-8bc8-c2e94232da5f · outbound

This paper cites Losparse: Structured compression of large language models based on low-rank and sparse approximation.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Losparse: Structured compression of large language models based on low-rank and sparse approximation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.541682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.564027Z digest=sha256:ddb010886094d961c1b7068976d561bc48036728b8706169e19050c72a12c5f1

Observation 08840d56-194c-45a0-9a03-329bef4349d9 · outbound

This paper cites The Microsoft Toolkit of Multi-Task Deep Neural Networks for Natural Language Understanding.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters The Microsoft Toolkit of Multi-Task Deep Neural Networks for Natural Language Understanding

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-08T13:44:00.955687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.567152Z digest=sha256:53420ed9f3eb6481881b0355d2996214d65ce6d361df7796562ad687f20779b3

Observation 9bbb9f5b-f553-4d78-9461-5702cce38c02 · outbound

This paper cites LLM-QAT: Data-Free Quantization Aware Training for Large Language Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters LLM-QAT: Data-Free Quantization Aware Training for Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.570594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.570594Z digest=sha256:1587e4f200fd5b52c60d080b892efa1eaa0f61d547257d88b7767ea725a75952

Observation a1aac243-d292-44e6-8a38-1538e7bab617 · outbound

This paper cites MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.574242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.574242Z digest=sha256:919ae2d50ba447631160b46b176d160720519c41b06ac2954d1b8d3fcd3f96d6

Observation c7c1be55-0710-45b8-8297-1609f64d69ef · outbound

This paper cites Deja vu: Contextual sparsity for efficient llms at inference time.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Deja vu: Contextual sparsity for efficient llms at inference time

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.532044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.577593Z digest=sha256:cb6d362e1e42526f6288e3616e51650b1adc600237241f3e5235a9052215927e

Observation d63f4650-1139-4eed-848c-01ff37af7fe7 · outbound

This paper cites The flan collection: Designing data and methods for effective instruction tuning.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters The flan collection: Designing data and methods for effective instruction tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.580675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.580675Z digest=sha256:b068e657967e6de024a7d82946a4f6a3929727232d2181d9adc4c2426c40bd05

Observation f1f98760-770e-4885-b5f7-eb28ddd2e9b4 · outbound

This paper cites Fineweb-edu, May 2024.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Fineweb-edu, May 2024

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.518017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.583970Z digest=sha256:408dd371304c5751b3241d6d37ad1fbf3338cca1c65c29e23ed4f6f019fb95d9

Observation 029c6bc1-782e-44a0-aa33-f43bad5966ed · outbound

This paper cites Llm-pruner: On the structural pruning of large language models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Llm-pruner: On the structural pruning of large language models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.587134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.587134Z digest=sha256:5769d743a59212fa18f01fe921d4b9f6bb36d6c0d66cedbaa7210fcd909c93ad

Observation 4357665d-8a38-4c5f-ac05-eba6b630ff83 · outbound

This paper cites Peft: State-of-the-art parameter-efficient fine-tuning methods.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Peft: State-of-the-art parameter-efficient fine-tuning methods

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.590146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.590146Z digest=sha256:78649b63f94cfb621491aa21a72f2c32663a90ad9af54a8d4bfe0b65d8ffcb44

Observation b1b78aea-63f9-44a5-ac99-833b591b03f4 · outbound

This paper cites ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.593433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.593433Z digest=sha256:7d0778a3e0cbd5351fcf1b816ad56378518538952e9aa289f27f46f8d5d5112d

Observation 8c46f36c-57ce-453b-8ee9-57271195c5ea · outbound

This paper cites Orca: Progressive Learning from Complex Explanation Traces of GPT-4.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Orca: Progressive Learning from Complex Explanation Traces of GPT-4

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.598069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.598069Z digest=sha256:f9a448ab2f640f38a95e55dd97b710bb89711a84e37163fdc6de9b982d857bb2

Observation 2c0e37d9-602e-4b4e-b89a-2eae112783c5 · outbound

This paper cites Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.498518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.601574Z digest=sha256:5d0b56798c78e1727f5771aea250266c15f94036aea49735c2d124dba86baf7e

Observation cd3188e7-a885-4cf6-adbb-99ab332cd029 · outbound

This paper cites The lambada dataset, Aug 2016.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters The lambada dataset, Aug 2016

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.604742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.604742Z digest=sha256:977fb5a54d640fa21fcdb2dee7346a13d3fda3a44cd92fda139935cc3ab696d2

Observation 73e2852b-3fd7-438a-bbd9-86f2ff6cf807 · outbound

This paper cites Hovy, Pamela Forner, \'A lvaro Rodrigo, Richard F.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Hovy, Pamela Forner, \'A lvaro Rodrigo, Richard F

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.482413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.608043Z digest=sha256:88546ed64ed16bf2b419c47113bcfa872a1fccb8b9931c59d78c21b8e1315463

Observation 94a895b8-0d03-49ae-a33b-828a0dc37633 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.610889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.610889Z digest=sha256:077ba85950f7e9a21e88f36552bc77fea29fdf6309b50b1117ffae6cbd48ce03

Observation 9ae97c1d-32cb-4a28-bd85-2fe2b0b93c53 · outbound

This paper cites Instruction Tuning with GPT-4.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Instruction Tuning with GPT-4

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.614144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.614144Z digest=sha256:a059593f699fb0e40b6b2f6a8cbfa8c97fa4716e32997407070123d1c4bc6792

Observation 9f48c874-de2e-4c5d-af16-d8b5f722c1f3 · outbound

This paper cites Efficiently scaling transformer inference.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Efficiently scaling transformer inference

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.617414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.617414Z digest=sha256:cb54a306d4ea61c710ec8b26f9a38ea9c86db65cb6be873a6fb837da4eaf280b

Observation 05bb520b-de63-45e1-ac97-2cefe282825b · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.620603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.620603Z digest=sha256:a5073345d345ee7b7f3edb16dea1f3c03a4f487447fc78eaa64be3e0971a4905

Observation 7d118554-e8be-428f-b59f-ac40539283d7 · outbound

This paper cites FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.623850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.623850Z digest=sha256:8733e9bf084c5b24ae15c4f9e160b2f8687d9555191761e912eee2a65619f519

Observation 2bc13ed9-f377-440b-89eb-7d0e1d2795b8 · outbound

This paper cites S-LoRA: Serving Thousands of Concurrent LoRA Adapters.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters S-LoRA: Serving Thousands of Concurrent LoRA Adapters

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.627303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.627303Z digest=sha256:583a5178c896ec9db16af87df026dc0c69f92a1775c0562dda156720188f16f8

Observation 4429da89-7e1a-4f88-8de7-731eb9f950dc · outbound

This paper cites Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.630570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.630570Z digest=sha256:e409c280571700c0c29aa4236d2229b8bca2741f7346bb96a609102d10c5e280

Observation 79f7021f-7e3c-4f95-ad0f-c03563279973 · outbound

This paper cites Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.467507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.634440Z digest=sha256:9252a620fb4e9a6e7b7ebb620c35eb3da39abd27972410eddbe0e3beeb1b59c1

Observation d0863e9f-5c08-4595-a736-0465ac3e2d42 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters A Simple and Effective Pruning Approach for Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.638205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.638205Z digest=sha256:4e5a2193a1a91ae9c67365cacea73f031cadfb837fa624b3afcfc8fb923dc5c7

Observation 83636a88-52f1-4891-b66d-98e554e2fb56 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.641561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.641561Z digest=sha256:2a66da02f75a1136356c11896a71cba85439592579eeedb68ba978ec283bb09a

Observation aa25504b-1c32-4c06-8ab1-ac8be1fea274 · outbound

This paper cites KroneckerBERT: Learning Kronecker Decomposition for Pre-trained Language Models via Knowledge Distillation.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters KroneckerBERT: Learning Kronecker Decomposition for Pre-trained Language Models via Knowledge Distillation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.645377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.645377Z digest=sha256:b721c3ad0b3868e40ae43b59fbd7cb5cc49adbf3fd27283203b4884ecc0479fa

Observation ba86a792-cd7d-4434-8f78-c84ca18c9055 · outbound

This paper cites C ommonsense QA : A question answering challenge targeting commonsense knowledge.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters C ommonsense QA : A question answering challenge targeting commonsense knowledge

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.648906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.648906Z digest=sha256:fbbaa3cd02fe5282a25547b8fe63cff3bf345f330e483c476d964833ed876989

Observation 4414b662-01c4-436b-a618-9f0b6af1aae3 · outbound

This paper cites Multi-Domain Neural Machine Translation.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Multi-Domain Neural Machine Translation

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-08-08T13:44:00.837719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.652312Z digest=sha256:ba94793d1e706d50fcc21178add91e1be10e246df81d8c6bbe2bfaefc46ccc9c

Observation 6025f9a1-8fe8-4a9d-88e3-8abc8d210e5f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Gemini: A Family of Highly Capable Multimodal Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.655703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.655703Z digest=sha256:17f9026bac6e11ed07595a2d3c51984350e3094c75392c6ea78cb65af7042035

Observation 878a148a-0bcb-4fc8-86d7-4e755cb41522 · outbound

This paper cites Baby Llama: knowledge distillation from an ensemble of teachers trained on a small dataset with no performance penalty.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Baby Llama: knowledge distillation from an ensemble of teachers trained on a small dataset with no performance penalty

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.659142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.659142Z digest=sha256:55cf541498794fe355c1d535657979b1399801f3696da7aa32c397bbf1e1d650

Observation f3aa82dc-5707-4ac0-a5d0-6748380455f8 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.662881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.662881Z digest=sha256:50f76b5304a89e147bf2670aa68514717bf668b2a50333ac786ce7b0d568d00a

Observation a189eff1-a16f-4fbc-9d37-3025366e41d4 · outbound

This paper cites Smith, Iz Beltagy, and Hannaneh Hajishirzi.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Smith, Iz Beltagy, and Hannaneh Hajishirzi

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.666585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.666585Z digest=sha256:90c5907e0dadeeccc6db1fe62bccd4ddf5515df7a3fd252e181f66f9803a6480

Observation 6090a7a2-414f-4013-9fd5-43057ea2c880 · outbound

This paper cites Liu, and Matt Gardner.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Liu, and Matt Gardner

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.452340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.669764Z digest=sha256:77d013816a981f4ba968afd36593a2873fe1571ad39e0609bfb5a655746ba062

Observation 72d3a5b1-ca63-4a16-bac2-3fb80da30ac2 · outbound

This paper cites Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.672927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.672927Z digest=sha256:9bce97119ad9a3fd5a82f8f71f5c3a53e3e18803b98c3dcd5a10494bf78f5241

Observation 197c667f-0f91-4cb0-9932-05469fef679c · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.676353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.676353Z digest=sha256:063da2bdca44c46c80a1d4b31e052c776a70f047109a4611802d021964ae7dc1

Observation 398d67a9-e006-4337-9adf-55117488af82 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Smoothquant: Accurate and efficient post-training quantization for large language models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.680017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.680017Z digest=sha256:3bb3e1d892851823c4a4851653b18d9d81b375641bc7464249c3d011f27505e5

Observation 006da0f2-6bb7-4c61-bd0d-254ff8a5779e · outbound

This paper cites Wizard LM : Empowering large pre-trained language models to follow complex instructions.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Wizard LM : Empowering large pre-trained language models to follow complex instructions

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.683256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.683256Z digest=sha256:9b048b9292689f51478ab31dce84381471686930b06037d691fc24716655b109

Observation a40591c7-eb13-4dae-9979-8ad86ed6550b · outbound

This paper cites Zeroquant: Efficient and affordable post-training quantization for large-scale transformers.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Zeroquant: Efficient and affordable post-training quantization for large-scale transformers

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.432369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.686408Z digest=sha256:d4bc91bedb6b2f1f74e7080339d2e01e01816f36d8e862cbbf19a3b90b6e2b4d

Observation d552fbc1-bcdd-4087-80e6-6d7e01384118 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters TinyLlama: An Open-Source Small Language Model

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.689555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.689555Z digest=sha256:773555987bfc167dec81494b85f09b36f3d6078e3c60c3aa90364263a01d0b1a

Observation e2ada6c8-054c-4cf5-b556-dffa5ae0fd7f · outbound

This paper cites AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.692882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.692882Z digest=sha256:d91b52f05b0cde10306df9bc8918bda83b532c4d960f17db4677534e45ce4731

Observation 6a0cda3f-0f67-4f00-b8db-30140497589f · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters OPT: Open Pre-trained Transformer Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.696184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.696184Z digest=sha256:e6c8086a73607180eedb9d715717784e13b1333e5b344b9d024e35f5479ace55

Observation 788b02db-186b-4043-aaa4-f3b04a13b843 · outbound

This paper cites Adaptive-precision framework for sgd using deep q-learning.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Adaptive-precision framework for sgd using deep q-learning

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.423105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.699834Z digest=sha256:3b5e8a1c6a75d0bcdfa51492bebb68cadb3c4829505420ea3ff647ded6858d31

Observation f457a825-98a0-48cd-883c-08043957a28e · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.413628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.702933Z digest=sha256:ba71eecc352dbd6ea7dd5c05856d7075c9db013290c48abcd85a4b54a96bb5e3

Observation 77948001-69c7-4cd3-8931-ccdf9ca04e35 · outbound

This paper cites Lima: Less is more for alignment.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Lima: Less is more for alignment

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.705824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.705824Z digest=sha256:6aa9e923b1380074cca9b031bc6a3769ea74971f9db4f88c29eefa1039b0530e

Observation 1e7fd59a-1c84-4197-b6b6-2ef9c4a72fd9 · outbound

This paper cites write newline.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters write newline

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.708804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.708804Z digest=sha256:07b53fe945729de8fc60bd64b3469656be1ff69e97806039591018e531d3965b

Observation f32fecc0-8bc8-4ca4-ab83-82106412c486 · outbound

This paper cites @esa (Ref.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters @esa (Ref

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.712425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.712425Z digest=sha256:09e0e0704a734c672f63d9911f0b6359a4f12b8983f8f4357b6d3270e6976560

Observation 669dfaa8-092f-47d8-970c-c4cf8d21571f · outbound

This paper cites an unresolved cited work.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Unresolved cited work

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.717308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.717308Z digest=sha256:915c370f038df5f5454b96c6e2b8f31997ee7d9933a6b7ea4e7a0f070130e372

Observation 555edb9b-f77f-4b11-a6a6-97964e650b49 · outbound

This paper cites an unresolved cited work.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Unresolved cited work

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.720633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.720633Z digest=sha256:cecd5e1ac6a8038f35b9e972793b85ca28ade8a704ae387e44ecebd1d47b6499

Pith citing papers

No inbound Pith citation observations are available.