Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:06:41.062180Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2506.03510.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:06:41.062180Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5556fa6b-2f0c-49df-abfd-fe95bc6ab418 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Slicegpt: Compress large language models by deleting rows and columns
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation aa1fee7e-3982-41fb-ad3d-30f40759b92d · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information • ARC-Challenge and ARC-Easy[Clark et al., 2018] are datasets composed of grade-school level multiple-choice science questions
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e01af565-2563-4ba0-ac88-66611601c497 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Pea-kd: Parameter-efficient and accurate knowledge distillation on bert
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 38237da4-63d0-476f-9c26-bc362fbc1804 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information This demonstrates that five candidates are sufficient to find the sublayer with the least importance score
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 662f402c-6d1d-4514-8715-854e9007b60e · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Sparsegpt: Massive language models can be accurately pruned in one-shot
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7a9db158-1b3b-4329-bf3c-fedc0359d932 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information OPTQ: Accurate quantization for generative pre-trained transformers
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0dd00334-b931-47a4-a8d4-f03df3027c78 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Lazyllm: Dynamic token pruning for efficient long context llm inference
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9c5207f5-51a4-4bd3-85b5-6cedc6725560 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information A frame- work for few-shot language model evaluation, 12
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 22e84dfc-d61e-4113-b2c8-148e3b09b1f1 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b75d018c-4b48-4a64-b762-152809afba74 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Falcon: lightweight and accurate convo- lution based on depthwise separable convolution
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 90598526-0a8a-4c41-b45d-74f3b64d6507 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 61086d0b-0683-4ab0-8507-943d620cd40c · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation dfd76e9a-163d-41eb-9838-4cd5703574b0 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5235280a-ba91-4152-8169-b2c466d0e57f · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information DDK: distilling domain knowledge for efficient large language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5634f1d0-04b7-40a3-bf6b-b7e174af23a0 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Llm-pruner: On the structural pruning of large lan- guage models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bbaf37f5-70f7-46e7-9ad7-22660095dc21 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Shortgpt: Layers in large language models are more redundant than you expect
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 12d06413-1a80-4661-852a-7e0e5163877e · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Pointer sentinel mixture models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0f0aae67-aac6-4589-b922-b40e8174a583 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Sensimix: Sensitivity-aware 8-bit index & 1-bit value mixed precision quantization for bert compression
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 69e25463-b4f2-48c4-a910-89cb9613309a · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Mixture-of-depths: Dynamically al- locating compute in transformer-based language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2e88c0cc-9fc4-499e-8b3a-59aee2d02567 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Exaone 3.0 7.8 b instruction tuned language model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8fa26fd2-3181-48d5-b2bb-ec9cf99d3a37 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Confident adaptive language modeling
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 010dc779-83e3-4e6f-acb5-d19675455bbf · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Omniquant: Omnidi- rectionally calibrated quantization for large language mod- els
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4b1d685d-2559-4158-89a4-b03b97987496 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Sleb: Streamlining llms through redundancy verification and elimination of transformer blocks
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2140cf69-46e2-4cc5-b1dc-37371f098a7d · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information A simple and effective pruning approach for large language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 97329a70-27f7-43d8-870c-50f274ae52ed · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Gemini: a family of highly capable multimodal models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 472f7a99-b653-4b6b-bb98-fe8f39d00ee4 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Accelerating llama infer- ence by enabling intermediate layer decoding via instruc- tion tuning with lite
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b6ec1720-b2c9-45df-8ca1-06f156af4020 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Attention is all you need
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 09eba7dd-b119-43f7-a49d-34a9e96886ab · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Outlier weighed layerwise sparsity (OWL): A missing secret sauce for pruning llms to high sparsity
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 10c44411-5b20-437a-9bd8-69aab4373dad · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Knowledge extraction with no observable data
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b9686ceb-7242-4416-a864-7bc9aead71ef · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Hellaswag: Can a ma- chine really finish your sentence? arXiv,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e09bf8af-4cc2-4796-a241-0ad8ba18e944 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Opt: Open pre-trained transformer language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e680e9dd-125f-4e8c-99b4-24799b1a1e9d · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Blockpruner: Fine- grained pruning for large language models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a6c6d7c0-a273-46d9-9568-efd467e9c860 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information LLM-Pruner cannot prune Llama-2 70B and Llama-3 8B, 70B since it does not support group query attention
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 34f0ca9a-7111-4823-8ce8-3ce34f438b41 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information We use a single GPU to evaluate Llama-2 7B, 13B and Llama-3 8B
Reference 1024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a8a4870c-8c8c-4888-b5eb-97a159a6e20b · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Qa-lora: Quantization-aware low-rank adaptation of large language models
Reference 2017
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6208b8cd-86ff-4936-a3f7-bf9d8d13b290 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Boolq: Exploring the surprising diffi- culty of natural yes/no questions
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c3660369-e8ee-4ff8-ba96-58a416a8c0a0 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information The llama 3 herd of models
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 262e6f9e-725a-4f8c-b8cd-7eca616d768b · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Language models are few-shot learners
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 76a1307c-49d8-4f7d-bccd-a6c255ac0ab9 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Flexround: Learnable rounding based on element-wise division for post-training quantiza- tion
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 75c45ef0-2e10-403a-b356-6c4eade2493a · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Palm: Scaling language modeling with pathways
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ee1a10d3-c294-4b29-834f-e99a59aef860 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Think you have solved question answer- ing? try arc, the ai2 reasoning challenge
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 08a4932e-0ad1-4204-ae89-03e97ba2dec3 · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Piqa: Reasoning about physical commonsense in natural language
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 83424917-ff2e-41e3-bc0e-ba0caa9b00ec · outbound
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Distillm: Towards streamlined distil- lation for large language models
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
No inbound Pith citation observations are available.