Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:48:53.872551Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 7 inbound Pith citation observations for arXiv:2502.06663.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:48:53.872551Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:57:41.273776Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T16:39:58.173800Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8f3ab860-3862-425a-b051-c27b1b780814 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 798e0fec-30ae-4efe-b2e4-00e3b0bada86 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6eacaf97-e309-42b5-96b8-72f4521e9fc6 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Once-for-All: Train One Network and Specialize it for Efficient Deployment
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ae18473-803c-4e45-a69f-3f33a544c34e · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4cfeaa97-fc74-46f0-8658-45dc07c9b08f · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 894b5e95-d578-428f-98b6-91591743197c · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 190f5ca2-4292-4134-be93-cc6086e7ec93 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9b68355-29d5-4fc4-950c-ae16daa6f3f1 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models The Llama 3 Herd of Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34a1062a-1c98-4e6b-add1-0222eca57003 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models OLMo: Accelerating the Science of Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ece9b58d-50be-440a-b617-d68de8f4da27 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecda896b-6e75-443b-8092-fc93d81d32bc · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe5e4031-22cd-457a-a7e5-9ab88f5c360a · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Scaling Laws for Neural Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bab0a534-e073-49c7-ad5c-4fdaa397ab83 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models DistiLLM: Towards Streamlined Distillation for Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f487d1e6-c35c-409e-938a-f6a0537a3302 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models DataComp-LM: In search of the next generation of training sets for language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f642e492-42a3-44e6-90c7-b93ca0ed93f5 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models DARTS: Differentiable Architecture Search
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4e02bb4-9866-440f-9ff9-1f30885a9f7d · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a3a09ec-7afe-49f6-b34f-d893f9a8f504 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models KVPruner: Structural Pruning for Faster and Memory-Efficient Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8d8470a-9652-44c4-8e86-59b7f1957ee7 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 529cbc2c-0263-490d-93fe-2dfcf4ae0751 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models A Simple and Effective Pruning Approach for Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c547a0a-7c88-4789-8cc9-243b8221fb93 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models PanGu-$\pi$ Pro:Rethinking Optimization and Architecture for Tiny Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fdeed727-1187-4614-b9eb-63ea10f9b19e · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e48a6e5c-c83a-4769-a84a-670bd1d3bd6b · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models RedPajama: an Open Dataset for Training Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5aba48e-cbcd-41a1-b271-16fb0f188a4e · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef81d76d-59cd-4edf-84ca-4660fd1f7634 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Nas-bert: Task-agnostic and adaptive-size bert com- pression with neural architecture search
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 72a9d564-1cbc-4249-bedd-9290e6313f56 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Qwen2 Technical Report
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2f6e10b-e5fa-4e2e-ae13-efc03a6cb1a6 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a58e704a-733f-433e-a9ef-3ad4aa58dc74 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6ad8578-5174-4626-a153-99844b7d91e6 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models OPT: Open Pre-trained Transformer Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59a97f15-5739-43e6-a926-b375282da3d0 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models The OBD only uses the second-order term in Eq.9, which applied the diagonal of the Hessian matrix for approximate calculation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 14fa2ed5-5772-49d4-9467-e63cb8ce02b0 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models In the case of GQA, cluster attention can be obtained through pruning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 020bf5b9-ccaa-4e71-81a8-f8c185639cdd · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models • Common Sense Reasoning: Follow most of recent works (Xia et al., 2023; Ma et al., 2023; Li et al., 2024), we apply the widely used lm-evaluation-harness package (Gao et al.,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 656c0315-c6e3-49f5-99ca-e62948459046 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models In EfficientLLM, the pruning ratio of 13 EfficientLLM hidden-size is smaller than attention heads and FFN intermediate channels driven by saliency
Reference 1989
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9124714c-a921-4b6e-989f-2aab1e64f92e · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Measuring Massive Multitask Language Understanding
Reference 1993
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cecafcd9-8a91-4ee3-a45d-bd68ef53d76c · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e0ee9c7-a947-4b34-864f-107dcb7c7a2c · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models LoRAShear: Efficient Large Language Model Structured Pruning and Knowledge Recovery
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7287ffb9-d81b-4697-98f5-767a630b1420 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models GPT-NeoX-20B: An Open-Source Autoregressive Language Model
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 313ce8fb-7d42-4e94-b702-473a0da0fd46 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models LLM Pruning and Distillation in Practice: The Minitron Approach
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e7ef76d-b74f-4dce-a9a0-03e4a6acc423 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e755fc7b-d8bf-40da-8d09-6ad0a50c03e8 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4398016a-ddfb-455e-8aa3-4384cb73cd52 · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models SliceGPT: Compress Large Language Models by Deleting Rows and Columns
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3876975-b4dc-4fce-8d8a-cbd91ebb2a7e · outbound
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccdee7e0-425a-4450-bed7-d00ea8d62d8b · inbound
Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe5c8862-117c-40f9-a117-ba5ddff78986 · inbound
PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e24c69d8-d99b-4368-9e81-201e45a515d0 · inbound
Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models
Reference 121
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37eedbe6-45ce-4eac-81ba-75c2ae95c56e · inbound
EGGS-PTP: An Expander-Graph Guided Structured Post-training Pruning Method for Large Language Models EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e73a38b8-909a-41d9-ab6f-06ffc7f20578 · inbound
SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ad8f0973-8cfa-46fa-ac74-2c6116ae904b · inbound
LongSpike: Fractional Order Spiking State Space Models for Efficient Long Sequence Learning EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b141b3e0-d1e4-420d-a508-b1237e209ec6 · inbound
IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.