Pith. sign in

Paper Citation Record · LEDGER

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models

As of 9 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 7 inbound Pith citation observations for arXiv:2502.06663.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06663 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:48:53.872551Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:57:41.273776Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:58.173800Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f3ab860-3862-425a-b051-c27b1b780814 · outbound

This paper cites GPT-4 Technical Report.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.713646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.713646Z digest=sha256:b0fc88b8e5b84248799c1030ed60bd0881cfb1b9e10841acb54a33884c8335ab

Observation 798e0fec-30ae-4efe-b2e4-00e3b0bada86 · outbound

This paper cites an unresolved cited work.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:48:54.302037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:48:53.854946Z digest=sha256:2138835d0b52d82416f527a38ee57493cc64c92df3af8c3854d5bcf513621b6f

Observation 6eacaf97-e309-42b5-96b8-72f4521e9fc6 · outbound

This paper cites Once-for-All: Train One Network and Specialize it for Efficient Deployment.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Once-for-All: Train One Network and Specialize it for Efficient Deployment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.736258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.736258Z digest=sha256:2078a32a6cfdbddd924e20e825b908aadd6cf65b01007c0657a3caaef2dfbc61

Observation 1ae18473-803c-4e45-a69f-3f33a544c34e · outbound

This paper cites an unresolved cited work.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:48:54.254028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:48:53.869090Z digest=sha256:fc75d2d5a68068efdacf54ffba7b6031d15f83b1bb56a7071407e79a79a07551

Observation 4cfeaa97-fc74-46f0-8658-45dc07c9b08f · outbound

This paper cites an unresolved cited work.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:48:54.325593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:48:53.745087Z digest=sha256:fb722ffc0cfe50bb7337da50a8d4c8ed58c4656cde39ae35ef9b6b6c009e9af7

Observation 894b5e95-d578-428f-98b6-91591743197c · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.752765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.752765Z digest=sha256:4411a71b91ded7f18bfa777f93629d9999fddcaef019f57f090c925dbeafe025

Observation 190f5ca2-4292-4134-be93-cc6086e7ec93 · outbound

This paper cites Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.757177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.757177Z digest=sha256:4827345a1db95b4f20ca5647178202d64cc6d294c766ce468d05239e4e073cf1

Observation a9b68355-29d5-4fc4-950c-ae16daa6f3f1 · outbound

This paper cites The Llama 3 Herd of Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.761083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.761083Z digest=sha256:c0f13d9a944e8fff967cc3f6e52917a2ce642cc65ea4780fc010c501bc5229ce

Observation 34a1062a-1c98-4e6b-add1-0222eca57003 · outbound

This paper cites OLMo: Accelerating the Science of Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models OLMo: Accelerating the Science of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.764890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.764890Z digest=sha256:44ee74cc48c99e08284049ec59d1c003abac87f1676d11732e1a95e48cfdb0f8

Observation ece9b58d-50be-440a-b617-d68de8f4da27 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.768926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.768926Z digest=sha256:262bea250e24088d2278341de06eb81f0d3389928d3f7223565442e1b825d4b9

Observation ecda896b-6e75-443b-8092-fc93d81d32bc · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.780997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.780997Z digest=sha256:642ea64bca6a659321f9e1d6b01345817bd5b0c3c6bb2033b096709ffe4784e4

Observation fe5e4031-22cd-457a-a7e5-9ab88f5c360a · outbound

This paper cites Scaling Laws for Neural Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Scaling Laws for Neural Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.784871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.784871Z digest=sha256:f1d476eaa3cb796e9dcb2aacf4748cbc83c59dcb50b44f0e3f831d0b40826f71

Observation bab0a534-e073-49c7-ad5c-4fdaa397ab83 · outbound

This paper cites DistiLLM: Towards Streamlined Distillation for Large Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models DistiLLM: Towards Streamlined Distillation for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.789335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.789335Z digest=sha256:b36ea47cc556436df4caadcf7b70cc65a7af78d45521af31b5188802087a2212

Observation f487d1e6-c35c-409e-938a-f6a0537a3302 · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models DataComp-LM: In search of the next generation of training sets for language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.792864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.792864Z digest=sha256:8d9b8fa2c7630a74ef244177ec93f0d50fe71c3bc09b0473914fc0faefdca553

Observation f642e492-42a3-44e6-90c7-b93ca0ed93f5 · outbound

This paper cites DARTS: Differentiable Architecture Search.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models DARTS: Differentiable Architecture Search

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.796680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.796680Z digest=sha256:42e1255a5cee39c83a415fe323ff93b9d55fb76b452f747b60b7a02aa1c4798a

Observation b4e02bb4-9866-440f-9ff9-1f30885a9f7d · outbound

This paper cites MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.800690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.800690Z digest=sha256:232e32afd197a3d56a14d84f90a0450abcb32989d0013085bfa2a42d6d649c37

Observation 8a3a09ec-7afe-49f6-b34f-d893f9a8f504 · outbound

This paper cites KVPruner: Structural Pruning for Faster and Memory-Efficient Large Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models KVPruner: Structural Pruning for Faster and Memory-Efficient Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.804404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.804404Z digest=sha256:a05d24efb47561a05f40263fe38c538d5f9eb222524b0def60d7f1e8a1d2accb

Observation e8d8470a-9652-44c4-8e86-59b7f1957ee7 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.807846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.807846Z digest=sha256:b08b94a9273e14d60bb39d7c4c52308bd96a4d087e5737e22f59be5a93daf957

Observation 529cbc2c-0263-490d-93fe-2dfcf4ae0751 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models A Simple and Effective Pruning Approach for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.815448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.815448Z digest=sha256:a8cbfda0cbe04e8ae692cf9ddde7b34af82bbc9f36ba422810300e9bed1356a9

Observation 0c547a0a-7c88-4789-8cc9-243b8221fb93 · outbound

This paper cites PanGu-$\pi$ Pro:Rethinking Optimization and Architecture for Tiny Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models PanGu-$\pi$ Pro:Rethinking Optimization and Architecture for Tiny Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:48:53.984095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:48:53.819732Z digest=sha256:c8e17bd50a62be5f29f20541aa1ec7486c5f2b360f8ec0d5d2422e1f85d67e84

Observation fdeed727-1187-4614-b9eb-63ea10f9b19e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.823338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.823338Z digest=sha256:46305f82e9c6c9c6c29ab9ea7c8982e3304c80df32a7c6c64b3bea4492f72ae1

Observation e48a6e5c-c83a-4769-a84a-670bd1d3bd6b · outbound

This paper cites RedPajama: an Open Dataset for Training Large Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models RedPajama: an Open Dataset for Training Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.827249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.827249Z digest=sha256:4d655f77f3f76d14bbb72090f56ca98571228856098e45cf7cb56e0f8a9a75ff

Observation a5aba48e-cbcd-41a1-b271-16fb0f188a4e · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.831312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.831312Z digest=sha256:2ad968c6be2a4e3dc65749cd71f6120fcf66d4e08b16a7401de247115a7cfb73

Observation ef81d76d-59cd-4edf-84ca-4660fd1f7634 · outbound

This paper cites Nas-bert: Task-agnostic and adaptive-size bert com- pression with neural architecture search.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Nas-bert: Task-agnostic and adaptive-size bert com- pression with neural architecture search

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:48:54.313868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:48:53.835395Z digest=sha256:dde8aef62a5e0e2e095761e128874064c2e78a630de6fb8cc9e7e00ffb2c7772

Observation 72a9d564-1cbc-4249-bedd-9290e6313f56 · outbound

This paper cites Qwen2 Technical Report.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Qwen2 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.839396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.839396Z digest=sha256:1ec473dd4fe720b8d7c10e4b208a5a4559c1ace9f34c0858902b31a43d423a3e

Observation f2f6e10b-e5fa-4e2e-ae13-efc03a6cb1a6 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.843460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.843460Z digest=sha256:5586057666ba74a77ee22e990cc423858972cba5f22026d4124262b7f5786293

Observation a58e704a-733f-433e-a9ef-3ad4aa58dc74 · outbound

This paper cites LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.847426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.847426Z digest=sha256:85304fa741e6b5537f3928db54d6f44395a5ea887961745fec0bf12e96a258ec

Observation f6ad8578-5174-4626-a153-99844b7d91e6 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models OPT: Open Pre-trained Transformer Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.851478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.851478Z digest=sha256:9f863bce2c36597650e0981e2188c90906c164e4067f815191f70b1c1b7c6a1a

Observation 59a97f15-5739-43e6-a926-b375282da3d0 · outbound

This paper cites The OBD only uses the second-order term in Eq.9, which applied the diagonal of the Hessian matrix for approximate calculation.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models The OBD only uses the second-order term in Eq.9, which applied the diagonal of the Hessian matrix for approximate calculation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:48:54.278528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:48:53.861524Z digest=sha256:bb4ddb207de1d510542eeeadb5e581fad979b8a5f56af4a46f2412f74d91f83f

Observation 14fa2ed5-5772-49d4-9467-e63cb8ce02b0 · outbound

This paper cites In the case of GQA, cluster attention can be obtained through pruning.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models In the case of GQA, cluster attention can be obtained through pruning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:48:54.266579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:48:53.865131Z digest=sha256:df9d3867bae195907e3f0de8bb661fec7a0e26c8d7e1229790c8c7b8b176234c

Observation 020bf5b9-ccaa-4e71-81a8-f8c185639cdd · outbound

This paper cites • Common Sense Reasoning: Follow most of recent works (Xia et al., 2023; Ma et al., 2023; Li et al., 2024), we apply the widely used lm-evaluation-harness package (Gao et al.,.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models • Common Sense Reasoning: Follow most of recent works (Xia et al., 2023; Ma et al., 2023; Li et al., 2024), we apply the widely used lm-evaluation-harness package (Gao et al.,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:48:54.243326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:48:53.872551Z digest=sha256:ae26ad628a650cbafe973bf056506f04eb7e1ead4692e5380a3bc8c36b05c54e

Observation 656c0315-c6e3-49f5-99ca-e62948459046 · outbound

This paper cites In EfficientLLM, the pruning ratio of 13 EfficientLLM hidden-size is smaller than attention heads and FFN intermediate channels driven by saliency.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models In EfficientLLM, the pruning ratio of 13 EfficientLLM hidden-size is smaller than attention heads and FFN intermediate channels driven by saliency

Reference 1989

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:48:54.290343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:48:53.858277Z digest=sha256:65e634a10943d3c4aca7dba0814c7bb59517ee0e49f3f9586be0f829581a4217

Observation 9124714c-a921-4b6e-989f-2aab1e64f92e · outbound

This paper cites Measuring Massive Multitask Language Understanding.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Measuring Massive Multitask Language Understanding

Reference 1993

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.776783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.776783Z digest=sha256:1923bb3f51e0a2b7cde2f97e650b75df00052436cc301784e2b108befb3a8eee

Observation cecafcd9-8a91-4ee3-a45d-bd68ef53d76c · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.748682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.748682Z digest=sha256:67a7364540d1385f5c583525d2527419a8d6628bbe5a6e5da61770f62739b11c

Observation 6e0ee9c7-a947-4b34-864f-107dcb7c7a2c · outbound

This paper cites LoRAShear: Efficient Large Language Model Structured Pruning and Knowledge Recovery.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models LoRAShear: Efficient Large Language Model Structured Pruning and Knowledge Recovery

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.740780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.740780Z digest=sha256:cd2aa7815669ecb46ef1bce601869526a14506522880d1069beb32084164ec4e

Observation 7287ffb9-d81b-4697-98f5-767a630b1420 · outbound

This paper cites GPT-NeoX-20B: An Open-Source Autoregressive Language Model.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.727799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.727799Z digest=sha256:5ea6141eec2f3894309829d69b44ad1662d9afb80a0551e1fdbdfaf113a51d9a

Observation 313ce8fb-7d42-4e94-b702-473a0da0fd46 · outbound

This paper cites LLM Pruning and Distillation in Practice: The Minitron Approach.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models LLM Pruning and Distillation in Practice: The Minitron Approach

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.811446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.811446Z digest=sha256:7d1f83a1c73555685301b87e246987944ae8d9ff99a6939ecce83687621f6988

Observation 8e7ef76d-b74f-4dce-a9a0-03e4a6acc423 · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.731978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.731978Z digest=sha256:0f6e44fa9a0e1004b98ebdfffa88e898b61d6f837fe3c4d78cb705949d894aa3

Observation e755fc7b-d8bf-40da-8d09-6ad0a50c03e8 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.718637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.718637Z digest=sha256:35965b771f6d1023789ab621aef246cb690dedc4431447ac7d7bfa496b88b2e0

Observation 4398016a-ddfb-455e-8aa3-4384cb73cd52 · outbound

This paper cites SliceGPT: Compress Large Language Models by Deleting Rows and Columns.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models SliceGPT: Compress Large Language Models by Deleting Rows and Columns

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.723438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.723438Z digest=sha256:848508a0ac5e5d8f2e0e61ff58d6aa6a23fe169210548a007b2e52e27ea15ada

Observation b3876975-b4dc-4fce-8d8a-cbd91ebb2a7e · outbound

This paper cites Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.772787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.772787Z digest=sha256:8c2f389a4310f8c6b9356c09102a3a507ff92c94f5b32371b3d6877b2b6052c6

Pith citing papers

Observation ccdee7e0-425a-4450-bed7-d00ea8d62d8b · inbound

Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation cites this paper.

Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:57:41.273776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:57:41.273776Z digest=sha256:2bab39a7975633ed0cf20743c6a55cf3a06f0d5c01fab112cc331b189e40d4ed

Observation fe5c8862-117c-40f9-a117-ba5ddff78986 · inbound

PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models cites this paper.

PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:05.522354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:51:05.522354Z digest=sha256:fbdc7b0ac87f7116b58cf5c58e22f699f5541ad195c3f3815ede9ffceb39ca7a

Observation e24c69d8-d99b-4368-9e81-201e45a515d0 · inbound

Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges cites this paper.

Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models

Reference 121

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:48.293267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:48.293267Z digest=sha256:0828058230b4cef83cf437f2957c8e968951c0247822818033a573cdf312b740

Observation 37eedbe6-45ce-4eac-81ba-75c2ae95c56e · inbound

EGGS-PTP: An Expander-Graph Guided Structured Post-training Pruning Method for Large Language Models cites this paper.

EGGS-PTP: An Expander-Graph Guided Structured Post-training Pruning Method for Large Language Models EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T21:08:18.721011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:08:18.721011Z digest=sha256:b337bf14dc832242c85e6a1a70ab732079fbfa00ab2d38bdfa7b55ec5dfebd48

Observation e73a38b8-909a-41d9-ab6f-06ffc7f20578 · inbound

SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba cites this paper.

SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:46:12.425193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T09:44:53.290259Z digest=sha256:b23f11dc3c1c00c0c9333ac544f8aeaf09b60bdb4359d07fe555c2177a10ce1f

Observation ad8f0973-8cfa-46fa-ac74-2c6116ae904b · inbound

LongSpike: Fractional Order Spiking State Space Models for Efficient Long Sequence Learning cites this paper.

LongSpike: Fractional Order Spiking State Space Models for Efficient Long Sequence Learning EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:48:21.554743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T07:26:58.608421Z digest=sha256:e1716c4376270079b0e87fff8f9c89418cdc4e274d3fef5c3e4d3ea9c21c5b68

Observation b141b3e0-d1e4-420d-a508-b1237e209ec6 · inbound

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation cites this paper.

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:58.175300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T00:19:49.071495Z digest=sha256:abec4881459d92aa5e3a7a65307a2b8ed35296efd5eaaccee01ff2ac0bdfc318