Pith. sign in

Paper Citation Record · LEDGER

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models

As of 10 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 7 inbound Pith citation observations for arXiv:2502.06663.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06663 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:48:53.872551Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:57:41.273776Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:58.173800Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f3ab860-3862-425a-b051-c27b1b780814 · outbound

This paper cites GPT-4 Technical Report.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.713646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.713646Z digest=sha256:b0fc88b8e5b84248799c1030ed60bd0881cfb1b9e10841acb54a33884c8335ab

Observation 798e0fec-30ae-4efe-b2e4-00e3b0bada86 · outbound

This paper cites an unresolved cited work.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:48:54.302037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:48:53.854946Z digest=sha256:90a941197756671f48cd9d874b83b09ce7a37cf76a7b1f985b592cf16fff3100

Observation 6eacaf97-e309-42b5-96b8-72f4521e9fc6 · outbound

This paper cites Once-for-All: Train One Network and Specialize it for Efficient Deployment.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Once-for-All: Train One Network and Specialize it for Efficient Deployment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.736258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.736258Z digest=sha256:a175fbadd20b2e7be309c61bbb153192b027009233fd5ce9567393fa56313c68

Observation 1ae18473-803c-4e45-a69f-3f33a544c34e · outbound

This paper cites an unresolved cited work.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:48:54.254028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:48:53.869090Z digest=sha256:d56bf1dc3ef7f5aafddcc01fab2caaee3cfd595906db9f30086affd22ccae1f1

Observation 4cfeaa97-fc74-46f0-8658-45dc07c9b08f · outbound

This paper cites an unresolved cited work.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:48:54.325593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:48:53.745087Z digest=sha256:acf6d898e2b8d192bf78cd9eb9c0d87adf1eeef0af5c103e8035c986e6307399

Observation 894b5e95-d578-428f-98b6-91591743197c · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.752765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.752765Z digest=sha256:4411a71b91ded7f18bfa777f93629d9999fddcaef019f57f090c925dbeafe025

Observation 190f5ca2-4292-4134-be93-cc6086e7ec93 · outbound

This paper cites Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.757177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.757177Z digest=sha256:7b62d613cbe3482cc7a9ad46ecde0676acd794e9f19873328dc7b61422645466

Observation a9b68355-29d5-4fc4-950c-ae16daa6f3f1 · outbound

This paper cites The Llama 3 Herd of Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.761083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.761083Z digest=sha256:c0f13d9a944e8fff967cc3f6e52917a2ce642cc65ea4780fc010c501bc5229ce

Observation 34a1062a-1c98-4e6b-add1-0222eca57003 · outbound

This paper cites OLMo: Accelerating the Science of Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models OLMo: Accelerating the Science of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.764890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.764890Z digest=sha256:44ee74cc48c99e08284049ec59d1c003abac87f1676d11732e1a95e48cfdb0f8

Observation ece9b58d-50be-440a-b617-d68de8f4da27 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.768926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.768926Z digest=sha256:262bea250e24088d2278341de06eb81f0d3389928d3f7223565442e1b825d4b9

Observation ecda896b-6e75-443b-8092-fc93d81d32bc · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.780997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.780997Z digest=sha256:1d9046bfcb83aff8a2f8b4b12a2da187995ca4515cf46a3c332b98d13afd8950

Observation fe5e4031-22cd-457a-a7e5-9ab88f5c360a · outbound

This paper cites Scaling Laws for Neural Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Scaling Laws for Neural Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.784871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.784871Z digest=sha256:f1d476eaa3cb796e9dcb2aacf4748cbc83c59dcb50b44f0e3f831d0b40826f71

Observation bab0a534-e073-49c7-ad5c-4fdaa397ab83 · outbound

This paper cites DistiLLM: Towards Streamlined Distillation for Large Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models DistiLLM: Towards Streamlined Distillation for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.789335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.789335Z digest=sha256:b36ea47cc556436df4caadcf7b70cc65a7af78d45521af31b5188802087a2212

Observation f487d1e6-c35c-409e-938a-f6a0537a3302 · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models DataComp-LM: In search of the next generation of training sets for language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.792864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.792864Z digest=sha256:8d9b8fa2c7630a74ef244177ec93f0d50fe71c3bc09b0473914fc0faefdca553

Observation f642e492-42a3-44e6-90c7-b93ca0ed93f5 · outbound

This paper cites DARTS: Differentiable Architecture Search.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models DARTS: Differentiable Architecture Search

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.796680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.796680Z digest=sha256:4fab685d3061dc377e46993ed9d953af4cccb5ac5596dae74f82784ee12d92c8

Observation b4e02bb4-9866-440f-9ff9-1f30885a9f7d · outbound

This paper cites MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.800690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.800690Z digest=sha256:232e32afd197a3d56a14d84f90a0450abcb32989d0013085bfa2a42d6d649c37

Observation 8a3a09ec-7afe-49f6-b34f-d893f9a8f504 · outbound

This paper cites KVPruner: Structural Pruning for Faster and Memory-Efficient Large Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models KVPruner: Structural Pruning for Faster and Memory-Efficient Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.804404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.804404Z digest=sha256:a05d24efb47561a05f40263fe38c538d5f9eb222524b0def60d7f1e8a1d2accb

Observation e8d8470a-9652-44c4-8e86-59b7f1957ee7 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.807846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.807846Z digest=sha256:b08b94a9273e14d60bb39d7c4c52308bd96a4d087e5737e22f59be5a93daf957

Observation 529cbc2c-0263-490d-93fe-2dfcf4ae0751 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models A Simple and Effective Pruning Approach for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.815448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.815448Z digest=sha256:a8cbfda0cbe04e8ae692cf9ddde7b34af82bbc9f36ba422810300e9bed1356a9

Observation 0c547a0a-7c88-4789-8cc9-243b8221fb93 · outbound

This paper cites PanGu-$\pi$ Pro:Rethinking Optimization and Architecture for Tiny Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models PanGu-$\pi$ Pro:Rethinking Optimization and Architecture for Tiny Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:48:53.984095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:48:53.819732Z digest=sha256:711a05263084de0391da1feddbd935667905bbb508717d262f64138309117fd9

Observation fdeed727-1187-4614-b9eb-63ea10f9b19e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.823338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.823338Z digest=sha256:46305f82e9c6c9c6c29ab9ea7c8982e3304c80df32a7c6c64b3bea4492f72ae1

Observation e48a6e5c-c83a-4769-a84a-670bd1d3bd6b · outbound

This paper cites RedPajama: an Open Dataset for Training Large Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models RedPajama: an Open Dataset for Training Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.827249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.827249Z digest=sha256:4d655f77f3f76d14bbb72090f56ca98571228856098e45cf7cb56e0f8a9a75ff

Observation a5aba48e-cbcd-41a1-b271-16fb0f188a4e · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.831312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.831312Z digest=sha256:2ad968c6be2a4e3dc65749cd71f6120fcf66d4e08b16a7401de247115a7cfb73

Observation ef81d76d-59cd-4edf-84ca-4660fd1f7634 · outbound

This paper cites Nas-bert: Task-agnostic and adaptive-size bert com- pression with neural architecture search.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Nas-bert: Task-agnostic and adaptive-size bert com- pression with neural architecture search

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:48:54.313868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:48:53.835395Z digest=sha256:8f2caf4da1a65cc8bb11c2197ee3b5c4703ee3ce4906e87772a6c37730c49fe1

Observation 72a9d564-1cbc-4249-bedd-9290e6313f56 · outbound

This paper cites Qwen2 Technical Report.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Qwen2 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.839396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.839396Z digest=sha256:1ec473dd4fe720b8d7c10e4b208a5a4559c1ace9f34c0858902b31a43d423a3e

Observation f2f6e10b-e5fa-4e2e-ae13-efc03a6cb1a6 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.843460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.843460Z digest=sha256:5586057666ba74a77ee22e990cc423858972cba5f22026d4124262b7f5786293

Observation a58e704a-733f-433e-a9ef-3ad4aa58dc74 · outbound

This paper cites LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.847426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.847426Z digest=sha256:b04e34636077bb5c789f588b4ebe99fe136fd1782ed14f84f9b6a5a35aab04da

Observation f6ad8578-5174-4626-a153-99844b7d91e6 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models OPT: Open Pre-trained Transformer Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.851478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.851478Z digest=sha256:9f863bce2c36597650e0981e2188c90906c164e4067f815191f70b1c1b7c6a1a

Observation 59a97f15-5739-43e6-a926-b375282da3d0 · outbound

This paper cites The OBD only uses the second-order term in Eq.9, which applied the diagonal of the Hessian matrix for approximate calculation.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models The OBD only uses the second-order term in Eq.9, which applied the diagonal of the Hessian matrix for approximate calculation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:48:54.278528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:48:53.861524Z digest=sha256:1d84eb9be009e0d7da469f77f05db0fac68923e1b87010b743cb2caee9e16b92

Observation 14fa2ed5-5772-49d4-9467-e63cb8ce02b0 · outbound

This paper cites In the case of GQA, cluster attention can be obtained through pruning.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models In the case of GQA, cluster attention can be obtained through pruning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:48:54.266579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:48:53.865131Z digest=sha256:3c88b5c831bbd7b1cd55e9f0f5ec3d269a3962d10eb39fdb793f2a9c96e13473

Observation 020bf5b9-ccaa-4e71-81a8-f8c185639cdd · outbound

This paper cites • Common Sense Reasoning: Follow most of recent works (Xia et al., 2023; Ma et al., 2023; Li et al., 2024), we apply the widely used lm-evaluation-harness package (Gao et al.,.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models • Common Sense Reasoning: Follow most of recent works (Xia et al., 2023; Ma et al., 2023; Li et al., 2024), we apply the widely used lm-evaluation-harness package (Gao et al.,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:48:54.243326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:48:53.872551Z digest=sha256:a836d3c2fc14c67c3d3f4c0c68a9f5b080684c8059c6ff80ddab619a1605e3bd

Observation 656c0315-c6e3-49f5-99ca-e62948459046 · outbound

This paper cites In EfficientLLM, the pruning ratio of 13 EfficientLLM hidden-size is smaller than attention heads and FFN intermediate channels driven by saliency.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models In EfficientLLM, the pruning ratio of 13 EfficientLLM hidden-size is smaller than attention heads and FFN intermediate channels driven by saliency

Reference 1989

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:48:54.290343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:48:53.858277Z digest=sha256:4c794f5e5cf4b991e415867040d06f85c072b716a6c513e11bf0303506770b66

Observation 9124714c-a921-4b6e-989f-2aab1e64f92e · outbound

This paper cites Measuring Massive Multitask Language Understanding.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Measuring Massive Multitask Language Understanding

Reference 1993

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.776783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.776783Z digest=sha256:dbfc395df0a5f8433ba23fe3d6c398f5f6c3c17a90af401249e997eb04be1385

Observation cecafcd9-8a91-4ee3-a45d-bd68ef53d76c · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.748682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.748682Z digest=sha256:67a7364540d1385f5c583525d2527419a8d6628bbe5a6e5da61770f62739b11c

Observation 6e0ee9c7-a947-4b34-864f-107dcb7c7a2c · outbound

This paper cites LoRAShear: Efficient Large Language Model Structured Pruning and Knowledge Recovery.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models LoRAShear: Efficient Large Language Model Structured Pruning and Knowledge Recovery

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.740780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.740780Z digest=sha256:cd2aa7815669ecb46ef1bce601869526a14506522880d1069beb32084164ec4e

Observation 7287ffb9-d81b-4697-98f5-767a630b1420 · outbound

This paper cites GPT-NeoX-20B: An Open-Source Autoregressive Language Model.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.727799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.727799Z digest=sha256:5ea6141eec2f3894309829d69b44ad1662d9afb80a0551e1fdbdfaf113a51d9a

Observation 313ce8fb-7d42-4e94-b702-473a0da0fd46 · outbound

This paper cites LLM Pruning and Distillation in Practice: The Minitron Approach.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models LLM Pruning and Distillation in Practice: The Minitron Approach

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.811446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.811446Z digest=sha256:7d1f83a1c73555685301b87e246987944ae8d9ff99a6939ecce83687621f6988

Observation 8e7ef76d-b74f-4dce-a9a0-03e4a6acc423 · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.731978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.731978Z digest=sha256:0f6e44fa9a0e1004b98ebdfffa88e898b61d6f837fe3c4d78cb705949d894aa3

Observation e755fc7b-d8bf-40da-8d09-6ad0a50c03e8 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.718637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.718637Z digest=sha256:35965b771f6d1023789ab621aef246cb690dedc4431447ac7d7bfa496b88b2e0

Observation 4398016a-ddfb-455e-8aa3-4384cb73cd52 · outbound

This paper cites SliceGPT: Compress Large Language Models by Deleting Rows and Columns.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models SliceGPT: Compress Large Language Models by Deleting Rows and Columns

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.723438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.723438Z digest=sha256:848508a0ac5e5d8f2e0e61ff58d6aa6a23fe169210548a007b2e52e27ea15ada

Observation b3876975-b4dc-4fce-8d8a-cbd91ebb2a7e · outbound

This paper cites Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.772787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.772787Z digest=sha256:8c2f389a4310f8c6b9356c09102a3a507ff92c94f5b32371b3d6877b2b6052c6

Pith citing papers

Observation ccdee7e0-425a-4450-bed7-d00ea8d62d8b · inbound

Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation cites this paper.

Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:57:41.273776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:57:41.273776Z digest=sha256:2bab39a7975633ed0cf20743c6a55cf3a06f0d5c01fab112cc331b189e40d4ed

Observation fe5c8862-117c-40f9-a117-ba5ddff78986 · inbound

PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models cites this paper.

PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:05.522354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:51:05.522354Z digest=sha256:fbdc7b0ac87f7116b58cf5c58e22f699f5541ad195c3f3815ede9ffceb39ca7a

Observation e24c69d8-d99b-4368-9e81-201e45a515d0 · inbound

Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges cites this paper.

Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models

Reference 121

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:48.293267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:48.293267Z digest=sha256:0828058230b4cef83cf437f2957c8e968951c0247822818033a573cdf312b740

Observation 37eedbe6-45ce-4eac-81ba-75c2ae95c56e · inbound

EGGS-PTP: An Expander-Graph Guided Structured Post-training Pruning Method for Large Language Models cites this paper.

EGGS-PTP: An Expander-Graph Guided Structured Post-training Pruning Method for Large Language Models EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T21:08:18.721011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:08:18.721011Z digest=sha256:dc34c6f90f2146cbf363e379aa540efe89085bbe68b175543e7dd11ac7037198

Observation e73a38b8-909a-41d9-ab6f-06ffc7f20578 · inbound

SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba cites this paper.

SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:46:12.425193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T09:44:53.290259Z digest=sha256:13e01f2321380523444b9ad43165bb05a8f276a1c7006ff6989b2cf84fa260a7

Observation ad8f0973-8cfa-46fa-ac74-2c6116ae904b · inbound

LongSpike: Fractional Order Spiking State Space Models for Efficient Long Sequence Learning cites this paper.

LongSpike: Fractional Order Spiking State Space Models for Efficient Long Sequence Learning EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:48:21.554743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T07:26:58.608421Z digest=sha256:e2e453373fb8eceefe6416c1919917c7578dbdbb5d8151d78fbe8ed28abbab1d

Observation b141b3e0-d1e4-420d-a508-b1237e209ec6 · inbound

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation cites this paper.

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:58.175300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T00:19:49.071495Z digest=sha256:abd22aa9161d6f7c46b64b4759a487a644a29fd3130fa88c5712f4ad21031d4e