{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:7UL53LZUN4TVGMXDYABD2IBTCR","short_pith_number":"pith:7UL53LZU","schema_version":"1.0","canonical_sha256":"fd17ddaf346f275332e3c0023d20331468f89b58bd69ce23886655b354268d4b","source":{"kind":"arxiv","id":"2310.06694","version":2},"attestation_state":"computed","paper":{"title":"Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.LG"],"primary_cat":"cs.CL","authors_text":"Danqi Chen, Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng","submitted_at":"2023-10-10T15:13:30Z","abstract_excerpt":"The popularity of LLaMA (Touvron et al., 2023a;b) and other recently emerged moderate-sized large language models (LLMs) highlights the potential of building smaller yet powerful LLMs. Regardless, the cost of training such models from scratch on trillions of tokens remains high. In this work, we study structured pruning as an effective means to develop smaller LLMs from pre-trained, larger models. Our approach employs two key techniques: (1) targeted structured pruning, which prunes a larger model to a specified target shape by removing layers, heads, and intermediate and hidden dimensions in "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2310.06694","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2023-10-10T15:13:30Z","cross_cats_sorted":["cs.AI","cs.LG"],"title_canon_sha256":"c7e7e4e016a85fe3bc5d69c3c68bfe91acc95bfc0dc1f07bbcf3a176eafbbaa4","abstract_canon_sha256":"465770da4f9dda524b3ca55409e3f0f1d05b7627faea2d87bdd0b5c5e91e29ea"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:06:43.461522Z","signature_b64":"eqFZ2U7YgKPySXN2OIaw4xzPVmNK33LJpeuy3V0Zw9SL69jn8jqYZP7qG1FJP+GH3LZHx4VzE7dRun+oZSYKCw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"fd17ddaf346f275332e3c0023d20331468f89b58bd69ce23886655b354268d4b","last_reissued_at":"2026-07-05T08:06:43.461034Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:06:43.461034Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.LG"],"primary_cat":"cs.CL","authors_text":"Danqi Chen, Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng","submitted_at":"2023-10-10T15:13:30Z","abstract_excerpt":"The popularity of LLaMA (Touvron et al., 2023a;b) and other recently emerged moderate-sized large language models (LLMs) highlights the potential of building smaller yet powerful LLMs. Regardless, the cost of training such models from scratch on trillions of tokens remains high. In this work, we study structured pruning as an effective means to develop smaller LLMs from pre-trained, larger models. Our approach employs two key techniques: (1) targeted structured pruning, which prunes a larger model to a specified target shape by removing layers, heads, and intermediate and hidden dimensions in "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2310.06694","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2310.06694/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2310.06694","created_at":"2026-07-05T08:06:43.461093+00:00"},{"alias_kind":"arxiv_version","alias_value":"2310.06694v2","created_at":"2026-07-05T08:06:43.461093+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2310.06694","created_at":"2026-07-05T08:06:43.461093+00:00"},{"alias_kind":"pith_short_12","alias_value":"7UL53LZUN4TV","created_at":"2026-07-05T08:06:43.461093+00:00"},{"alias_kind":"pith_short_16","alias_value":"7UL53LZUN4TVGMXD","created_at":"2026-07-05T08:06:43.461093+00:00"},{"alias_kind":"pith_short_8","alias_value":"7UL53LZU","created_at":"2026-07-05T08:06:43.461093+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":31,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2607.07557","citing_title":"PALS: Percentile-Aware Layerwise Sparsity for LLM Pruning","ref_index":15,"is_internal_anchor":true},{"citing_arxiv_id":"2605.24956","citing_title":"NITP: Next Implicit Token Prediction for LLM Pre-training","ref_index":54,"is_internal_anchor":false},{"citing_arxiv_id":"2606.28438","citing_title":"When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs","ref_index":149,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15491","citing_title":"Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs","ref_index":36,"is_internal_anchor":false},{"citing_arxiv_id":"2605.24956","citing_title":"NITP: Next Implicit Token Prediction for LLM Pre-training","ref_index":54,"is_internal_anchor":false},{"citing_arxiv_id":"2605.25134","citing_title":"Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate","ref_index":89,"is_internal_anchor":false},{"citing_arxiv_id":"2605.25344","citing_title":"A Hamiltonian-Inspired Local-Operator Ansatz for Slimming Large Language Models","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2605.28207","citing_title":"Pruning and Distilling Mixture-of-Experts into Dense Language Models","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2606.00535","citing_title":"DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation","ref_index":143,"is_internal_anchor":false},{"citing_arxiv_id":"2606.26538","citing_title":"CascadeFormer: Depth-Tapered Transformers Motivated by Gradient Fan-in Asymmetry","ref_index":39,"is_internal_anchor":false},{"citing_arxiv_id":"2411.11707","citing_title":"Federated Co-tuning Framework for Large and Small Language Models","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2501.05465","citing_title":"Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026)","ref_index":138,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14738","citing_title":"TAPIOCA: Why Task- Aware Pruning Improves OOD model Capability","ref_index":52,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08738","citing_title":"SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training","ref_index":68,"is_internal_anchor":false},{"citing_arxiv_id":"2605.17985","citing_title":"SAFE-SVD: Sensitivity-Aware Fidelity-Enforcing SVD for Physics Foundation Models","ref_index":28,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15491","citing_title":"Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs","ref_index":36,"is_internal_anchor":false},{"citing_arxiv_id":"2507.15640","citing_title":"Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training","ref_index":42,"is_internal_anchor":false},{"citing_arxiv_id":"2510.22767","citing_title":"TELL-TALE: Task Efficient LLMs with Task Aware Layer Elimination","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2511.14582","citing_title":"OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models","ref_index":50,"is_internal_anchor":false},{"citing_arxiv_id":"2512.22671","citing_title":"Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2309.03883","citing_title":"DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models","ref_index":93,"is_internal_anchor":false},{"citing_arxiv_id":"2602.01997","citing_title":"On the Limits of Layer Pruning for Generative Reasoning in Large Language Models","ref_index":31,"is_internal_anchor":false},{"citing_arxiv_id":"2409.04429","citing_title":"VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14350","citing_title":"Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling","ref_index":269,"is_internal_anchor":false},{"citing_arxiv_id":"2404.14294","citing_title":"A Survey on Efficient Inference for Large Language Models","ref_index":175,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/7UL53LZUN4TVGMXDYABD2IBTCR","json":"https://pith.science/pith/7UL53LZUN4TVGMXDYABD2IBTCR.json","graph_json":"https://pith.science/api/pith-number/7UL53LZUN4TVGMXDYABD2IBTCR/graph.json","events_json":"https://pith.science/api/pith-number/7UL53LZUN4TVGMXDYABD2IBTCR/events.json","paper":"https://pith.science/paper/7UL53LZU"},"agent_actions":{"view_html":"https://pith.science/pith/7UL53LZUN4TVGMXDYABD2IBTCR","download_json":"https://pith.science/pith/7UL53LZUN4TVGMXDYABD2IBTCR.json","view_paper":"https://pith.science/paper/7UL53LZU","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2310.06694&json=true","fetch_graph":"https://pith.science/api/pith-number/7UL53LZUN4TVGMXDYABD2IBTCR/graph.json","fetch_events":"https://pith.science/api/pith-number/7UL53LZUN4TVGMXDYABD2IBTCR/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/7UL53LZUN4TVGMXDYABD2IBTCR/action/timestamp_anchor","attest_storage":"https://pith.science/pith/7UL53LZUN4TVGMXDYABD2IBTCR/action/storage_attestation","attest_author":"https://pith.science/pith/7UL53LZUN4TVGMXDYABD2IBTCR/action/author_attestation","sign_citation":"https://pith.science/pith/7UL53LZUN4TVGMXDYABD2IBTCR/action/citation_signature","submit_replication":"https://pith.science/pith/7UL53LZUN4TVGMXDYABD2IBTCR/action/replication_record"}},"created_at":"2026-07-05T08:06:43.461093+00:00","updated_at":"2026-07-05T08:06:43.461093+00:00"}