Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:59:48.350745Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 1 inbound Pith citation observation for arXiv:2501.14713.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:59:48.350745Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T17:52:28.170298Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T17:52:28.290404Z
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 71a7b657-c5c8-4a52-b94c-c8fce5c2cb8b · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d050b94-7530-448f-85c4-29d3b92b4509 · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Head-wise Shareable Attention for Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b998c2a9-f8a3-46c5-b9b5-4bf5bdbb2583 · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Universal Transformers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03967287-d11a-4066-9a80-d4262baee57a · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 096671c5-da39-4457-9027-a9c865b69c54 · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing LoRA: Low-Rank Adaptation of Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 381fde8c-c300-43a5-a67b-859ea6a06d4c · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing The MiniPile Challenge for Data-Efficient Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c63f6358-52c3-4134-875a-b3b1b02e4097 · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing The Power of Scale for Parameter-Efficient Prompt Tuning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 022155da-5826-47fe-9228-d151cb30d5ff · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing DoRA: Weight-Decomposed Low-Rank Adaptation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0ff9be9-71d9-456f-a40a-8d06e0fb6c8e · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b52c128-d06d-4c9b-aaed-1b6e3fb89ff4 · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c53e008-1a12-4626-91d5-5a0d4f019fff · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Pointer Sentinel Mixture Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c14f7341-46bb-480e-975a-ee48ea735a50 · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Large Language Models: A Survey
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1bce8de-2d85-4448-ab4b-cd73becccce5 · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing A Comprehensive Overview of Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbf55e81-76b6-4a48-ac36-3cd3d18a12d4 · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing EE-Tuning: An Economical yet Scalable Solution for Tuning Early-Exit Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bcbf74b-1ddb-47e3-ad53-9f5b398d549c · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Subformer: Exploring Weight Sharing for Parameter Efficiency in Generative Transformers
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d04cfec8-1c97-4c4a-bb5d-bd871882b7ac · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Accelerating Transformer Inference for Translation via Parallel Decoding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38583a41-8ab5-4c1b-ae1e-2bf9da7864cb · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Lessons on Parameter Sharing across Layers in Transformers
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6cb0186-e725-4e98-bdb6-8b6a6d063d1b · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing LLaMA: Open and Efficient Foundation Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c69ce8be-58d2-4e11-bc9a-c3f7eba90d24 · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing LaCo: Large Language Model Pruning via Layer Collapse
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57c418fc-d422-42a0-9d67-ab4c55626e42 · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing TinyLlama: An Open-Source Small Language Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cc733c4-c108-4279-958e-4d34f1faaaf5 · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing OPT: Open Pre-trained Transformer Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1921a86-2b40-4166-ac1b-11252f00cfde · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing For zero-shot performance evaluations, we used the ARC-e, ARC-c (Clark et al., 2018), PIQA (Bisk et al., 2020), WinoGrande (Sakaguchi et al., 2021), and HellaSwag (Zellers et al.,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation eb18ec19-2fb0-4e74-a975-2ea1477988df · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing For perplexity perfor- mance evaluations, we used the validation MiniP- ile (Kaddour,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 697440f5-e2b3-4665-9975-ccd751f4a3d2 · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ac36b88b-c193-4622-8c60-352c07211b0c · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Layer Normalization
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3be2bc2e-8043-4b15-a740-b14a3c659549 · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 765a187d-65ab-4d0a-b874-08c7216758a3 · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0be7abd8-2ab4-4ee8-af43-a850c2b04bd8 · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 836ca620-f868-4a1c-aaaf-dc40be6eb857 · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Towards a Unified View of Parameter-Efficient Transfer Learning
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6909fb1c-f9a8-43b1-9b96-497726c1e759 · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Language model compression with weighted low-rank factorization
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6765384a-655a-40d2-b6ec-dd16d198f3b6 · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c4bb74e-6d21-456f-86f3-308bbb75b210 · outbound
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing SliceGPT: Compress Large Language Models by Deleting Rows and Columns
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9364cb06-1227-4de1-ad60-37f0966a8470 · inbound
CALR: Corrective Adaptive Low-Rank Decomposition for Efficient Large Language Model Layer Compression FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.