Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T14:25:06.182146Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 12 inbound Pith citation observations for arXiv:2412.12094.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T14:25:06.182146Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T00:38:48.057219Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
20 of 20 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation 0f0ff43e-151c-41d4-92bb-2a5d8e419bca · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator The Falcon Series of Open Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f865bea-b69d-4fb8-9556-181f5aa2d572 · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator Additionally, feed-forward networks with ReLU activation can effectively represent any piecewise linear function
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5989b33f-4e24-40b9-a089-2a708f03521c · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61972174-c251-47d3-a52a-ab1ebddcfd21 · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a716a03f-1bf4-4da6-a402-865a87e69930 · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator Chen, G., Xia, L., and Huang, C
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18a6fb13-35f6-4da5-8002-6a5fb40ffaa9 · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40a2fefc-d6f1-4bee-b9a2-6e587108eace · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 788d97a0-b99e-46e8-9771-83c0ae06f2c4 · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 79da381c-6552-44c3-b2ac-5d207f879b15 · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator PyramidKV (Zhang et al., 2024)
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0407bef2-3476-4576-a892-bfd3b702f81d · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator ” and “?
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation eeb6ea4c-4082-4ae6-b4a7-3f7e7fa568a3 · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator (n−1) d−1X i=0 δ−i :δ: (n−1) d−1X i=0 δ−i +δ −d+1 −δ # , u⊤Zk ∈
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4f79ad32-015f-4cf5-ad94-1a0d541c79ae · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator 4 initial tokens are kept
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2ad808ad-71b5-426d-8d50-b350bd63c349 · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator 32 initial tokens are kept
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation dcae64ab-e48c-4d7b-98f6-450c255989ef · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1039abf-a3ff-4e7b-ace6-453cbbbb4998 · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator Training Verifiers to Solve Math Word Problems
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86b3e9f2-bd5e-4078-bcbe-a9110222a9d2 · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator The Llama 3 Herd of Models
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fedc5b42-97d8-43f8-b498-74959d10411c · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator The results in Table 17 show that FixLLM has a significant gap compared to SepLLM in both mathematical logical reasoning and knowledge-based reasoning capabili- ties
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9a80c46f-697b-45df-bcf6-9d1a3f52d890 · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator Longformer: The Long-Document Transformer
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0727550-c477-4fe7-a605-2f220933b70f · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe622c78-ea51-4154-a854-9c00766ff892 · outbound
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71b2eebb-bf39-443f-b1bf-e96dcb0c4954 · inbound
A Survey on Large Language Model Acceleration based on KV Cache Management SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Reference 136
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc318184-6f14-4f85-aeb2-b1463aee789f · inbound
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a5f70ddc-f502-4087-85e2-06912880a329 · inbound
GEM: Empowering LLM for both Embedding Generation and Language Understanding SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f496848-d855-4471-8142-2c892c297b6b · inbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a8464b1-0285-4290-bf96-5dec339b3a40 · inbound
EARN: Efficient Inference Acceleration for LLM-based Generative Recommendation by Register Tokens SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e3cbe33-ce91-4136-8db3-9c084ff0c855 · inbound
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57454bc2-d061-44aa-ae76-7f3e17e17fe3 · inbound
CaliDrop: KV Cache Compression with Calibration SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97e2590e-7e60-4785-bac5-541b06db62e6 · inbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4761a781-8aaa-4876-8e24-731d6cab01ea · inbound
LightThinker++: From Reasoning Compression to Memory Management SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8ac373e4-d79f-4568-91fd-1f58389d87dd · inbound
SAGE: Selective Attention-Guided Extraction for Token-Efficient Document Indexing SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8d5cab32-d60d-4afa-8d74-00ed049ff9fc · inbound
What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ce8edf70-bd41-4f56-8b59-cb6c2af227a6 · inbound
Metaphor Tracer: A Theory-Informed Analysis of Hidden States SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.