Pith. sign in

Paper Citation Record · LEDGER

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

As of 14 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 12 inbound Pith citation observations for arXiv:2412.12094.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12094 v6

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:25:06.182146Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:38:48.057219Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 0f0ff43e-151c-41d4-92bb-2a5d8e419bca · outbound

This paper cites The Falcon Series of Open Language Models.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator The Falcon Series of Open Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T14:25:06.023014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:25:06.023014Z digest=sha256:74ba69c56d493c60600e7883ce38752fa1870aef5eb4d9856d27d0d33730502e

Observation 2f865bea-b69d-4fb8-9556-181f5aa2d572 · outbound

This paper cites Additionally, feed-forward networks with ReLU activation can effectively represent any piecewise linear function.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator Additionally, feed-forward networks with ReLU activation can effectively represent any piecewise linear function

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:25:06.732832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T14:25:06.158626Z digest=sha256:41f6569df77154f0c5b5bce2a5e986f42e7a71db44b07b3a22285ed5fcb661eb

Observation 5989b33f-4e24-40b9-a089-2a708f03521c · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:25:06.063912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:25:06.063912Z digest=sha256:b7c01736b215709f84ae74bda1980863d7b945decb0c6a790dfd2d078b52d874

Observation 61972174-c251-47d3-a52a-ab1ebddcfd21 · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T14:25:06.085460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:25:06.085460Z digest=sha256:d950815602d38aabcefe821ee8767bcaf0ab4a8cc424b8e7b703eec42905ada7

Observation a716a03f-1bf4-4da6-a402-865a87e69930 · outbound

This paper cites Chen, G., Xia, L., and Huang, C.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator Chen, G., Xia, L., and Huang, C

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T14:25:06.036755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:25:06.036755Z digest=sha256:afac0397583b1ad7efa70649b9a4d60c0440830c1555dd2fe00c810217b5a864

Observation 18a6fb13-35f6-4da5-8002-6a5fb40ffaa9 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T14:25:06.100416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:25:06.100416Z digest=sha256:ce31da4897bc159238e75df9ebce3a6885bddf20c653e8ebef5a051cad4aa75b

Observation 40a2fefc-d6f1-4bee-b9a2-6e587108eace · outbound

This paper cites an unresolved cited work.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:25:06.855705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T14:25:06.118633Z digest=sha256:b7b9011e26887939225740abaa1264d4d2a3023a64eacc11603495b88778bb10

Observation 788d97a0-b99e-46e8-9771-83c0ae06f2c4 · outbound

This paper cites an unresolved cited work.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:25:06.830308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T14:25:06.123921Z digest=sha256:e3c8c5e03212ffce022b7bd6b4f90c9f2306d5be0e0c22c9352c5cb01042bcec

Observation 79da381c-6552-44c3-b2ac-5d207f879b15 · outbound

This paper cites PyramidKV (Zhang et al., 2024).

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator PyramidKV (Zhang et al., 2024)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:25:06.798914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T14:25:06.130816Z digest=sha256:78b4c8fc38cb1ccb94851825ea3adc8c66a801276025661abc661b0e78f041c7

Observation 0407bef2-3476-4576-a892-bfd3b702f81d · outbound

This paper cites ” and “?.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator ” and “?

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:25:06.776540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T14:25:06.137167Z digest=sha256:1b7c7dc03bd820440528cc1aacb2073385502e30322969c2ea17e33edf0b47e6

Observation eeb6ea4c-4082-4ae6-b4a7-3f7e7fa568a3 · outbound

This paper cites (n−1) d−1X i=0 δ−i :δ: (n−1) d−1X i=0 δ−i +δ −d+1 −δ # , u⊤Zk ∈.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator (n−1) d−1X i=0 δ−i :δ: (n−1) d−1X i=0 δ−i +δ −d+1 −δ # , u⊤Zk ∈

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:25:06.708598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T14:25:06.167887Z digest=sha256:4f8db95a90167ab069cfde9752d85a0142e0aecb150c1045033d37f141e032d2

Observation 4f79ad32-015f-4cf5-ad94-1a0d541c79ae · outbound

This paper cites 4 initial tokens are kept.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator 4 initial tokens are kept

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:25:06.674486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T14:25:06.176673Z digest=sha256:2a13114daebec4eb1c54352f09b3f35d54ce542a1de1bb044c83373fcb1d7c83

Observation 2ad808ad-71b5-426d-8d50-b350bd63c349 · outbound

This paper cites 32 initial tokens are kept.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator 32 initial tokens are kept

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:25:06.641891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T14:25:06.182146Z digest=sha256:a096983450ad048bc397208bd73a782ceed624209a578784c56dfcebfae3f9fb

Observation dcae64ab-e48c-4d7b-98f6-450c255989ef · outbound

This paper cites PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T14:25:06.072675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:25:06.072675Z digest=sha256:29fa0c71b00ad7795641ae1744eff3e841d936854104a38ec93e5e4738bbac1c

Observation a1039abf-a3ff-4e7b-ace6-453cbbbb4998 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator Training Verifiers to Solve Math Word Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T14:25:06.049120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:25:06.049120Z digest=sha256:f0ffe9e224484773906aaec0ecc1c326b3242f502f8329efc49440b2ed0f8650

Observation 86b3e9f2-bd5e-4078-bcbe-a9110222a9d2 · outbound

This paper cites The Llama 3 Herd of Models.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator The Llama 3 Herd of Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T14:25:06.057530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:25:06.057530Z digest=sha256:6e2145d3a1c615689c793c80628f4ce83450be6bfa35df201d86f4905ac90dae

Observation fedc5b42-97d8-43f8-b498-74959d10411c · outbound

This paper cites The results in Table 17 show that FixLLM has a significant gap compared to SepLLM in both mathematical logical reasoning and knowledge-based reasoning capabili- ties.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator The results in Table 17 show that FixLLM has a significant gap compared to SepLLM in both mathematical logical reasoning and knowledge-based reasoning capabili- ties

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:25:06.755414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T14:25:06.149825Z digest=sha256:087f5a5860fb6da4ebbe519f4dd79685e6641e7ee4609026d016c683b097c1ad

Observation 9a80c46f-697b-45df-bcf6-9d1a3f52d890 · outbound

This paper cites Longformer: The Long-Document Transformer.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator Longformer: The Long-Document Transformer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T14:25:06.029478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:25:06.029478Z digest=sha256:c2f0859b0ee0b4e2ec0332dfedaf69758075b51d20000f2db5963d693e75b142

Observation b0727550-c477-4fe7-a605-2f220933b70f · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T14:25:06.042723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:25:06.042723Z digest=sha256:b89a99b460b250df8e9d9409149e26608e4489cfc5464df0653bac5f21f1adf7

Observation fe622c78-ea51-4154-a854-9c00766ff892 · outbound

This paper cites SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T14:25:06.112967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:25:06.112967Z digest=sha256:4d16eb1e3051d493fdb1c92f1c99d410585d2935c58e0ee345c7f58876a0981b

Pith citing papers

Observation 71b2eebb-bf39-443f-b1bf-e96dcb0c4954 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 136

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:48.057219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:48.057219Z digest=sha256:907aa19b4429e9211a7023eaac0db6c1dbec9cbbf80f0052acea5a35dd6e7a8d

Observation dc318184-6f14-4f85-aeb2-b1463aee789f · inbound

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention cites this paper.

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:46:30.212281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-16T23:46:29.975858Z digest=sha256:2f3a60e4f2e8ee460dc25c1e0d08d1d414af788ea4e1f1a3a3a499a65b774570

Observation a5f70ddc-f502-4087-85e2-06912880a329 · inbound

GEM: Empowering LLM for both Embedding Generation and Language Understanding cites this paper.

GEM: Empowering LLM for both Embedding Generation and Language Understanding SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:51.014693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:50:51.014693Z digest=sha256:87b87b2ace476a3496fe4bc5f403ae09adb8cabd4e6c990575e80c36262d5b18

Observation 8f496848-d855-4471-8142-2c892c297b6b · inbound

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning cites this paper.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.572234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.572234Z digest=sha256:7d0d4a76550d428eaa999db9f7b838766156ea326e0290b7338fd5db0b39bf1a

Observation 1a8464b1-0285-4290-bf96-5dec339b3a40 · inbound

EARN: Efficient Inference Acceleration for LLM-based Generative Recommendation by Register Tokens cites this paper.

EARN: Efficient Inference Acceleration for LLM-based Generative Recommendation by Register Tokens SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:18:46.625475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:18:46.625475Z digest=sha256:d0896b26b8608360859eeff2b279d4627fe0e889c89dbd8691a602560f2b34a9

Observation 7e3cbe33-ce91-4136-8db3-9c084ff0c855 · inbound

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference cites this paper.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.728594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.728594Z digest=sha256:38424040608dd1d6e4e1943500d562b4930b6af58e430b04ccec1f5ee18daf13

Observation 57454bc2-d061-44aa-ae76-7f3e17e17fe3 · inbound

CaliDrop: KV Cache Compression with Calibration cites this paper.

CaliDrop: KV Cache Compression with Calibration SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.675892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.675892Z digest=sha256:61ae788cd108b816062d8fc16ae019f842fce050cf3b036fb9a335f3998bd873

Observation 97e2590e-7e60-4785-bac5-541b06db62e6 · inbound

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference cites this paper.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:00.999081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:00.999081Z digest=sha256:f3addae1829983ba40289261dfeb4e5060b94b24df396a58cfe001d7cb10acb9

Observation 4761a781-8aaa-4876-8e24-731d6cab01ea · inbound

LightThinker++: From Reasoning Compression to Memory Management cites this paper.

LightThinker++: From Reasoning Compression to Memory Management SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:28:01.786972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T17:25:28.432170Z digest=sha256:7b133946567debd5f0e1a69e0f1e6a32d92cfe61b20216283e7d5648cd2f7837

Observation 8ac373e4-d79f-4568-91fd-1f58389d87dd · inbound

SAGE: Selective Attention-Guided Extraction for Token-Efficient Document Indexing cites this paper.

SAGE: Selective Attention-Guided Extraction for Token-Efficient Document Indexing SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:32:52.129872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T08:32:02.222528Z digest=sha256:58b6bb1272253ebb348a8202d65996a4d86223bab54b79200853869e2d536f3c

Observation 8d5cab32-d60d-4afa-8d74-00ed049ff9fc · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.291190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:fd28b3f7b107b91f642df16f673df32024b61a5f5dbb6f2adda9d5144213b6a0

Observation ce8edf70-bd41-4f56-8b59-cb6c2af227a6 · inbound

Metaphor Tracer: A Theory-Informed Analysis of Hidden States cites this paper.

Metaphor Tracer: A Theory-Informed Analysis of Hidden States SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-31T07:30:05.476997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:30:05.476997Z digest=sha256:211b0776e84f2fb3e5ce623dafac793de70271b166c0feb93e2b5cf40a7eb979