Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T14:43:05.950819Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2510.02361.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T14:43:05.950819Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a8d18167-a913-4ac9-8e41-3ad25b30b9fc · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b12ba76-acc2-47d8-acce-1d2451347a53 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference @esa (Ref
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ddfb705-3488-4297-90fb-a31656d009c8 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 104e38e2-3542-411b-83df-f7867d428aa5 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fce07af1-9bac-48c6-81b5-1468dad88988 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Longbench: A bilingual, multitask benchmark for long context understanding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd80f98a-0981-4b7d-8f6a-dd7684312707 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Longformer: The Long-Document Transformer
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd5af0bb-12cc-4238-80e5-4125e1dd2577 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Transformers to ssms: Distilling quadratic knowledge to subquadratic models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 016852ee-4fae-4590-a967-73bee6d5cdc2 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Piqa: Reasoning about physical commonsense in natural language
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8154630b-1f14-4887-8f2f-9687e91c15b0 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97e2590e-7e60-4785-bac5-541b06db62e6 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a67f7d59-ed6d-46f7-953b-5e800d5b977a · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Vicuna: An open-source chatbot impressing gpt-4 with 90\ See https://vicuna
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf8ad47e-48b9-4ba7-b266-9f84a9d0e709 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1807d04c-ee86-40d6-83ab-575bb89913cd · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2854c181-e0d5-4004-9549-182896572482 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference The llama 3 herd of models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 467f76a0-70ca-4e9f-b95a-9809793fafd2 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Knowledge Distillation: A Survey
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26dea23d-bc8e-4bf0-8e9e-80576bde32da · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Measuring massive multitask language understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ef0cbf8-c6b6-4e6d-8e69-2abb581633a0 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Distilling the Knowledge in a Neural Network
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37af2e96-0f4c-4d2e-a8dc-53101bfb4ccc · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48e79267-803c-42cf-97bf-795be4731117 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Wise: Weak-supervision-guided step-by-step explanations for multimodal llms in image classification
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58855a08-501a-41a4-915e-7e471fddf205 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fcc620d-c2e4-4979-8758-ff28057a29d9 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Sequence-level knowledge distillation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd1c0ab8-89ca-4471-956a-5a7afbdc86fa · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference MiniMax-01: Scaling Foundation Models with Lightning Attention
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0598c0a0-a870-4d7f-ac25-43779cbf6c5d · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Snapkv: LLM knows what you are looking for before generation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f3ad922-f04f-4d20-8792-028b5c857272 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff01ec78-3682-4291-ab6f-743adea4f501 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Fineweb-edu: the finest collection of educational content, 2024
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ed79671-7dcb-4b84-950c-0656a590b46d · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference MoBA: Mixture of Block Attention for Long-Context LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef20dec0-da35-4c6a-8799-1313eea66e3e · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Linearizing Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91117444-c746-4609-94ff-a3eaca973aaa · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Llama 3 model card
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4377e9ba-f304-42bb-a6f4-dcda090b67a5 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Can a suit of armor conduct electricity? A new dataset for open book question answering
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b7b5af8-f332-45fb-88f8-8611e640120e · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Instruction Tuning with GPT-4
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbc559e8-3dd0-4421-a00d-f667281827dd · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference RWKV: Reinventing RNNs for the Transformer Era
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54089480-e5a5-4bcd-878d-4df82f66fb35 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c08777be-abd3-4ce6-bf49-45963a88d071 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Compressive Transformers for Long-Range Sequence Modelling
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71965019-297d-4cc7-a3e8-4b88bf4b4f8e · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Policy Distillation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5eaea05-fb50-4f36-a4d3-eb0384861fa3 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference P y SBD : Pragmatic sentence boundary disambiguation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f9ff6e1-1742-4186-ae24-e1e5f3ae8460 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Winogrande: An adversarial winograd schema challenge at scale
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c621803b-ddf1-4a43-bee0-8915fb00f999 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcfcb007-c3dc-42b4-8492-5ab33576c729 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Social iqa: Commonsense reasoning about social interactions
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f969937e-b197-4c30-a92d-16ec3a3360e2 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Instruction-Tuning LLMs for Event Extraction with Annotation Guidelines
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20adb174-0089-4902-b1f2-e9c63581619a · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Retentive Network: A Successor to Transformer for Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e7a825f-9af7-4054-bda5-771f5493eae2 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Commonsenseqa: A question answering challenge targeting commonsense knowledge
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9b34afe-ae15-4f3a-8d9a-79e59a02d169 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4acb60d5-cb70-443f-81d8-02722dc0d81b · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Stanford alpaca: An instruction-following llama model, 2023
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8ce3ebc-30e2-413a-b6ee-4b4359cf4d68 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Qwen2.5: A party of foundation models, September 2024
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85bb8371-85ae-436e-9948-b4c75b4e5cf5 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Attention is all you need
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3147b43f-bfb3-45b2-9b87-4a2eeddd7343 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Unshackling Context Length: An Efficient Selective Attention Approach through Query-Key Compression
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dd7f531a-c428-423f-8e94-936bd33ced52 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference The mamba in the llama: Distilling and accelerating hybrid models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdc4b953-2b0c-4dcb-ad3c-a0beca44416d · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Liu, and Matt Gardner
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e262e16b-7752-4a2a-817d-1a3ab3e0cc1f · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Efficient streaming language models with attention sinks
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a1b167d-f5a6-4487-9c26-bd4d168218c2 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Pyramidinfer: Pyramid KV cache compression for high-throughput LLM inference
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75b48b7c-4374-4c56-bcaf-eb1b40817ffb · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Native sparse attention: Hardware-aligned and natively trainable sparse attention
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32d47e24-cf6a-4ab0-bf9e-96dcccc5e6b0 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f0b6726-7019-4b05-92df-f50d6e2c8a91 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Event temporal relation extraction based on retrieval-augmented on llms
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5503ea8f-c7cd-4436-8c9e-7b7c4c89d781 · outbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Barrett, Zhangyang Wang, and Beidi Chen
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.