Pith. sign in

Paper Citation Record · LEDGER

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference

As of 18 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2510.02361.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.02361 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:43:05.950819Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved52
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a8d18167-a913-4ac9-8e41-3ad25b30b9fc · outbound

This paper cites write newline.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:59.614005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:42:59.614005Z digest=sha256:a5647f240e43bed4901ce4029326c730b8a5191b46ade75fe8a91b41d33d1384

Observation 1b12ba76-acc2-47d8-acce-1d2451347a53 · outbound

This paper cites @esa (Ref.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference @esa (Ref

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:59.768694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:42:59.768694Z digest=sha256:70c4fcf0b3f53d11f2326c53c7dc9234b9f074e5262037eb79c468be225f3a57

Observation 3ddfb705-3488-4297-90fb-a31656d009c8 · outbound

This paper cites an unresolved cited work.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:59.893251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:42:59.893251Z digest=sha256:818f202735f8ba0fcbcfb2c28abba4845784d60f0fac3973bb9c7f181177ee28

Observation 104e38e2-3542-411b-83df-f7867d428aa5 · outbound

This paper cites an unresolved cited work.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:00.066367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:00.066367Z digest=sha256:c65d6a53aed1036dc4aa11b2a81d0f10d83cd9ac2ab80769670d91cc3a2c79dd

Observation fce07af1-9bac-48c6-81b5-1468dad88988 · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Longbench: A bilingual, multitask benchmark for long context understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:00.189486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:00.189486Z digest=sha256:fc9a9e42a22096aaea045811e5542dfa970e92d8a0bc7a001f9e16baa716215d

Observation bd80f98a-0981-4b7d-8f6a-dd7684312707 · outbound

This paper cites Longformer: The Long-Document Transformer.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Longformer: The Long-Document Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:00.368727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:00.368727Z digest=sha256:a693de601ceee8a2571c27baee29fa604bffbf449a6548eb5cabfc4ac8bd7d1d

Observation dd5af0bb-12cc-4238-80e5-4125e1dd2577 · outbound

This paper cites Transformers to ssms: Distilling quadratic knowledge to subquadratic models.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Transformers to ssms: Distilling quadratic knowledge to subquadratic models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:00.540648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:00.540648Z digest=sha256:219b490478be1e390c339edc1e3fcd9e959b5e2f40db4d5c5f0fd3e27dde9f63

Observation 016852ee-4fae-4590-a967-73bee6d5cdc2 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Piqa: Reasoning about physical commonsense in natural language

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:00.673999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:00.673999Z digest=sha256:439331ac0e3800f625096a68d6734d315fa1766f2bcc0734b8d877eb0a2fe066

Observation 8154630b-1f14-4887-8f2f-9687e91c15b0 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:00.828806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:00.828806Z digest=sha256:fccc63cec6ef7753d8206b2704ab7e6e71d32829241ebac2bee5d1aa1eaae7f7

Observation 97e2590e-7e60-4785-bac5-541b06db62e6 · outbound

This paper cites SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:00.999081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:00.999081Z digest=sha256:48d84aefe9522dc2008d1766a9e4882c4153042673035412c8bc1a1c8e654210

Observation a67f7d59-ed6d-46f7-953b-5e800d5b977a · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90\ See https://vicuna.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Vicuna: An open-source chatbot impressing gpt-4 with 90\ See https://vicuna

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:01.105639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:01.105639Z digest=sha256:b6296cdc6b7747568ab3cc50b50f3d00cfc660050cf17bdeb974ed2b71998651

Observation cf8ad47e-48b9-4ba7-b266-9f84a9d0e709 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:01.240790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:01.240790Z digest=sha256:7536a93e844b5f9c37c819e1d6fcda7e1125d290dbdb3d8f712ad2035684b758

Observation 1807d04c-ee86-40d6-83ab-575bb89913cd · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:01.350545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:01.350545Z digest=sha256:9cb290cc25c7ccf41e8d34723d868e7ea0b41f614c4c8669e018b6727070f1ff

Observation 2854c181-e0d5-4004-9549-182896572482 · outbound

This paper cites The llama 3 herd of models.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference The llama 3 herd of models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:01.458419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:01.458419Z digest=sha256:ed0131fefec22cc88635c73dd0f0ed4462ea098b32cda4183770d9d68f42fa82

Observation 467f76a0-70ca-4e9f-b95a-9809793fafd2 · outbound

This paper cites Knowledge Distillation: A Survey.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Knowledge Distillation: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:01.544923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:01.544923Z digest=sha256:c526b5fa753fb600baa306b0ca87ef19e245341c215dd3f96e741476823b2265

Observation 26dea23d-bc8e-4bf0-8e9e-80576bde32da · outbound

This paper cites Measuring massive multitask language understanding.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Measuring massive multitask language understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:01.668092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:01.668092Z digest=sha256:ce101d1cb328f3b84abc8f8e3bd1732b35e05b4580781ee1662f15a7ad733491

Observation 0ef0cbf8-c6b6-4e6d-8e69-2abb581633a0 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Distilling the Knowledge in a Neural Network

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:01.774895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:01.774895Z digest=sha256:92f93758e138e17ef5be9f977928ab23d8889a81dae6da64c87d2c42c0ef7feb

Observation 37af2e96-0f4c-4d2e-a8dc-53101bfb4ccc · outbound

This paper cites Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:01.955236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:01.955236Z digest=sha256:c81ab56b8809a9cb9cd38c6031861db00fbdb73f340d97baab9d48425d4a827a

Observation 48e79267-803c-42cf-97bf-795be4731117 · outbound

This paper cites Wise: Weak-supervision-guided step-by-step explanations for multimodal llms in image classification.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Wise: Weak-supervision-guided step-by-step explanations for multimodal llms in image classification

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:02.086066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:02.086066Z digest=sha256:a50778ed46609699332fb0eb002e7ca79bb064e7d3b111db533e9ea023cb4663

Observation 58855a08-501a-41a4-915e-7e471fddf205 · outbound

This paper cites an unresolved cited work.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:02.248626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:02.248626Z digest=sha256:304cdce028db6d92f4636651c7c27c09e7f3e4d80ea913712666aaf92d310663

Observation 8fcc620d-c2e4-4979-8758-ff28057a29d9 · outbound

This paper cites Sequence-level knowledge distillation.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Sequence-level knowledge distillation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:02.402138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:02.402138Z digest=sha256:7b930135c40e44fae63d8de03eac6898b5ef729dd1c63c1c356020e61a7a1ea8

Observation bd1c0ab8-89ca-4471-956a-5a7afbdc86fa · outbound

This paper cites MiniMax-01: Scaling Foundation Models with Lightning Attention.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference MiniMax-01: Scaling Foundation Models with Lightning Attention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:02.535332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:02.535332Z digest=sha256:56f040a19d19dae4f56848277dc98b60ee4d6293b0a6c12cd399266a70f3388a

Observation 0598c0a0-a870-4d7f-ac25-43779cbf6c5d · outbound

This paper cites Snapkv: LLM knows what you are looking for before generation.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Snapkv: LLM knows what you are looking for before generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:02.642104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:02.642104Z digest=sha256:e8779ef14a510e85e7874ab0473a7debb67ccb27b21d691e5931c4dee6f151c7

Observation 0f3ad922-f04f-4d20-8792-028b5c857272 · outbound

This paper cites RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:02.786626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:02.786626Z digest=sha256:4044789c7532859dd1efde73be591443712b12a58fd720a2887bf3de9028859c

Observation ff01ec78-3682-4291-ab6f-743adea4f501 · outbound

This paper cites Fineweb-edu: the finest collection of educational content, 2024.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Fineweb-edu: the finest collection of educational content, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:02.886560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:02.886560Z digest=sha256:29eeea499545ddf2f3cd2d2a9c57a8616252152a2ce3d80dd5830cdb2d8e355f

Observation 5ed79671-7dcb-4b84-950c-0656a590b46d · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:02.987885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:02.987885Z digest=sha256:55cedab4cbd79c9b3a0bf648f754c387c15a1a5d5bfe032733af24eee6f4ac7e

Observation ef20dec0-da35-4c6a-8799-1313eea66e3e · outbound

This paper cites Linearizing Large Language Models.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Linearizing Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:03.125292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:03.125292Z digest=sha256:11aba97d98ee79bd51fa8d0713368d99623909244d0ffd9fac75dd10f0755ade

Observation 91117444-c746-4609-94ff-a3eaca973aaa · outbound

This paper cites Llama 3 model card.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Llama 3 model card

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:03.191713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:03.191713Z digest=sha256:4a0a9730226ef3e9235da4c16869c9762fa63bbf4055f75d8d1c4c72b56184c1

Observation 4377e9ba-f304-42bb-a6f4-dcda090b67a5 · outbound

This paper cites Can a suit of armor conduct electricity? A new dataset for open book question answering.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Can a suit of armor conduct electricity? A new dataset for open book question answering

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:03.332107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:03.332107Z digest=sha256:6c3df812be41e97b328cb028595ab8a9452c3124685eea222f54e71e2ba89f14

Observation 4b7b5af8-f332-45fb-88f8-8611e640120e · outbound

This paper cites Instruction Tuning with GPT-4.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Instruction Tuning with GPT-4

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:03.438399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:03.438399Z digest=sha256:a0f19ed97bffb3230ceaade538df0cc26945ae37b30a433952814fc0dd941053

Observation fbc559e8-3dd0-4421-a00d-f667281827dd · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference RWKV: Reinventing RNNs for the Transformer Era

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:03.509806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:03.509806Z digest=sha256:3b4ed19e74d8ec3906d34f3e82e4d11c7978b89e2449372133556712d31be1b2

Observation 54089480-e5a5-4bcd-878d-4df82f66fb35 · outbound

This paper cites Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:03.587581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:03.587581Z digest=sha256:b132fcb67232d529fbf1e65f84fe51169b97465ee4f97dfb791034bcc9d92818

Observation c08777be-abd3-4ce6-bf49-45963a88d071 · outbound

This paper cites Compressive Transformers for Long-Range Sequence Modelling.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Compressive Transformers for Long-Range Sequence Modelling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:03.696218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:03.696218Z digest=sha256:8a526b13f481db9248e982b2db0a9fac2d28b2c48620070d64ed600d644cafa4

Observation 71965019-297d-4cc7-a3e8-4b88bf4b4f8e · outbound

This paper cites Policy Distillation.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Policy Distillation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:03.837743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:03.837743Z digest=sha256:29a78fed4f4c10b47e31e7251022f0996bc3a5f580c3574d6cf2f3dcc1ef50dc

Observation a5eaea05-fb50-4f36-a4d3-eb0384861fa3 · outbound

This paper cites P y SBD : Pragmatic sentence boundary disambiguation.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference P y SBD : Pragmatic sentence boundary disambiguation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:03.965654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:03.965654Z digest=sha256:351129d9f669aa257888270583f7336a22bf8235aefaccaad9b056eb9dc04e33

Observation 3f9ff6e1-1742-4186-ae24-e1e5f3ae8460 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Winogrande: An adversarial winograd schema challenge at scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:04.122792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:04.122792Z digest=sha256:ae2c4961f56713cb2690b5e6b03e62426ada9bb3351fa634d7eea42f909e04f4

Observation c621803b-ddf1-4a43-bee0-8915fb00f999 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:04.255288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:04.255288Z digest=sha256:651df13a5b261bbdbaaaa94792cb568d5e6b6fb141a9372906fcf067e578f4f3

Observation bcfcb007-c3dc-42b4-8492-5ab33576c729 · outbound

This paper cites Social iqa: Commonsense reasoning about social interactions.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Social iqa: Commonsense reasoning about social interactions

Reference 39

Resolution
malformed identifier
no resolver link, observed 2026-08-04T14:43:04.369336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:04.369336Z digest=sha256:84f9c88d49bace93915f70f6645e501f4cc0dc3f2fb03094bdc767fa440d6d36

Observation f969937e-b197-4c30-a92d-16ec3a3360e2 · outbound

This paper cites Instruction-Tuning LLMs for Event Extraction with Annotation Guidelines.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Instruction-Tuning LLMs for Event Extraction with Annotation Guidelines

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:04.514760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:04.514760Z digest=sha256:ad9223ebb874e428ae35e9ffb23d63e8168e53d6810f6bbb3e18b738cc356036

Observation 20adb174-0089-4902-b1f2-e9c63581619a · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Retentive Network: A Successor to Transformer for Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:04.645205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:04.645205Z digest=sha256:ead8eebf38062fa8eac6d39935bb7ac49876c53dcc4500327751721d3ad1e01f

Observation 8e7a825f-9af7-4054-bda5-771f5493eae2 · outbound

This paper cites Commonsenseqa: A question answering challenge targeting commonsense knowledge.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Commonsenseqa: A question answering challenge targeting commonsense knowledge

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:04.730271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:04.730271Z digest=sha256:2a2057f07bf813e1b6facc58c860c06a408c5824d47295b92da37a854f1a1754

Observation c9b34afe-ae15-4f3a-8d9a-79e59a02d169 · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:04.849425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:04.849425Z digest=sha256:8697b3a9f3631f9f2d9a1a8ee85e214dbb3799a8733e826faacb0830fed49fc9

Observation 4acb60d5-cb70-443f-81d8-02722dc0d81b · outbound

This paper cites Stanford alpaca: An instruction-following llama model, 2023.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Stanford alpaca: An instruction-following llama model, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:04.966387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:04.966387Z digest=sha256:8c76db08cee90990a51df7877ae49a32873a83aa9d2c376e84da06ceda9d2a86

Observation e8ce3ebc-30e2-413a-b6ee-4b4359cf4d68 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Qwen2.5: A party of foundation models, September 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:05.023640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:05.023640Z digest=sha256:635175a2894cccc4e882192c1114a27cddea8d541ed86d52b3ead3bb7da11880

Observation 85bb8371-85ae-436e-9948-b4c75b4e5cf5 · outbound

This paper cites Attention is all you need.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Attention is all you need

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:05.115564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:05.115564Z digest=sha256:5e58630aed96ea0659113e476f177a6af77df0ebdf96f0456a4f26f25b87d16c

Observation 3147b43f-bfb3-45b2-9b87-4a2eeddd7343 · outbound

This paper cites Unshackling Context Length: An Efficient Selective Attention Approach through Query-Key Compression.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Unshackling Context Length: An Efficient Selective Attention Approach through Query-Key Compression

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-04T14:43:25.804741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-04T14:43:05.227833Z digest=sha256:518c60c0c1a3485e2583ef24ccbffda06e068e972185deb67ae7c7041b1757b1

Observation dd7f531a-c428-423f-8e94-936bd33ced52 · outbound

This paper cites The mamba in the llama: Distilling and accelerating hybrid models.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference The mamba in the llama: Distilling and accelerating hybrid models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:05.312199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:05.312199Z digest=sha256:e44d6ad28ee3f113510357009c78d577d7e53e7627dedfd9a7e66262f6c00054

Observation fdc4b953-2b0c-4dcb-ad3c-a0beca44416d · outbound

This paper cites Liu, and Matt Gardner.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Liu, and Matt Gardner

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:05.422514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:05.422514Z digest=sha256:78089743093c91c40ace81fce73cbd28048b6f9124bcc1ee49365f092c791033

Observation e262e16b-7752-4a2a-817d-1a3ab3e0cc1f · outbound

This paper cites Efficient streaming language models with attention sinks.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Efficient streaming language models with attention sinks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:05.551726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:05.551726Z digest=sha256:0ff6054e79bc268a0117eca9aaabe7cd5477b8902d5f5b869e2c29bec070ed68

Observation 1a1b167d-f5a6-4487-9c26-bd4d168218c2 · outbound

This paper cites Pyramidinfer: Pyramid KV cache compression for high-throughput LLM inference.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Pyramidinfer: Pyramid KV cache compression for high-throughput LLM inference

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:05.674264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:05.674264Z digest=sha256:0c58511a0c4acb93a5caca22853a68b8675c07d3c32568f7f7f07ccab5f951de

Observation 75b48b7c-4374-4c56-bcaf-eb1b40817ffb · outbound

This paper cites Native sparse attention: Hardware-aligned and natively trainable sparse attention.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Native sparse attention: Hardware-aligned and natively trainable sparse attention

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:05.742499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:05.742499Z digest=sha256:a70800544c648c18584cfae848096982a02b307307e6b6dec40d2506e1609328

Observation 32d47e24-cf6a-4ab0-bf9e-96dcccc5e6b0 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:05.810453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:05.810453Z digest=sha256:f0f323b8fc58141109c99b27713d6b626e657e9bb8aa25c0ecd9208db302b475

Observation 1f0b6726-7019-4b05-92df-f50d6e2c8a91 · outbound

This paper cites Event temporal relation extraction based on retrieval-augmented on llms.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Event temporal relation extraction based on retrieval-augmented on llms

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:05.888909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:05.888909Z digest=sha256:0ef7c341091edaed530b33b445c30d55134c0842be08e3b3291e4ee75494f2f4

Observation 5503ea8f-c7cd-4436-8c9e-7b7c4c89d781 · outbound

This paper cites Barrett, Zhangyang Wang, and Beidi Chen.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference Barrett, Zhangyang Wang, and Beidi Chen

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:05.950819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:05.950819Z digest=sha256:97391b29641b51c8cc29c3b23e0fe5d6a0df7580f3b524e5f18e6290adef7174

Pith citing papers

No inbound Pith citation observations are available.