Pith. sign in

Paper Citation Record · LEDGER

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization

As of 8 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 2 inbound Pith citation observations for arXiv:2505.19586.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19586 v2

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:16:42.984625Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:45:12.553567Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T17:55:13.328095Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0009c9d3-4a69-441a-84cb-52e148f9e73e · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:46.286667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:16:38.963314Z digest=sha256:d95a9c3af1374b83b6f1755d75f76d5143450411fbd7e90c2a495d38b932dfc6

Observation e61ec203-ee94-4a34-acf9-b5da69754873 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:46.145966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:16:39.013330Z digest=sha256:39004cf6f542eafc3f7de3bda117ca4b5468ff4ba999a0bab381b551d4062830

Observation 56da942a-5538-4a0f-801a-95e2e1113a05 · outbound

This paper cites GPT-4 Technical Report.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:39.105730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:39.105730Z digest=sha256:ceb9bfccbd1ba88d35b9a61e1efa0ecd281f9de58d71409ef1bc9a26772fd38a

Observation bff1fdef-e70e-4e89-ae0f-f8f661f4b18d · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:39.190085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:39.190085Z digest=sha256:5ee480a8da6d87f151483f20b8d60ae68255e76757a48c9e4d4d255a8c390ad9

Observation 3da24f0b-5d7d-4529-b2cc-5e6e828b937e · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:39.266790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:39.266790Z digest=sha256:039c99098f1e7abfe0d2efdb770f15b624a7d96763c282633ec835d14bcd3308

Observation b0700a1d-b0be-4d48-a286-c8dc4035002d · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:45.978753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:16:39.358085Z digest=sha256:ee8a3c61623e925527d614aa58ce2b497197424b864ca7b14692ae4d1e07313a

Observation c358e741-9b03-4389-adaa-4897f3af5a9c · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:39.438404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:39.438404Z digest=sha256:e2989e6ea1ed3f9d15f3e8c5069b97adb84c405e872270ff226b875ecf225267

Observation 7917efbd-44f6-4c9b-adca-9457f47d1654 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:39.530091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:39.530091Z digest=sha256:f8f17f435d9c7c94ee5afe8d2b47493e283c0190dfa2504eb15f48cf94d0a654

Observation 050c77f7-e884-40c9-8c1c-5b1a9e8b8769 · outbound

This paper cites The Llama 3 Herd of Models.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:39.619435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:39.619435Z digest=sha256:2577fd2bcb360ad1f5f9890925c2fa04cf75e2a5b9618e941ebce89fcdee0cda

Observation be925673-3d0b-4f96-ba40-04959777713c · outbound

This paper cites Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:39.736272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:39.736272Z digest=sha256:85fab4fd147c7ca6071928a528e843dcfc0c855ab3b68a17def2442475aaf21f

Observation e5b8f505-a5b4-486f-add0-f3ca2a95bc1c · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:45.756652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:16:39.882527Z digest=sha256:978bf2120ba6268f5833288063c5bcc8c720562794c1142d72df77413178b6b1

Observation 75d60ba6-1429-42b5-9726-e7d47c12acd8 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:39.957230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:39.957230Z digest=sha256:3fd0c019795c328c527cfeb6c58c2f111c76ab317d62f1e40033fc4887306438

Observation 1d546562-dc02-475e-9b47-73266fb0ab03 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.059217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.059217Z digest=sha256:a9a38dc0cbf3489220aec362b5a439eb0cdfd7af8a2eb6d8f268c3c5f99cb371

Observation 4d09a508-089b-47bb-a5f2-97700cf23728 · outbound

This paper cites Abdi, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Abdi, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.185668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.185668Z digest=sha256:404ee9fb599a220f25d7c3a45dea2c7c4af8799342730af29ac3e03a06a55a0e

Observation 7e68ccd2-51c8-47b1-a411-fa6b04f267a7 · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.300614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.300614Z digest=sha256:ab60fead8fc2138e2eabfd0f16a239097eb45081170a9d6d1e3e963dd87e9691

Observation 83393d5b-049a-4ca5-bb72-3ce241205685 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.399838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.399838Z digest=sha256:988ef71b7448dea71efb21335cfde48999add680eddc0f36b6cd569e66339acc

Observation 0ba4e595-50ed-412f-a995-e05f8a7aeaea · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.533731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.533731Z digest=sha256:bff11eb344f1a2551956f2a4e94c3c71bd153054eeea1fa2cdcf003a99d811ef

Observation 9c533e52-2d32-4895-af8c-64dc29d1e3eb · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.602920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.602920Z digest=sha256:7b73e8176819352030d487e72362be6acbdac293d2d70be1c34ec764899580fa

Observation 2a066f05-082b-49cb-b01b-6ee0622ab53b · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.733388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.733388Z digest=sha256:7367fa294ce1890f801ae56a17724401f5947d11263fad77fe77b2ea84de3ec1

Observation fd7275c9-816b-45ff-96c9-3afbc57e709d · outbound

This paper cites DeepSeek-V3 Technical Report.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization DeepSeek-V3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.840599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.840599Z digest=sha256:41f9d870928574af0934a562decb2d2b62db32b7aa86f9e1f47408ee91933078

Observation 9a04388b-1d84-45e3-b2ab-1a747ca1e589 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:45.426399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:16:40.929427Z digest=sha256:d26ba0a2cfbe781d2982ee473499aad84e3f5635e7acaeaff34103116071aea8

Observation 8cd1bf44-9bf8-486c-96b6-084e41673e6f · outbound

This paper cites RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.998220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.998220Z digest=sha256:b61fce5f02e517222a0397417df7d0c067231b2295629f3f14d213491ad0360e

Observation f66f4c85-15b0-439d-8ff4-9ce75c957d13 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:45.178131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:16:41.097932Z digest=sha256:e7567f82f0e5f156ed7deb3c2af8db6923c9719cc1f76e6720d898eed06f0bfc

Observation 663a592d-2f60-412c-b8b6-82a99ede768d · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:44.909157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:16:41.196983Z digest=sha256:75bf4f131ca192679afeeeb61ce554cefa017fb235e3aa78fb6115932509b177

Observation a079a4d3-391c-45f9-b92e-54d21c51123b · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:44.754722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:16:41.296524Z digest=sha256:77fe9120818cc95245fb0a52ed5fde4655c97f6335c231f0e6230ddbaacb474e

Observation 1a8fe988-6039-4b96-bfdf-7a2269b3c748 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:44.592866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:16:41.372786Z digest=sha256:30d87ce7889575b4b7ce0beadbb7a7fdadc543ef9866af8f3b054777a269f9a4

Observation 52e8d0bc-9369-454a-a3b1-c7267ef526cd · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:44.434457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:16:41.464856Z digest=sha256:3064190929e4cd3cf7eda2faec78ad7c3d1366d333b8837ade39a250088443d5

Observation 83f3bebf-fe43-41b5-ae85-56d670b907ed · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:44.304156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:16:41.553509Z digest=sha256:9d220c8428bded6fa88efa005144521a7876a9918118ea447070278a94b1103c

Observation b100b888-b2b6-4049-839b-082a781e0eac · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization LLaMA: Open and Efficient Foundation Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:41.649120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:41.649120Z digest=sha256:b852f76d3b03cad52561566506ed2c524e9ed0db542793b8d00268ec483730d8

Observation 868b6744-535c-4cc2-8824-f3fd5b1dad29 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:44.083353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:16:41.769077Z digest=sha256:91928d441b03bd822463a58da00d99fb757ea0c94a2576f726985e97f82b00bf

Observation c4e447d9-8d09-4019-8676-3864b2e16da2 · outbound

This paper cites Deep Graph Library: A Graph-Centric, Highly-Performant Package for Graph Neural Networks.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Deep Graph Library: A Graph-Centric, Highly-Performant Package for Graph Neural Networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:41.882697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:41.882697Z digest=sha256:141f34ef14c4ba4fad37137ed6448b545f5359a0ea648530872938f721fe0f72

Observation 7924559a-cd26-4bca-b90b-b394ea4b6781 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:43.891966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:16:41.976279Z digest=sha256:d3faa105ecf9e273e55e0a9ef04d92501ce31940ef2d21b3a0728dfb7abea91a

Observation 0fa3dd3c-f37c-4e1f-a563-75bc8879d05f · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:43.696181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:16:42.074404Z digest=sha256:53b812112c9094190c36e2b5d59b98245ac552b42dbb18d5fa18ce48a574a2fc

Observation e3a1e404-48ed-442b-abc8-6981acdd4d24 · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:42.177025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:42.177025Z digest=sha256:21c077306a250b315ca352657ec57274c89d447b7666e84b97bfdba3b22a5eb3

Observation 8bae7c40-8cad-4de8-b326-6a4e6348b460 · outbound

This paper cites Post-Training Sparse Attention with Double Sparsity.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Post-Training Sparse Attention with Double Sparsity

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:42.302352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:42.302352Z digest=sha256:c4f6589cbc14bd3a63460fdf3b51049834d1c825eae69a2685b0cafcec4bb57a

Observation d06da1e1-b521-4fa2-8675-91a428ba9431 · outbound

This paper cites PQCache: Product Quantization-based KVCache for Long Context LLM Inference.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:42.391215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:42.391215Z digest=sha256:4b90006cad3b69a8cfaef53e29433cb02b15ffcec0de110870220822ccc93f7d

Observation 32ac6e75-29cb-4ae5-8a7b-7d439bd44850 · outbound

This paper cites Dovetail: A CPU/GPU Heterogeneous Speculative Decoding for LLM inference.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Dovetail: A CPU/GPU Heterogeneous Speculative Decoding for LLM inference

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:42.476185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:42.476185Z digest=sha256:5bbf9314230ddf96485cfc635dbf9ccd6b182c1be9fe99643656290ff3692c96

Observation 11eb4041-968b-4b28-860f-0d80e8b8c153 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:42.573762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:42.573762Z digest=sha256:da3b508fb1c1dfd414bd7c4b848ecc800e6719822cf677b1a2357c211328122b

Observation ceee1b52-8eca-48c7-8c1e-6345ac942997 · outbound

This paper cites LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:42.683728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:42.683728Z digest=sha256:fcbd16ba5b0df5ed8f3b0207226f284df730cdef158c794c2653f9cdd04a09b1

Observation ef746939-1c9f-47a3-9895-200ca4c889e3 · outbound

This paper cites Barrett, Zhangyang Wang, and Beidi Chen.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Barrett, Zhangyang Wang, and Beidi Chen

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:43.494267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:16:42.793624Z digest=sha256:e6e08fbf2b12f568748ccfb4efb905be9445b2c9598e6b584d6381403c656cb0

Observation ba9f6c0b-ec83-4895-932f-6a81fdad01d7 · outbound

This paper cites online" 'onlinestring :=.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization online" 'onlinestring :=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:42.896424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:42.896424Z digest=sha256:7cde98661081c333c8a3378cd98bcc7ce548dd1cf2079c34990b2c0d3c41e2fd

Observation 5a44742e-afbb-46f6-8f28-6abae5b91aea · outbound

This paper cites write newline.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization write newline

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:42.984625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:42.984625Z digest=sha256:ce8c8dc330c6921592b3bd2165af70efc27ad763a324e2e77b8d04468ea876de

Pith citing papers

Observation c81e1b24-b203-471f-98e7-755364f2c47d · inbound

LOOM-Scope: a comprehensive and efficient LOng-cOntext Model evaluation framework cites this paper.

LOOM-Scope: a comprehensive and efficient LOng-cOntext Model evaluation framework TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:12.553567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:45:12.553567Z digest=sha256:54b326591af5f03acd315b332b937a3c4bc786c5bfcd2957f597c1fdb98396fc

Observation 06792c92-3127-4413-b66c-d107532327c0 · inbound

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference cites this paper.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:55:13.365091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:55:10.372511Z digest=sha256:68eb7bc02d19013278101a00915dce3de5b3a30f07934a1966c4df40d0b5d3ac