Pith. sign in

Paper Citation Record · LEDGER

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

As of 18 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 16 inbound Pith citation observations for arXiv:2507.07400.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07400 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:45:42.411640Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:58:39.204380Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ddc67721-7ff7-45d2-a005-486c7bc7f317 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows React: Synergizing reasoning and acting in language models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.222956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.222956Z digest=sha256:f87036d5f0881036c7928aec3f44a71e43ad6b8fc12f299cea67f870e1077433

Observation a60f3c9d-5f10-40f5-95ac-be25f82fab4d · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.227937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.227937Z digest=sha256:4400bc42d53054251605b89078ef7247e7ea6e6dc66e386e0fe70fe0aeb2eee8

Observation 55921259-8778-4990-a660-09036a7a2743 · outbound

This paper cites MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.233314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.233314Z digest=sha256:c99543a39123e8ad99be246071ace0b8d67f86338088bc1fe192fe9d655d3906

Observation 7914f024-1158-4497-b3e8-1fb12c7a18a9 · outbound

This paper cites Camel: Communicative agents for" mind" exploration of large language model society.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Camel: Communicative agents for" mind" exploration of large language model society

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.238391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.238391Z digest=sha256:aa51fafd6595f989c8b0658755608dbb033fe9732f450ae774a6262f64dc1226

Observation d96f7d36-be33-47db-9b86-ca3418007a30 · outbound

This paper cites PEER: Expertizing Domain-Specific Tasks with a Multi-Agent Framework and Tuning Methods.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows PEER: Expertizing Domain-Specific Tasks with a Multi-Agent Framework and Tuning Methods

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:45:42.772387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:45:42.243742Z digest=sha256:559c205b8174132da54de132c3db39c8e914fbf231ae4f1741856c5106beec1b

Observation 90d6b685-6b40-4beb-9549-a1b1c9b8e75f · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.248910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.248910Z digest=sha256:c88a5d585a8dd4831a9eba43faf3778785755fcb8ac0d6f3c43656b979a07829

Observation f7eb983f-f5fe-46e0-b19b-ab1d3674e4a6 · outbound

This paper cites Gptswarm: Language agents as optimizable graphs.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Gptswarm: Language agents as optimizable graphs

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:43.022975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:45:42.255938Z digest=sha256:e400015073d1af5f4ed8499c7453164b0e082a36c30514cc32e269db7f948825

Observation bc9def31-20cf-4e04-a077-65e5c6cd55e1 · outbound

This paper cites AFlow: Automating Agentic Workflow Generation.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows AFlow: Automating Agentic Workflow Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.261460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.261460Z digest=sha256:d77686058cc671be1cbde262576941d25b6c1e8e4bab7fc1cfbddc4ce74fdaa2

Observation 7c2ce70a-94c3-4c42-affe-fe97b883e641 · outbound

This paper cites Very Large-Scale Multi-Agent Simulation in AgentScope.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Very Large-Scale Multi-Agent Simulation in AgentScope

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.266656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.266656Z digest=sha256:4bc51109bd607d579ce1bff8aa56ed551871d3b6ca1e0e39a4f9a247bcdbb454

Observation ce19083b-5ba9-4f51-b18a-a88d0bb19648 · outbound

This paper cites Cognify: Supercharging Gen-AI Workflows With Hierarchical Autotuning.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Cognify: Supercharging Gen-AI Workflows With Hierarchical Autotuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.272015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.272015Z digest=sha256:135b51b75e0f906df82f1ba148a29d707b45d0cf121f493767b3d031028bdd2d

Observation 96a5f89f-c5d8-4ad5-abee-3a93b3096963 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Efficient memory management for large language model serving with pagedattention

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.276704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.276704Z digest=sha256:a8a698e066963aefde3bf1527f836665191a38c3ff32f8ada292a2b72b0e3f98

Observation bda1cd92-185a-472c-9c29-957d9edb573a · outbound

This paper cites Gonzalez, Clark W.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Gonzalez, Clark W

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.996931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:45:42.281347Z digest=sha256:d295c67018de9734aeb61f1cc430d19718ab038213bf2e21ecac970946c7fabf

Observation c7c4aa97-e930-41cd-bf27-9bdc2e4bce7c · outbound

This paper cites TensorRT-LLM.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows TensorRT-LLM

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.982123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:45:42.286228Z digest=sha256:2d4ce5eef698b7c9db8776a4f546761af4c814dde213aa4bf2feb434473bc9ff

Observation 687a0462-43bb-40c4-8c89-d89cf734eba8 · outbound

This paper cites Automatic Prefix Caching.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Automatic Prefix Caching

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.966950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:45:42.290936Z digest=sha256:5cf5b5bc9a96f3cdded8c300e42bfb92f7301f6df07b738eec2adb925eaf0446

Observation 74408f15-ecf0-4635-b034-f7e1da15f41e · outbound

This paper cites RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.296114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.296114Z digest=sha256:ccb3f78c9dd604c5cb223d883e2d6c5866343e3f3610ece5767e04c319078a94

Observation fb353792-c84a-45fa-98b5-ff09cd5c5a15 · outbound

This paper cites {Cost-Efficient} large language model serving for multi- turn conversations with {CachedAttention}.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows {Cost-Efficient} large language model serving for multi- turn conversations with {CachedAttention}

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.951768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:45:42.302064Z digest=sha256:d6909f67049e7394a6b9a484a3a391915025a27d15e426088adc74ea02db83c2

Observation 6426a51c-213d-4d06-b5d7-2c56a4c8a539 · outbound

This paper cites Generative agents: Interactive simulacra of human behavior.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Generative agents: Interactive simulacra of human behavior

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.306995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.306995Z digest=sha256:972459aaafd2247e55c9452c8029f32c6cb696eeb40a7047efa7ad033def357b

Observation 96815315-538e-4beb-b77a-fa09f0deda9a · outbound

This paper cites A survey on large language model based autonomous agents.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows A survey on large language model based autonomous agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.311794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.311794Z digest=sha256:276229cd90c639640b1e6847fac28ddb27a0ca8c1497d6c38964e34643988ca6

Observation 8cf33364-7348-4f83-9dcf-2c133442380a · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.316209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.316209Z digest=sha256:7076e2c79fa23c1b7716741be06aed492f1e0023bba5c4e84b85a4b9d752f257

Observation 6ec7bd6f-7b16-49e6-be9b-08466e82c288 · outbound

This paper cites Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.320701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.320701Z digest=sha256:bb13e357042343d0dd2d351560722e3d6c1a5400326f3c9b8b7345c4fd27a51d

Observation 007a7681-db51-4dd4-9127-da403a12d622 · outbound

This paper cites Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-Collaboration.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-Collaboration

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.325909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.325909Z digest=sha256:40bdda2d264a2083840e2eef52e47735daf24e6f1fb2823262555a8d5499b01b

Observation efa75a61-7f40-4e16-8f3e-6bc7ec6816df · outbound

This paper cites MAGE: A Multi-Agent Engine for Automated RTL Code Generation.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows MAGE: A Multi-Agent Engine for Automated RTL Code Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.330477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.330477Z digest=sha256:9ad695c3778b73571e41d983b0438fe540b6cb0bc3d8794ba003c25994daa4da

Observation 0af19f0d-9703-4ccb-8a29-fe0a9137a7e4 · outbound

This paper cites ChatDev: Communicative Agents for Software Development.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows ChatDev: Communicative Agents for Software Development

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.334764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.334764Z digest=sha256:d4de392d7c0565eb5553e12191f50faaee079eb109889b1095ee3729d6b1f36d

Observation 071ae83c-0559-41d4-9882-9ab8bc5e9250 · outbound

This paper cites Improv- ing factuality and reasoning in language models through multiagent debate.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Improv- ing factuality and reasoning in language models through multiagent debate

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.339249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.339249Z digest=sha256:a121dda766035a61f4af766af19ce23aee5ac530a1bb41b030b27f47de17e859

Observation c39ff46b-ee40-459a-8bda-3d5cb2e53ba3 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Accelerating Large Language Model Decoding with Speculative Sampling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.344014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.344014Z digest=sha256:470c8e0461831630bb709887b4ad21b3d73126c16a27dada7d8e102badea2b23

Observation 5a0484f4-020a-4727-8d2e-9342846f4c11 · outbound

This paper cites Fast inference from transformers via speculative decoding.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Fast inference from transformers via speculative decoding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.348481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.348481Z digest=sha256:ae7c89326c3bf30b88d7adf7cb700e0bd10d34baaa51c08b62c1eca0a69d30a5

Observation 3ba568b7-6fb8-4c60-afe1-3312768270a2 · outbound

This paper cites Prompt lookup decoding, November 2023.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Prompt lookup decoding, November 2023

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.894668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:45:42.352902Z digest=sha256:ec419b40156634e512e01c1d34e92d9c06119ac4beff723173691164025fcdc1

Observation ef4c5fae-3d8f-48b2-a03c-032fc212b8eb · outbound

This paper cites Efficient streaming language models with attention sinks.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Efficient streaming language models with attention sinks

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.879053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:45:42.357195Z digest=sha256:178b3d09775cecb47680fb514316c8f7afefbaa50d767219358fd22ef98836f5

Observation cf829135-f430-4acc-8754-43b503bed880 · outbound

This paper cites Efficiently Scaling LLM Reasoning with Certaindex.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Efficiently Scaling LLM Reasoning with Certaindex

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.361717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.361717Z digest=sha256:56f07f20176a75017e964f7e7ec4ecaddbf6b5fc008baed1e889cc42f60bf55b

Observation 218cb1c7-c682-42ad-a1f2-aa2bae8309cc · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based} generative models.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Orca: A distributed serving system for {Transformer-Based} generative models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.366575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.366575Z digest=sha256:01a298e29cfe778f5825ef2dd47398c8c9a0b67e057f2159a7d1ddb92a317a0b

Observation 8fc98626-c13a-466f-aff9-12bf892ad6e8 · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Fast Distributed Inference Serving for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.370910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.370910Z digest=sha256:8c948e1725ff7ccc2a0c3d8cab9b17fbec652a804ca13cd04664806fcbdbb89b

Observation 66d38722-36b3-4010-bb3f-fde70f928c6d · outbound

This paper cites Andes: Defining and Enhancing Quality-of-Experience in LLM-Based Text Streaming Services.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Andes: Defining and Enhancing Quality-of-Experience in LLM-Based Text Streaming Services

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.375175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.375175Z digest=sha256:16c9e9ff67af3ca85c8a1ee42e35200520e21e56421f6c99577cdb7055f3785e

Observation 39572001-5ffb-47af-b551-e1bd8474c506 · outbound

This paper cites Stateful large language model serving with pensieve.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Stateful large language model serving with pensieve

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.853980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:45:42.379599Z digest=sha256:5a2ba9280db0cff1cb7b47d02c1fbd441c1f9d38ed66d35ca53ad293d937262d

Observation 519ae05f-2cce-46bd-92ab-866c62479560 · outbound

This paper cites InferCept: Efficient Intercept Support for Augmented Large Language Model Inference.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows InferCept: Efficient Intercept Support for Augmented Large Language Model Inference

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.384225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.384225Z digest=sha256:ac5887de4d456c3c592c52de0a4bec826c671348e63a013d5804263afaf2f665

Observation f300aba2-02cf-48da-a93b-cbe612a86f26 · outbound

This paper cites Autellix: An Efficient Serving Engine for LLM Agents as General Programs.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Autellix: An Efficient Serving Engine for LLM Agents as General Programs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.388683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.388683Z digest=sha256:e2a7f4ce8aaeefa8e283f929d550249f7bf991cfb96b19c53d7059d83ee256ef

Observation 99c684cf-6b44-4846-b990-b15ee6860f83 · outbound

This paper cites Parrot: Efficient serving of {LLM-based} applications with semantic variable.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Parrot: Efficient serving of {LLM-based} applications with semantic variable

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.837952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:45:42.393131Z digest=sha256:ec5112d8ad0f114138cca144e92a14d2a504a21f22e93ef8e27872545dbeb7d6

Observation 72716aa1-a196-4ab0-9289-d329a85fc8d1 · outbound

This paper cites LangGraph.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows LangGraph

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.821121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:45:42.397557Z digest=sha256:b37a55e1329abc366fdb7630691d3fb3170184c2d3ff9c5cc300ecbbb9344193

Observation 9e15dc4e-6d66-4013-92a9-70b26b4662fe · outbound

This paper cites Building effective agents.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Building effective agents

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.805802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:45:42.402250Z digest=sha256:e7f8187534a176014c0fd100fa18613e50b3fdefae60552d0ed8786f89807614

Observation 13f354ff-166e-4bb5-ace0-ee2b53f11b46 · outbound

This paper cites AgentScope: A Flexible yet Robust Multi-Agent Platform.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows AgentScope: A Flexible yet Robust Multi-Agent Platform

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.406791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.406791Z digest=sha256:e5c5bf54d655ff4fb691b6420a1cfac301d1e209ad894f7df7d312266cfb4f69

Observation 887322ec-44c8-4927-a54c-ed5fc257900b · outbound

This paper cites Cut the Crap: An Economical Communication Pipeline for LLM-based Multi-Agent Systems.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Cut the Crap: An Economical Communication Pipeline for LLM-based Multi-Agent Systems

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.411640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.411640Z digest=sha256:443b9478c15d2b5f145a7e0c727b023166d965fce64b964b060f56820619aa42

Pith citing papers

Observation 26d3d1de-fff5-4a5c-9e1a-38e9331db713 · inbound

Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving cites this paper.

Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T22:20:51.838309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:20:51.838309Z digest=sha256:ea8800de733084c6da2dfab7cb5d77d39907433da5f4c582855304ef59e09296

Observation 8504fea4-ac1a-4c6f-ad49-df4739e91feb · inbound

ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache cites this paper.

ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:25:46.907699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:16:49.292491Z digest=sha256:15c76a0e00054c56120484eccf020e05163b14c844d1e69cde4810ff32290536

Observation 9bd72f02-e53d-4bdd-a705-6e85b90aa4a2 · inbound

Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling cites this paper.

Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:36:36.686395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T06:31:36.776819Z digest=sha256:df7c49ca7fac7618b3eedec877e89b97b2035b8b4cd7a6f0ee48d50438407b03

Observation ed36350b-6157-449b-bb3f-bb64896b7591 · inbound

PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference cites this paper.

PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:34.256580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T03:55:32.585495Z digest=sha256:6098ca50a30813afe32bb58d1928af5443087709bbc148552bec417d63e0ee2d

Observation d380eb5e-e830-48b5-bdc7-32b31d024ce3 · inbound

AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving cites this paper.

AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:46:13.127977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T14:16:46.241126Z digest=sha256:af69441745dd2f2fa7536ee0e85169909edb529aeb44e313e82363dc429386e7

Observation 903503ab-0c69-45a8-bbe7-430b6ddc9c2a · inbound

PRISM: Fast Online LLM Serving via Scheduling-Memory Co-design cites this paper.

PRISM: Fast Online LLM Serving via Scheduling-Memory Co-design KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:26.474892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T00:53:50.578344Z digest=sha256:25644e7517a146e67b38cc9aa86b415c4845b0932e6f497b1ce5cbdc6d168b2a

Observation 3d13b212-72be-4504-8766-0d483000950c · inbound

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference cites this paper.

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-19T22:12:51.215622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T22:09:15.782104Z digest=sha256:d3d597eb448a724376c47bbb561e6cf36e02c101697b5eabc515b3f6ac29e983

Observation d7ec5bef-2697-4fe9-be9a-fa4b5f8e50a3 · inbound

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches cites this paper.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:09:07.016867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:83c114940372dc3ce35ca727fb82bbbf680acaf02883d8be8009d5313f2ed336

Observation 061ca800-78f4-4b34-b6dd-832b565c2b15 · inbound

VineLM: Trie-Based Fine-Grained Control for Agentic Workflows cites this paper.

VineLM: Trie-Based Fine-Grained Control for Agentic Workflows KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T23:56:22.405561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:56:22.405561Z digest=sha256:af2e79888e3ba9cf9c091147724e864754b126e8526af3bc9427d73e8bd3d5a0

Observation 22dab550-d3c8-4bb7-9865-251ed0bc1eaf · inbound

Resident KV Claims: A Conformance Contract for Future Reuse under Active KV Pressure cites this paper.

Resident KV Claims: A Conformance Contract for Future Reuse under Active KV Pressure KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:24:45.202521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T14:18:35.209571Z digest=sha256:8e86568dc6f9bccc98c4bce1cea1020e81874e0f17412a353791497431317a48

Observation 57a31b54-9b25-48a8-92a6-c6e5d950bfc7 · inbound

VikingMem: A Memory Base Management System for Stateful LLM-based Applications cites this paper.

VikingMem: A Memory Base Management System for Stateful LLM-based Applications KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:33:13.635254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T07:30:42.104455Z digest=sha256:9e596097bd881ea5b192aa67ea6299944e37b9c1b09a97297a97fc22de3ac7be

Observation aec54df1-77a5-4946-b6ae-20d28c31d373 · inbound

Streaming Communication in Multi-Agent Reasoning cites this paper.

Streaming Communication in Multi-Agent Reasoning KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:06:47.918375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T06:24:07.052856Z digest=sha256:eb87c164dbd1b11cb544b4b57cc5fbdcf66deb3f84a9477b6c539ef467dc7f8c

Observation 4ae630f1-8f27-4dba-8663-26525168fd5d · inbound

Streaming Communication in Multi-Agent Reasoning cites this paper.

Streaming Communication in Multi-Agent Reasoning KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T04:58:39.204380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:58:39.204380Z digest=sha256:1775faeb405d7fd990f3972c1789153d4bfed4f6d028e2d8516eec4d4acca2e3

Observation 4e8cb190-cd0a-4d88-8fe9-f9124eebd6a0 · inbound

SmoothAgent: Efficient Long-Horizon LLM-Based Agent Serving with Lookahead Context Engineering cites this paper.

SmoothAgent: Efficient Long-Horizon LLM-Based Agent Serving with Lookahead Context Engineering KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:14.694422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-02T17:19:53.560039Z digest=sha256:e3ef61d8cc40d545ee90628c0297b587bd0775b0ec4f028a560910d9e4f065ab

Observation bd742d16-48fb-4143-9404-57f04d3397a2 · inbound

Workload-Aware Caching for Multi-Agent Systems cites this paper.

Workload-Aware Caching for Multi-Agent Systems KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T11:22:57.020958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:22:57.020958Z digest=sha256:6e8b0bfa1d4648dedd8cb00d5900803bbb89ba18347fe2132ffae2c59783fd66

Observation 0a65ee2b-195e-4223-ba1d-95c47b4382bf · inbound

Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework cites this paper.

Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T14:14:12.063684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:14:12.063684Z digest=sha256:870007eb638ea84417dcba1b95678490df6e317b2a09d31d261899c97118927e