Pith. sign in

Paper Citation Record · LEDGER

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure

As of 10 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 0 inbound Pith citation observations for arXiv:2608.06007.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06007 v1

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:23:08.569863Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

88 of 88 outbound references displayed

  • verified exact2
  • verified fuzzy51
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7266694b-2167-460e-a9a2-c6a1461ddcae · outbound

This paper cites https://www.kimi.com/blog/kimi- k3.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://www.kimi.com/blog/kimi- k3

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.210099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.210099Z digest=sha256:61d4d3ded574984a5f7dfccd09d71ac984f215ba49807a6afe44a535ee73b585

Observation c5640c60-ddb8-46c0-820b-289b6a800bd5 · outbound

This paper cites ServerlessLLM: Low-Latency serverless inference for large language models.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure ServerlessLLM: Low-Latency serverless inference for large language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.215746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.215746Z digest=sha256:7b43313c14dce3197f07c7e16acae8fcf7dec85fb0f74310e91cf7d97ddeaa48

Observation d373c80b-505e-4110-9072-7ef34e468468 · outbound

This paper cites Deep- flow: Serverless large language model serving at scale.arXiv e-prints, pages arXiv–2501, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Deep- flow: Serverless large language model serving at scale.arXiv e-prints, pages arXiv–2501, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.220987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.220987Z digest=sha256:c9fb64fb9e26ee6c79e046735ff308b35338c9d4f32a4da02dc009478db8c6c7

Observation 059a9267-99f7-4671-b484-7768bf48ee96 · outbound

This paper cites Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.225722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.225722Z digest=sha256:fa943dcdcfe09722863a49f5aa625779d5ad69b3451a133a84196218f8b25557

Observation 7bb633c6-8d07-43e5-a398-f51152382d11 · outbound

This paper cites an unresolved cited work.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.230680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.230680Z digest=sha256:cda7fb58592b9cb478925b322c649c0d6ca07fd939d4461bbdca635050d8de43

Observation e4436a8a-eac5-4d58-a312-6f25afa6d915 · outbound

This paper cites BlitzScale: Fast and Live Large Model Autoscaling with O (1) Host Caching.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure BlitzScale: Fast and Live Large Model Autoscaling with O (1) Host Caching

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.234958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.234958Z digest=sha256:d8c95aac4de041276c6cd5110e39f1b3eb15bb34f61e458afe1ae22761dc6794

Observation 4c447d41-9049-42d6-a526-314478418a4a · outbound

This paper cites Hydraserve: Minimizing cold start latency for serverless llm serving in public clouds.arXiv preprint arXiv:2502.15524, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Hydraserve: Minimizing cold start latency for serverless llm serving in public clouds.arXiv preprint arXiv:2502.15524, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.239382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.239382Z digest=sha256:05ce9a577dc755a5e26b992549c0c6e1eafe304e1923faae0ebe3a777e489ae9

Observation cff16d1b-78d6-4d53-afdf-1a465b2f75e0 · outbound

This paper cites https://lmsys.org/blog/2025-12-10-rfork/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://lmsys.org/blog/2025-12-10-rfork/

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.243347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.243347Z digest=sha256:acb3bda915b851e48ce09d5eddade003f621b4f7798e11b57106217071594ffd

Observation 9012b892-fc8d-46fe-9e9a-83c5ef0d592a · outbound

This paper cites Mooncake: Trad- ing more storage for less computation—a KVCache-centric architecture for serving LLM chatbot.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Mooncake: Trad- ing more storage for less computation—a KVCache-centric architecture for serving LLM chatbot

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.247723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.247723Z digest=sha256:3b17cbdeca55c0f670204c2d57cb3e6365bf39062e7426896ffc0c4400967f90

Observation 214672d2-ef54-4bbc-9f95-a8ecd3432886 · outbound

This paper cites Dualmap: Enabling both cache affinity and load bal- ancing for distributed LLM serving.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Dualmap: Enabling both cache affinity and load bal- ancing for distributed LLM serving

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.252135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.252135Z digest=sha256:6925ac275773d6076a3a34ba0f6da0af67ecd59d93e3ad4ceacd7190561c33e6

Observation 530077e3-3828-4a3b-8d88-67308194f45b · outbound

This paper cites Lmcache: An efficient kv cache layer for enterprise-scale llm inference.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Lmcache: An efficient kv cache layer for enterprise-scale llm inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.256927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.256927Z digest=sha256:0e7df6161913733b9178808a2fc0dbebf4542d4368e14ebba392bbf10937ae70

Observation 1bab6d33-92c7-4dee-8e00-57df7e7504a8 · outbound

This paper cites https://lmsys.org/blog/2025-09-10-sglang-hicache/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://lmsys.org/blog/2025-09-10-sglang-hicache/

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.261289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.261289Z digest=sha256:c48121f64e069f542c55932153f464b4878ffa71b563985e131435e1a7b6fd9a

Observation 42f290ec-4062-48e3-bfca-971efb32b231 · outbound

This paper cites Stateful large language model serving with pensieve.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Stateful large language model serving with pensieve

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.647674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.265567Z digest=sha256:577ca166b8d2a0795dd642c4a39359aadc53a5dae21e2c649724ad046506ea22

Observation 9194a7a9-2f73-4e9c-b28f-1f56071a2766 · outbound

This paper cites Cacheblend: Fast large language model serving for rag with cached knowledge fusion.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Cacheblend: Fast large language model serving for rag with cached knowledge fusion

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.635789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.270101Z digest=sha256:88f00958d2853b00debd0ae6d32aeed17d3920fbbbb156ec76ce6ceb6782d29f

Observation da5d784e-d1e4-4906-82b2-e611454f3ffb · outbound

This paper cites Cost-Efficient large language model serving for multi-turn conversations with Cache- dAttention.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Cost-Efficient large language model serving for multi-turn conversations with Cache- dAttention

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.624659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.274263Z digest=sha256:10b3b3bc0ee21ef1c6e410d8c0dea32dec23ffc2a8f06a56d022cf8d4851b4bc

Observation 7b20c028-7b9f-4073-9b77-231851ecc7da · outbound

This paper cites {ByteCheckpoint}: A unified checkpointing system for large foundation model development.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure {ByteCheckpoint}: A unified checkpointing system for large foundation model development

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.613146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.278745Z digest=sha256:281bc05c11efc542f83c1eb43d1b1f025f6c505c41806ada7bcebc49b26b9153

Observation 4662b785-a21c-4aa6-a755-bec13fc799eb · outbound

This paper cites https://github.com/MoonshotAI/chec kpoint-engine.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://github.com/MoonshotAI/chec kpoint-engine

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.601572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.282891Z digest=sha256:db99cac92c68d147af154e3de9cbdace522b7f5f5dfc7137fbc64681408e09fd

Observation 31c3bf0a-5caf-4897-8dfd-3bc84f726e13 · outbound

This paper cites https://vllm.ai/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://vllm.ai/

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.589655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.286808Z digest=sha256:ed8c6cdb1e6d978219afa01a903e5bc5d763f05cae13156163d381916adcbdb9

Observation 4559eff5-182a-4fde-b2db-6844d6b4eef1 · outbound

This paper cites Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583, 2024.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.292156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.292156Z digest=sha256:b2c7d6256f713c13dd1b4ea066ec3f118e2656a4000c44519343c4bfe8d67083

Observation f1a612b6-281d-44a8-966a-de5cd8b8260d · outbound

This paper cites https://nvidia.github.io/TensorRT-LLM/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://nvidia.github.io/TensorRT-LLM/

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.571245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.296384Z digest=sha256:ba3c685bd81cca63ceb12359e5a8dfdf027b5b092ca2f76b0bda649a292b4cca

Observation 88da8772-4852-4076-aeed-3d06e9d38c5c · outbound

This paper cites an unresolved cited work.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:23:10.559915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.300372Z digest=sha256:067ded218f674282988fcc5334aee719ab410753532d82e13e6897c9393f0f3f

Observation db842daf-c122-4ea2-933e-7cb220a4e955 · outbound

This paper cites https://redis.io/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://redis.io/

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.546071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.304346Z digest=sha256:ab0b5054d04ce5dd6b99a520e886b7cbcd73934ccece672362c12ea1dfe905ca

Observation e67a7a63-1040-4cf5-a8c1-e12cedf1315d · outbound

This paper cites Vineyard: Optimizing data sharing in data- intensive analytics.Proc.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Vineyard: Optimizing data sharing in data- intensive analytics.Proc

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.534517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.308377Z digest=sha256:895b03ff8c1ea354a8ee95ce3d3b26926b2692881a8f03da7df9c9be124400da

Observation 45e262c5-30d2-4c6f-be79-3d1fc74ab155 · outbound

This paper cites https://github.com/ray-project/plasma.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://github.com/ray-project/plasma

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.522032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.312295Z digest=sha256:e23913fbed6fffd66671726e9c8cdcc0c873e96661e75b24e9a2e1375757106c

Observation 8436dcad-3029-4bdf-bb28-e3e864b50b22 · outbound

This paper cites Ray: A distributed framework for emerging ai applications.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Ray: A distributed framework for emerging ai applications

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.509673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.315974Z digest=sha256:2999b6bc86d562c3d2e668a68826c9028184f9a7f2c15050cba4e9d174a22dbe

Observation 76dc0904-c3ee-4e49-a3cb-3e5b4a919e09 · outbound

This paper cites Resilient distributed datasets: A Fault-Tolerant abstraction for In-Memory cluster computing.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Resilient distributed datasets: A Fault-Tolerant abstraction for In-Memory cluster computing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.497933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.320121Z digest=sha256:6cd0a3f9050500008467e4e85bcb5f6f2607853fc2f29398adb4dc37f61ad279

Observation 0d236e8d-f09c-4621-b24c-a6a9d6cc73ed · outbound

This paper cites Serverless computing: Design, implementation, and performance.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Serverless computing: Design, implementation, and performance

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.324667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.324667Z digest=sha256:3b3aff7231383519e1a25094c12c54e5b0e8e77fcd188ffb70931b50bd65c6d5

Observation b4725148-3cf4-48f3-bcf2-2e285e89f710 · outbound

This paper cites Infless: a native serverless system for low-latency, high-throughput inference.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Infless: a native serverless system for low-latency, high-throughput inference

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.478626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.328764Z digest=sha256:71f7f68fe420139f0aa600254d54132f446a54a9b6e01b8aa9f0215b09084e94

Observation 61af6026-a968-4494-9880-29304696fc9d · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.332432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.332432Z digest=sha256:71eae65c47b3355c355756be66a6d3e04e071c9389a8ade933c35a323159154d

Observation d04194b5-e4b6-438e-8cae-de4d804b89f0 · outbound

This paper cites Language models are few-shot learners.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Language models are few-shot learners

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.336296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.336296Z digest=sha256:583e8a78a37d78426078df7eefd97d5fbe40353746850b2ce205d72c0cffa9bf

Observation 8c2c0de8-44a0-435c-8316-63e889a005a8 · outbound

This paper cites Flexkv: Flexible index offloading for memory-disaggregated key-value store.arXiv preprint arXiv:2512.16148, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Flexkv: Flexible index offloading for memory-disaggregated key-value store.arXiv preprint arXiv:2512.16148, 2025

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-08-07T18:23:09.340739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.339917Z digest=sha256:98a8041b107e22606b14c403b565b2fbfba41e6b5fbac8d193ebac655532579c

Observation 770044c6-56af-42e6-a0b9-808c69d64444 · outbound

This paper cites In USENIX OSDI, 2024.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure In USENIX OSDI, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.450965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.344903Z digest=sha256:4a2918b88c9fc908ec7a931d463b3ba4987b57b97a08fa5fcef00d75e8cf7756

Observation c0f22ded-c796-44de-a7ef-351c921e3fd7 · outbound

This paper cites Splitwise: Efficient genera- tive llm inference using phase splitting.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Splitwise: Efficient genera- tive llm inference using phase splitting

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.438704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.348995Z digest=sha256:1f25e213cf75530b946fd858f53d38ac0eba1d3ad3a9ab834ccd40d74caf2e6a

Observation ea77d351-5772-4605-b59d-d2342eac07e3 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.353271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.353271Z digest=sha256:3849753606447abc9b9f67cc7ab93bbc5fcbe6b3f9ff4043e8849288b5fb9e19

Observation 60fe65b5-c7d9-437f-aa3f-c6cb4b316c79 · outbound

This paper cites D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.357426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.357426Z digest=sha256:fe179b3805769a9eb1cd1964415a74b4a5fc1cc6b3c1edd8c1d1bddfce8e756d

Observation e3d048ff-38d0-4f2e-9257-ef5544fc095f · outbound

This paper cites Prompt cache: Modular attention reuse for low-latency inference.Proceedings of Machine Learning and Systems, 6:325–338, 2024.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Prompt cache: Modular attention reuse for low-latency inference.Proceedings of Machine Learning and Systems, 6:325–338, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.425673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.361591Z digest=sha256:ef459ccb18ac78610577e1e5db5b2ced31ece3d987c8e8084cd77ddb1cc6edaf

Observation bcc27414-be25-4ed7-90ec-d706b5fad965 · outbound

This paper cites Cachegen: Kv cache compression and streaming for fast large language model serving.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Cachegen: Kv cache compression and streaming for fast large language model serving

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.414206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.365396Z digest=sha256:8a8a8f79c53cc678b9bd2fe290af76dddc53ac555cb5e2ca2c837c9ee921d606

Observation 1a8aac5a-b716-4e1a-b453-3e151a60595c · outbound

This paper cites Ragcache: Efficient knowledge caching for retrieval-augmented generation.ACM Transactions on Computer Sys- tems, 44(1):1–27, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Ragcache: Efficient knowledge caching for retrieval-augmented generation.ACM Transactions on Computer Sys- tems, 44(1):1–27, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.401403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.369036Z digest=sha256:701e20c94f5a04d8ffc46ede5e88f81aa0f059789d3c30e3e31714f5b0488cef

Observation 573cad29-4ab0-4ad7-ae10-34ce8d5c21d4 · outbound

This paper cites Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.373466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.373466Z digest=sha256:f4c20239c885ecb6318185ee34e37e605d406593a576dc0cc1805216641e59d1

Observation 254654b9-7c4a-46cd-95fe-2db36adfceca · outbound

This paper cites Dist checkpointing package.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Dist checkpointing package

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.386334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.378063Z digest=sha256:2a3afb747ff910a0033c374763d141b507ffa54a40546ada7328cd85cf794e88

Observation 69313a8f-f463-45a8-ac16-3a10cb5bdf78 · outbound

This paper cites Getting started with Distributed Check- point (DCP).

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Getting started with Distributed Check- point (DCP)

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.373829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.382598Z digest=sha256:232c18d8183c913074e96fff100389c774bcab963a59cafb2f179d26e704943a

Observation 5f164233-bc05-4186-942a-ec60c49025d7 · outbound

This paper cites Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.386378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.386378Z digest=sha256:55cc1318f3cec55166e4ba6bd523ee411f077341b229999b6a86627b80f3309d

Observation 9cfffabc-cb58-4e9a-b75f-bc7b8d852320 · outbound

This paper cites Simple is better: Multiplication may be all you need for llm request scheduling.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Simple is better: Multiplication may be all you need for llm request scheduling

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.359485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.390858Z digest=sha256:e752bd448fd7fb15b33e41b21d78cf02ff007cd9f0148fc55baa94a2d72e28a2

Observation 8a62c2a8-f310-4c44-aae6-6933689baeb9 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Efficient memory management for large language model serving with pagedattention

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.347609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.394615Z digest=sha256:eee4d7be3d11f66c79559d5ab3a997f696861fbf891934eef4fb3f97eaa746ee

Observation 48fdc530-c428-411c-87e2-112cddc9bfd4 · outbound

This paper cites https://docs.vllm.ai/en/stable/desig n/prefix_caching/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://docs.vllm.ai/en/stable/desig n/prefix_caching/

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.334793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.398238Z digest=sha256:3a8917fcf6ff2445e17ee7451cacdf720fd3192f94db24d7ea35aaceec5a7f07

Observation 2ec7bafd-e84c-4950-a29a-579d9f6908fa · outbound

This paper cites https://lmsys.org/blog/2024-01-17-sglang/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://lmsys.org/blog/2024-01-17-sglang/

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.322790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.401764Z digest=sha256:817d29cd1bd99212fed7d0e6c75151fdbd270ed3b9df565695b55d63045e848b

Observation aed7ed80-0a54-4021-9bad-50150fca6b34 · outbound

This paper cites Pie: A pro- grammable serving system for emerging llm applications.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Pie: A pro- grammable serving system for emerging llm applications

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.310698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.405393Z digest=sha256:76cd0426dce576b7e54e9ebf71a5cdb84144ffc02767bee43c69e83be74b971d

Observation 2d7fe69a-4765-4fa5-a54c-4bd2caefbe3d · outbound

This paper cites Chain-of-thought prompt- ing elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Chain-of-thought prompt- ing elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.409426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.409426Z digest=sha256:8e5917323c15760cfba39d3946db1502a9a9e04d99dd99c8094426d6e1cb2f15

Observation 4c63f9c8-ebc7-45d6-aa2d-1fb0fd3500ba · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.412959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.412959Z digest=sha256:ab8fdf6efe7e983be3c096a28d7a4c7bd5df911c55154bee6296c3983d668d03

Observation 60382737-0556-4911-ab3d-b66a51c7f2ec · outbound

This paper cites Training language models to follow instruc- tions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Training language models to follow instruc- tions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.282711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.416940Z digest=sha256:487301f1823b8679098ddbe8d7bb761d22dc8426fa90b04df5a89913a7064ac6

Observation 4c583cb5-6425-44b7-82ee-ab34222e266e · outbound

This paper cites Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.421353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.421353Z digest=sha256:57496d41a19dc81ee72d9e2e22fdb2709ec1a48a17c97bf51fa0b8d337a548db

Observation f8c235cd-4da4-4fe4-88dc-0cf04bcc8159 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.425755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.425755Z digest=sha256:4b5a08af323ae3c3eee555c302b122ce54e8d91279520e76ac4bccdda171628d

Observation 2847fa79-e643-40e6-a65c-fb56839f6f7e · outbound

This paper cites Totrl: Unlock llm tree-of-thoughts reasoning potential through puzzles solving.arXiv preprint arXiv:2505.12717, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Totrl: Unlock llm tree-of-thoughts reasoning potential through puzzles solving.arXiv preprint arXiv:2505.12717, 2025

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.430349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.430349Z digest=sha256:70c094a410d3f4169929cddbc619488890780c23e2e3b7d2e660b03dc6482d1b

Observation 50a88b26-30cb-4888-a5af-cec5627c30a6 · outbound

This paper cites Using name-based map- pings to increase hit rates.IEEE/ACM Transactions on networking, 6(1):1–14, 2002.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Using name-based map- pings to increase hit rates.IEEE/ACM Transactions on networking, 6(1):1–14, 2002

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.269755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.434044Z digest=sha256:d570b72a5552bfc8932142efa9d83da4198f799220026d0b1fe3922d7f468ad9

Observation 006bb2b3-6cfd-4972-844b-26a05685227f · outbound

This paper cites Paxos made simple.ACM SIGACT News (Distributed Computing Column) 32, 4 (Whole Number 121, December 2001), pages 51–58, 2001.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Paxos made simple.ACM SIGACT News (Distributed Computing Column) 32, 4 (Whole Number 121, December 2001), pages 51–58, 2001

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.438822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.438822Z digest=sha256:691fca910d4f993546c30607d49ff06f2c3a997405531e85b4eb47c5822b3b44

Observation 32b178f7-8443-4441-b96b-3e37cc30c25a · outbound

This paper cites In search of an understandable consensus algorithm.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure In search of an understandable consensus algorithm

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.249413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.442644Z digest=sha256:0b02a8fbef13864583266cdc310aa5f9ad5d1d84d5e9fac1b1bc3022bb277a03

Observation bbf2b854-54d3-4aaa-bb1d-fe5f9036b578 · outbound

This paper cites Chain replication for sup- porting high throughput and availability.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Chain replication for sup- porting high throughput and availability

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.238009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.446537Z digest=sha256:4aa289f3299477b3e91080cb5e6a3bdcacf04e4bdfcc92aefd3fe981ffaac469

Observation ff7f2fae-e973-42ba-8053-ca9eff8b8ff9 · outbound

This paper cites Object storage on craq: High- throughput chain replication for read-mostly workloads.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Object storage on craq: High- throughput chain replication for read-mostly workloads

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.224973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.450954Z digest=sha256:5e6467fe11b57c43c92edb4137472d8007e4248c0c93c1de70c4d959a65b0550

Observation 9168acd8-ed5d-4b95-8f5f-2019e44874c8 · outbound

This paper cites https://duckdb.org/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://duckdb.org/

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.211272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.454631Z digest=sha256:52f93122041766db8584ef9da75ac5fd21416ddc77ced6a01b13993ecc51e0f0

Observation 53bf552d-21dd-4d45-a1ec-41026a92b6c8 · outbound

This paper cites mtcp: a highly scalable user-level tcp stack for multicore systems.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure mtcp: a highly scalable user-level tcp stack for multicore systems

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.198701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.459162Z digest=sha256:d431abfa40f4cfc80e93cabfeb15c54a5b4e172f70bd0b87b7ac82758f07fb49

Observation 6a330889-e78c-40b5-9722-42ead742e5d1 · outbound

This paper cites https://github.com/juicedata/juicefs.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://github.com/juicedata/juicefs

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.186247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.462909Z digest=sha256:3e4889a46eeee7807f3bc942b76dc2a82d77cc98b449bf213261bf93d49257c0

Observation e5e4bb20-5e0d-4601-ba22-8341ef673049 · outbound

This paper cites https://github.com/scitix/InstantTensor.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://github.com/scitix/InstantTensor

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.172944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.466659Z digest=sha256:2be552d03387ebda745652441d299b83ab7be4ac7b5acfe769da0671557a4e02

Observation 3e039b24-fd2d-4026-b25b-8d5d812f2fed · outbound

This paper cites Long- bench: A bilingual, multitask benchmark for long context understand- ing.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Long- bench: A bilingual, multitask benchmark for long context understand- ing

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.160803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.470342Z digest=sha256:d543676ccdafb9c0634618ade1cfc8e321cee3bbff358f9fe014c7429857348c

Observation 39ad49e7-c9ab-40df-ac0f-262f08e9d101 · outbound

This paper cites Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.474373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.474373Z digest=sha256:d01e69fc41688d4fa06fc94867a336cabf812a7e3b3aea01f760af46afa88796

Observation 22a3cb25-dbb3-44a9-83a9-ba3209d9e25d · outbound

This paper cites https://docs.sglang.io/docs/advanced_feature s/sgl_model_gateway.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://docs.sglang.io/docs/advanced_feature s/sgl_model_gateway

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.147703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.478209Z digest=sha256:e436e38bf19587dcc7790accb1b727b88d21d51fba597b37eb05581c55867924

Observation f82df5cf-a17a-497c-a3eb-179e3270f03c · outbound

This paper cites The power of two choices in randomized load balancing.IEEE transactions on parallel and distributed systems, 12(10):1094–1104, 2002.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure The power of two choices in randomized load balancing.IEEE transactions on parallel and distributed systems, 12(10):1094–1104, 2002

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.135851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.481823Z digest=sha256:7c76b3ba2766a6709abd03b547aac6be2665feaa4dce8f41955844c96dbe05f3

Observation 916f0139-d982-41e4-9dda-0c918187b03b · outbound

This paper cites https://huggingf ace.co/datasets/SWE-Gym/OpenHands-Sampled-Trajectories.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://huggingf ace.co/datasets/SWE-Gym/OpenHands-Sampled-Trajectories

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.123576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.486166Z digest=sha256:fa6c949655d2e3869f58dd9e4c17e919eb11888166cd04cdadf7e039ca274c43

Observation ad3735ae-f21e-46a5-9c1d-c929ea663c71 · outbound

This paper cites Statistical analysis of a telephone call center: A queueing-science perspective.Journal of the American statistical association, 100(469):36–50, 2005.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Statistical analysis of a telephone call center: A queueing-science perspective.Journal of the American statistical association, 100(469):36–50, 2005

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.109945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.489786Z digest=sha256:0383373d32d16fc10d2c665d3eb12bb0240602fbb626d24ec018e5b288814c2a

Observation 50970dc1-da77-4f6d-87a2-d011f79d8682 · outbound

This paper cites A poissonian explanation for heavy tails in e-mail communi- cation.Proceedings of the National Academy of Sciences, 105(47):18153– 18158, 2008.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure A poissonian explanation for heavy tails in e-mail communi- cation.Proceedings of the National Academy of Sciences, 105(47):18153– 18158, 2008

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.092229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.493268Z digest=sha256:9e1754a5fc2379071f39232e6a2f64c558d61799a280479eaf6bbc074ef42636

Observation 9da05baf-c805-45e5-b197-927e82cb80a7 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.497271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.497271Z digest=sha256:747a38353f6d1fd4fc22429056472b7722373f4263a412f8c16e1a4628456b44

Observation 7c371ba0-aef4-4f2f-9ba9-cc684d7c79ce · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.501915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.501915Z digest=sha256:24f9f9a2b088fea20931ff37cedd73904aab4151ac77f5e6200ea2e16ee4f68d

Observation 3f1f042e-f93e-4b13-97de-451d75f5d386 · outbound

This paper cites Alpa: Automating inter-and{Intra-Operator} par- allelism for distributed deep learning.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Alpa: Automating inter-and{Intra-Operator} par- allelism for distributed deep learning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.070297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.506285Z digest=sha256:4d8903244deb08f1785e875a20a9522a06dc1a70bb899b1470bdfcefc5d12598

Observation ec3db6c5-8618-41dd-a0e0-fc2975811352 · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.056029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.510490Z digest=sha256:9de54fcb930a1c17a836dcf4bddce1cd17fc1c1a569f11a37c1f93ba7a9ea2f5

Observation ce5ad13f-c44f-4996-b5d5-3d1c1068e11c · outbound

This paper cites Loongserve: Efficiently serving long-context large lan- guage models with elastic sequence parallelism.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Loongserve: Efficiently serving long-context large lan- guage models with elastic sequence parallelism

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.041553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.514615Z digest=sha256:07b6eb1b35deede3db35b6a39a4ecb9baa5e518dc82014c779bc2ee7ca7f84ea

Observation 321a3a82-8a7f-431a-b288-9d475621f8dd · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Fast Distributed Inference Serving for Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.520249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.520249Z digest=sha256:b91097a34f12c539df00cc813936b391f0a3676e5805b44081913b4312da3f51

Observation d44d2940-594a-4da8-9e92-b240f4b8893d · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based}generative models.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Orca: A distributed serving system for {Transformer-Based}generative models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.028543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.524500Z digest=sha256:fcd82a77ef260baf4b9fb533d86ead7b00d8911aac2783a387a7df971bf6558a

Observation e5b698e6-323f-4355-87d5-cc3ac62d84e0 · outbound

This paper cites Taming{Throughput-Latency} tradeoff in{LLM} inference with {Sarathi-Serve}.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Taming{Throughput-Latency} tradeoff in{LLM} inference with {Sarathi-Serve}

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.016099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.528420Z digest=sha256:47f672f8935fcc407dabbdd5309c2383a59b3e0891d274680325051aab154968

Observation a47f28ee-3dff-4fe0-b73d-64c8f230a1b1 · outbound

This paper cites {InfiniGen}: Efficient generative inference of large language mod- els with dynamic{KV}cache management.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure {InfiniGen}: Efficient generative inference of large language mod- els with dynamic{KV}cache management

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.999320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.532162Z digest=sha256:7cb9317ac25a650eef33098ae6766459576640573d4afefdc0d4ef1f7a444aae

Observation d7f6d99a-bfde-4c43-b0d1-93c0571d3a74 · outbound

This paper cites Jenga: Effective memory management for serving llm with heterogeneity.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Jenga: Effective memory management for serving llm with heterogeneity

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.985186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.535995Z digest=sha256:f532329c2b84611eda182a89324be44fa8b887f5d40a4be4b68e5968459815e2

Observation 47d81dd6-1a6f-458a-91d8-0c0eb459dba6 · outbound

This paper cites BitNet: Scaling 1-bit Transformers for Large Language Models.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure BitNet: Scaling 1-bit Transformers for Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.539810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.539810Z digest=sha256:f6502bc775ef6a11e84d37ed2a4993db5100a02ea608ce6ea5da2c83a03a4f95

Observation 1ec8b77c-2dc6-4430-a65d-ae0ace596689 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of machine learning and systems, 6:87–100, 2024.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of machine learning and systems, 6:87–100, 2024

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.544105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.544105Z digest=sha256:acbf6c3495d7de7645a0d90e21586505fddfb7351815d45508ce35af8b094444

Observation e727258c-6057-4d0d-b6fc-f4007d5afa6c · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io- awareness.Advances in neural information processing systems, 35:16344–16359, 2022.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Flashattention: Fast and memory-efficient exact attention with io- awareness.Advances in neural information processing systems, 35:16344–16359, 2022

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.965148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.547722Z digest=sha256:ade89ca7f92ba80fe3884b7c8c01e18128d5f1ac045aed1f6554fa2cc6eeaeb9

Observation 84c4db5f-97ed-4335-a15a-5f09fc966fd5 · outbound

This paper cites FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.551553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.551553Z digest=sha256:2df112478fabc98ee4a4e1ad4e17b904996e0ef219e8566e935c9856ca291580

Observation 097379e1-7500-43d0-8074-29ede2225e43 · outbound

This paper cites https://gith ub.com/langchain-ai/langchain.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://gith ub.com/langchain-ai/langchain

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.953397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.555354Z digest=sha256:43ed6b99911fe28c5a4f471ccdaffc07380b61cc8030fbdb9015a213e68f1c6c

Observation da08e753-aeb8-4041-b346-72bca9005c74 · outbound

This paper cites https://www.langflow.org/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://www.langflow.org/

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.940054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.558727Z digest=sha256:7b032d1ab056825d42e2c9fdae22905ed91c6239b8acd971193a815e7e38bf4d

Observation 829b080f-c9c2-465c-af2a-a92e10744480 · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.562330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.562330Z digest=sha256:68dbe3653ecba9a93395b59826e1e6278d978d147ce8ccc1d5f1de386fb3d39a

Observation a982b658-0abc-40b2-bc57-01279d88e8c3 · outbound

This paper cites Dspy: Compiling declarative 17 language model calls into state-of-the-art pipelines.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Dspy: Compiling declarative 17 language model calls into state-of-the-art pipelines

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.928405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.566354Z digest=sha256:c80090e0e794d67a097e238a078e7851a7cf26714a564d6e02bd2007f89a9035

Observation d6368581-12ab-416e-8120-5c3d69cb5437 · outbound

This paper cites A System for Microserving of LLMs.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure A System for Microserving of LLMs

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-08-07T18:23:08.608861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.569863Z digest=sha256:36c183a276d3e1f21e429de5e3b54896836d47b5c4e7f6331faf6a7dc7b547ad

Pith citing papers

No inbound Pith citation observations are available.