Pith. sign in

Paper Citation Record · LEDGER

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure

As of 10 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 0 inbound Pith citation observations for arXiv:2608.06007.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06007 v1

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:23:08.569863Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

88 of 88 outbound references displayed

  • verified exact2
  • verified fuzzy51
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7266694b-2167-460e-a9a2-c6a1461ddcae · outbound

This paper cites https://www.kimi.com/blog/kimi- k3.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://www.kimi.com/blog/kimi- k3

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.210099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.210099Z digest=sha256:4f7b9f10b983ba780b9f6238a016f158f13ede963a7e36a4e014ff94358db964

Observation c5640c60-ddb8-46c0-820b-289b6a800bd5 · outbound

This paper cites ServerlessLLM: Low-Latency serverless inference for large language models.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure ServerlessLLM: Low-Latency serverless inference for large language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.215746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.215746Z digest=sha256:b3404c8dbde10b6926018e33aef7e8a7bf7f0492f8cd5ba0bfb6daae36714e1a

Observation d373c80b-505e-4110-9072-7ef34e468468 · outbound

This paper cites Deep- flow: Serverless large language model serving at scale.arXiv e-prints, pages arXiv–2501, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Deep- flow: Serverless large language model serving at scale.arXiv e-prints, pages arXiv–2501, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.220987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.220987Z digest=sha256:d999ba554e695ffe0376bee321c7f429c8c224dff493903416464fecebcae96d

Observation 059a9267-99f7-4671-b484-7768bf48ee96 · outbound

This paper cites Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.225722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.225722Z digest=sha256:a6aa9680bf8060853d6460086f3c48c8f938d539df05418fb59d694b6d4f4979

Observation 7bb633c6-8d07-43e5-a398-f51152382d11 · outbound

This paper cites an unresolved cited work.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.230680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.230680Z digest=sha256:55b57d0b24a84e9f4727d210688d609ea0d342610c2649edf73be440e598d20e

Observation e4436a8a-eac5-4d58-a312-6f25afa6d915 · outbound

This paper cites BlitzScale: Fast and Live Large Model Autoscaling with O (1) Host Caching.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure BlitzScale: Fast and Live Large Model Autoscaling with O (1) Host Caching

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.234958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.234958Z digest=sha256:81b2ce68c63bcbbd86a060b38e96b434471d8b8f9c49eb5c4c2e62d407d7a909

Observation 4c447d41-9049-42d6-a526-314478418a4a · outbound

This paper cites Hydraserve: Minimizing cold start latency for serverless llm serving in public clouds.arXiv preprint arXiv:2502.15524, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Hydraserve: Minimizing cold start latency for serverless llm serving in public clouds.arXiv preprint arXiv:2502.15524, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.239382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.239382Z digest=sha256:972a83cdc88b97136a952c9ad31e77add15fdc025899ed7cec6b1c6e9a31b78c

Observation cff16d1b-78d6-4d53-afdf-1a465b2f75e0 · outbound

This paper cites https://lmsys.org/blog/2025-12-10-rfork/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://lmsys.org/blog/2025-12-10-rfork/

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.243347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.243347Z digest=sha256:f0193464eda4d41ce6d519990aeba516ddca4da6e243412d9a37c282ab6b9f19

Observation 9012b892-fc8d-46fe-9e9a-83c5ef0d592a · outbound

This paper cites Mooncake: Trad- ing more storage for less computation—a KVCache-centric architecture for serving LLM chatbot.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Mooncake: Trad- ing more storage for less computation—a KVCache-centric architecture for serving LLM chatbot

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.247723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.247723Z digest=sha256:c20a6769b20615ead0463a060624a2983389c48867b0b323b6755b5a1b33289a

Observation 214672d2-ef54-4bbc-9f95-a8ecd3432886 · outbound

This paper cites Dualmap: Enabling both cache affinity and load bal- ancing for distributed LLM serving.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Dualmap: Enabling both cache affinity and load bal- ancing for distributed LLM serving

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.252135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.252135Z digest=sha256:8305f35f577c3a3d551789fc3e8af8db07085aa321b95e9208e876ba4efe310b

Observation 530077e3-3828-4a3b-8d88-67308194f45b · outbound

This paper cites Lmcache: An efficient kv cache layer for enterprise-scale llm inference.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Lmcache: An efficient kv cache layer for enterprise-scale llm inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.256927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.256927Z digest=sha256:7de7ce2bd987a3f8774451de961ecbcff01b75f6ac2316132924a259c5f58790

Observation 1bab6d33-92c7-4dee-8e00-57df7e7504a8 · outbound

This paper cites https://lmsys.org/blog/2025-09-10-sglang-hicache/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://lmsys.org/blog/2025-09-10-sglang-hicache/

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.261289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.261289Z digest=sha256:3d15f6cc0d462353709022a870a7b6aae5a3c63a11f2499c853d9f6e8190a4cd

Observation 42f290ec-4062-48e3-bfca-971efb32b231 · outbound

This paper cites Stateful large language model serving with pensieve.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Stateful large language model serving with pensieve

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.647674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.265567Z digest=sha256:73ce489d0e99d55aa6475a0c56130a5edd91e1a323dc48fec4216e4d09a8a5aa

Observation 9194a7a9-2f73-4e9c-b28f-1f56071a2766 · outbound

This paper cites Cacheblend: Fast large language model serving for rag with cached knowledge fusion.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Cacheblend: Fast large language model serving for rag with cached knowledge fusion

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.635789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.270101Z digest=sha256:64a2c51f3db97be742eb394acc5b7002925b0f05e7fed6ee1a28e7bb3ace23f9

Observation da5d784e-d1e4-4906-82b2-e611454f3ffb · outbound

This paper cites Cost-Efficient large language model serving for multi-turn conversations with Cache- dAttention.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Cost-Efficient large language model serving for multi-turn conversations with Cache- dAttention

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.624659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.274263Z digest=sha256:5131cf7c56eb3368650518b6400d47e17233e96630647c32f2b5518f870d9328

Observation 7b20c028-7b9f-4073-9b77-231851ecc7da · outbound

This paper cites {ByteCheckpoint}: A unified checkpointing system for large foundation model development.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure {ByteCheckpoint}: A unified checkpointing system for large foundation model development

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.613146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.278745Z digest=sha256:0f8be5784ac0b3a952dd74b454870beb5802f7a70bb1d8b6717d01aa50d079f4

Observation 4662b785-a21c-4aa6-a755-bec13fc799eb · outbound

This paper cites https://github.com/MoonshotAI/chec kpoint-engine.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://github.com/MoonshotAI/chec kpoint-engine

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.601572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.282891Z digest=sha256:4b972f55aac4dc22f95bafe2e8e08eb8bdbc87bb095cfe0285752bb764ddaf60

Observation 31c3bf0a-5caf-4897-8dfd-3bc84f726e13 · outbound

This paper cites https://vllm.ai/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://vllm.ai/

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.589655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.286808Z digest=sha256:0fcfc0bc8cfbee70c47bc93ef8a2648d7eb2878cdf51aea6cc09f84adec931d6

Observation 4559eff5-182a-4fde-b2db-6844d6b4eef1 · outbound

This paper cites Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583, 2024.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.292156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.292156Z digest=sha256:38cdd215075aca3a444a1e77fa6f9cbcef0e78984f4f582b5ff6f6be7626c90c

Observation f1a612b6-281d-44a8-966a-de5cd8b8260d · outbound

This paper cites https://nvidia.github.io/TensorRT-LLM/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://nvidia.github.io/TensorRT-LLM/

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.571245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.296384Z digest=sha256:952e126e2abd09db8c594c18a217204b429fe72f99b1c5194e02c3ce42b5ab5e

Observation 88da8772-4852-4076-aeed-3d06e9d38c5c · outbound

This paper cites an unresolved cited work.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:23:10.559915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.300372Z digest=sha256:8d6f041fa38a4065a84c86ace12874e756a254e49c1cc7a2ed5a02bdc6758953

Observation db842daf-c122-4ea2-933e-7cb220a4e955 · outbound

This paper cites https://redis.io/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://redis.io/

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.546071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.304346Z digest=sha256:6751471d99cc5bbe80fead754ea22ce77e2f6a7e6a7e09dda795e8b35c6f4847

Observation e67a7a63-1040-4cf5-a8c1-e12cedf1315d · outbound

This paper cites Vineyard: Optimizing data sharing in data- intensive analytics.Proc.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Vineyard: Optimizing data sharing in data- intensive analytics.Proc

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.534517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.308377Z digest=sha256:0aeea247c7c09b0b29a9cd73312202551a5ad42947df3c0db163a27453dab228

Observation 45e262c5-30d2-4c6f-be79-3d1fc74ab155 · outbound

This paper cites https://github.com/ray-project/plasma.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://github.com/ray-project/plasma

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.522032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.312295Z digest=sha256:b956921cb6e119348ba339301efa97511d1fa1c68edfa58e7b6854a06beebaca

Observation 8436dcad-3029-4bdf-bb28-e3e864b50b22 · outbound

This paper cites Ray: A distributed framework for emerging ai applications.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Ray: A distributed framework for emerging ai applications

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.509673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.315974Z digest=sha256:a31f481b05175e3de38604943c27868042b0e2b4b0c5602385d3c46933457128

Observation 76dc0904-c3ee-4e49-a3cb-3e5b4a919e09 · outbound

This paper cites Resilient distributed datasets: A Fault-Tolerant abstraction for In-Memory cluster computing.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Resilient distributed datasets: A Fault-Tolerant abstraction for In-Memory cluster computing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.497933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.320121Z digest=sha256:f57ff1139b27f5d22e1110b88d764ff13538a6ed5147820af02959ecb4d71e46

Observation 0d236e8d-f09c-4621-b24c-a6a9d6cc73ed · outbound

This paper cites Serverless computing: Design, implementation, and performance.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Serverless computing: Design, implementation, and performance

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.324667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.324667Z digest=sha256:53e2e3e77f663f95565871eee5352f08512b8cdf718675d923b49f8ce316e676

Observation b4725148-3cf4-48f3-bcf2-2e285e89f710 · outbound

This paper cites Infless: a native serverless system for low-latency, high-throughput inference.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Infless: a native serverless system for low-latency, high-throughput inference

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.478626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.328764Z digest=sha256:c366b4a243cc64d9e998a7cd55022f98c136b97f334fb557f505cb6f5b73cded

Observation 61af6026-a968-4494-9880-29304696fc9d · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.332432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.332432Z digest=sha256:6732ba4f08474a9655ad8418db6056c31899f79c9eca19ed5b2a5f0eacb8e1f1

Observation d04194b5-e4b6-438e-8cae-de4d804b89f0 · outbound

This paper cites Language models are few-shot learners.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Language models are few-shot learners

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.336296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.336296Z digest=sha256:1d8d741e4588f85a72efb8e108c4f9afe976fa34052c6ff798e51b33bacfcc14

Observation 8c2c0de8-44a0-435c-8316-63e889a005a8 · outbound

This paper cites Flexkv: Flexible index offloading for memory-disaggregated key-value store.arXiv preprint arXiv:2512.16148, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Flexkv: Flexible index offloading for memory-disaggregated key-value store.arXiv preprint arXiv:2512.16148, 2025

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-08-07T18:23:09.340739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.339917Z digest=sha256:80b0cf2c642d362dcd981eff6f3c9e6118e4d878acca6d170cf2886eea860903

Observation 770044c6-56af-42e6-a0b9-808c69d64444 · outbound

This paper cites In USENIX OSDI, 2024.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure In USENIX OSDI, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.450965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.344903Z digest=sha256:079cb0bc78382d4dc69a7cbf02242f1fc6c3455779011dc83512c2f3f4051fe2

Observation c0f22ded-c796-44de-a7ef-351c921e3fd7 · outbound

This paper cites Splitwise: Efficient genera- tive llm inference using phase splitting.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Splitwise: Efficient genera- tive llm inference using phase splitting

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.438704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.348995Z digest=sha256:8ecd1817537f71e2744de971feec6169b6757e9d22d42af425090cdbb70d13cb

Observation ea77d351-5772-4605-b59d-d2342eac07e3 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.353271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.353271Z digest=sha256:ac796e551d8dcbe0557f8de206ace2e7dc37e3465b2773ffc22ce662de4f4c2c

Observation 60fe65b5-c7d9-437f-aa3f-c6cb4b316c79 · outbound

This paper cites D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.357426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.357426Z digest=sha256:4312c72b7177e81b0340582299b7eca20aadd6c444317344bf298c7248d88eae

Observation e3d048ff-38d0-4f2e-9257-ef5544fc095f · outbound

This paper cites Prompt cache: Modular attention reuse for low-latency inference.Proceedings of Machine Learning and Systems, 6:325–338, 2024.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Prompt cache: Modular attention reuse for low-latency inference.Proceedings of Machine Learning and Systems, 6:325–338, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.425673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.361591Z digest=sha256:01a3f775f8b64e7f5edffdd98ea66d9c04c06bbc60a29de4c773406632cd5264

Observation bcc27414-be25-4ed7-90ec-d706b5fad965 · outbound

This paper cites Cachegen: Kv cache compression and streaming for fast large language model serving.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Cachegen: Kv cache compression and streaming for fast large language model serving

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.414206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.365396Z digest=sha256:de4013e1be576e9cf23d862ca0985f8534e3227d47ec6084e592fc05113e1620

Observation 1a8aac5a-b716-4e1a-b453-3e151a60595c · outbound

This paper cites Ragcache: Efficient knowledge caching for retrieval-augmented generation.ACM Transactions on Computer Sys- tems, 44(1):1–27, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Ragcache: Efficient knowledge caching for retrieval-augmented generation.ACM Transactions on Computer Sys- tems, 44(1):1–27, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.401403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.369036Z digest=sha256:27d397580d6d0e968ad6cfa760aa910a807331f11e204c07257c5a7b71f178c6

Observation 573cad29-4ab0-4ad7-ae10-34ce8d5c21d4 · outbound

This paper cites Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.373466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.373466Z digest=sha256:645ea2e14d9465877a5388b6315ca6b80671377b5f0330f1e3fa666b5284fa76

Observation 254654b9-7c4a-46cd-95fe-2db36adfceca · outbound

This paper cites Dist checkpointing package.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Dist checkpointing package

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.386334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.378063Z digest=sha256:7ad734cc8dabed48d9488168ec6647d43c26b977d745428682226016b33a96df

Observation 69313a8f-f463-45a8-ac16-3a10cb5bdf78 · outbound

This paper cites Getting started with Distributed Check- point (DCP).

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Getting started with Distributed Check- point (DCP)

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.373829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.382598Z digest=sha256:d3d9b099a1fd251260c0f48c44ba917aabe9cf3dec919057d7c8a1133ef7ec85

Observation 5f164233-bc05-4186-942a-ec60c49025d7 · outbound

This paper cites Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.386378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.386378Z digest=sha256:35c46a5fc86b6c8adf0426fa5c1a03db751253a858acdffe4c8731288ce0f012

Observation 9cfffabc-cb58-4e9a-b75f-bc7b8d852320 · outbound

This paper cites Simple is better: Multiplication may be all you need for llm request scheduling.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Simple is better: Multiplication may be all you need for llm request scheduling

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.359485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.390858Z digest=sha256:32486c34f2fc76ef40ff5d49a2a5244b7d90e2c33e64960b28422c2cdc7b5f4a

Observation 8a62c2a8-f310-4c44-aae6-6933689baeb9 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Efficient memory management for large language model serving with pagedattention

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.347609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.394615Z digest=sha256:02f1f751befa71ae383d19ad4720ef771aa72128c23a54e6b1f4d56e55cf66e1

Observation 48fdc530-c428-411c-87e2-112cddc9bfd4 · outbound

This paper cites https://docs.vllm.ai/en/stable/desig n/prefix_caching/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://docs.vllm.ai/en/stable/desig n/prefix_caching/

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.334793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.398238Z digest=sha256:cd78bbbd35f8afea7b8f37527322bc6d58c1378b3f41b8048de07a075b1843c7

Observation 2ec7bafd-e84c-4950-a29a-579d9f6908fa · outbound

This paper cites https://lmsys.org/blog/2024-01-17-sglang/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://lmsys.org/blog/2024-01-17-sglang/

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.322790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.401764Z digest=sha256:68e2bf1cec2a681e94b8248f8ba11088e939252f38f488a6cbb7ac96bb2237e7

Observation aed7ed80-0a54-4021-9bad-50150fca6b34 · outbound

This paper cites Pie: A pro- grammable serving system for emerging llm applications.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Pie: A pro- grammable serving system for emerging llm applications

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.310698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.405393Z digest=sha256:a1dc84e11df66ef4b177d0c5a334446c60b1e989663cbcc9f06fded2b2558234

Observation 2d7fe69a-4765-4fa5-a54c-4bd2caefbe3d · outbound

This paper cites Chain-of-thought prompt- ing elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Chain-of-thought prompt- ing elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.409426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.409426Z digest=sha256:893eef6b663436ae10ade394ebb391dd307730eb0242068bb679968114405f3e

Observation 4c63f9c8-ebc7-45d6-aa2d-1fb0fd3500ba · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.412959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.412959Z digest=sha256:a533a13579749b4fd73f5d1cb4b51d34ebed90f393c134c0faf37a498bc96d81

Observation 60382737-0556-4911-ab3d-b66a51c7f2ec · outbound

This paper cites Training language models to follow instruc- tions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Training language models to follow instruc- tions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.282711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.416940Z digest=sha256:9cd46a55c7a507ce744b1f9cf938aa5139465e700ff011dbddd6096f59885666

Observation 4c583cb5-6425-44b7-82ee-ab34222e266e · outbound

This paper cites Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.421353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.421353Z digest=sha256:908ec94ab3d04faa0e5edfd98c2f274b3088131c8f0560b38034e5175ec2421d

Observation f8c235cd-4da4-4fe4-88dc-0cf04bcc8159 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.425755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.425755Z digest=sha256:e37551f8ba03afa66441c2923c3eb0075f63abd8b4916b3a008e6b19dc76d165

Observation 2847fa79-e643-40e6-a65c-fb56839f6f7e · outbound

This paper cites Totrl: Unlock llm tree-of-thoughts reasoning potential through puzzles solving.arXiv preprint arXiv:2505.12717, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Totrl: Unlock llm tree-of-thoughts reasoning potential through puzzles solving.arXiv preprint arXiv:2505.12717, 2025

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.430349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.430349Z digest=sha256:e9ccb3af57a1015e7a832a00fab90e27d10d39b30cc5b1c0a345f7022eea52c7

Observation 50a88b26-30cb-4888-a5af-cec5627c30a6 · outbound

This paper cites Using name-based map- pings to increase hit rates.IEEE/ACM Transactions on networking, 6(1):1–14, 2002.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Using name-based map- pings to increase hit rates.IEEE/ACM Transactions on networking, 6(1):1–14, 2002

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.269755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.434044Z digest=sha256:2615ace8d5018383bba58e0a82078867611d76f7f714180533667b9727f45b3c

Observation 006bb2b3-6cfd-4972-844b-26a05685227f · outbound

This paper cites Paxos made simple.ACM SIGACT News (Distributed Computing Column) 32, 4 (Whole Number 121, December 2001), pages 51–58, 2001.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Paxos made simple.ACM SIGACT News (Distributed Computing Column) 32, 4 (Whole Number 121, December 2001), pages 51–58, 2001

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.438822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.438822Z digest=sha256:5d9be7bb3649e3fc17c7a1089cb86bbf2b7ebcb3583718771a6875875d14ba73

Observation 32b178f7-8443-4441-b96b-3e37cc30c25a · outbound

This paper cites In search of an understandable consensus algorithm.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure In search of an understandable consensus algorithm

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.249413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.442644Z digest=sha256:b262d7c4ec44501ebefe5a828ccac06ee3470ac71442547bb9a82b224f748b79

Observation bbf2b854-54d3-4aaa-bb1d-fe5f9036b578 · outbound

This paper cites Chain replication for sup- porting high throughput and availability.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Chain replication for sup- porting high throughput and availability

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.238009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.446537Z digest=sha256:bdfa95fcd82a6d5e4bc2bfed9013d47cd5339194a99e5af38d0495b69283f417

Observation ff7f2fae-e973-42ba-8053-ca9eff8b8ff9 · outbound

This paper cites Object storage on craq: High- throughput chain replication for read-mostly workloads.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Object storage on craq: High- throughput chain replication for read-mostly workloads

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.224973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.450954Z digest=sha256:a2414c28bde763ae94fb09ad7fd5c55202951df93127db70cddeee23f5f468ba

Observation 9168acd8-ed5d-4b95-8f5f-2019e44874c8 · outbound

This paper cites https://duckdb.org/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://duckdb.org/

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.211272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.454631Z digest=sha256:75688b60a12d9850b36dfe373b039154317d469b4a707619db923411bc9e7383

Observation 53bf552d-21dd-4d45-a1ec-41026a92b6c8 · outbound

This paper cites mtcp: a highly scalable user-level tcp stack for multicore systems.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure mtcp: a highly scalable user-level tcp stack for multicore systems

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.198701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.459162Z digest=sha256:54f081708ee6b618753b27bac508d3ece5a9ea80ded50697c84ecf5741a9beea

Observation 6a330889-e78c-40b5-9722-42ead742e5d1 · outbound

This paper cites https://github.com/juicedata/juicefs.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://github.com/juicedata/juicefs

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.186247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.462909Z digest=sha256:5ef8dfe51e4fec3d22d1ccc6cec147e80588778482d572ff1f8d11b36536c288

Observation e5e4bb20-5e0d-4601-ba22-8341ef673049 · outbound

This paper cites https://github.com/scitix/InstantTensor.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://github.com/scitix/InstantTensor

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.172944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.466659Z digest=sha256:43727e5e54ab8655d129407cd79c6967ad807b0b0972fdddb93caae6dc272a0c

Observation 3e039b24-fd2d-4026-b25b-8d5d812f2fed · outbound

This paper cites Long- bench: A bilingual, multitask benchmark for long context understand- ing.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Long- bench: A bilingual, multitask benchmark for long context understand- ing

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.160803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.470342Z digest=sha256:7704de91f55401a036ab691e20145eaf8c494370491a06fbeb11327538a6ecd8

Observation 39ad49e7-c9ab-40df-ac0f-262f08e9d101 · outbound

This paper cites Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.474373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.474373Z digest=sha256:5264bfcae369ffda7552d49e92834fcb42be5efccc2b7ef1aa2406ac4e9f19a2

Observation 22a3cb25-dbb3-44a9-83a9-ba3209d9e25d · outbound

This paper cites https://docs.sglang.io/docs/advanced_feature s/sgl_model_gateway.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://docs.sglang.io/docs/advanced_feature s/sgl_model_gateway

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.147703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.478209Z digest=sha256:636307ff5abbc6cd613f509756435f840839cee345d27ba77fd0c2757dcfdd20

Observation f82df5cf-a17a-497c-a3eb-179e3270f03c · outbound

This paper cites The power of two choices in randomized load balancing.IEEE transactions on parallel and distributed systems, 12(10):1094–1104, 2002.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure The power of two choices in randomized load balancing.IEEE transactions on parallel and distributed systems, 12(10):1094–1104, 2002

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.135851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.481823Z digest=sha256:ff068918ef77de24fb0df69ea2d2424ad0de1867f7e128c2dca9fef6dda89afe

Observation 916f0139-d982-41e4-9dda-0c918187b03b · outbound

This paper cites https://huggingf ace.co/datasets/SWE-Gym/OpenHands-Sampled-Trajectories.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://huggingf ace.co/datasets/SWE-Gym/OpenHands-Sampled-Trajectories

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.123576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.486166Z digest=sha256:31458e179d2eb75b5963d3c68e6dcc179ff9ecd315e458feccbf7539d599e176

Observation ad3735ae-f21e-46a5-9c1d-c929ea663c71 · outbound

This paper cites Statistical analysis of a telephone call center: A queueing-science perspective.Journal of the American statistical association, 100(469):36–50, 2005.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Statistical analysis of a telephone call center: A queueing-science perspective.Journal of the American statistical association, 100(469):36–50, 2005

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.109945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.489786Z digest=sha256:2400fc5262b535d1375ef4a9085c4d517a3a6d0ff549112b221e9a7f038bac17

Observation 50970dc1-da77-4f6d-87a2-d011f79d8682 · outbound

This paper cites A poissonian explanation for heavy tails in e-mail communi- cation.Proceedings of the National Academy of Sciences, 105(47):18153– 18158, 2008.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure A poissonian explanation for heavy tails in e-mail communi- cation.Proceedings of the National Academy of Sciences, 105(47):18153– 18158, 2008

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.092229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.493268Z digest=sha256:0ef3b40d9de1c7ad2a0f76f9de9ca774351efbe8fc36b6d4626ebae3fd791b9a

Observation 9da05baf-c805-45e5-b197-927e82cb80a7 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.497271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.497271Z digest=sha256:ac8af704008747c1796a2d6d0edb5b00d536f636b8bf42f723453b3271dcb3f6

Observation 7c371ba0-aef4-4f2f-9ba9-cc684d7c79ce · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.501915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.501915Z digest=sha256:337e6cff8006d331398bccbe85dd77e3722d0418a62bd610d8f947643d018448

Observation 3f1f042e-f93e-4b13-97de-451d75f5d386 · outbound

This paper cites Alpa: Automating inter-and{Intra-Operator} par- allelism for distributed deep learning.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Alpa: Automating inter-and{Intra-Operator} par- allelism for distributed deep learning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.070297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.506285Z digest=sha256:136952be6d987aa1fd490bc2f7c80373163f8a6e371c4df18d327036b02c21a6

Observation ec3db6c5-8618-41dd-a0e0-fc2975811352 · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.056029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.510490Z digest=sha256:36ead62fbf7f7c37d0517a9c654c844caf3f4226cca1d407f5ff2b353902b102

Observation ce5ad13f-c44f-4996-b5d5-3d1c1068e11c · outbound

This paper cites Loongserve: Efficiently serving long-context large lan- guage models with elastic sequence parallelism.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Loongserve: Efficiently serving long-context large lan- guage models with elastic sequence parallelism

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.041553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.514615Z digest=sha256:83e5dfcd4dd72d9c1ab6be1e1fa70f7ac208dcc1971f0cc1f0677e3318566312

Observation 321a3a82-8a7f-431a-b288-9d475621f8dd · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Fast Distributed Inference Serving for Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.520249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.520249Z digest=sha256:c7026bb6e9c870970c4424b143c0b5802e37d6790fa4d94425c670130e4cbb2d

Observation d44d2940-594a-4da8-9e92-b240f4b8893d · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based}generative models.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Orca: A distributed serving system for {Transformer-Based}generative models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.028543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.524500Z digest=sha256:7e3ddf71016cec18beb28c1f2fd532993422c3b08df634515064dac836a4289f

Observation e5b698e6-323f-4355-87d5-cc3ac62d84e0 · outbound

This paper cites Taming{Throughput-Latency} tradeoff in{LLM} inference with {Sarathi-Serve}.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Taming{Throughput-Latency} tradeoff in{LLM} inference with {Sarathi-Serve}

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.016099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.528420Z digest=sha256:7a29d101670cf4c1dadadb0e12e6231398bd5961325f9572fde405dfa2ef52bd

Observation a47f28ee-3dff-4fe0-b73d-64c8f230a1b1 · outbound

This paper cites {InfiniGen}: Efficient generative inference of large language mod- els with dynamic{KV}cache management.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure {InfiniGen}: Efficient generative inference of large language mod- els with dynamic{KV}cache management

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.999320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.532162Z digest=sha256:e6f5aa55ad86afaa00fbae9f312251546fd31ed6fc221f914b740ae4c9e80885

Observation d7f6d99a-bfde-4c43-b0d1-93c0571d3a74 · outbound

This paper cites Jenga: Effective memory management for serving llm with heterogeneity.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Jenga: Effective memory management for serving llm with heterogeneity

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.985186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.535995Z digest=sha256:7137b3ca2fa70f6737516c3510ec39ca4cb1602ea74b7320fb6407f2621a325f

Observation 47d81dd6-1a6f-458a-91d8-0c0eb459dba6 · outbound

This paper cites BitNet: Scaling 1-bit Transformers for Large Language Models.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure BitNet: Scaling 1-bit Transformers for Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.539810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.539810Z digest=sha256:7023c0fe832643310008019981eb41dfa68f6b4f059ff3468aa12c5f5c4d5e7a

Observation 1ec8b77c-2dc6-4430-a65d-ae0ace596689 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of machine learning and systems, 6:87–100, 2024.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of machine learning and systems, 6:87–100, 2024

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.544105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.544105Z digest=sha256:f7f5d0155331d6c780ac81107ea83e03ee473b07192a96956aebb1884eafd647

Observation e727258c-6057-4d0d-b6fc-f4007d5afa6c · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io- awareness.Advances in neural information processing systems, 35:16344–16359, 2022.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Flashattention: Fast and memory-efficient exact attention with io- awareness.Advances in neural information processing systems, 35:16344–16359, 2022

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.965148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.547722Z digest=sha256:10daf979e90bc056f96038cb7a0fb7c83ae2b529dcb60bf7ec1c7c4d7c59cd26

Observation 84c4db5f-97ed-4335-a15a-5f09fc966fd5 · outbound

This paper cites FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.551553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.551553Z digest=sha256:6f241eb2e383b542db7177f28c06694d9b09363221a05d29fa94ec02a5d4271f

Observation 097379e1-7500-43d0-8074-29ede2225e43 · outbound

This paper cites https://gith ub.com/langchain-ai/langchain.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://gith ub.com/langchain-ai/langchain

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.953397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.555354Z digest=sha256:1f7c3b97e6a5ba066be4ed57b16cc2386880c5139cd2abc1460a3646aff96ab2

Observation da08e753-aeb8-4041-b346-72bca9005c74 · outbound

This paper cites https://www.langflow.org/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://www.langflow.org/

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.940054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.558727Z digest=sha256:9a0e015c6e3b72205611d970387525b6588b5bf87cb8e4920b81b38b7a4e425e

Observation 829b080f-c9c2-465c-af2a-a92e10744480 · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.562330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.562330Z digest=sha256:ae6c9702feea87dc6ade7a48dc0fd8ac3f97efe4008be152e6ce007ad0925fcf

Observation a982b658-0abc-40b2-bc57-01279d88e8c3 · outbound

This paper cites Dspy: Compiling declarative 17 language model calls into state-of-the-art pipelines.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Dspy: Compiling declarative 17 language model calls into state-of-the-art pipelines

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.928405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.566354Z digest=sha256:09713f39b710bda236ac09464d6e274babd79fc6b008d458af69768cd20f1c26

Observation d6368581-12ab-416e-8120-5c3d69cb5437 · outbound

This paper cites A System for Microserving of LLMs.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure A System for Microserving of LLMs

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-08-07T18:23:08.608861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T18:23:08.569863Z digest=sha256:c5eaa23f5a3a5c8f0551d37080564d8f17e741b72d1ff483df39c1f53c5704b8

Pith citing papers

No inbound Pith citation observations are available.