Pith. sign in

Paper Citation Record · LEDGER

Speeding up Model Loading with fastsafetensors

As of 13 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 4 inbound Pith citation observations for arXiv:2505.23072.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23072 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:58:07.427165Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T13:52:43.151100Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 1eb807ef-31d1-4c1d-9bdf-4ecf60c58ed7 · outbound

This paper cites Introducing ChatGPT,.

Speeding up Model Loading with fastsafetensors Introducing ChatGPT,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.656524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:03.832875Z digest=sha256:f187415e121d1b526620fece6d16572129c1bc92e4f4d93a1252cc8bb1cd1324

Observation f5b86204-bff4-4e3f-ac2b-9f89a0b3f10e · outbound

This paper cites Introducing Gemini: our largest and most capable AI model,.

Speeding up Model Loading with fastsafetensors Introducing Gemini: our largest and most capable AI model,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.492922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:03.902873Z digest=sha256:dfc4b0a1ec6175adba8ed6873a68e6cb5901ac0208e4a6b68dffb5426e4424b9

Observation 9a68980f-964e-44a0-9507-846aeffe203f · outbound

This paper cites Granite Code Models: A Family of Open Foundation Models for Code Intelligence,.

Speeding up Model Loading with fastsafetensors Granite Code Models: A Family of Open Foundation Models for Code Intelligence,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.366624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:03.973943Z digest=sha256:3c39902e992050ac9d6f21cb1020095adc088ded127bbd7e4e1d361cdb63174a

Observation 25defce5-5c40-4fd9-a40e-73806d7154ac · outbound

This paper cites The Shift from Models to Compound AI Systems,.

Speeding up Model Loading with fastsafetensors The Shift from Models to Compound AI Systems,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.176639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:04.077312Z digest=sha256:b18c90a16563fc336acd7757d53425c1d8ecbc0308589a1ab7f989a69cff35d4

Observation 87dc08d5-980d-4bd8-b7d9-83a130d7e7d3 · outbound

This paper cites FlashAttention: Fast and memory-efficient exact attention with IO-awareness,.

Speeding up Model Loading with fastsafetensors FlashAttention: Fast and memory-efficient exact attention with IO-awareness,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.059280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:04.145099Z digest=sha256:0710327abb4d59d32e49bf2ffde721fb6221e41648737b5de142e23c472206fc

Observation fd73d497-5dc8-42c4-9eb6-1b2a8d87bdc7 · outbound

This paper cites FlashAttention-2: Faster attention with better parallelism and work partitioning,.

Speeding up Model Loading with fastsafetensors FlashAttention-2: Faster attention with better parallelism and work partitioning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.895787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:04.214691Z digest=sha256:330cc93a9b3980f8b5030e4633feec42bc424fa1d2d1ddf8c056f1efbd662ec4

Observation 12cc850f-9c15-4005-87b1-9ae7c7d25ae2 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention,.

Speeding up Model Loading with fastsafetensors Efficient Memory Management for Large Language Model Serving with PagedAttention,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.758834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:04.304079Z digest=sha256:590a7755224e20a697951f3c3082ee229d1fd262ae6bc6e53b49d03c912f4554

Observation ae9808a8-4a8a-421d-9b8b-1cd773232cd5 · outbound

This paper cites Orca: A Distributed Serving System for Transformer-Based Generative Models,.

Speeding up Model Loading with fastsafetensors Orca: A Distributed Serving System for Transformer-Based Generative Models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.611414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:04.350083Z digest=sha256:823dc32cc28d58dfcc819751fa0a419ca133e069063e7cf8da7dedf89b4610cc

Observation e30d29f1-ab5e-4cdb-9cc2-d96cb0e427d8 · outbound

This paper cites Accelerating Production LLMs with Combined Token/Embedding Speculators.

Speeding up Model Loading with fastsafetensors Accelerating Production LLMs with Combined Token/Embedding Speculators

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.425862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.425862Z digest=sha256:59719724779f5403534c06e68fc0ff9699fe33ef71aae10b963daf219f6f77af

Observation e64cecdc-7e37-422b-9e32-862ec995d3f4 · outbound

This paper cites DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving.

Speeding up Model Loading with fastsafetensors DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.496281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.496281Z digest=sha256:df763e7f4cb39019ba138edc757eb3f869fc8e8d6ed1b0ba220585b0c179e9fc

Observation c2718510-a082-46bb-aa53-00fac5a49f0d · outbound

This paper cites Taming throughput-latency tradeoff in llm inference with sarathi-serve,.

Speeding up Model Loading with fastsafetensors Taming throughput-latency tradeoff in llm inference with sarathi-serve,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.472887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:04.561087Z digest=sha256:4eca9e714b57ed1a18a617f3d119f5e3a710661fb5f7dc191b38b3f1248671d1

Observation 910f4ac6-e5ae-4023-9856-08bdf6a695fc · outbound

This paper cites Decrease PyTorch Model Load Times with CoreWeave’s Tensorizer,.

Speeding up Model Loading with fastsafetensors Decrease PyTorch Model Load Times with CoreWeave’s Tensorizer,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.309831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:04.639769Z digest=sha256:52cdb1780ac529c9c982d4ec53aacf59587b104cf0681dad4534c1055d63e43f

Observation cda45ea4-2e7e-4ab7-952f-e4009f79e25f · outbound

This paper cites ServerlessLLM: Low-Latency Serverless Inference for Large Language Models.

Speeding up Model Loading with fastsafetensors ServerlessLLM: Low-Latency Serverless Inference for Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.729125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.729125Z digest=sha256:eaf1a89284131d90e51ad2f0e1da69f391c619e9d4e17626318d8150227bd5f1

Observation 2b1adefe-1e6a-4e8c-8fe1-fb2f309c00e3 · outbound

This paper cites Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM.

Speeding up Model Loading with fastsafetensors Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.791744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.791744Z digest=sha256:6cf6ee0f8fea85ae118a8a23b80aff784ff14dcd130e2f4853066c5361fa36a3

Observation 89ee66e9-bbba-4b12-ab18-81eab9f57679 · outbound

This paper cites Safetensors,.

Speeding up Model Loading with fastsafetensors Safetensors,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.115727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:04.876060Z digest=sha256:ccbd931241e71e4e318a77d81e66bfcd47f14cc9f80b32b21b3f713da7e71bd1

Observation bc3dc6c4-e620-4184-8408-f68174e86e7b · outbound

This paper cites Models – Hugging Face,.

Speeding up Model Loading with fastsafetensors Models – Hugging Face,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.934865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:04.943400Z digest=sha256:6669ad3b6575aebee33ec581c519072c4a81717897873c87abeb6f5cda795d9c

Observation 8718fe95-bf6e-4e3d-8bfd-c7ac573d3059 · outbound

This paper cites pickle — Python object Serialization,.

Speeding up Model Loading with fastsafetensors pickle — Python object Serialization,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.697345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:05.012673Z digest=sha256:fd7198f780ed6d25e2924eeadacc19bb4edc3d3694c25cf1c0a954c86c8e2dff

Observation a4160a11-7b76-4e06-ac30-f8ef8dab91ce · outbound

This paper cites Pytorch,.

Speeding up Model Loading with fastsafetensors Pytorch,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.287909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:05.160118Z digest=sha256:a93a408eaae826c7ba57d5f6153a5812523a9b337538b9db60d4339a262fd101

Observation d5df08d8-f0c5-4082-a97e-39e272121d0e · outbound

This paper cites Available: https://docs .python.org/3/library/pickle.html.

Speeding up Model Loading with fastsafetensors Available: https://docs .python.org/3/library/pickle.html

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.472708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:05.086172Z digest=sha256:47327a5bf13211f4a15b751cae52618b2e1de42acbdaec8459dd1b76e6cd0b51

Observation eac066d1-55ad-4914-9134-18f059fbd4c8 · outbound

This paper cites ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning.

Speeding up Model Loading with fastsafetensors ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.322950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.322950Z digest=sha256:ea53a93b681fbb8b1594969c396fd3e1da904859c74d086cf6f7917cd7966c0a

Observation 1c0e8d69-88b5-4a1d-9c5e-c55ee495787a · outbound

This paper cites TensorFlow: A System for Large-Scale Machine Learning,.

Speeding up Model Loading with fastsafetensors TensorFlow: A System for Large-Scale Machine Learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.162396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:05.256431Z digest=sha256:f045ea3c2d7ef8a72a25491c69f7c85b9a64ff73a0705df8ee72cf5aa5c4e4b6

Observation f87c9321-eb6f-4360-a83c-21e94e87b89e · outbound

This paper cites ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development,.

Speeding up Model Loading with fastsafetensors ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.974543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:05.583944Z digest=sha256:067cec75600d879479af37221bb4dc882024ec97363355fd4b04a051c162dd3a

Observation da9d9610-51ee-4e8f-8d8d-fd5c57aa3058 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel,.

Speeding up Model Loading with fastsafetensors PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.415353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.415353Z digest=sha256:6cc17e55238d56aada95344f8a9b331b2c2a71624957c33ae274718ff30d9368

Observation f5fdf0db-d9a9-45f0-a3a6-9d4f12f43a1b · outbound

This paper cites NVIDIA Magnum IO GPUDirect Storage Design Guide,.

Speeding up Model Loading with fastsafetensors NVIDIA Magnum IO GPUDirect Storage Design Guide,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.563051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:05.795124Z digest=sha256:4b8d4e215d7d9923f9941ecaf0aa1198750fdf46ec5bd28591292b22a715680a

Observation ea280cc0-7b24-4930-8870-e6faefdd52ae · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Speeding up Model Loading with fastsafetensors Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.994728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.994728Z digest=sha256:79a1f3889f35d233adf75ca44194b4e7816a612f4b8be2420d6efa7a41fdf746

Observation 17d7f8ea-1c41-42aa-b6e3-5e2dfc00f514 · outbound

This paper cites ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development.

Speeding up Model Loading with fastsafetensors ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.667602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.667602Z digest=sha256:8852496a39f682ab0595e768a88f309bd6ab1d766798e5f0fa2017bdee4ec0e0

Observation bcd5b2ae-657d-4b94-bf1a-6bd50878fd60 · outbound

This paper cites Welcome to DLPack’s documentation! — DLPack 0.6.0 documentation,.

Speeding up Model Loading with fastsafetensors Welcome to DLPack’s documentation! — DLPack 0.6.0 documentation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.782963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:05.733164Z digest=sha256:71c5b9163fb6c10434c8170b7e3a65fd234a66d78f7ef346828b631ca9c247a6

Observation 44d43244-a860-49cb-a83a-d032e69233e5 · outbound

This paper cites huggingface/text-generation-inference: Large Language Model Text Generation Inference,.

Speeding up Model Loading with fastsafetensors huggingface/text-generation-inference: Large Language Model Text Generation Inference,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.177797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:06.274780Z digest=sha256:92d471467a0f03cfcd11f24854a4f328de19078bd8ccad578e769fe45f27fb22

Observation 6356c0c6-767c-49d9-842a-31481fc4d10e · outbound

This paper cites vllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for LLMs,.

Speeding up Model Loading with fastsafetensors vllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for LLMs,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.966023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:06.378960Z digest=sha256:9f7a5e012379f68b34062616c6f7aa4f7f9a9e36b2cf17e2085f7c8e19d1cd31

Observation 1145e70a-359e-48e6-99fa-8aa855cee725 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs,.

Speeding up Model Loading with fastsafetensors SGLang: Efficient Execution of Structured Language Model Programs,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:06.449698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:06.449698Z digest=sha256:5fc9e88d4243f2e471cf7b2dea2445ad1c90723dbfe7236d706c229b18601d4f

Observation 11caba6e-a8c6-410f-a6e8-0f52e01635ec · outbound

This paper cites The Falcon Series of Open Language Models.

Speeding up Model Loading with fastsafetensors The Falcon Series of Open Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:06.100629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:06.100629Z digest=sha256:0619d7195f5a738251cb0935d9f85aa5597a39ac1ce8efb93c535a1c493474ff

Observation 911f8b35-c10d-40d5-bf03-ef139478ad66 · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

Speeding up Model Loading with fastsafetensors BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:06.183400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:06.183400Z digest=sha256:93d7309a935209a242eb732bb491ca499ce8a007f80ae04f67eb42977615d689

Observation 0471cae3-d345-484f-a461-0339ac493343 · outbound

This paper cites PEP 703 — Making the Global Interpreter Lock Optional in CPython,.

Speeding up Model Loading with fastsafetensors PEP 703 — Making the Global Interpreter Lock Optional in CPython,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.485679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:06.692824Z digest=sha256:4cc47e74841ed0b9da5f8e7a0abdd3c19503909eb4e2254d2abba7573ac0ee2d

Observation beddeef5-3e13-41b1-9354-49ff7190c4d2 · outbound

This paper cites SPIN: Seam- less Operating System Integration of Peer-to-Peer DMA Between SSDs and GPUs,.

Speeding up Model Loading with fastsafetensors SPIN: Seam- less Operating System Integration of Peer-to-Peer DMA Between SSDs and GPUs,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.306983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:06.791585Z digest=sha256:3d73e88321b4aeed9ed0902aaf4e2b41cd4cf549b362c0bb1e9cf4ec4bba81dc

Observation 55590a15-344c-4e65-857b-a27540e8f8c1 · outbound

This paper cites How beneficial is peer-to-peer DMA?.

Speeding up Model Loading with fastsafetensors How beneficial is peer-to-peer DMA?

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.178224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:06.863437Z digest=sha256:354de2fc2dda3637d8cd92ac517181d238d72538c02489683ccdddfce0a0d0eb

Observation 3ee878dd-2212-47c6-ba76-acdff335049c · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Speeding up Model Loading with fastsafetensors SGLang: Efficient Execution of Structured Language Model Programs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:06.513769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:06.513769Z digest=sha256:8f659872953082d58ad6f1f1ab7d455b44ef41094e27db38f39b49b5ccab7b55

Observation bee2d220-7aeb-4748-819b-4dc82321b264 · outbound

This paper cites Rapid Data Pre- Processing with NVIDIA DALI,.

Speeding up Model Loading with fastsafetensors Rapid Data Pre- Processing with NVIDIA DALI,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.800106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:06.584427Z digest=sha256:8e6fc4c4b48c1aa572107c1dc86c9b51826b9f30068be05f2db22e631d86e5b0

Observation b828a218-c17b-406d-b38d-c4422a76126c · outbound

This paper cites Accelerate AI and ML workloads with OCI, NVIDIA Magnum IO GPUDirect Storage, and IBM Storage Scale,.

Speeding up Model Loading with fastsafetensors Accelerate AI and ML workloads with OCI, NVIDIA Magnum IO GPUDirect Storage, and IBM Storage Scale,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.649672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:06.654325Z digest=sha256:64e1a1e7644a6958cca521967c4205134adfdf5dba4c12a39bebd38222b28359

Observation 84dd1e58-3951-47ed-a478-0fe6cd855632 · outbound

This paper cites FP8 Formats for Deep Learning.

Speeding up Model Loading with fastsafetensors FP8 Formats for Deep Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:07.253100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:07.253100Z digest=sha256:e703f4d998bd35a0e02f51b3182a0036d62055254d4f5ac57037a4e2d95e2fa6

Observation 151eca64-8839-4a13-bc73-a709b13deba7 · outbound

This paper cites Efficient Post-training Quantization with FP8 Formats.

Speeding up Model Loading with fastsafetensors Efficient Post-training Quantization with FP8 Formats

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:07.344437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:07.344437Z digest=sha256:795948628915690bbdf351191b9a714967c8eb8178b51345d6719ec82e3c56aa

Observation 8514445a-46b1-414c-9313-606aeec79f5a · outbound

This paper cites FP8 Quantization: The Power of the Exponent.

Speeding up Model Loading with fastsafetensors FP8 Quantization: The Power of the Exponent

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:07.427165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:07.427165Z digest=sha256:4f86ceb3cfa9fc498586953364c8eb9b0e8cd49b78af14aa675c35d5e388bfa5

Observation 0ce17f57-773f-4db8-a49a-e450f646dbc0 · outbound

This paper cites Column Cache: Buffer Cache for Columnar Storage on HDFS,.

Speeding up Model Loading with fastsafetensors Column Cache: Buffer Cache for Columnar Storage on HDFS,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.008794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:06.941011Z digest=sha256:38918e57969247ebce93adae995d6514b956c3269b44eccf4c623ab9a90dac8a

Observation 4525d906-06ac-473e-a464-382a9ac9027a · outbound

This paper cites Teraheap: Reducing memory pressure in managed big data frameworks,.

Speeding up Model Loading with fastsafetensors Teraheap: Reducing memory pressure in managed big data frameworks,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:07.862265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:07.049645Z digest=sha256:1471592f8670673f1065205a8a8b34cef0c657ab6e60d60ce78f40d9460a51e2

Observation fc358a47-a979-4eb3-9fd4-c52b5fc52e0b · outbound

This paper cites Accelerating multilingual applications with in-memory array sharing,.

Speeding up Model Loading with fastsafetensors Accelerating multilingual applications with in-memory array sharing,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:07.752961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:07.153431Z digest=sha256:f8cbd3e7594faba0b88639359fba4a124487bf765d6bb57332fa03e0260d4284

Observation 5d4bbd67-4c41-4ccf-be99-f739761dd62b · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Speeding up Model Loading with fastsafetensors PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.511142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.511142Z digest=sha256:84e8003c636740be1a483a9aed3020d053440ae84bb5778a50ec30a755a2b123

Observation 465ebe04-efce-4d36-a082-c9493210a02f · outbound

This paper cites Granite Code Models: A Family of Open Foundation Models for Code Intelligence.

Speeding up Model Loading with fastsafetensors Granite Code Models: A Family of Open Foundation Models for Code Intelligence

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.032742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.032742Z digest=sha256:68703bede2070775aa280e4bb8bb51d17c237e2fa40685a3e16b901d1c40f8b5

Observation 842d19cb-7c3f-4cbd-a9ff-095b167ac77c · outbound

This paper cites Available: https://docs .nvidia.com/gpudirect-storage/ design-guide/index.html.

Speeding up Model Loading with fastsafetensors Available: https://docs .nvidia.com/gpudirect-storage/ design-guide/index.html

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.368234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T12:58:05.911032Z digest=sha256:ffb3adee0b71881036d5735d416aed751381a15ba348ed803820ebaf0ebc1074

Pith citing papers

Observation 42ecb65a-73c0-4442-bca7-7388a2662378 · inbound

RTP-LLM: High-Performance Alibaba LLM Inference Engine cites this paper.

RTP-LLM: High-Performance Alibaba LLM Inference Engine Speeding up Model Loading with fastsafetensors

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:52:49.146323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T23:52:40.763228Z digest=sha256:3a5127ada225f62ea045dcc6c66642c8c24d527966623dbeef0d61fb473dea44

Observation 7d5d1c20-ccf9-4905-a4bf-701289fb01aa · inbound

The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing cites this paper.

The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing Speeding up Model Loading with fastsafetensors

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-26T22:10:09.429560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-26T06:39:04.161585Z digest=sha256:f745a1aa91ca97b576b5461207a9a2457e6bc3b234892146c3c1b6d2ac59f769

Observation c7989ea8-90ee-4b21-9d22-1c174033d54c · inbound

The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing cites this paper.

The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing Speeding up Model Loading with fastsafetensors

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T07:20:36.786149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:20:36.786149Z digest=sha256:0fa02148ed26c0a4e5a8a3be15180de44094dcf759cd9ec1506017c5a89c4e9a

Observation 28bcb40f-0c63-43a1-96ad-41b4b97a7e89 · inbound

InstantInfer: Enabling Fast LLM Cold Start with Communicating Finite Automata cites this paper.

InstantInfer: Enabling Fast LLM Cold Start with Communicating Finite Automata Speeding up Model Loading with fastsafetensors

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T13:52:43.151100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:52:43.151100Z digest=sha256:860885a196070fd734fbd17c324d77a4aa28ae4889de13c5d4e1193edb1b1f11