Pith. sign in

Paper Citation Record · LEDGER

Speeding up Model Loading with fastsafetensors

As of 8 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 4 inbound Pith citation observations for arXiv:2505.23072.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23072 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:58:07.427165Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T13:52:43.151100Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 1eb807ef-31d1-4c1d-9bdf-4ecf60c58ed7 · outbound

This paper cites Introducing ChatGPT,.

Speeding up Model Loading with fastsafetensors Introducing ChatGPT,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.656524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:03.832875Z digest=sha256:f4299e856e1ce9a5b3ffbaebdb46d2bd0302d1d12f742bad2a8fb4431965b0e9

Observation f5b86204-bff4-4e3f-ac2b-9f89a0b3f10e · outbound

This paper cites Introducing Gemini: our largest and most capable AI model,.

Speeding up Model Loading with fastsafetensors Introducing Gemini: our largest and most capable AI model,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.492922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:03.902873Z digest=sha256:ce9e9b821c70515f39d1b580173ee922f7b9c81766713fa628e89db4f93bd404

Observation 9a68980f-964e-44a0-9507-846aeffe203f · outbound

This paper cites Granite Code Models: A Family of Open Foundation Models for Code Intelligence,.

Speeding up Model Loading with fastsafetensors Granite Code Models: A Family of Open Foundation Models for Code Intelligence,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.366624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:03.973943Z digest=sha256:412839a1e8e88da69cc0d2fedd27732b2fb3230576c6fd2107f6ed08e9a55790

Observation 25defce5-5c40-4fd9-a40e-73806d7154ac · outbound

This paper cites The Shift from Models to Compound AI Systems,.

Speeding up Model Loading with fastsafetensors The Shift from Models to Compound AI Systems,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.176639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:04.077312Z digest=sha256:422389401a05a12a9f72be9308e97f56949ffcb7cb86af1d16ba2fde2a177296

Observation 87dc08d5-980d-4bd8-b7d9-83a130d7e7d3 · outbound

This paper cites FlashAttention: Fast and memory-efficient exact attention with IO-awareness,.

Speeding up Model Loading with fastsafetensors FlashAttention: Fast and memory-efficient exact attention with IO-awareness,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.059280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:04.145099Z digest=sha256:f64b9c839991832e650b514c325ab385bb406d7cb16b896f80e9656605fd6f24

Observation fd73d497-5dc8-42c4-9eb6-1b2a8d87bdc7 · outbound

This paper cites FlashAttention-2: Faster attention with better parallelism and work partitioning,.

Speeding up Model Loading with fastsafetensors FlashAttention-2: Faster attention with better parallelism and work partitioning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.895787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:04.214691Z digest=sha256:a7af91db532e15903c695986c64be07da7f711579cd5e119d22d7509ea0886ef

Observation 12cc850f-9c15-4005-87b1-9ae7c7d25ae2 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention,.

Speeding up Model Loading with fastsafetensors Efficient Memory Management for Large Language Model Serving with PagedAttention,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.758834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:04.304079Z digest=sha256:b1b0a3044c604f55b116cb6f769a712fb3def25f0bd5df250f515905ccca6bd1

Observation ae9808a8-4a8a-421d-9b8b-1cd773232cd5 · outbound

This paper cites Orca: A Distributed Serving System for Transformer-Based Generative Models,.

Speeding up Model Loading with fastsafetensors Orca: A Distributed Serving System for Transformer-Based Generative Models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.611414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:04.350083Z digest=sha256:aee03da95c77f792fd95a78ccb9339fd80ed482afd53cab00b9d9505cc5680b8

Observation e30d29f1-ab5e-4cdb-9cc2-d96cb0e427d8 · outbound

This paper cites Accelerating Production LLMs with Combined Token/Embedding Speculators.

Speeding up Model Loading with fastsafetensors Accelerating Production LLMs with Combined Token/Embedding Speculators

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.425862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.425862Z digest=sha256:fc4174164d1b611c5b9f0c5c24b4d31f06d408db9612560a50a85448f7c3590b

Observation e64cecdc-7e37-422b-9e32-862ec995d3f4 · outbound

This paper cites DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving.

Speeding up Model Loading with fastsafetensors DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.496281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.496281Z digest=sha256:852a159a11a4658f888d66f6200826eaf23a4be9cd991b1e322567327f9c1d35

Observation c2718510-a082-46bb-aa53-00fac5a49f0d · outbound

This paper cites Taming throughput-latency tradeoff in llm inference with sarathi-serve,.

Speeding up Model Loading with fastsafetensors Taming throughput-latency tradeoff in llm inference with sarathi-serve,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.472887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:04.561087Z digest=sha256:8964457bc21cf3f6e167b5e060434c098e88bc040fe52359b057d7921ac3f9f7

Observation 910f4ac6-e5ae-4023-9856-08bdf6a695fc · outbound

This paper cites Decrease PyTorch Model Load Times with CoreWeave’s Tensorizer,.

Speeding up Model Loading with fastsafetensors Decrease PyTorch Model Load Times with CoreWeave’s Tensorizer,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.309831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:04.639769Z digest=sha256:3a71ab8f279b9b4f10bedda237b3a9481303b91b5829b560ef644800686d321d

Observation cda45ea4-2e7e-4ab7-952f-e4009f79e25f · outbound

This paper cites ServerlessLLM: Low-Latency Serverless Inference for Large Language Models.

Speeding up Model Loading with fastsafetensors ServerlessLLM: Low-Latency Serverless Inference for Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.729125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.729125Z digest=sha256:52b604b399b3e0144ea1db16f72eaa42c0eae0da00d58c64dbf8a89a57c9a2aa

Observation 2b1adefe-1e6a-4e8c-8fe1-fb2f309c00e3 · outbound

This paper cites Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM.

Speeding up Model Loading with fastsafetensors Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.791744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.791744Z digest=sha256:43bc62804ea4cacc4b6f14b124542d2de88fd108a900d37b8c98b0f2f7bca9b3

Observation 89ee66e9-bbba-4b12-ab18-81eab9f57679 · outbound

This paper cites Safetensors,.

Speeding up Model Loading with fastsafetensors Safetensors,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.115727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:04.876060Z digest=sha256:1b3240b12315bf10151c116c64e06bec382d08303318530403b587d4d3f819c1

Observation bc3dc6c4-e620-4184-8408-f68174e86e7b · outbound

This paper cites Models – Hugging Face,.

Speeding up Model Loading with fastsafetensors Models – Hugging Face,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.934865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:04.943400Z digest=sha256:daa690389bbb6e0e3e9b2884bba9041d92d69f47d8ef5305955116d701fdd37f

Observation 8718fe95-bf6e-4e3d-8bfd-c7ac573d3059 · outbound

This paper cites pickle — Python object Serialization,.

Speeding up Model Loading with fastsafetensors pickle — Python object Serialization,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.697345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:05.012673Z digest=sha256:17c2bd71f594fd6641058f02ffc51f9720fe6f64118d48361cc8cee59d1d3693

Observation a4160a11-7b76-4e06-ac30-f8ef8dab91ce · outbound

This paper cites Pytorch,.

Speeding up Model Loading with fastsafetensors Pytorch,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.287909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:05.160118Z digest=sha256:2ef8985443840c04471e78b8d4b679ea3f883052dd16ec101e9522e2bc3f89b2

Observation d5df08d8-f0c5-4082-a97e-39e272121d0e · outbound

This paper cites Available: https://docs .python.org/3/library/pickle.html.

Speeding up Model Loading with fastsafetensors Available: https://docs .python.org/3/library/pickle.html

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.472708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:05.086172Z digest=sha256:56063b0bd7abbf4624af50ab1fb4b0ef36e54e29d49388fbf79f6e8e8d985cd8

Observation eac066d1-55ad-4914-9134-18f059fbd4c8 · outbound

This paper cites ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning.

Speeding up Model Loading with fastsafetensors ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.322950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.322950Z digest=sha256:4b56f6d05fb2bfb7cdc9f6dd9e478ca3409c9e712bacbe81b9b04d7aaa23adf6

Observation 1c0e8d69-88b5-4a1d-9c5e-c55ee495787a · outbound

This paper cites TensorFlow: A System for Large-Scale Machine Learning,.

Speeding up Model Loading with fastsafetensors TensorFlow: A System for Large-Scale Machine Learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.162396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:05.256431Z digest=sha256:56004293419850cec6677ba98f3b0875b478b2077bb054150b537e1bd6287982

Observation f87c9321-eb6f-4360-a83c-21e94e87b89e · outbound

This paper cites ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development,.

Speeding up Model Loading with fastsafetensors ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.974543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:05.583944Z digest=sha256:3e94184f9cbbc176134c9aadd5ceecd38c469073152943d3b8b53fd8a2543aef

Observation da9d9610-51ee-4e8f-8d8d-fd5c57aa3058 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel,.

Speeding up Model Loading with fastsafetensors PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.415353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.415353Z digest=sha256:307cb0bb5e1888c2b8877b219373b1d5e0efbcff3a37ffcf78fd8de75ead3dee

Observation f5fdf0db-d9a9-45f0-a3a6-9d4f12f43a1b · outbound

This paper cites NVIDIA Magnum IO GPUDirect Storage Design Guide,.

Speeding up Model Loading with fastsafetensors NVIDIA Magnum IO GPUDirect Storage Design Guide,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.563051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:05.795124Z digest=sha256:ff9e4eaf4dd9a4465d0e162e7bb1cce23c55e5f26b7dd9f5a8adc9bc0c3a3b35

Observation ea280cc0-7b24-4930-8870-e6faefdd52ae · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Speeding up Model Loading with fastsafetensors Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.994728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.994728Z digest=sha256:b989171803e6fef1d6f70a636e0979b7817eb7f429f9f5c32753439682d3c4ab

Observation 17d7f8ea-1c41-42aa-b6e3-5e2dfc00f514 · outbound

This paper cites ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development.

Speeding up Model Loading with fastsafetensors ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.667602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.667602Z digest=sha256:7b137d47b2bf467327a48ec1073b9dcaf024dc5190c68cc5f8aad6e3b1af25b6

Observation bcd5b2ae-657d-4b94-bf1a-6bd50878fd60 · outbound

This paper cites Welcome to DLPack’s documentation! — DLPack 0.6.0 documentation,.

Speeding up Model Loading with fastsafetensors Welcome to DLPack’s documentation! — DLPack 0.6.0 documentation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.782963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:05.733164Z digest=sha256:4a5913abf80b231dbcea1f299e00e0c836592bf42445f656afcd000bed0579d3

Observation 44d43244-a860-49cb-a83a-d032e69233e5 · outbound

This paper cites huggingface/text-generation-inference: Large Language Model Text Generation Inference,.

Speeding up Model Loading with fastsafetensors huggingface/text-generation-inference: Large Language Model Text Generation Inference,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.177797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:06.274780Z digest=sha256:b47417b16800b849e77cb956695e78ec5a4ee7e693162ec2358fc35776d6c6fb

Observation 6356c0c6-767c-49d9-842a-31481fc4d10e · outbound

This paper cites vllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for LLMs,.

Speeding up Model Loading with fastsafetensors vllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for LLMs,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.966023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:06.378960Z digest=sha256:5ad511e96d1654656cd0d6bf691ad2797ffb3b9cd5c5cda9cbbb14f8ee441cf3

Observation 1145e70a-359e-48e6-99fa-8aa855cee725 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs,.

Speeding up Model Loading with fastsafetensors SGLang: Efficient Execution of Structured Language Model Programs,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:06.449698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:06.449698Z digest=sha256:ab9926cadf00680f2378172640d38a2206fcf813e84102b075f86151cc3f29ce

Observation 11caba6e-a8c6-410f-a6e8-0f52e01635ec · outbound

This paper cites The Falcon Series of Open Language Models.

Speeding up Model Loading with fastsafetensors The Falcon Series of Open Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:06.100629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:06.100629Z digest=sha256:90316b39a5b9b1370a74a412fa81e007a2b7413a09cc8c8549b6f8d501584dde

Observation 911f8b35-c10d-40d5-bf03-ef139478ad66 · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

Speeding up Model Loading with fastsafetensors BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:06.183400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:06.183400Z digest=sha256:6e92665d35202131badb1de0b55c1bc79cf2b4c0d386aa8a74169328319d9e9c

Observation 0471cae3-d345-484f-a461-0339ac493343 · outbound

This paper cites PEP 703 — Making the Global Interpreter Lock Optional in CPython,.

Speeding up Model Loading with fastsafetensors PEP 703 — Making the Global Interpreter Lock Optional in CPython,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.485679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:06.692824Z digest=sha256:b627592b8e9e0c0441efc759d5a7ef98cd22091700486c8d083ab939d4acc80e

Observation beddeef5-3e13-41b1-9354-49ff7190c4d2 · outbound

This paper cites SPIN: Seam- less Operating System Integration of Peer-to-Peer DMA Between SSDs and GPUs,.

Speeding up Model Loading with fastsafetensors SPIN: Seam- less Operating System Integration of Peer-to-Peer DMA Between SSDs and GPUs,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.306983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:06.791585Z digest=sha256:25b49a7a31d45a87645157fd93a687402d2a5bf03e7bcad79999aeb1ddd7e4f4

Observation 55590a15-344c-4e65-857b-a27540e8f8c1 · outbound

This paper cites How beneficial is peer-to-peer DMA?.

Speeding up Model Loading with fastsafetensors How beneficial is peer-to-peer DMA?

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.178224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:06.863437Z digest=sha256:a97a1200c3d18994b5174e8a01697f54e6503fcfd6088d42dc72bd0e871b8442

Observation 3ee878dd-2212-47c6-ba76-acdff335049c · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Speeding up Model Loading with fastsafetensors SGLang: Efficient Execution of Structured Language Model Programs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:06.513769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:06.513769Z digest=sha256:48a5333f26df6ca5f5f12b249137c8afdba1c247c7e996cf75a6d017947b534d

Observation bee2d220-7aeb-4748-819b-4dc82321b264 · outbound

This paper cites Rapid Data Pre- Processing with NVIDIA DALI,.

Speeding up Model Loading with fastsafetensors Rapid Data Pre- Processing with NVIDIA DALI,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.800106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:06.584427Z digest=sha256:e0f4f5851aaa72f56f71744f5bdc17229e763c682b08092c9af99e7b29cef2c0

Observation b828a218-c17b-406d-b38d-c4422a76126c · outbound

This paper cites Accelerate AI and ML workloads with OCI, NVIDIA Magnum IO GPUDirect Storage, and IBM Storage Scale,.

Speeding up Model Loading with fastsafetensors Accelerate AI and ML workloads with OCI, NVIDIA Magnum IO GPUDirect Storage, and IBM Storage Scale,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.649672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:06.654325Z digest=sha256:bf9be565350a0b3a8c79e2afb339d15cb5c87a3bd771e09264bd28fc2bc1ae0e

Observation 84dd1e58-3951-47ed-a478-0fe6cd855632 · outbound

This paper cites FP8 Formats for Deep Learning.

Speeding up Model Loading with fastsafetensors FP8 Formats for Deep Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:07.253100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:07.253100Z digest=sha256:42b1c8a8a1f5b480bd6424b6a76bf1145371ae9681b2b7b104b8e2c1477bfd9b

Observation 151eca64-8839-4a13-bc73-a709b13deba7 · outbound

This paper cites Efficient Post-training Quantization with FP8 Formats.

Speeding up Model Loading with fastsafetensors Efficient Post-training Quantization with FP8 Formats

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:07.344437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:07.344437Z digest=sha256:c78f275cc2e7305bd8813596db8f9d5645c36b195b43620459312812e6a4c490

Observation 8514445a-46b1-414c-9313-606aeec79f5a · outbound

This paper cites FP8 Quantization: The Power of the Exponent.

Speeding up Model Loading with fastsafetensors FP8 Quantization: The Power of the Exponent

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:07.427165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:07.427165Z digest=sha256:7d513af2f488358c33661ae164d9e068fee9afbe951363b85c16ea9dffe450a7

Observation 0ce17f57-773f-4db8-a49a-e450f646dbc0 · outbound

This paper cites Column Cache: Buffer Cache for Columnar Storage on HDFS,.

Speeding up Model Loading with fastsafetensors Column Cache: Buffer Cache for Columnar Storage on HDFS,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.008794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:06.941011Z digest=sha256:31ed1c19470ebe6ec6b5b039cd3f98def854a7bdfeffc8ff96d67372219e7ae9

Observation 4525d906-06ac-473e-a464-382a9ac9027a · outbound

This paper cites Teraheap: Reducing memory pressure in managed big data frameworks,.

Speeding up Model Loading with fastsafetensors Teraheap: Reducing memory pressure in managed big data frameworks,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:07.862265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:07.049645Z digest=sha256:b77d47dbd1842f35123023b8ee6970fb9cdabc5ad91b86940e98a1cbbc1ae10e

Observation fc358a47-a979-4eb3-9fd4-c52b5fc52e0b · outbound

This paper cites Accelerating multilingual applications with in-memory array sharing,.

Speeding up Model Loading with fastsafetensors Accelerating multilingual applications with in-memory array sharing,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:07.752961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:07.153431Z digest=sha256:e720a31392ae623fe808be59763b1626cbacfb8fc8718db6664ebc3a66a90530

Observation 5d4bbd67-4c41-4ccf-be99-f739761dd62b · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Speeding up Model Loading with fastsafetensors PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.511142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.511142Z digest=sha256:d974906431420f39b9920868dab110b0a8fdb7781804f8ca3ea8b41d6e1d9581

Observation 465ebe04-efce-4d36-a082-c9493210a02f · outbound

This paper cites Granite Code Models: A Family of Open Foundation Models for Code Intelligence.

Speeding up Model Loading with fastsafetensors Granite Code Models: A Family of Open Foundation Models for Code Intelligence

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.032742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.032742Z digest=sha256:f9d868589a08bd3b28b20858f19ad2b7783ffee46311dcaffcf584b2db0bcefb

Observation 842d19cb-7c3f-4cbd-a9ff-095b167ac77c · outbound

This paper cites Available: https://docs .nvidia.com/gpudirect-storage/ design-guide/index.html.

Speeding up Model Loading with fastsafetensors Available: https://docs .nvidia.com/gpudirect-storage/ design-guide/index.html

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.368234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:58:05.911032Z digest=sha256:b30dc6d443f7a68dcb7110f14f76c7c1336c9b857db9b6af309fb5359401383f

Pith citing papers

Observation 42ecb65a-73c0-4442-bca7-7388a2662378 · inbound

RTP-LLM: High-Performance Alibaba LLM Inference Engine cites this paper.

RTP-LLM: High-Performance Alibaba LLM Inference Engine Speeding up Model Loading with fastsafetensors

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:52:49.146323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T23:52:40.763228Z digest=sha256:aa7b5acc5b1dfeec978d62ef387433b23b8b13f4f90a4b2496b6cf9d2344bf53

Observation 7d5d1c20-ccf9-4905-a4bf-701289fb01aa · inbound

The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing cites this paper.

The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing Speeding up Model Loading with fastsafetensors

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-26T22:10:09.429560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T06:39:04.161585Z digest=sha256:48ac8d834f4cf6380b9ef7de9dec8c1effbe54f20daf7f3c55b6a6291d36d3e3

Observation c7989ea8-90ee-4b21-9d22-1c174033d54c · inbound

The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing cites this paper.

The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing Speeding up Model Loading with fastsafetensors

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T07:20:36.786149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:20:36.786149Z digest=sha256:93d0038d050d82fc40b097199b3a135f052bdd685b61e252dc316d3133bcbd3f

Observation 28bcb40f-0c63-43a1-96ad-41b4b97a7e89 · inbound

InstantInfer: Enabling Fast LLM Cold Start with Communicating Finite Automata cites this paper.

InstantInfer: Enabling Fast LLM Cold Start with Communicating Finite Automata Speeding up Model Loading with fastsafetensors

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T13:52:43.151100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:52:43.151100Z digest=sha256:f40eb2f65ce97722a06166c63cea596c65095e10641514fd16354c3730d2824a