Pith. sign in

Paper Citation Record · LEDGER

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving

As of 10 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2607.29575.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.29575 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:22:54.961186Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1283f76c-0bbf-4207-80e6-721c88cf4e48 · outbound

This paper cites The rapid adoption of generative ai,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving The rapid adoption of generative ai,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.621032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.621032Z digest=sha256:d2ecadeeffc7d976acfd6b9a78885b12dd04cd2df6249d932b13aadd1136a9f8

Observation 71750f8e-e65d-4dfc-8fc7-5c8e831faba4 · outbound

This paper cites The adoption of chatgpt,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving The adoption of chatgpt,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.687853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.687853Z digest=sha256:600d2ec8c2f903ef9bace23b795851bdcdd24c7fe6eaf081210caf625785354f

Observation 5e3aec83-1a24-4388-adde-a9bb58b1fdf0 · outbound

This paper cites Quantifying large language model usage in scientific papers,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Quantifying large language model usage in scientific papers,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.737457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.737457Z digest=sha256:ea9ce2a94538c1fc7f36d12085103cb3f10dad7ec6098f38757ef6e7356ce851

Observation d7a31b72-eb9d-4ed6-8138-bb0287ef2850 · outbound

This paper cites Llumnix: Dynamic scheduling for large language model serving,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Llumnix: Dynamic scheduling for large language model serving,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.844613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.844613Z digest=sha256:8b7942910d87f29049f97376c2f4a9627137b4c909525fc50c6d86a76f346951

Observation a627426c-a4f4-4206-a394-04d805de12ae · outbound

This paper cites Sageserve: Optimizing llm serving on cloud data centers with forecast aware auto-scaling,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Sageserve: Optimizing llm serving on cloud data centers with forecast aware auto-scaling,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.935319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.935319Z digest=sha256:ebb0de39f1cdcd6f55c8c929e77554370252bb0f938799550d47d3666fdebcda

Observation dc0fee93-9e4d-46c1-8b28-cfa5f828b8ae · outbound

This paper cites Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.032744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.032744Z digest=sha256:5bce84b6436affb26ce508acbd7facc2f47bfbe8e4ec34babfd95c64e2356398

Observation 3d7e53ee-93d8-47cb-a734-76e6e4e5c880 · outbound

This paper cites Efficiently scaling transformer inference,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Efficiently scaling transformer inference,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.082197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.082197Z digest=sha256:bb317a47ddc09139685ca1f7018b76f0f320fd856f0a5120f822e75eb09d9830

Observation dbd5fb7e-2239-4859-89da-77361ba9beb1 · outbound

This paper cites Efficient llm inference: Bandwidth, compute, synchronization, and capacity are all you need,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Efficient llm inference: Bandwidth, compute, synchronization, and capacity are all you need,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.167267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.167267Z digest=sha256:00d02c97b3e64d60170302b1f0c71382696510c8aaf36c9e4089da69e38581e6

Observation 8401ac48-13a1-42e6-96f7-d75a09241a32 · outbound

This paper cites Mind the memory gap: Unveiling gpu bot- tlenecks in large-batch llm inference,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Mind the memory gap: Unveiling gpu bot- tlenecks in large-batch llm inference,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.228945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.228945Z digest=sha256:52125893ca279ec27837edcae15e09232806d0bd6220d120d8d2602e70566721

Observation e9be6525-2ed8-4deb-baa1-20852599966e · outbound

This paper cites Llmvisor: A real-time latency attribution model for multi-tenant llm serving,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Llmvisor: A real-time latency attribution model for multi-tenant llm serving,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.312448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.312448Z digest=sha256:1c2ebd55ea7f5f32432d24fea14af3147d4d8b2a253fe60bc4088a22e60464a4

Observation 169dd382-9cd3-48bf-8c43-d32d50983289 · outbound

This paper cites Predicting llm inference latency: A roofline-driven ml method,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Predicting llm inference latency: A roofline-driven ml method,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.386461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.386461Z digest=sha256:3155df0272f75a642e0fcaa0ca63a05176752ad67a4c9a5b1cdb1d78a81e9e61

Observation faa6fa12-8937-465d-87bf-2c10a68eb48a · outbound

This paper cites Language mod- els are few-shot learners,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Language mod- els are few-shot learners,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.529800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.529800Z digest=sha256:eac3469d80846e230d33a24652d251dac9fa3bf0c080d9ff88e76926f0eab719

Observation 06875560-041e-47bd-bca5-4e0a529eaf44 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving LLaMA: Open and Efficient Foundation Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.616647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.616647Z digest=sha256:7097680db3fe1d173ba14a3e34c97626ee74849a485e4212d02a4899b3917c3e

Observation ab551368-fb6b-46ed-b8d7-789c77352651 · outbound

This paper cites Attention is all you need,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Attention is all you need,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.703801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.703801Z digest=sha256:fb2baf6a8cd99a6cf297d00969ad3f81299b53be80122fe06e021d79abf7c997

Observation de2897af-d4a1-4f69-8e86-b899daadc481 · outbound

This paper cites Orca: A distributed serving system for transformer-based generative models,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Orca: A distributed serving system for transformer-based generative models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.769107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.769107Z digest=sha256:51960f0acf480c3f9ffb6f3f3d0874e9822646d4f2366bf677bf9e1bc12830a5

Observation 56f5eeae-c9e5-433a-b6c4-22b4f42723c9 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Efficient memory management for large language model serving with pagedattention,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.843887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.843887Z digest=sha256:30c3c78f0895aae61dabac992bdaff0a5ea244d6299affccc40fbbdffbd11172

Observation 3aa6c78b-0554-4beb-a73e-df1f537fc2e2 · outbound

This paper cites TensorRT-LLM,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving TensorRT-LLM,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.928323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.928323Z digest=sha256:dc16a2d4be360194ff1746d02676ede843d08e30ade27ac6254687a83b6acc21

Observation 94854dc3-b780-4314-b2a6-ef0ecc871228 · outbound

This paper cites DeepSpeed-MII,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving DeepSpeed-MII,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.030125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.030125Z digest=sha256:f01015865a2c2c3bb59258eadffd960e684f877c2948bab24636d701717248ec

Observation f142fcee-c690-41b7-a722-d974ca81131e · outbound

This paper cites Slora: Scalable serving of thousands of lora adapters,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Slora: Scalable serving of thousands of lora adapters,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.117570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.117570Z digest=sha256:01f0cc629bea25d08fcd4579cc3d9f603eeab81752d3d47dd69116de4008c578

Observation a133c8c9-ab5f-42b6-a05b-570a8d2e5d44 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.202838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.202838Z digest=sha256:b443929bdc9b94d973d2db95ac1e8dbf4f9c12b5512aa62ce28b235650fe844b

Observation 46ded418-0c69-4282-b2fd-5ca8d11f8edc · outbound

This paper cites Taming{Throughput-Latency}tradeoff in{LLM}inference with{Sarathi-Serve},.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Taming{Throughput-Latency}tradeoff in{LLM}inference with{Sarathi-Serve},

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.305385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.305385Z digest=sha256:198f22597997d0f55d9a1e86f1c3ae09b9dff33753f7addec979ada387b2485e

Observation e556d816-ff08-40d6-835e-f20286af23c8 · outbound

This paper cites Fast inference from transform- ers via speculative decoding,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Fast inference from transform- ers via speculative decoding,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.428587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.428587Z digest=sha256:e041010c1e785bf51d973008745f0fa25ed9d6fa4c778dbd5f334a5e673244df

Observation 11947e0b-a77a-40a4-9b7f-51eeb4fa4fc2 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Accelerating Large Language Model Decoding with Speculative Sampling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.550502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.550502Z digest=sha256:f14e9d257d30d0462da7d507763e9d78f6ad8296407a6ef1c251b58abff96f3d

Observation f0c8d29b-ce4a-4423-b2d6-fc9afdaae304 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Fast Transformer Decoding: One Write-Head is All You Need

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.668526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.668526Z digest=sha256:3ef13a3f2c730e662cd588dd1a2a8026bd2f53235ae69be4704d64c09cebf45a

Observation 4fff4c06-15c0-4d6c-82a7-e8a0c5e61c97 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.767380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.767380Z digest=sha256:57a123639c5644f446c0eb45cbba17e78016ac331e766a1539eea2f1960ecf3d

Observation 8a9d34a5-5e25-41b3-b4f1-f9dde91eb23f · outbound

This paper cites Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.847735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.847735Z digest=sha256:6ffcb83ea279c1d4723c860f111302aa91864e27ccc59de5f9f188913809771a

Observation 6f35459a-2052-4824-bc59-7fc008a098fc · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.915456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.915456Z digest=sha256:ebc5e83641dac7eb1f732856b61049c13e9b1737ecaa52a2f205c2927f5e7ee1

Observation ec390e4f-1a26-45ca-b8ef-c96e3928b659 · outbound

This paper cites Vidur: A large-scale simulation framework for llm inference,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Vidur: A large-scale simulation framework for llm inference,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.024548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.024548Z digest=sha256:a2b9b40d33da2c422f9ce1ad209ad218a1e7825ecd9090868b0a891ffc856d1b

Observation 9474e2ac-8847-448d-a2c5-33784756ba3a · outbound

This paper cites Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.136193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.136193Z digest=sha256:628d1cdf2d5b682f8cdbeb6844c668c63ed968c3587b207d8a90c821bec72e83

Observation e3a09dfd-6bd0-4c5c-b9c0-dad1f0f4702a · outbound

This paper cites Llmcompass: Enabling efficient hardware design for large language model inference,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Llmcompass: Enabling efficient hardware design for large language model inference,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.209163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.209163Z digest=sha256:39dae6503134ef18d78cef00e205099f414cc5a86d9f5f6270e6d7fa28f42ccf

Observation 8d956b91-494c-4617-83d1-4b357831ad48 · outbound

This paper cites Amali: An analytical model for accurately modeling llm inference on modern gpus,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Amali: An analytical model for accurately modeling llm inference on modern gpus,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.321060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.321060Z digest=sha256:edf2a71907577c827c9264eb13ad4756cf6566b30f274e8a6eafd58ceb102147

Observation 46db8890-3e1f-4963-9f18-6cc14c0f8fef · outbound

This paper cites Fairness in serving large language models,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Fairness in serving large language models,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.393342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.393342Z digest=sha256:f5ea1d137d25ea473e33cfc434689636cd4347ba6e2bc358c05f7c79f45eae04

Observation ef1c6471-f486-4f1d-a3b2-dc42413c0c53 · outbound

This paper cites Clean sharegpt dataset,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Clean sharegpt dataset,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.458412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.458412Z digest=sha256:88ffd1a55446a0924119af561dc9ea1d936f68fe0865f45d4e5db5dfda3b6fc3

Observation 6b314849-88b5-42a6-910a-381b7b4b23f0 · outbound

This paper cites Mistral 7B.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Mistral 7B

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.533081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.533081Z digest=sha256:986e8f46113b24138e1f4ee6f941c7ab250c751c2deea40d446d5f2cbf06b50f

Observation a2a6e250-6339-4122-b15f-b4bac850f715 · outbound

This paper cites Granite Code Models: A Family of Open Foundation Models for Code Intelligence.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Granite Code Models: A Family of Open Foundation Models for Code Intelligence

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.615902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.615902Z digest=sha256:2b6ca15efc6703caa650b0e5a7b48bd4b94738d7ee015479c3e37d3a4da8eeb9

Observation ab58eaea-b3aa-4790-bff5-1bc003501ea7 · outbound

This paper cites Opt: Open pre-trained transformer language models,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Opt: Open pre-trained transformer language models,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.707629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.707629Z digest=sha256:4e7cc5d4593d5706ce0dd89a31d26a876c1123d73c58e3dfa27301bbdd1c44ff

Observation 543d0be1-ba38-436c-9a08-fd300a83dc8a · outbound

This paper cites Qwen2 Technical Report.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Qwen2 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.849854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.849854Z digest=sha256:2eacb26f52193061d91e184e5206b05ef2ddf1c287d9876fc17542d9d6d3b320

Observation c589949a-a181-4c18-b42c-a44f50ab1084 · outbound

This paper cites Ai and memory wall,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Ai and memory wall,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.961186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.961186Z digest=sha256:b37dfd2a1a230bfbd1bc0ef6c9d81c567f7d8c86304c22d144504aff45040279

Observation faa8d4de-c4d2-4dc7-919c-4814191e17ec · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving OPT: Open Pre-trained Transformer Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.775813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.775813Z digest=sha256:4aebf121c75d6537e6e0610822db8c17419242c2ae38bfb1f1838edf03bcb6f4

Pith citing papers

No inbound Pith citation observations are available.