Pith. sign in

Paper Citation Record · LEDGER

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving

As of 8 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2607.29575.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.29575 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:22:54.961186Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1283f76c-0bbf-4207-80e6-721c88cf4e48 · outbound

This paper cites The rapid adoption of generative ai,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving The rapid adoption of generative ai,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.621032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.621032Z digest=sha256:dba9fd6d1663ffb3f638d786fd0fbf8e82845f37b7e49885dacfc86f56f86a84

Observation 71750f8e-e65d-4dfc-8fc7-5c8e831faba4 · outbound

This paper cites The adoption of chatgpt,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving The adoption of chatgpt,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.687853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.687853Z digest=sha256:19f60a1dafb0187fc56c454388f391b2f8308bb0becb4f92bf92c535d958dbab

Observation 5e3aec83-1a24-4388-adde-a9bb58b1fdf0 · outbound

This paper cites Quantifying large language model usage in scientific papers,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Quantifying large language model usage in scientific papers,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.737457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.737457Z digest=sha256:a6631e6924a68c670a0544ef85bb84ca3be8d7f73cde8f44caa74c3331ffe977

Observation d7a31b72-eb9d-4ed6-8138-bb0287ef2850 · outbound

This paper cites Llumnix: Dynamic scheduling for large language model serving,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Llumnix: Dynamic scheduling for large language model serving,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.844613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.844613Z digest=sha256:cb0de26d5a6bcb37027643a875426f1ddc9cba523f0563337b9b12e4eb0c50e1

Observation a627426c-a4f4-4206-a394-04d805de12ae · outbound

This paper cites Sageserve: Optimizing llm serving on cloud data centers with forecast aware auto-scaling,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Sageserve: Optimizing llm serving on cloud data centers with forecast aware auto-scaling,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.935319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.935319Z digest=sha256:ecb1490c3b73e23ee95925fc80b0fb4a1d465decdef3f4f40875747047e0f5de

Observation dc0fee93-9e4d-46c1-8b28-cfa5f828b8ae · outbound

This paper cites Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.032744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.032744Z digest=sha256:e5b36542fdebfa12c62aec49e11543a3d6671c08950db36efd0953c0d6ca13c2

Observation 3d7e53ee-93d8-47cb-a734-76e6e4e5c880 · outbound

This paper cites Efficiently scaling transformer inference,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Efficiently scaling transformer inference,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.082197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.082197Z digest=sha256:5d807d954eb7dea7554c14f4c44d0fdf27419512f8d56a0a5ff4d6796b4171f5

Observation dbd5fb7e-2239-4859-89da-77361ba9beb1 · outbound

This paper cites Efficient llm inference: Bandwidth, compute, synchronization, and capacity are all you need,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Efficient llm inference: Bandwidth, compute, synchronization, and capacity are all you need,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.167267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.167267Z digest=sha256:f08cfc1272afd5b0050593cccf48ab1ba5df7cc51fc9579ec947f372a1dcad56

Observation 8401ac48-13a1-42e6-96f7-d75a09241a32 · outbound

This paper cites Mind the memory gap: Unveiling gpu bot- tlenecks in large-batch llm inference,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Mind the memory gap: Unveiling gpu bot- tlenecks in large-batch llm inference,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.228945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.228945Z digest=sha256:ccc048716519a1644f122b4b936c7d5f7ab08c9e638ea00ea496efa2f8b8009e

Observation e9be6525-2ed8-4deb-baa1-20852599966e · outbound

This paper cites Llmvisor: A real-time latency attribution model for multi-tenant llm serving,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Llmvisor: A real-time latency attribution model for multi-tenant llm serving,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.312448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.312448Z digest=sha256:4727a9ae87d9c0f97ddbbbbbd37bf0a24b0dcd5eaac804c6a5fea4ec9005f065

Observation 169dd382-9cd3-48bf-8c43-d32d50983289 · outbound

This paper cites Predicting llm inference latency: A roofline-driven ml method,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Predicting llm inference latency: A roofline-driven ml method,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.386461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.386461Z digest=sha256:7c5a796d437ef68319b585a6a122a7a2ecc57c4fd9e039b17772f634d70df01e

Observation faa6fa12-8937-465d-87bf-2c10a68eb48a · outbound

This paper cites Language mod- els are few-shot learners,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Language mod- els are few-shot learners,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.529800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.529800Z digest=sha256:2c0f9a85324309a13e925b410db939b11eb5060499981f18c462ded0f89cecde

Observation 06875560-041e-47bd-bca5-4e0a529eaf44 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving LLaMA: Open and Efficient Foundation Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.616647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.616647Z digest=sha256:80b08e2d82556f6939fd70450784f63e22d528288f8a5e9705fc7457d5087755

Observation ab551368-fb6b-46ed-b8d7-789c77352651 · outbound

This paper cites Attention is all you need,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Attention is all you need,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.703801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.703801Z digest=sha256:54aae1b0256a50b2561d4b741eabcc8fc42ee6bbb04d068d02243df04c0424fc

Observation de2897af-d4a1-4f69-8e86-b899daadc481 · outbound

This paper cites Orca: A distributed serving system for transformer-based generative models,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Orca: A distributed serving system for transformer-based generative models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.769107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.769107Z digest=sha256:0bc8e7dd12003ec520113814bb101405b04307e8869934178274b454c781e08f

Observation 56f5eeae-c9e5-433a-b6c4-22b4f42723c9 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Efficient memory management for large language model serving with pagedattention,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.843887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.843887Z digest=sha256:9ae5ac04dbe99d3e59f4abdf52dc45a4c5b602d39b209dc30b6a0584f9b3e9f1

Observation 3aa6c78b-0554-4beb-a73e-df1f537fc2e2 · outbound

This paper cites TensorRT-LLM,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving TensorRT-LLM,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.928323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.928323Z digest=sha256:b71d0e91ea626c912a9f5f5b30dfdd1edcf4cf3fb78c922274416f532f842dcf

Observation 94854dc3-b780-4314-b2a6-ef0ecc871228 · outbound

This paper cites DeepSpeed-MII,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving DeepSpeed-MII,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.030125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.030125Z digest=sha256:9f1837d3c81e74ef4615a5482b9530f43a04cc1bf2bc9aca223b2610b1f846aa

Observation f142fcee-c690-41b7-a722-d974ca81131e · outbound

This paper cites Slora: Scalable serving of thousands of lora adapters,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Slora: Scalable serving of thousands of lora adapters,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.117570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.117570Z digest=sha256:2b95ee2cc055e262d984c12a236421412e0ffcb56a08b498b6c1b6277099738d

Observation a133c8c9-ab5f-42b6-a05b-570a8d2e5d44 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.202838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.202838Z digest=sha256:7c9ce3a153b36b55d744b398d80d8799becfa912b306da8a00a0becbd00a899e

Observation 46ded418-0c69-4282-b2fd-5ca8d11f8edc · outbound

This paper cites Taming{Throughput-Latency}tradeoff in{LLM}inference with{Sarathi-Serve},.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Taming{Throughput-Latency}tradeoff in{LLM}inference with{Sarathi-Serve},

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.305385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.305385Z digest=sha256:3c704f143a162ed10b41457847e43bc8ff78e040107522c3742fdecd43ccdfa7

Observation e556d816-ff08-40d6-835e-f20286af23c8 · outbound

This paper cites Fast inference from transform- ers via speculative decoding,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Fast inference from transform- ers via speculative decoding,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.428587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.428587Z digest=sha256:f70960257d146220f525cd7dd10a8f9ee09db2dd159eecdabaefceb9386c74ee

Observation 11947e0b-a77a-40a4-9b7f-51eeb4fa4fc2 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Accelerating Large Language Model Decoding with Speculative Sampling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.550502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.550502Z digest=sha256:f46dcd4707b4752349389cd0b4cb65a6486124305bd947e34d8bb029feb37c70

Observation f0c8d29b-ce4a-4423-b2d6-fc9afdaae304 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Fast Transformer Decoding: One Write-Head is All You Need

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.668526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.668526Z digest=sha256:d2f7b05a96d3e29497a5b3e3a436b3bae9a045843ca21a7726c8bf1c2f062ff9

Observation 4fff4c06-15c0-4d6c-82a7-e8a0c5e61c97 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.767380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.767380Z digest=sha256:6b2d7bc574edd07fb4bf5201aa5091f3600749543fd147e9b0e71081149133ab

Observation 8a9d34a5-5e25-41b3-b4f1-f9dde91eb23f · outbound

This paper cites Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.847735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.847735Z digest=sha256:39df24dfc20831d0a33a751d94d0a919c136a3874b48a66cd11d251e312e5c8f

Observation 6f35459a-2052-4824-bc59-7fc008a098fc · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.915456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.915456Z digest=sha256:52fa58e724eb3e20d33245d592bc7db2cd6a0f3c14b0e8b0e52af2b474815740

Observation ec390e4f-1a26-45ca-b8ef-c96e3928b659 · outbound

This paper cites Vidur: A large-scale simulation framework for llm inference,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Vidur: A large-scale simulation framework for llm inference,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.024548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.024548Z digest=sha256:24005005afa492f12a9eafaa3ede0a57bd319c48269e9065f4ab7b6571b7f820

Observation 9474e2ac-8847-448d-a2c5-33784756ba3a · outbound

This paper cites Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.136193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.136193Z digest=sha256:7bc77a8168ddf0cef7008a317fa0332726f1e8b71bcd56cce48f636c6b8db044

Observation e3a09dfd-6bd0-4c5c-b9c0-dad1f0f4702a · outbound

This paper cites Llmcompass: Enabling efficient hardware design for large language model inference,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Llmcompass: Enabling efficient hardware design for large language model inference,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.209163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.209163Z digest=sha256:93e2a8b4bd27862acbced39647313d584725bb20dd1ba835d3dc29bb8a83801f

Observation 8d956b91-494c-4617-83d1-4b357831ad48 · outbound

This paper cites Amali: An analytical model for accurately modeling llm inference on modern gpus,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Amali: An analytical model for accurately modeling llm inference on modern gpus,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.321060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.321060Z digest=sha256:6c59a37193bb2a4cdb3cf1d001c7a01bd5695c90b47852d9a286f6d8f95117aa

Observation 46db8890-3e1f-4963-9f18-6cc14c0f8fef · outbound

This paper cites Fairness in serving large language models,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Fairness in serving large language models,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.393342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.393342Z digest=sha256:fffaf9023ce5ce34430db08354e912399a2e5d727d547c69920606f023ad3c0f

Observation ef1c6471-f486-4f1d-a3b2-dc42413c0c53 · outbound

This paper cites Clean sharegpt dataset,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Clean sharegpt dataset,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.458412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.458412Z digest=sha256:9382b68f98f028b5c65ac7d9038aecda7360a06e42485087011a967f3183e187

Observation 6b314849-88b5-42a6-910a-381b7b4b23f0 · outbound

This paper cites Mistral 7B.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Mistral 7B

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.533081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.533081Z digest=sha256:ee4d58cc2a027bd0b481756e1bdf4a8f5b73fe2d91f43f6358c5742cb9d8ae3c

Observation a2a6e250-6339-4122-b15f-b4bac850f715 · outbound

This paper cites Granite Code Models: A Family of Open Foundation Models for Code Intelligence.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Granite Code Models: A Family of Open Foundation Models for Code Intelligence

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.615902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.615902Z digest=sha256:4bbab17417f619924350c6be91f53b607a2e60fa196b133838ac88285f5d9314

Observation ab58eaea-b3aa-4790-bff5-1bc003501ea7 · outbound

This paper cites Opt: Open pre-trained transformer language models,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Opt: Open pre-trained transformer language models,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.707629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.707629Z digest=sha256:a3eff42c3987bedb39a361684974cf3177bb59c578702aeb86ef09b087d62765

Observation 543d0be1-ba38-436c-9a08-fd300a83dc8a · outbound

This paper cites Qwen2 Technical Report.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Qwen2 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.849854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.849854Z digest=sha256:57f1f0a75958bb62fcd76851b17152d5d8464f7ac2b05aa11e51278c75c13dd2

Observation c589949a-a181-4c18-b42c-a44f50ab1084 · outbound

This paper cites Ai and memory wall,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Ai and memory wall,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.961186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.961186Z digest=sha256:e795014b293e0f3ab15f66c60799b711474314781717b06424b5f16b5952ce6e

Observation faa8d4de-c4d2-4dc7-919c-4814191e17ec · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving OPT: Open Pre-trained Transformer Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.775813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.775813Z digest=sha256:6227f0f81120a98b5cbd20a73a0fb9e626f786bc5c6c0246772518978d407ba2

Pith citing papers

No inbound Pith citation observations are available.