Pith. sign in

Paper Citation Record · LEDGER

Topology-aware Preemptive Scheduling for Co-located LLM Workloads

As of 19 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2411.11560.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11560 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:28:01.641782Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f64c548-f65d-4eba-b4dd-645c130c0fbc · outbound

This paper cites A survey on large language models: Applications, challenges, limitations, and practical usage.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads A survey on large language models: Applications, challenges, limitations, and practical usage

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T18:28:01.428717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:28:01.428717Z digest=sha256:23cf97ca1f5ee02d42c6ce275aa66ac0b4da8e043c073f8795deb28e36f6ece7

Observation b3303931-8436-4e1c-b87e-27fcafcf2e7a · outbound

This paper cites Efficient Training of Large Language Models on Distributed Infrastructures: A Survey.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Efficient Training of Large Language Models on Distributed Infrastructures: A Survey

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T18:28:01.433698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:28:01.433698Z digest=sha256:793a895e412839bff6de8a3d8cefd8b70b055a8c6790fb42a98f62cb6d6c2b29

Observation decb8d3c-e512-4ab7-a628-a57eb2b26359 · outbound

This paper cites BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T18:28:01.438721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:28:01.438721Z digest=sha256:243afe39b665eac4cc9218eec4b238509b93909eedba660a0e3f940198fcf365

Observation 1838389c-2088-4752-ae94-a643ce2ca9f1 · outbound

This paper cites Gödel: Unified large-scale resource management and scheduling at bytedance.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Gödel: Unified large-scale resource management and scheduling at bytedance

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:02.306668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.443368Z digest=sha256:e814cb4a13e5647a244eac026b5004b35a6e575bfdf00420af3340b26c3758be

Observation 686d5a79-898d-46f3-9724-e2a7a3495913 · outbound

This paper cites Topology-aware gpu scheduling for learning workloads in cloud environments.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Topology-aware gpu scheduling for learning workloads in cloud environments

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:02.291237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.447617Z digest=sha256:743237015e2df4dd03d7ddc12e501814329ecd2e1bab6e43a35307a330ca78d3

Observation cb7af90d-5efa-4522-b89b-27aadcf7a300 · outbound

This paper cites Numa (non-uniform memory access): An overview: Numa becomes more common because memory controllers get close to execution units on microprocessors.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Numa (non-uniform memory access): An overview: Numa becomes more common because memory controllers get close to execution units on microprocessors

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:02.275245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.452715Z digest=sha256:a368b327ae167b01a118d65a53fbaea2bc5201b83248454667449996ab9c6590

Observation c2a87714-bfe2-4b13-bb56-e6dce973d287 · outbound

This paper cites M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T18:28:01.457638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:28:01.457638Z digest=sha256:4b60d9e6571c4f9c04fef8c54e3f240f8c19706e3d698cda19dd4666762eee62

Observation b274c18f-8aa3-4049-8c3b-ef85184042f6 · outbound

This paper cites Fastertransformer.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Fastertransformer

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:02.261255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.462536Z digest=sha256:954331719d402bec221a07acf3af3944033fff12979bda9cc627785feb6d267e

Observation d72f1f46-7198-4ed8-b7e6-929dac152349 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Efficient memory management for large language model serving with pagedattention

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T18:28:01.466924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:28:01.466924Z digest=sha256:e74808e2476544f2ae2e1d696792c4e9ee1eb2a5cede81f45a80e90d5c686d69

Observation 8059a009-b02e-4a5e-90ab-5b5462e55ffc · outbound

This paper cites an unresolved cited work.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:28:02.236920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.471222Z digest=sha256:938b59b727fd7fadf4dd46eacafd4d36f0feb689bcf97dd2a0334aeef7830eb1

Observation 55bbb4f4-33c7-4400-9a72-8dedbf17fc3d · outbound

This paper cites Huggingface text generation inference.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Huggingface text generation inference

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:02.223059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.476060Z digest=sha256:0aab42c3f64a7afd79f99dd57d9c538a2d27e0d595f8d6d265bb113bf391cb43

Observation 8cf4ed1a-5c32-4f90-8fc4-6828229232e4 · outbound

This paper cites Deepspeed inference.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Deepspeed inference

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:02.206174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.481104Z digest=sha256:4588cd2e1d973a65888a9b48eb16b4bab1eaabe25114e101194a00e54492370b

Observation 24929857-98da-480d-bb69-5d8d84af3978 · outbound

This paper cites Tensorrt-llm.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Tensorrt-llm

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:02.187894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.486113Z digest=sha256:964d76a57fe9be158595f4c63e48a35e6caaf1a2169ab082afc87337a3556d1c

Observation c33301c8-4278-4565-8476-81bae1b4adad · outbound

This paper cites Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T18:28:01.491752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:28:01.491752Z digest=sha256:19cfa75bc8ec37909862c85625a5cd96e6d456b8807a3a12407b51e596120961

Observation 35e1fa7c-29a6-48ae-8bf0-d7501de9fdf6 · outbound

This paper cites Kubernetes topology manager moves to beta.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Kubernetes topology manager moves to beta

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:02.170281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.497305Z digest=sha256:7b9b6fe3f81ec61ff3c5ffb524c0e8d6745db54dd9ef3aa975f39d9f05d6eaa5

Observation 4a4273a1-70b4-4d0e-a9fb-d136a6eafbe2 · outbound

This paper cites Pod priority and preemption.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Pod priority and preemption

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:02.151560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.502510Z digest=sha256:267cf597c9aa78e53c60ba9d97a404fb96af88f304dba0ebb3ed6d68bb38c5af

Observation 56641ff5-55ad-4d69-a0ed-47f142ee90b4 · outbound

This paper cites Godel scheduler: a unified scheduler for online and offline tasks.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Godel scheduler: a unified scheduler for online and offline tasks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:02.132252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.507707Z digest=sha256:a4d74075c70ed0838eff8f1cb6d7f8c1f9683d8db463a087255c5fba25c2a0e2

Observation 65743995-cd8e-4356-bc95-b0d38e2e1028 · outbound

This paper cites Daemonset.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Daemonset

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:02.112382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.513084Z digest=sha256:bcc57a4e138f12f0a8937349997a2e18e337672dcaa02e7a146eefb94b5637da

Observation 38571f2a-df86-4258-95d3-9677b53151bd · outbound

This paper cites Kubernetes without kubelet.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Kubernetes without kubelet

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:02.092846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.518354Z digest=sha256:ae4be05f243ab85f97300c197976d6218ccee700991bb2f47e36cfad2d6e7793

Observation 60eb0db1-4c97-4a32-ab61-2449af02ad9f · outbound

This paper cites Control topology management policies on a node.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Control topology management policies on a node

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:02.076029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.523308Z digest=sha256:bb0fd7592da0c4b0a1d1e54b0fc47ea81776c4508efa12022d2e0a1f1b706455

Observation 88a88c86-aed8-4e04-8cf0-847c994707ee · outbound

This paper cites Towards {GPU} utilization prediction for cloud deep learning.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Towards {GPU} utilization prediction for cloud deep learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:02.060218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.528634Z digest=sha256:86338b055f10faa551f9a392becb2489cf9e392c3db39d69d9f64276526e071f

Observation b1a8bceb-c2a4-40f6-8530-0d45e9a5ee3f · outbound

This paper cites Horus: Interference-aware and prediction-based scheduling in deep learning systems.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Horus: Interference-aware and prediction-based scheduling in deep learning systems

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:02.046202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.533781Z digest=sha256:06170f35e639cd54c3c633a979c3dc47cd6db2f558e746a3bb83d3b3d25d13ff

Observation 925f7790-266f-4f84-b56e-a683690cf4c0 · outbound

This paper cites Beware of fragmentation: Scheduling {GPU-Sharing} workloads with fragmentation gradient descent.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Beware of fragmentation: Scheduling {GPU-Sharing} workloads with fragmentation gradient descent

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:02.031649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.539007Z digest=sha256:24522b6cfd9062d5cb9452594aed8757e8f83161c017598bc9dbf629a6b15713

Observation b9e132b9-40af-44c1-bf05-587c08e6e83b · outbound

This paper cites {HiveD}: Sharing a {GPU} cluster for deep learning with guarantees.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads {HiveD}: Sharing a {GPU} cluster for deep learning with guarantees

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:02.016593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.543941Z digest=sha256:1ba865beea678fce43affa4d8c3d40dd3d4384d3ab9f8890497a64d6890e2f9a

Observation a71a4c0e-220e-4f0c-b4f4-ff8a56479a6f · outbound

This paper cites Supporting gpu sharing in cloud environments with a transparent runtime consolidation framework.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Supporting gpu sharing in cloud environments with a transparent runtime consolidation framework

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:02.001582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.549509Z digest=sha256:8ede92d073c9e91028c9e7e7049d8517f043be6517d2723c20ab17ec1f17b8b1

Observation 748ddaea-6be4-49c5-ac6c-c15e16774fd0 · outbound

This paper cites Fine-grained gpu sharing primitives for deep learning applications.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Fine-grained gpu sharing primitives for deep learning applications

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:01.987696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.554580Z digest=sha256:ebf7cdab8569ac81f8132788a8dbc51193722d7b94813989ed2228dba8d79c36

Observation 8223a453-87fd-43c1-93e8-110117107960 · outbound

This paper cites Gpushare: Fair-sharing middleware for gpu clouds.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Gpushare: Fair-sharing middleware for gpu clouds

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:01.973274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.560039Z digest=sha256:9b3ad341b48218a421adbc2764b3fcb197730a6de658123ddb635cf7f9ed76f9

Observation 4506a771-c34a-4ba0-9258-99d951452b22 · outbound

This paper cites Nvidia cloud native technologies: Gpu sharing.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Nvidia cloud native technologies: Gpu sharing

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:01.959345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.565070Z digest=sha256:b68f506d452187746e323b246909968100f3f15159e83d145a67cd9ecc979cd9

Observation 4c09c855-110c-49b5-b9a1-3475a4e1f3c7 · outbound

This paper cites Advanced features in ibm power8 systems.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Advanced features in ibm power8 systems

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:01.945397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.570281Z digest=sha256:3182a399146a068f4158b90c8e178709c80627beadd5dce851f2c2273f17cd9a

Observation 1fb6f2b0-c3d3-45af-a511-191befec4d7f · outbound

This paper cites Performance evaluation of the nvidia tesla p100: Our directive-based partitioning and pipelining vs.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Performance evaluation of the nvidia tesla p100: Our directive-based partitioning and pipelining vs

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:01.931958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.575734Z digest=sha256:d7f8f77f0d0bb2ebbe90a6ab1640ccde1921beeb2a9f79e3a60e174693bfe3d8

Observation deeb195f-5236-4900-8171-18ad8d71598f · outbound

This paper cites Topology-aware scheduling framework for microservice applications in cloud.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Topology-aware scheduling framework for microservice applications in cloud

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:01.917790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.580861Z digest=sha256:c8f4d831cee2af1794df215df2564152d7dd5cc794a57b06347279e715108351

Observation dbf852d4-d6ba-44a7-a2ef-b31305b0d091 · outbound

This paper cites Katalyst core.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Katalyst core

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:01.901739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.586803Z digest=sha256:bbf6ff5420164a9dea546810954ed0b5318653247b22fc2f9ed88d5175e292c9

Observation 93263d25-13c5-437c-9a69-5016325fcf88 · outbound

This paper cites Topology-aware resource allocation for data-intensive workloads.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Topology-aware resource allocation for data-intensive workloads

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:01.884085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.592516Z digest=sha256:26520729e2375f6a4ee46f1370e714938c32a41128d10fc437087b79a603469c

Observation 6f4c2c72-2123-4979-afa7-c94df6b035d9 · outbound

This paper cites Towards topology aware pre-emptive job scheduling with deep reinforcement learning.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Towards topology aware pre-emptive job scheduling with deep reinforcement learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:01.865992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.597592Z digest=sha256:abaf45d72f16649112bf8986123feaf2204062cee3a2a012e10451a670e8fee0

Observation 792f1e0a-0f86-4e97-a7ec-91a733a6150d · outbound

This paper cites Microsecond-scale preemption for concurrent {GPU- accelerated}{DNN} inferences.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Microsecond-scale preemption for concurrent {GPU- accelerated}{DNN} inferences

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:01.849163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.603019Z digest=sha256:d404b0156419064cf2083a687d839e20f727fec3bf2aacbd7fd3c570b2a93dbd

Observation cc0d2b9d-a5d2-457a-9492-ffaf42566c25 · outbound

This paper cites Efficiently programming large language models using sglang.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Efficiently programming large language models using sglang

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:01.830860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.608437Z digest=sha256:6287f9877e4fa3f4db9544f407d13d2dcaf7b400248b787f07fdc67a6cc89805

Observation d548ec38-5922-4dd3-826e-3af1d996ebbd · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based} generative models.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Orca: A distributed serving system for {Transformer-Based} generative models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T18:28:01.614459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:28:01.614459Z digest=sha256:3e80bbc4de80ebb601e72221d4096e9bcff3265bff36ee8498b107afaf77f41c

Observation 9ff1a8a3-098a-48de-9bd6-1c601afba6c4 · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Fast Distributed Inference Serving for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T18:28:01.620434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:28:01.620434Z digest=sha256:03f8d8661f0acef5b66212adaadefb9649a4d4cda37b759d09aa999a722e9ccb

Observation 799a9f9b-e866-443b-a8f8-1e266cf3c4c9 · outbound

This paper cites Bert loses patience: Fast and robust inference with early exit.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Bert loses patience: Fast and robust inference with early exit

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:01.801769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.625813Z digest=sha256:682b0fa63c31b177af81cdd1ae366dcf988b6884c6d7f42f5d2458656846741d

Observation 2ef5b5fb-8604-47f4-9f99-e25ed9042c7d · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T18:28:01.631009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:28:01.631009Z digest=sha256:b8b452c9196c1926b5a392fa22e501060af32ecc259d6ed10f7246a5f88488fe

Observation a7a7f4e6-21d2-45c6-bc1d-78087adeda08 · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Sparsegpt: Massive language models can be accurately pruned in one-shot

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T18:28:01.636569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:28:01.636569Z digest=sha256:2038038eb7f16c47aa22bd7326bb6e4a3a7a18f79436d0f578541f06c2f557b8

Observation 563b8ecf-5b3a-47ba-af67-92051c370136 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads Awq: Activation-aware weight quantization for on-device llm compression and acceleration

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:28:01.763314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:28:01.641782Z digest=sha256:65dec574eedaf3c294bae5ca73b6fa7c14aac5b2f6547de3872464cb6223bc5c

Pith citing papers

No inbound Pith citation observations are available.