Pith. sign in

Paper Citation Record · LEDGER

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs

As of 10 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 0 inbound Pith citation observations for arXiv:2607.23115.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.23115 v1

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:40:12.723229Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

86 of 86 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved83
  • parse uncertain2
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e81fbd40-63ff-4fc2-bd55-05114131724b · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 1

Resolution
parse uncertain
no resolver link, observed 2026-08-01T03:40:10.999156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:10.999156Z digest=sha256:d6dce7c08075ee502a387be13c1623ef7b4a4c92c5d30647f73919e9cbbec677

Observation 4da32b21-7b4a-4b7e-b44f-f20dc4ec28b5 · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.005260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.005260Z digest=sha256:394a14e94dc25ee988edcbc0d8d077a1cdffeab68c3f9a7b5f2c9cbcf36759df

Observation a5039eec-7bfa-4d8e-85b5-96275f47ad23 · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.010282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.010282Z digest=sha256:f08afb5c6f5e3e4a685423b9f4fcc54e787359fb9fc2a6ea427de4a8d548768a

Observation ee08d559-9968-4f12-9b7c-842efcc8f3e2 · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.016717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.016717Z digest=sha256:388366ed6e2f4eea5a97347047a821c8556365078bd348ca122f38f6cadc3015

Observation 75dcedb9-39d9-49cb-8803-3983dd7e0a82 · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.021558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.021558Z digest=sha256:6c4a0be5453b19679f1a7b4c0bb6d63c5f34428cf9704b6bf738e5d59358e099

Observation acf32578-7480-4dd9-8ee9-a728a93cabdc · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.026430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.026430Z digest=sha256:6de59a1fd2351f93e5b3efd4bb58c5f6e4bd86dfb142d13fd1aab830a0cf5853

Observation 398ec4d2-4f66-46a8-8a2b-aebc9527be8d · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 7

Resolution
parse uncertain
no resolver link, observed 2026-08-01T03:40:11.031727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.031727Z digest=sha256:f62074ef5af92685a369aa1a5f1e783a524938400bcf636218825626357920d6

Observation d6d493f9-cf75-428d-9f6b-79d4780f6d6c · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.035939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.035939Z digest=sha256:5e8646a11c16311eb550e04cc7da3bfd81ef3dbc355a8cf1572f34d8a00c6cfd

Observation f91212a2-8ee2-4d32-b5d3-de5ef048aaf6 · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.040074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.040074Z digest=sha256:36412f9e719ec03bdc177564bca6fdf2e9ee3c33ead429d64af3fbf58fc12176

Observation faf139c7-f5ae-4003-a1d8-05b1a98d601e · outbound

This paper cites Scissionlite: Accelerating distributed deep learning with lightweight data compression for iiot.IEEE Transac- tions on Industrial Informatics, 20(10):11950–11960, 2024.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Scissionlite: Accelerating distributed deep learning with lightweight data compression for iiot.IEEE Transac- tions on Industrial Informatics, 20(10):11950–11960, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.044193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.044193Z digest=sha256:55abf059964ab8d27ae4ad0122149ef5fcc8e9d28ed6fc59b95f4743a3d817c8

Observation cfd4fa3d-553e-4145-a395-0831c5726fd3 · outbound

This paper cites Tooth: Toward optimal balance of video {QoE} and redundancy cost by {Fine- Grained}{FEC} in cloud gaming streaming.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Tooth: Toward optimal balance of video {QoE} and redundancy cost by {Fine- Grained}{FEC} in cloud gaming streaming

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.048531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.048531Z digest=sha256:e0a7301315ed781bf0dc4d624c5a6bf433f85df4dd8f38c3256c9adfe772bcf8

Observation a98f0dba-9b93-43cf-930d-3a588ffe6c00 · outbound

This paper cites Crux: Gpu-efficient communication scheduling for deep learning training.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Crux: Gpu-efficient communication scheduling for deep learning training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.053431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.053431Z digest=sha256:ed24513b5c892f509389c20bea90582d4d8174b39cc3703b6a7ca11c8590a1ea

Observation 7aba9b77-b011-46a8-995d-5a2d2eb23de6 · outbound

This paper cites Eva: Cost- efficient cloud-based cluster scheduling.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Eva: Cost- efficient cloud-based cluster scheduling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.057720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.057720Z digest=sha256:4b1160c841173ce75e7420f6f355ef43c75906853427fb44b3c3f5ccd44405cc

Observation 5af2baa5-b736-4b88-9b16-8515905d1d06 · outbound

This paper cites Kernel oper- ations on the gpu, with autodiff, without memory over- flows.Journal of Machine Learning Research, 22(74):1– 6, 2021.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Kernel oper- ations on the gpu, with autodiff, without memory over- flows.Journal of Machine Learning Research, 22(74):1– 6, 2021

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.061940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.061940Z digest=sha256:00d529dd5eb3c2d6bad9dffd79b6d768095477c9e79b2e9baa2e571a4ac72467

Observation 7bbd9110-e6f2-46af-85ee-526807ff3686 · outbound

This paper cites Remote procedure call as a managed system service.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Remote procedure call as a managed system service

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.066423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.066423Z digest=sha256:03107063d0dee54b18ef64b2dc9a708b6fec9b0ba99eb68733290e5752fccdd7

Observation 591726b1-51c1-47dc-b966-fdb425e9b7e1 · outbound

This paper cites Multiplexing dynamic deep learn- ing workloads with slo-awareness in gpu clusters.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Multiplexing dynamic deep learn- ing workloads with slo-awareness in gpu clusters

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.070517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.070517Z digest=sha256:fc3b3bbf65e95e0c4061149a5ae6157a99424682aa29da595574c6e42dee6044

Observation 45221d94-c5ea-4cb3-a447-c93037871049 · outbound

This paper cites {GRACE}:{Loss- Resilient}{Real-Time} video through neural codecs.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs {GRACE}:{Loss- Resilient}{Real-Time} video through neural codecs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.074609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.074609Z digest=sha256:8e7c3caff72bbe1b86807d330a61b3589cb3d1b82b900ca7a15ce3a41047009b

Observation 7379fdeb-82e1-4252-8bb5-d5414e62298a · outbound

This paper cites Flashattention: Fast and memory- efficient exact attention with io-awareness.Advances in neural information processing systems, 35:16344– 16359, 2022.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Flashattention: Fast and memory- efficient exact attention with io-awareness.Advances in neural information processing systems, 35:16344– 16359, 2022

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.078564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.078564Z digest=sha256:b52356ad524b0d72cbc906faef74495da78b1126b057955b050576afffc8cb66

Observation e4812f51-687c-4fc5-a585-03122aa6c47f · outbound

This paper cites Oneadapt: Fast adapta- tion for deep learning applications via backpropagation.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Oneadapt: Fast adapta- tion for deep learning applications via backpropagation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.083135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.083135Z digest=sha256:80d9df10dc837457c8e86dfaeaab08be3207ad5ab9e675e852ff9ca18d4c2475

Observation 85a1258d-f65f-47ea-86ae-05737e63e0c4 · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.087432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.087432Z digest=sha256:772aff593cd9a95498b42a31e7821088e10852b5ece4468fa3313415d701e039

Observation 6b7fc549-6faf-42cb-82e9-bc59ca50a6b4 · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.091621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.091621Z digest=sha256:07dffdd7ae78b912afa24e3f5724f7379b5e66d69dc882d14dd7babb4720fcb7

Observation 379c6fbb-fec8-47e1-9d20-3021728b2615 · outbound

This paper cites Dgsf: Disaggregated gpus for serverless functions.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Dgsf: Disaggregated gpus for serverless functions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.095749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.095749Z digest=sha256:9d03b3cf9798ffc06115a5853f1c4047eaf99b976a6deb95cd3fcce1adef8de1

Observation 818656c9-a51e-4238-b955-4067ca201da4 · outbound

This paper cites Rdma over ethernet for distributed training at meta scale.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Rdma over ethernet for distributed training at meta scale

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.099856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.099856Z digest=sha256:a48bc67433dd41b3f85f486f7d0869a9c2506fdac0151fe9cf68a97e45e15e97

Observation d63b93d0-c65c-4979-947e-2eec9e855e89 · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.104095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.104095Z digest=sha256:8f9887308f79e259a9291dfe864da207080ae48662eec11e04e50ef8cfbdb822

Observation ee927be3-9751-4415-8937-0d7a837b8199 · outbound

This paper cites A gpgpu transparent virtualization component for high performance computing clouds.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A gpgpu transparent virtualization component for high performance computing clouds

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.108698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.108698Z digest=sha256:43386764d9745f5010650ca74b80e58330271e06774369d089070fb8e0182697

Observation ab9ae254-bd5c-408d-9972-c3ec21c2066f · outbound

This paper cites Shin, Yibo Zhu, Myeongjae Jeon, Junjie Qian, Hongqiang Liu, and Chuanxiong Guo.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Shin, Yibo Zhu, Myeongjae Jeon, Junjie Qian, Hongqiang Liu, and Chuanxiong Guo

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.113329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.113329Z digest=sha256:963ed220969b0dfa97b268e3da8ed34d4ea240e14c6da83aa9a0a77f670ebea8

Observation 9e82ee4d-322c-478b-b98b-eddd16370dcd · outbound

This paper cites Kace: Kernel-aware colocation for effi- cient gpu spatial sharing.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Kace: Kernel-aware colocation for effi- cient gpu spatial sharing

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.120130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.120130Z digest=sha256:f5ee3d2eb868315c5104b07ccee2a2931e2c07c87b27427cd51d61d9c9af7176

Observation b4ca0d1e-f0ff-4e5f-a691-c506cc9959ad · outbound

This paper cites Microsecond-scale preemption for concurrent GPU-accelerated DNN inferences.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Microsecond-scale preemption for concurrent GPU-accelerated DNN inferences

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.125670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.125670Z digest=sha256:4a2d857fb063ab1900bb8b5d613637bd1676584cd9007f97e4f7b5f03c004f1a

Observation 834d32cb-73f6-4e79-96b7-042367f4f229 · outbound

This paper cites Multi-agent collaborative infer- ence via dnn decoupling: Intermediate feature compres- sion and edge learning.IEEE Transactions on Mobile Computing, 22(10):6041–6055, 2023.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Multi-agent collaborative infer- ence via dnn decoupling: Intermediate feature compres- sion and edge learning.IEEE Transactions on Mobile Computing, 22(10):6041–6055, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.130167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.130167Z digest=sha256:99f74a978a8bc84cf9cbfbe773a7b75f0593cdb26a1c2c11e323ee52f7fc6e51

Observation 6a4fff8d-2f02-44bf-a2e0-5806d25bd8c6 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.134612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.134612Z digest=sha256:ddf3a37fcd844f843daa24484ca6cd030d1f6d273c578e865106fed170dc57d5

Observation 6af4123f-65db-4204-9fe1-e3ee783c1dfa · outbound

This paper cites In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23), pages 87–101, 2023.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23), pages 87–101, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.138890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.138890Z digest=sha256:e82d6ce260256ef4cfb6fff07ff8426e0ed2870d3181d0c2fabb49d20696b1a4

Observation 895f8f53-7860-40f3-8e5f-32e8a2ff6f53 · outbound

This paper cites Charlie Hu, Xiaojun Lin, and Nan Deng.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Charlie Hu, Xiaojun Lin, and Nan Deng

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.143422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.143422Z digest=sha256:7cba61524d5397deaa013b413e42569cc8cb4b61cce0070c46b54626ea8ac69d

Observation a12eb9f5-097c-4c74-ade9-57c9a012db8c · outbound

This paper cites A house united within itself: Slo-awareness for on-premises containerized ml inference clusters via faro.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A house united within itself: Slo-awareness for on-premises containerized ml inference clusters via faro

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.147743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.147743Z digest=sha256:ac114193721bcddcea815c25a8a8f26a9598f90e5dc263ebfd39d4ba8e35da10

Observation 817605af-c388-4aec-bca1-2b9a568f9d27 · outbound

This paper cites Mor- ley Mao.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Mor- ley Mao

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.152064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.152064Z digest=sha256:97793497c0ceae3c2d21281bde4c0dadf0e36d75ecac4938be431dde168b3b69

Observation 6d512668-de1e-402f-84e4-697d9c7ef85a · outbound

This paper cites Deepum: Tensor migration and prefetching in unified memory.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Deepum: Tensor migration and prefetching in unified memory

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.156399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.156399Z digest=sha256:42be4542541adf4f5579ede68a4f51ffba6f9ee22e7f112fc5b3440cb15f03fb

Observation b6ab778c-3b3f-4b23-b486-80fd42e4dd23 · outbound

This paper cites A neural-network- based realization of in-network computation for the in- ternet of things.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A neural-network- based realization of in-network computation for the in- ternet of things

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.161088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.161088Z digest=sha256:5f174f88e90d330265d5c56a6036221cc17eaa771d95909dd120ae841e1518eb

Observation 5ee45b07-34d1-42ca-b61d-e5cee2fea889 · outbound

This paper cites {SuperServe}:{Fine-Grained} inference serving for unpredictable workloads.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs {SuperServe}:{Fine-Grained} inference serving for unpredictable workloads

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.165512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.165512Z digest=sha256:ca558f74b73321422e7bfeccc23819c0ea620fac0556896ddf3fbc2f08f4fc06

Observation 54e34d1c-f630-4ea1-b661-51f6b00f7317 · outbound

This paper cites A survey on in-network computing: Programmable data plane and technology specific applications.IEEE Communications Surveys & Tutorials, 25(1):701–761, 2023.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A survey on in-network computing: Programmable data plane and technology specific applications.IEEE Communications Surveys & Tutorials, 25(1):701–761, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.169742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.169742Z digest=sha256:4b276ecfd6f3c5895e5b30757495e51498ee4036fb968e3f3e43679bac1c5b94

Observation d438c829-33e9-49ed-9b5d-8fc735a80211 · outbound

This paper cites Navigator: Dynamic multi-kernel scheduling to improve gpu per- formance.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Navigator: Dynamic multi-kernel scheduling to improve gpu per- formance

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.173949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.173949Z digest=sha256:6f2a0c6c01a7ff844e028d1e45de59fdc8bc3b231a2ea64b96eef52b60c492a8

Observation 57aef14b-de47-4fa5-a3d9-552485dd5783 · outbound

This paper cites Efficient memory manage- ment for large language model serving with pagedatten- tion.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Efficient memory manage- ment for large language model serving with pagedatten- tion

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.178278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.178278Z digest=sha256:a0a1b912383418267093d77734e19feb7f340a28632258cdc00137663562e89c

Observation 2784287b-4418-4329-bb95-87bea9ede7f6 · outbound

This paper cites Forecasting gpu performance for deep learning train- ing and inference.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Forecasting gpu performance for deep learning train- ing and inference

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.182444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.182444Z digest=sha256:4abed5260214950397e2e9d8d4452f1cadda31a2c421e1a98af196fd1e36f48f

Observation 93854fcd-8bc9-495f-8ba3-0e35228838d0 · outbound

This paper cites A survey on large language model acceleration based on kv cache management, 2025.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A survey on large language model acceleration based on kv cache management, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.186651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.186651Z digest=sha256:9f3ae3db0382fe8ee7d7c3717e4d0ff9688a2d41234e2957a1c16a6e410477ef

Observation 59f34cb6-8c73-46c1-bac8-e7ddd46290ff · outbound

This paper cites THC: Accelerating distributed deep learning using ten- sor homomorphic compression.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs THC: Accelerating distributed deep learning using ten- sor homomorphic compression

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.190669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.190669Z digest=sha256:fffebfaca3edf1bc250ced91a06ca1d8c9cf12b36f638c7cb9f69652ea87b76b

Observation 8832f8da-a903-4fa7-a36e-7e876239de19 · outbound

This paper cites Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.195265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.195265Z digest=sha256:332f9f5f2b8bc12407cb6b50750df554b82035572e31966f1be63efe5993f1d4

Observation 4f757e3e-db2e-4b27-86fc-7c7bf30891d0 · outbound

This paper cites Incbricks: To- ward in-network computation with an in-network cache.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Incbricks: To- ward in-network computation with an in-network cache

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.199805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.199805Z digest=sha256:8bc0c3f4934a856265a0226e28ff3c85c92ff77133c2587baf1e036b61f5c217

Observation 9cc3512f-8635-4f4a-a61d-ab588b7d3118 · outbound

This paper cites Cachegen: Kv cache compression and streaming for fast large lan- guage model serving.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Cachegen: Kv cache compression and streaming for fast large lan- guage model serving

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.204037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.204037Z digest=sha256:07e813bbb7387363fb332bcb51a17f00ee8ac9a4133d27b56e3d7f3f517d1aa7

Observation 74a2a691-484f-422e-ad78-ae137fa5c5cf · outbound

This paper cites A convnet for the 2020s.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A convnet for the 2020s

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.209596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.209596Z digest=sha256:107b4d003dfa4c788f2206cd7de720bd80e9a323747d49e4b3ad67c440dcb0d4

Observation d225d4c4-314e-4207-a6cf-e764f6337ada · outbound

This paper cites A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.216782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.216782Z digest=sha256:147e17902b0713c5e08613d244082ca587ce1091605774717ebf5af5fc1bdbd6

Observation 5147f1f6-46fc-4e50-8ab4-02134dfaec9b · outbound

This paper cites A survey of storage systems in the rdma era.IEEE Transac- tions on Parallel and Distributed Systems, 33(12):4395– 4409, 2022.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A survey of storage systems in the rdma era.IEEE Transac- tions on Parallel and Distributed Systems, 33(12):4395– 4409, 2022

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.228099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.228099Z digest=sha256:f0ac4bc1ce1cb6ba00e612f473ed932c070cc9a1d43833a8f47336969cd741e8

Observation 803fc393-b9d1-4a27-9bda-af2fae79b699 · outbound

This paper cites Skyserve: Serving ai mod- els across regions and clouds with spot instances.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Skyserve: Serving ai mod- els across regions and clouds with spot instances

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.240104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.240104Z digest=sha256:db62d28043505a9a99ab34e4b382c419a71185baba8f7bc3db9fbd4469465189

Observation 3f7e2269-9f65-4a92-9f83-c07e307cbf42 · outbound

This paper cites Efficient scheduling policies for Microsecond-Scale tasks.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Efficient scheduling policies for Microsecond-Scale tasks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.252541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.252541Z digest=sha256:575f7d37ffe494c3bc7fe3f15f13c5d727938102c08ca50507e9233808bcb6fe

Observation 5bf9d1f6-3277-4920-99de-e1685840341c · outbound

This paper cites To- ward performance-portable petsc for gpu-based exascale systems.Parallel Computing, 108:102831, 2021.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs To- ward performance-portable petsc for gpu-based exascale systems.Parallel Computing, 108:102831, 2021

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.266736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.266736Z digest=sha256:d178a3efc61a700787f6fd1d9217cd691ff2a69554c77c2b887f4aa2d15e3560

Observation 5220bb3a-f75b-48e1-ad6c-dcef1b447bc0 · outbound

This paper cites Porting warpx to gpu- accelerated platforms.Parallel Computing, 108:102833, 2021.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Porting warpx to gpu- accelerated platforms.Parallel Computing, 108:102833, 2021

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.285770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.285770Z digest=sha256:aaaf2b080cf02448e3c4e221eee0573a0d55d92caf07090bba745b615686b2c9

Observation ba9cc23b-1df0-4c96-a5f0-166389d8f973 · outbound

This paper cites Jellyfish: Timely inference serving for dynamic edge networks.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Jellyfish: Timely inference serving for dynamic edge networks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.314644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.314644Z digest=sha256:aaf10d5677e1e4cfbff046cfb0be51c6c7731e89b779ee38a89ebbcd60b505f3

Observation 0bf2f227-ab26-4dc9-9b0a-8a3d888cce04 · outbound

This paper cites Bringing umap closer to the speed of light with gpu acceleration.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Bringing umap closer to the speed of light with gpu acceleration

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.330029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.330029Z digest=sha256:8d01af8ef54c6c59381177522f1f8a83c0bff28f8c91bff5efffea52f11ad848

Observation e333a6ea-f011-4074-9c60-244140e14361 · outbound

This paper cites Nvidia nvswitch: The world’s highest- bandwidth on-node switch.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Nvidia nvswitch: The world’s highest- bandwidth on-node switch

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.337088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.337088Z digest=sha256:c5dfb79d2c8b6cad44e6926095b575012b8dafbf8bf0f3021410bafde8cc6a44

Observation e97e016e-0e66-4ac1-a005-0ec3c96f368a · outbound

This paper cites Gemel: Model merging for memory-efficient,real-time video analytics at the edge.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Gemel: Model merging for memory-efficient,real-time video analytics at the edge

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.344509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.344509Z digest=sha256:e0324d1febf82bad63cf4876c86909a52b18e81c7e7614918a7e6c78fc434757

Observation ba9ad0f5-3866-4e5e-8ef5-184567811aed · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.352874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.352874Z digest=sha256:b6fce356c9404f828e65f96c61329a5c487fd7cafd24b785c179ebaafa85f7ea

Observation 5e56c7da-083f-4b4d-9639-22c3682c3071 · outbound

This paper cites {CASSINI}:{Network-Aware} job scheduling in machine learning clusters.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs {CASSINI}:{Network-Aware} job scheduling in machine learning clusters

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.364647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.364647Z digest=sha256:ee5610dc6808334f2c95644c9ea70299e6694f5aa79a67cba54cc5d9f041844c

Observation b84179bc-e235-4d4c-ba79-78842ad7de83 · outbound

This paper cites Sentinel: Efficient tensor migration and allocation on heterogeneous memory systems for deep learning.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Sentinel: Efficient tensor migration and allocation on heterogeneous memory systems for deep learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.372960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.372960Z digest=sha256:a92079fe6065e9d895e68044f8469e51716ab67f1ceb5e778a78f88cc4172e41

Observation 2900bced-7c98-4149-8e37-75d00f579e1d · outbound

This paper cites Enabling large dynamic neural net- work training with learning-based memory manage- ment.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Enabling large dynamic neural net- work training with learning-based memory manage- ment

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.379969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.379969Z digest=sha256:893f3f7e680aae2967f5af3ef3d1864502576c7c2df8d862441ad898407d7d66

Observation 748e008d-6c97-475a-b8f2-1ed5e89acc24 · outbound

This paper cites {Cloud-LoRa}: Enabling cloud radio access {LoRa} networks using reinforce- ment learning based {Bandwidth-Adaptive} compres- sion.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs {Cloud-LoRa}: Enabling cloud radio access {LoRa} networks using reinforce- ment learning based {Bandwidth-Adaptive} compres- sion

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.387513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.387513Z digest=sha256:f44ebb05d469f0f067d662f1c88d21f575869e811050a0a6ba35867be3eb37cd

Observation b275ef95-bf4f-448f-9b27-da8cd39c9bc1 · outbound

This paper cites Exploiting simultaneous communications to accelerate data parallel distributed deep learning.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Exploiting simultaneous communications to accelerate data parallel distributed deep learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.394196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.394196Z digest=sha256:83a7b00df52dcb751adb6fd071231df5a7acc294f2f9ef5ab4973d2c090e8f9c

Observation 693cf1f8-96ac-43b0-b349-607ef9f9dd10 · outbound

This paper cites Orion: Interference-aware, fine-grained gpu sharing for ml ap- plications.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Orion: Interference-aware, fine-grained gpu sharing for ml ap- plications

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.400877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.400877Z digest=sha256:f638aeb9b256829cf6cb307dc7f036707a04edecea48bc3aff6f59e8f9ceaedb

Observation 1e32dc8d-be47-41bc-89af-2f2e08ef8975 · outbound

This paper cites gremote: Cloud ren- dering on gpu resource pool based on api-forwarding.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs gremote: Cloud ren- dering on gpu resource pool based on api-forwarding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.416566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.416566Z digest=sha256:60d7640f53a6be88867f5819a6aaf0b2606a45c080322e3a49a1563f184a8b4b

Observation 188d3cf4-d525-4a75-9005-8c8caff3d0ab · outbound

This paper cites Lammps-a flexible simu- lation tool for particle-based materials modeling at the atomic, meso, and continuum scales.Computer physics communications, 271:108171, 2022.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Lammps-a flexible simu- lation tool for particle-based materials modeling at the atomic, meso, and continuum scales.Computer physics communications, 271:108171, 2022

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.440322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.440322Z digest=sha256:64479c9e261dbd783542b3ca6f5a00aa287df759c2844d0e2709f9601b46bd05

Observation d9075125-c546-4c9d-8edd-367ac0d72b9c · outbound

This paper cites Plssvm—parallel least squares support vector machine.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Plssvm—parallel least squares support vector machine

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.468036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.468036Z digest=sha256:b954c0e361b2823b849549d272b02067cb734ba9b2cba4d722030ef92290274d

Observation 5771c190-475e-4cae-95be-a852315c3cea · outbound

This paper cites Aqua: Network-accelerated memory offloading for llms in scale-up gpu domains.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Aqua: Network-accelerated memory offloading for llms in scale-up gpu domains

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.522658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.522658Z digest=sha256:413a923a67478f6bbd326281d4ecfe5f76b7d179908bb795046b1c661ee7c17f

Observation 9ecc8377-7758-414f-9506-64ac79e2ed9a · outbound

This paper cites Coflow scheduling for llm training.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Coflow scheduling for llm training

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.576555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.576555Z digest=sha256:2d640c405064031a8b3d63a46b26e4ecfbbdb8908f65ce8af9bb29169d43bd60

Observation 158cf9d6-e008-4c7c-a59e-73473353b2b2 · outbound

This paper cites Characterizing Network Requirements for GPU API Remoting in AI Applications.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Characterizing Network Requirements for GPU API Remoting in AI Applications

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.630913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.630913Z digest=sha256:cc3492e25638d88283d4b18ac7e4ad53382bb9c92536f83f4794001aca733885

Observation b7884337-5363-4d05-b8a1-1ae263e751ab · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.675274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.675274Z digest=sha256:aef2a9985d9807aaeb1c2bc8ea19d13c035504c7a966d6ebddfb55093ce28803

Observation b25a7077-5fa2-4b4e-b832-c3e0c4d9b323 · outbound

This paper cites Beware of fragmentation: Scheduling {GPU-Sharing} workloads with fragmentation gradient descent.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Beware of fragmentation: Scheduling {GPU-Sharing} workloads with fragmentation gradient descent

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.755151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.755151Z digest=sha256:f1ed21c2f8cd773cbdd2c9d172b17ad54e56cf70d4e738267ed9896d776d439b

Observation ea9fd065-ddb7-48c9-a4c8-7b2ef3390bcb · outbound

This paper cites Transparent {GPU} sharing in container clouds for deep learning workloads.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Transparent {GPU} sharing in container clouds for deep learning workloads

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.798443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.798443Z digest=sha256:dba673433ca99fc5f01f9d0c7bc8f53935a7fb9aab2622a27f34f22ab4c4b672

Observation e46584d6-1c14-4058-bc99-72391d454c51 · outbound

This paper cites Yan, and Junchen Jiang.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Yan, and Junchen Jiang

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.872455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.872455Z digest=sha256:b87beb91e752bdfa9cd745958e0759b7a2dbc7e80516718143d0076ec17f792d

Observation 04c59993-3ce9-4598-a883-0405c0d8b2a1 · outbound

This paper cites Efficient tensor offloading for large deep-learning model training based on compute express link.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Efficient tensor offloading for large deep-learning model training based on compute express link

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.914851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.914851Z digest=sha256:a4349a830a4ded24f534944b7c981641ea1a479f47fdeffa3725e0a610cda115

Observation f0699efe-20a1-4124-a9e0-7c8f547898d5 · outbound

This paper cites Infless: a native serverless system for low-latency, high- throughput inference.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Infless: a native serverless system for low-latency, high- throughput inference

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.019245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.019245Z digest=sha256:38cf206395fd8f4b0298b846ab71678671695f1792f3a047e5ae68bf3d3044f6

Observation e87e1bfe-f4ef-4c12-85e7-b6ea4ff2bef8 · outbound

This paper cites Deep compressive offloading: Speeding up neural net- work inference by trading edge computation for network latency.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Deep compressive offloading: Speeding up neural net- work inference by trading edge computation for network latency

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.078779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.078779Z digest=sha256:ef0488b70a2fa8a9194f1a8a503e402bd8a3b9d3737506917b72938127446d07

Observation d9dce3a7-97b3-4d3a-8f78-165427374f2d · outbound

This paper cites Horus: granular in-network task sched- uler for cloud datacenters.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Horus: granular in-network task sched- uler for cloud datacenters

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.186797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.186797Z digest=sha256:a83dbac6128bb2d65468b4a0104e0c01301470105899182822911e19d160f881

Observation 0525d198-f573-46b9-bfc1-ec895eb25d05 · outbound

This paper cites Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.308277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.308277Z digest=sha256:7da1b843fe88d84f85acf68805596cf8887d5b3f443cd02af021d1ff798af1c8

Observation 32543cfe-c61d-4c10-80df-6a54ead58c8c · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.369125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.369125Z digest=sha256:a141e4a9662dc24e1dae42224efc76cfe1d582e0deeb062dbc6651cb8c379062

Observation 435783a6-c85b-4a06-95aa-502e3d4c7815 · outbound

This paper cites Expel: Llm agents are ex- periential learners.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Expel: Llm agents are ex- periential learners

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.439101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.439101Z digest=sha256:6b46d2d3995d616e62a150a4c2831052c68b3110c9b10ccc728c996d82631753

Observation 9ba649ac-fb75-41ad-98e5-0ec9e7e9ca01 · outbound

This paper cites Efficient {Direct-Connect} topologies for collective communications.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Efficient {Direct-Connect} topologies for collective communications

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.509184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.509184Z digest=sha256:01f640d615cee1c93c8db4478ae22a746207c6f070b6d25a6781141fde9a727a

Observation da20ac6b-e61e-4283-b465-2ecc048e4ac1 · outbound

This paper cites Tally: Non-intrusive performance isolation for concur- rent deep learning workloads.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Tally: Non-intrusive performance isolation for concur- rent deep learning workloads

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.594996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.594996Z digest=sha256:896856851d40ba6dfd6ca4ffb47db9dc371295a61908dd33adec34e1bba7a526

Observation 49ea237c-f11d-4484-a700-a0a96ebfd31c · outbound

This paper cites Sglang: Efficient execution of structured language model pro- grams.Advances in neural information processing sys- tems, 37:62557–62583, 2024.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Sglang: Efficient execution of structured language model pro- grams.Advances in neural information processing sys- tems, 37:62557–62583, 2024

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.665111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.665111Z digest=sha256:56f94ea4ce6ea58cf385fd0084698d27f09829758d0a9b8198611c54864c955c

Observation ab65e9fa-816f-4f3a-afe8-2b559f56f520 · outbound

This paper cites Kernelet: High- throughput gpu kernel executions with dynamic slicing and scheduling.IEEE Transactions on Parallel and Distributed Systems, 25(6):1522–1532, 2014.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Kernelet: High- throughput gpu kernel executions with dynamic slicing and scheduling.IEEE Transactions on Parallel and Distributed Systems, 25(6):1522–1532, 2014

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.683698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.683698Z digest=sha256:83734b1674f4f9b374833f284d5f00a61d2a73732ed1e4fb4e26f16d6dd3b3ae

Observation dbee7ac1-5724-4102-8cf5-d0e06cce693b · outbound

This paper cites Megascale-infer: Efficient mixture-of-experts model serving with disag- gregated expert parallelism.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Megascale-infer: Efficient mixture-of-experts model serving with disag- gregated expert parallelism

Reference 86

Resolution
malformed identifier
no resolver link, observed 2026-08-01T03:40:12.723229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.723229Z digest=sha256:7065da8b1ffc50f104a6fee0b030f65ce8859e0ad57e008c472660d9433daa29

Pith citing papers

No inbound Pith citation observations are available.