Pith. sign in

Paper Citation Record · LEDGER

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models

As of 9 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2607.13093.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13093 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T06:42:19.865695Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 704996cd-c2f2-4eff-93fe-991f9522c05a · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.125500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.125500Z digest=sha256:d13b792c502996ecd3ed9cba9dd5c73181adfc713d9194cf920b1ada7ebf844f

Observation a244b28b-a60f-4a5a-8424-34b6eba53ff3 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.165524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.165524Z digest=sha256:05b7b8ca99b0c69edcb78f00b7c89d2438f0c57dc0f3ece2306018eaade26de2

Observation dbf50e4c-e461-40a9-8de9-8898dd6a5caa · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.219713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.219713Z digest=sha256:a780d2711e088de0209750233dc5f7967c7df29c8bea32b48094b5a6e678a9c8

Observation 59aafbdb-ea91-4138-abbc-039fc7e78970 · outbound

This paper cites ORCA: A distributed serving system for transformer-based generative models.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models ORCA: A distributed serving system for transformer-based generative models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.254849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.254849Z digest=sha256:ef0419c1c5f7cf4aa94ec086dbdced08dec2c46adbcd2f3a70ddd4a4bc76534c

Observation 9ec4aa2a-05b0-4e4d-9fa1-fbca4f3b12e6 · outbound

This paper cites DeepSpeed-Inference: Enabling efficient inference of transformer models at unprecedented scale.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models DeepSpeed-Inference: Enabling efficient inference of transformer models at unprecedented scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.308814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.308814Z digest=sha256:5397f4daadb35733aba84aaae1b43bd5d4d615438c2bc9ca7294b532935f0b76

Observation 7a3237f2-631f-4397-abe2-2bd07717c460 · outbound

This paper cites Neurosurgeon: Collaborative intelligence between the cloud and mobile edge.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Neurosurgeon: Collaborative intelligence between the cloud and mobile edge

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.386010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.386010Z digest=sha256:43fa00ea7a7e4c90d67097d61431ca850d5ad73861bf91647923c2c6a388b72d

Observation a2c8ebfc-c67a-4f6c-8050-c791406b6084 · outbound

This paper cites Split computing and early exiting for deep learning applications: Survey and research challenges.ACM Computing Surveys, 2022.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Split computing and early exiting for deep learning applications: Survey and research challenges.ACM Computing Surveys, 2022

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.418410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.418410Z digest=sha256:ebd406ab94cfea05b08f8827484db98a75d8d9d87d19a7550d384ab0f8a2d187

Observation be4ee5b6-3c18-4bc6-ace7-ce68744ba00a · outbound

This paper cites DNN surgery: Accelerating DNN inference on the edge through layer partitioning.IEEE Transactions on Parallel and Distributed Systems, 2023.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models DNN surgery: Accelerating DNN inference on the edge through layer partitioning.IEEE Transactions on Parallel and Distributed Systems, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.469966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.469966Z digest=sha256:5d6297592e4bc0ffdddd4fba984aea8952f6821e0549c760536b04bba3bacb8e

Observation 5a7f8188-0c65-4f6f-be70-bde979f3536d · outbound

This paper cites Pipeedge: Pipeline parallelism for large-scale model inference on heterogeneous edge devices.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Pipeedge: Pipeline parallelism for large-scale model inference on heterogeneous edge devices

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.527834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.527834Z digest=sha256:9db94268f0c01e51843c030893049d839664ffea590e2a6829feab63739e1fd8

Observation a3e7d1d7-cc21-4428-a652-4b0f664a237c · outbound

This paper cites Hybrid SLM and LLM for edge-cloud collaborative inference.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Hybrid SLM and LLM for edge-cloud collaborative inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.597558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.597558Z digest=sha256:a236872515e5ef3566ff44c8c192f624161b20f9c9c394521bb188e4c523c685

Observation 33614f3b-2045-4f8a-b6b4-4898256a8cef · outbound

This paper cites Jupiter: Fast and resource-efficient collaborative inference of generative LLMs on edge devices.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Jupiter: Fast and resource-efficient collaborative inference of generative LLMs on edge devices

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.654481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.654481Z digest=sha256:fc12dde9d705cf82db3ef4dcc908431f344d5648aa2794598a427c0af4edb066

Observation 50603c84-6301-4705-b54b-b447c97d11af · outbound

This paper cites DSSD: Efficient Edge-Device LLM Deployment and Collaborative Inference via Distributed Split Speculative Decoding.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models DSSD: Efficient Edge-Device LLM Deployment and Collaborative Inference via Distributed Split Speculative Decoding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.700801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.700801Z digest=sha256:056d6600fddf61bed2f9f2f05c06c911ef48ca1aeb8c716debebd25314ad08c3

Observation cbd50d06-12a4-4179-8465-1e9474670daa · outbound

This paper cites CE-CoLLM: Efficient and adaptive large language models through cloud-edge collaboration.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models CE-CoLLM: Efficient and adaptive large language models through cloud-edge collaboration

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.760422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.760422Z digest=sha256:6e2940427e05df095eccf284bf532524c4a68c4f8f104574711558789f833d7c

Observation c13880c8-dbdb-41ce-bce7-e66b06bc695c · outbound

This paper cites Edgeshard: Efficient LLM inference via collaborative edge computing.IEEE Internet of Things Journal, 2024.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Edgeshard: Efficient LLM inference via collaborative edge computing.IEEE Internet of Things Journal, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.840945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.840945Z digest=sha256:b1f14a5df8c6c0ae4df736467857a595ea8266eebe6fb44991fc8b1d4c5c4157

Observation a7e64c04-0db6-4448-8b9c-9944903863d4 · outbound

This paper cites Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.915097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.915097Z digest=sha256:5484afaf27066c38e826e7fedf4898d82051b4f5c0add9c0dcf09f55aefbc1c0

Observation 2094fa64-e73e-443d-b654-c8744521eb64 · outbound

This paper cites Fast inference from transformers via speculative decoding.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Fast inference from transformers via speculative decoding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.997029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.997029Z digest=sha256:2b66ec3034a99df022d1f46f316059595bfc044933d5dc83e0785028551a26c1

Observation 8dbfbe5d-4166-42a5-a20c-708869b7b6ef · outbound

This paper cites Accelerating large language model decoding with speculative sampling, 2023.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Accelerating large language model decoding with speculative sampling, 2023

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.090265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.090265Z digest=sha256:07270d8ee63b759b2d70d4387a8f70d31925ac583f492ad2e464544c1eb27163

Observation e31208b4-04d7-439c-be8e-37d33de933a9 · outbound

This paper cites Specinfer: Accelerating generative large language model serving with tree-based speculative inference and verification.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Specinfer: Accelerating generative large language model serving with tree-based speculative inference and verification

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.143946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.143946Z digest=sha256:16dbbbc7b69af74a5e969ed904005738cf221d54c8fa3aab08ead30e5ae1cd0c

Observation b0df2741-fa12-4b04-ab32-22b12618180a · outbound

This paper cites Sequoia: Scalable, robust, and hardware-aware speculative decoding.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Sequoia: Scalable, robust, and hardware-aware speculative decoding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.237006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.237006Z digest=sha256:20b14e2d3453f7af69ed0e1880984b549d0e34bdfe24c708482cde0ae4596f47

Observation bba04103-6a0d-43b9-b4c3-73db6ab2f825 · outbound

This paper cites Lee, Deming Chen, and Tri Dao.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Lee, Deming Chen, and Tri Dao

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.304835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.304835Z digest=sha256:9362ec92867ba5becac59d3e151f139e80e16425022c2896184e4c5d35d79331

Observation b204fa25-a9b6-4561-ba1f-334675d9b30d · outbound

This paper cites EAGLE-2: Faster inference of language models with dynamic draft trees.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models EAGLE-2: Faster inference of language models with dynamic draft trees

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.363212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.363212Z digest=sha256:447c65407fb0a04f31d80927b4e0ad490b7503a47271909d087a5e76d0626bae

Observation 457a302d-9d3d-4119-ba8a-fa8b5a00785f · outbound

This paper cites Break the sequential dependency of LLM inference using lookahead decoding.arXiv preprint, 2024.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Break the sequential dependency of LLM inference using lookahead decoding.arXiv preprint, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.391598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.391598Z digest=sha256:a59a3b41aaafc83bfb84428a006e03c09f4d68346225a6b616942fc5896815f4

Observation 20f5b314-7564-4fa3-a578-c7ab3c0859ad · outbound

This paper cites Reddi, and Felix Yu.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Reddi, and Felix Yu

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.475568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.475568Z digest=sha256:f8eba2f2d97f09d7919bd47b53e3d3445c96aa41c87434c1e2b1a1afcca5f3e6

Observation a752c88b-db56-45de-bc72-a3f4fe9e945e · outbound

This paper cites Layerskip: Enabling early exit inference and self-speculative decoding.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Layerskip: Enabling early exit inference and self-speculative decoding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.542612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.542612Z digest=sha256:bd63b9ccf7984224842e1adcbf049e95adf7b2ecc9b1aa40a671660a419bec2e

Observation 7c4cab50-7dea-4036-aa84-f11339118f41 · outbound

This paper cites Draft & verify: Lossless large language model acceleration via self-speculative decoding.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Draft & verify: Lossless large language model acceleration via self-speculative decoding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.615010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.615010Z digest=sha256:677acdad9876197aa1c2fc9839d71b9544711d88c2585edc51af5d8d56f2a5ad

Observation 23008ec9-d20b-479b-b1dc-63915ef94f0b · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.702536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.702536Z digest=sha256:d41b1e1cffdef82a99e2e53816d2969713b1493558e7b3fee334c78b26d4af1a

Observation f66487df-1cd8-46ad-9931-9f26f325fdfb · outbound

This paper cites QLoRA: Efficient finetuning of quantized LLMs.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models QLoRA: Efficient finetuning of quantized LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.739903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.739903Z digest=sha256:832e8031e0f92309a2a2b1bf6a4b01ef6c99d5f8bbd5c2069627981cac9d5ff2

Observation 8e950cda-cbcb-45d0-8480-7af5fa2c85d0 · outbound

This paper cites SplitLoRA: A Split Parameter-Efficient Fine-Tuning Framework for Large Language Models.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models SplitLoRA: A Split Parameter-Efficient Fine-Tuning Framework for Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.876672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.876672Z digest=sha256:1a93b7664abf598c328733cdbcfefd8dcaa1cb06a5345607a4fa68650ef50fa0

Observation 4798fc88-828b-416f-ae75-d69c30355b48 · outbound

This paper cites Improving LoRA in Privacy-preserving Federated Learning.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Improving LoRA in Privacy-preserving Federated Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:17.098363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:17.098363Z digest=sha256:b79307dfaaa7aaead72fe246f7741d51b838bc1d3acc36568c1c4f8a5f73f36a

Observation 9b9bd033-a26f-4707-8e85-b49ebb328fd1 · outbound

This paper cites On the implicit relation between low-rank adaptation and differential privacy.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models On the implicit relation between low-rank adaptation and differential privacy

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:17.249674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:17.249674Z digest=sha256:2ac065fd561b7b24c309a3d61c998cc53178d61fbce24ff7caccdea6fa3b4679

Observation 87219b76-0a36-4fad-a0be-3076474b4077 · outbound

This paper cites Is split learning privacy-preserving for fine-tuning large language models?IEEE Transactions on Big Data, 2024.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Is split learning privacy-preserving for fine-tuning large language models?IEEE Transactions on Big Data, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:17.384853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:17.384853Z digest=sha256:d0e1c2bcf41d23a4b289719bb20765cbcd60460f0880192550b380908e275dd9

Observation 8a2ee5c0-71e1-41ee-ac6f-15ae965c3402 · outbound

This paper cites SLDP-LoRA: A privacy-preserving split learning framework with low-rank adaptation.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models SLDP-LoRA: A privacy-preserving split learning framework with low-rank adaptation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:17.521863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:17.521863Z digest=sha256:cb9209ecf19fb9dead871af65bfc75a26114e33c10c54cc9fc2ee414287e5291

Observation c36f73c1-69bf-41a8-95ca-810241f11245 · outbound

This paper cites Split-and-Denoise: Protect large language model inference with local differential privacy.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Split-and-Denoise: Protect large language model inference with local differential privacy

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:17.678536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:17.678536Z digest=sha256:86dfdc9ae154cf2f9a7c1572fb882c3a59b249309e8ac65b87c3583a52c35cda

Observation 1483a785-638b-4235-aaee-269cc19c2a66 · outbound

This paper cites Santos, et al.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Santos, et al

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:17.851965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:17.851965Z digest=sha256:960b25e767f8bf27ad8c7a82dc9e56292308ef9cf8a0e790e976d16251e3d24e

Observation 2dc0dd2e-5e5f-4118-8583-d065b60aefcd · outbound

This paper cites GPTQ: Accurate post-training quantization for generative pre-trained transformers.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models GPTQ: Accurate post-training quantization for generative pre-trained transformers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.076697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.076697Z digest=sha256:58d434d41ce8e65c7f1fc4c37a963c72a0161de7521a57528aa218157a1b117d

Observation 67ea7582-74b1-4a21-be37-a25d1d408574 · outbound

This paper cites AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.194038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.194038Z digest=sha256:3715aa734702afea7758bb776dea410e9fdfe52f5690bb57caef8fd352d35ca3

Observation 143f02b4-ff8e-46cf-a862-004b59d02e04 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Smoothquant: Accurate and efficient post-training quantization for large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.257172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.257172Z digest=sha256:feb924ff6763c2ff55c68aade5352db8e10d4f7de27b42682740d8eb52f9ddc2

Observation 7c1bc7fc-10c5-48be-88d3-35dbd3e50c00 · outbound

This paper cites A simple and effective pruning approach for large language models.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models A simple and effective pruning approach for large language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.321759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.321759Z digest=sha256:b7252422e8a0e354b48b4bdc12d7ac168148759a95ddd24e24f2a12d67059b21

Observation 58f287d4-07cc-4f10-bbab-3abaead8bc27 · outbound

This paper cites Minillm: Knowledge distillation of large language models.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Minillm: Knowledge distillation of large language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.401946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.401946Z digest=sha256:614d6b1f1403ef03fea265e0b11ba2ada847915033cfa9a3f269a6bc266ea5ef

Observation 62123e57-7a91-474e-9325-b83a21c203bc · outbound

This paper cites Efficient LLM inference on CPUs.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Efficient LLM inference on CPUs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.479747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.479747Z digest=sha256:b399a50d05a30f29c76359872f3dac668b5b9a4875c5a02960a2b3ccbf7e0d55

Observation 715331ba-a009-48c1-a933-2b3ebe623b87 · outbound

This paper cites H2O: Heavy-hitter oracle for accurate KV cache compression.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models H2O: Heavy-hitter oracle for accurate KV cache compression

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.566986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.566986Z digest=sha256:a57e0154c80f03c77fa2528f53969400960aee1a7148fb268ad3cf9b77c8faae

Observation 487eca1a-41f9-4c73-9bdd-671413163a11 · outbound

This paper cites Streamingllm: Efficient streaming language models with attention sinks.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Streamingllm: Efficient streaming language models with attention sinks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.679844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.679844Z digest=sha256:4f43036f23f875105106d0804e24941b7f3161508a0bc85dafbea1842f7f06a3

Observation 81176a82-252d-409c-ad5d-184fe5b7b48b · outbound

This paper cites Recommendation for block cipher modes of operation: Ga- lois/counter mode (gcm) and gmac.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Recommendation for block cipher modes of operation: Ga- lois/counter mode (gcm) and gmac

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.791543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.791543Z digest=sha256:bf4c6cc4502a5617cfd1011172f44a90c5e0b81d32f7150bef9888f5fca76691

Observation 13e26c63-4711-4bdf-bec6-4480245fafdd · outbound

This paper cites BOLT: Privacy-preserving, accurate and efficient inference for transformers.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models BOLT: Privacy-preserving, accurate and efficient inference for transformers

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.867792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.867792Z digest=sha256:0283c444d12b4500d3913fe6776f37b50d42dc3d9684ad38accdf609b30ccead

Observation e05b9057-df30-4519-a95a-a42138371044 · outbound

This paper cites THE-X: Privacy-preserving transformer inference with homomorphic encryption.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models THE-X: Privacy-preserving transformer inference with homomorphic encryption

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.969333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.969333Z digest=sha256:ca0cfe056feec81167b5d9a76649cd55f900867eb572ff927eaf2471f28e20ce

Observation 5369bbcc-5885-4905-9e6a-cf607f584081 · outbound

This paper cites Xing, et al.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Xing, et al

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:19.083020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:19.083020Z digest=sha256:52304311ac98077b5191c5a20197ee6546c37fae0279b9ec3c86ff1ae978fa33

Observation 223d5a16-0e04-49b4-8373-640f5860107b · outbound

This paper cites Secureinfer: Heterogeneous TEE-GPU architecture for privacy- critical LLM tensors.arXiv preprint arXiv:2510.19979, 2025.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Secureinfer: Heterogeneous TEE-GPU architecture for privacy- critical LLM tensors.arXiv preprint arXiv:2510.19979, 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:19.164171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:19.164171Z digest=sha256:fa1d7e25d05ab1c8b5390a9945f3e7a49bb706ad3047892d6f6129f690008899

Observation 31bfd236-2284-4e7d-970a-f89636aeddf0 · outbound

This paper cites ONNX Runtime: Cross-platform, high performance machine learning inferencing and training accelerator.https://onnxruntime.ai/.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models ONNX Runtime: Cross-platform, high performance machine learning inferencing and training accelerator.https://onnxruntime.ai/

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:19.238337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:19.238337Z digest=sha256:3a690f91f390a820cf72639356d30df05de58891c58241b540eaee9199eb71f7

Observation 02d45801-cac2-4a28-b676-337714cc26d7 · outbound

This paper cites Text Revealer: Private Text Reconstruction via Model Inversion Attacks against Transformers.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Text Revealer: Private Text Reconstruction via Model Inversion Attacks against Transformers

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:19.309412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:19.309412Z digest=sha256:6747250c5eacedaabbf7470f36e0c48f5b165fd4c6d4582940c98f2a7b697da4

Observation 555f580a-0501-459e-bbc7-4be1bc3da1d2 · outbound

This paper cites ML-Doctor: Holistic risk assessment of inference attacks against machine learning models.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models ML-Doctor: Holistic risk assessment of inference attacks against machine learning models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:19.410428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:19.410428Z digest=sha256:f2c09d1e4e4c9d47daac69646cdbb4f78e467b79be821e56300a3d33df95ea33

Observation b3620fcb-bd28-4f09-92f5-16b6b7356dd0 · outbound

This paper cites Privacy side channels in machine learning systems.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Privacy side channels in machine learning systems

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:19.565403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:19.565403Z digest=sha256:edb263f261db07835c40be9f4e8429eb1c4f30ccdb80725998d267f56725d9d2

Observation f26177c0-e00f-4545-a5fe-dece69615396 · outbound

This paper cites Secure multi-party computation for machine learning: A survey.IEEE Communications Surveys & Tutorials, 2024.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Secure multi-party computation for machine learning: A survey.IEEE Communications Surveys & Tutorials, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:19.737748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:19.737748Z digest=sha256:482e078426fc1e9bbaf467b51cdcb0862f9293d2b1635e3d9a60a4ed3ba5be08

Observation a2cb37e2-5ef6-48ae-8838-00b7e3d2b366 · outbound

This paper cites an unresolved cited work.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:19.865695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:19.865695Z digest=sha256:7533c7d1be81c48968297689163c6d6b335b5275196dcc484bfb374ef2e3c3e6

Pith citing papers

No inbound Pith citation observations are available.