Pith. sign in

Paper Citation Record · LEDGER

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices

As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2512.21835.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.21835 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:04:37.517635Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 347964fc-61d7-4a83-a328-98418ecb40bf · outbound

This paper cites DeepSeek-V3 Technical Report.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices DeepSeek-V3 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:33.761205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:33.761205Z digest=sha256:363e8d320c3f0014a60cfaf6265c78d4e37d8372704fdd05b02d86f2ff0b3aba

Observation 7f067646-16b2-4be3-8163-2e9422b6aab5 · outbound

This paper cites Improving language understanding by generative pre-training,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Improving language understanding by generative pre-training,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:33.825769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:33.825769Z digest=sha256:4230d69f583817d8b2aeefbf342eaaf6e88bc131538b5d8bc81ef430508dff6e

Observation 07d77fe5-ce41-4b45-a2df-b5a7371ec808 · outbound

This paper cites Sasha: creative goal-oriented reasoning in smart homes with large language models,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Sasha: creative goal-oriented reasoning in smart homes with large language models,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:33.884490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:33.884490Z digest=sha256:0ac39458ac4af751c7764eb531e2322a3d53d1184b66ec46c6f9475a3ab90238

Observation 9ec56baf-10b9-4b35-8efb-e63026edc1ae · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Gemini Robotics: Bringing AI into the Physical World

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:33.995628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:33.995628Z digest=sha256:730c76aacb0684c275e9d178351fde2b3a8b8a09667f520afdb51e4dbc9eec83

Observation d1158706-503f-43b3-a1c1-f6546f429006 · outbound

This paper cites Privacy Inference Attacks and Defenses in Cloud-based Deep Neural Network: A Survey.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Privacy Inference Attacks and Defenses in Cloud-based Deep Neural Network: A Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:34.103587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:34.103587Z digest=sha256:1fceb11361e46406daaf6458552b9e8b1ea4ce1355b088ec2edaa78b37e22edc

Observation 09d28748-6dee-4d5b-8e0f-faa145f3da81 · outbound

This paper cites Privacy-preserving federated learning for transportation mode prediction based on personal mobility data,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Privacy-preserving federated learning for transportation mode prediction based on personal mobility data,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:34.160746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:34.160746Z digest=sha256:622d5a4c98c3f3d00e203ac5670ac31aa43220d03d03e2be7ad80f9ab70ae06b

Observation 9663b91e-ecee-461f-ace8-78b01e93401a · outbound

This paper cites Re- search on medical data storage and sharing model based on blockchain,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Re- search on medical data storage and sharing model based on blockchain,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:34.208909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:34.208909Z digest=sha256:4b152befd210beeb48f2a0ef826eb05ad565bc1997759c2c136a0bd3b2ed5689

Observation 2c5b9380-7a76-4f85-9f18-62f21d860edf · outbound

This paper cites {ServerlessLLM}:{Low-Latency}serverless inference for large language models,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices {ServerlessLLM}:{Low-Latency}serverless inference for large language models,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:34.313116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:34.313116Z digest=sha256:9de6491f7d73d2fc7700a9609b572a9bb05f6723eb60d99f42cffc6ac49bc33b

Observation 9a95d7e2-e5e6-4848-b422-fa5bd7b96aa3 · outbound

This paper cites The llama 3 herd of models,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices The llama 3 herd of models,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:34.422921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:34.422921Z digest=sha256:62384583653be533f72ee05fa5c44e43caabebe5c22d151607378101315e6b69

Observation 08d95cc8-6c3c-4188-a13b-991ba6520659 · outbound

This paper cites Nvidia jetson xavier,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Nvidia jetson xavier,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:34.482417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:34.482417Z digest=sha256:98f37cc955410fe83fce81a56118eb3e69b3f5312d9aff4a64d36e42b4d4779f

Observation bce7fa31-24ac-4b45-979c-dddb123b5dbb · outbound

This paper cites Kvquant: Towards 10 million context length llm inference with kv cache quantization,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Kvquant: Towards 10 million context length llm inference with kv cache quantization,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:34.537868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:34.537868Z digest=sha256:22a6ba226167bc291fc54c544009072d7db9787d96eaaac29866bce438be2400

Observation 8c40f42f-b3c2-42b6-ac17-82841a3ccebe · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:34.644179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:34.644179Z digest=sha256:a1722e57d75d39719f56f40a7e2878e0d5ab0b7ec052b9abc888d607bbd5c4f3

Observation 65590f5b-6841-4d13-a596-b45f86c62864 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Qlora: Efficient finetuning of quantized llms,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:34.746272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:34.746272Z digest=sha256:88f540ae04e2689218c3add6f68c7f8980c9b8984697ac0218f122e65fc85e13

Observation 1222fb62-48b8-4dcb-9636-2ab067d195f6 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:34.853561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:34.853561Z digest=sha256:2e168e78ec72c3030b7e33fb2187b745bb0bffbf7f58e73fcd093c41d0726766

Observation 107578cc-9dea-4a79-8380-29dc03b4f203 · outbound

This paper cites Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:34.962065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:34.962065Z digest=sha256:b18c1445b20bf3d4cbb3dab77476ec2a1d9080ee5b892aa0f4bca11db95aed45

Observation 9f627b4a-b326-4e4e-bfb1-8f4a4da11880 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices MiniLLM: On-Policy Distillation of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:35.077208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:35.077208Z digest=sha256:9cb06ffd89c93bb1ee4ca6185deabe563a259209b3b94c36252ade20d02d4178

Observation 4aad829c-bdec-4f0c-bc3e-20f7b6c5e14e · outbound

This paper cites Less is more: Task-aware layer-wise distillation for language model compression,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Less is more: Task-aware layer-wise distillation for language model compression,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:35.184824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:35.184824Z digest=sha256:a1fbea081f71c46b44371267e5b34e25a8146d028b6f14d996e7c5d555c3fdcb

Observation fe458615-5a4c-4155-8bd1-677c24c9fcaf · outbound

This paper cites Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:35.292557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:35.292557Z digest=sha256:e00c65d64940148a791e193c6419153d64aa103dd00734f3308b6258f6866709

Observation e30c8108-2404-47a1-849f-5cd8af5bafd9 · outbound

This paper cites Resource-aware federated self-supervised learn- ing with global class representations,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Resource-aware federated self-supervised learn- ing with global class representations,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:35.355574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:35.355574Z digest=sha256:e4a7c39b19a34ac3003259491a84133c4a492737a2c093a17faa08ddba032e0f

Observation 6868cb3d-a0ed-43ad-8b16-8f1ed472c0f7 · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Sparsegpt: Massive language models can be accurately pruned in one-shot,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:35.463293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:35.463293Z digest=sha256:faf25e962b83511acb752d7d7fc321e6c5d0634edf82a6297dc4ab5c41ccab99

Observation 426fa491-f22c-41a2-a4c0-9414821500d4 · outbound

This paper cites Unity is power: Semi-asynchronous collaborative training of large-scale models with structured pruning in resource-limited clients,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Unity is power: Semi-asynchronous collaborative training of large-scale models with structured pruning in resource-limited clients,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:35.521177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:35.521177Z digest=sha256:8e21600a317d94eb432f0a7ce39372e286451024a2de58990a9454c48172dc62

Observation 6bee5a92-c87e-4a35-89dd-8fa390fecfbf · outbound

This paper cites SliceGPT: Compress Large Language Models by Deleting Rows and Columns.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices SliceGPT: Compress Large Language Models by Deleting Rows and Columns

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:35.628892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:35.628892Z digest=sha256:c766ab3b5180baa3f3fd44173b8bdfc71ea20e24371167f33b9000e1afff6e16

Observation 6634eded-36bc-4e47-b5fd-8f02d0ac0ebc · outbound

This paper cites Llm-pruner: On the structural pruning of large language models,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Llm-pruner: On the structural pruning of large language models,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:35.688378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:35.688378Z digest=sha256:1d408f959902c5fb6d9ebb67b7baced8af4406517e11ad37193124b8b61a9413

Observation 276674d8-e370-4885-8b8b-a7961b363594 · outbound

This paper cites Generalization-aware distributed minimax optimization for large-scale models on resource-limited devices,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Generalization-aware distributed minimax optimization for large-scale models on resource-limited devices,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:35.743614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:35.743614Z digest=sha256:bc7c3443a3cb484643a93a704f55d1bdb6644433fdaff28a034618d6db262c60

Observation d5284734-b987-4907-9914-2d0fc84c0a64 · outbound

This paper cites Theoretical convergence guaranteed resource-adaptive federated learning with mixed heterogeneity,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Theoretical convergence guaranteed resource-adaptive federated learning with mixed heterogeneity,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:35.789583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:35.789583Z digest=sha256:06e70fe7eb9226f2ba3e64eb619498791c849db1ea30a315794d28a7fe925d6e

Observation b9348d74-8ef6-4472-b5c2-f4051b35c602 · outbound

This paper cites Credit risk analysis using machine and deep learning models,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Credit risk analysis using machine and deep learning models,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:35.860427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:35.860427Z digest=sha256:1c3b240d1bdb79596ecdb848b1b958d1b6bcf68badfaa0c829be71047c98db88

Observation eab24cbf-13ad-42b4-b6dc-79d22740a56a · outbound

This paper cites Medical image analysis using deep learning algorithms,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Medical image analysis using deep learning algorithms,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:35.893519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:35.893519Z digest=sha256:5eb04f690ea86e9cdab5e45d29da499cc8406cf68c5c0ea48c87a3ccb2459d58

Observation 28cfe570-720f-4106-830b-2767489d3dd0 · outbound

This paper cites Pipeedge: Pipeline parallelism for large-scale model inference on heterogeneous edge devices,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Pipeedge: Pipeline parallelism for large-scale model inference on heterogeneous edge devices,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:35.963974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:35.963974Z digest=sha256:b485c3bf01a18f49edd141a6bac21abdfc6087a66f3fe1b026ef22ca15805765

Observation 0336e8dd-c341-4c79-9508-165368199780 · outbound

This paper cites Edgeshard: Efficient llm inference via collaborative edge computing,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Edgeshard: Efficient llm inference via collaborative edge computing,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:35.992124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:35.992124Z digest=sha256:3b7b727b16efdf7a527b3c05b058caf8d1d59303af7e3834f0c09696e5c92e02

Observation 77669c8c-c35c-431b-8035-9531a73be352 · outbound

This paper cites Galaxy: A resource-efficient collaborative edge ai system for in-situ transformer inference,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Galaxy: A resource-efficient collaborative edge ai system for in-situ transformer inference,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:36.010971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:36.010971Z digest=sha256:694a0a86fed104e996f3adba3cceda88b79892482f83c97c74ecd82b221f6512

Observation 28889b35-fdbd-408b-b30a-7d1d1259c442 · outbound

This paper cites Tpi-llm: Serving 70b-scale llms efficiently on low-resource mobile devices,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Tpi-llm: Serving 70b-scale llms efficiently on low-resource mobile devices,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:36.063253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:36.063253Z digest=sha256:3cf93abc91ef7a4e4e4e7e5dd7429c0734dbaf44708345a0c19237e1f056a7f0

Observation 9f461e3b-507e-4f4b-83d4-94c9134b0a53 · outbound

This paper cites Band: coordinated multi-dnn inference on heterogeneous mobile processors,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Band: coordinated multi-dnn inference on heterogeneous mobile processors,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:36.168116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:36.168116Z digest=sha256:fb28a09fbde92d311fa3d5fbe5fee5a2c0c6a8d73797c60b0a1860927824d907

Observation e8fb5fd0-5b98-40a6-935e-8707d523343b · outbound

This paper cites Coedge: Cooperative dnn inference with adaptive workload partitioning over heterogeneous edge devices,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Coedge: Cooperative dnn inference with adaptive workload partitioning over heterogeneous edge devices,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:36.277666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:36.277666Z digest=sha256:26a2a6657cefbf8dbc4c06bf35e186b32d15413524f0821e842f0eb46b9297d7

Observation b9d5f377-fc33-4055-b2df-052b3ab73c85 · outbound

This paper cites Model parallelism optimization for distributed inference via decoupled cnn structure,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Model parallelism optimization for distributed inference via decoupled cnn structure,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:36.441328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:36.441328Z digest=sha256:f07b00e867b872b854f5c4e904e31e6795094118990aba4245a318af551c43c8

Observation 25b0704f-d791-4f6b-b628-100d4179953b · outbound

This paper cites When the edge meets transformers: Distributed inference with transformer models,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices When the edge meets transformers: Distributed inference with transformer models,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:36.589273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:36.589273Z digest=sha256:d2a8cfeac0bf15679c19d765db3a06ead744458d7916c1e854b528399aba6adc

Observation 416b80cd-ab38-49ad-ae6c-0d35398c08e0 · outbound

This paper cites Jupiter: Fast and resource-efficient collaborative inference of generative llms on edge devices,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Jupiter: Fast and resource-efficient collaborative inference of generative llms on edge devices,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:36.741062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:36.741062Z digest=sha256:2e595347819bbc5c7340365c7deb86069c71b1086603eb4313fded234550e617

Observation 076bfe50-5ba5-4d3a-b77c-fb6b595a5533 · outbound

This paper cites Sequence parallelism: Long sequence training from system perspective,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Sequence parallelism: Long sequence training from system perspective,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:36.865777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:36.865777Z digest=sha256:7f926cee0f011113c2a776b0f4c73fe329aa456fecf82bec8846fc215aacf585

Observation 748c3806-7c81-4d1d-b196-73247132d685 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:36.975700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:36.975700Z digest=sha256:6391593ff9d173000c3ab23bb3cdf4d51d4e085ba99b58fcebc791f22bc5c4bb

Observation b052bd9a-b3da-46cd-9cb8-bac4877d8eae · outbound

This paper cites Pipedream: Generalized pipeline parallelism for dnn training,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Pipedream: Generalized pipeline parallelism for dnn training,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:37.104365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:37.104365Z digest=sha256:712e6af795a43f3c9934c47eece961961dd0273c462cbd21dcfe038c85e46d2c

Observation 399f08c9-c621-4d54-9d73-0928a71ac8d6 · outbound

This paper cites Gpipe: Efficient training of giant neu- ral networks using pipeline parallelism,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Gpipe: Efficient training of giant neu- ral networks using pipeline parallelism,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:37.216489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:37.216489Z digest=sha256:31ff47094d41edf0cc9fa1c1aeac379d09f241a5ba975f87a95dfa4fa024f180

Observation a3df890f-a88d-4903-a1f1-4aef2fa3fd01 · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:37.369007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:37.369007Z digest=sha256:834ad87f08d45d3308facd7e967fdebc4b7f393349f588ae1f46f4a48663616b

Observation ba45872c-6f51-40ae-bda4-b3e3d980f3ab · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Flexgen: High-throughput generative inference of large language models with a single gpu,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:37.412504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:37.412504Z digest=sha256:76df49588e20abce2c851d874fe8749dc97b24aef942a79f7030ac2690b966a9

Observation 838c0edd-4585-4dcc-bfac-e0a520bf3eb4 · outbound

This paper cites Linux advanced routing & traffic control howto,.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Linux advanced routing & traffic control howto,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:37.472113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:37.472113Z digest=sha256:4c5477b78f18835f26c5cf6b1f8883fd3331d190e1df8c0642058ec3f8170def

Observation 9a6f689e-e879-40b9-a078-8d8160822cbf · outbound

This paper cites Qwen3 Technical Report.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Qwen3 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:37.517635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:37.517635Z digest=sha256:5baec94aaee4af44dde5373e3f9c1189d20cfcd7fd8167679a7307e978ed9d38

Pith citing papers

No inbound Pith citation observations are available.