Pith. sign in

Paper Citation Record · LEDGER

Deploying Foundation Model Powered Agent Services: A Survey

As of 17 August 2026, this Paper Citation Record lists 100 of 300 outbound references and 2 inbound Pith citation observations for arXiv:2412.13437.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.13437 v1

Coverage vector

measured 100 of 300 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:09:45.699402Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:47:39.121487Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T14:06:38.131407Z

Reference resolution

100 of 300 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6fa4800a-d355-49fe-bf6e-e2b61fe09361 · outbound

This paper cites On the Opportunities and Risks of Foundation Models,.

Deploying Foundation Model Powered Agent Services: A Survey On the Opportunities and Risks of Foundation Models,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.779007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.779007Z digest=sha256:aa851910fe7b1499a4a4a9c11be7ed9193b407d160edbe27538a8b67b3480c79

Observation 37ebb9b6-bf7b-485a-9235-9e5908e39e7c · outbound

This paper cites The Rise and Potential of Large Language Model Based Agents: A Survey,.

Deploying Foundation Model Powered Agent Services: A Survey The Rise and Potential of Large Language Model Based Agents: A Survey,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.792843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.792843Z digest=sha256:dbf2f94edb226513285314db38dfdbcd0aaa27baaa3c7fc63ef9387a08681463

Observation 1efdd7fd-6506-48ec-b5c9-c8293d551286 · outbound

This paper cites 107 Up-to-Date ChatGPT Statistics & User Numbers,.

Deploying Foundation Model Powered Agent Services: A Survey 107 Up-to-Date ChatGPT Statistics & User Numbers,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.802780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.802780Z digest=sha256:1a7a2cce1d5c30a38ba116367b74f483f76da4dc61c6d922ab8ed6d7a8ca2a90

Observation 33052f6b-0d67-47ae-988f-9b3c6c585b04 · outbound

This paper cites A Survey on Hardware Accelerators for Large Language Models,.

Deploying Foundation Model Powered Agent Services: A Survey A Survey on Hardware Accelerators for Large Language Models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.810366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.810366Z digest=sha256:e0efe6197681515ca253000dd0d291c652c2e7d042a6ddd3df4a35b4035de9cf

Observation 95954845-4e77-4541-b10c-dcaa0a59abfb · outbound

This paper cites A Survey on Scheduling Techniques in Computing and Network Convergence,.

Deploying Foundation Model Powered Agent Services: A Survey A Survey on Scheduling Techniques in Computing and Network Convergence,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.818604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.818604Z digest=sha256:1775b9f88ec863f3550592d2a63f65c2f4038277e7c57aba2762db6d379ea3c0

Observation d2a7a6b2-3ea7-4d25-83dd-17da27509e8c · outbound

This paper cites Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems,.

Deploying Foundation Model Powered Agent Services: A Survey Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.826199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.826199Z digest=sha256:69d7932db0f5723617361dcc721f5c6013b411672978ff4848ee4198e1eb79df

Observation 8c2023fa-f709-4452-8055-df0d414ebf2d · outbound

This paper cites Unleashing the Power of Edge-Cloud Generative AI in Mobile Net- works: A Survey of AIGC Services,.

Deploying Foundation Model Powered Agent Services: A Survey Unleashing the Power of Edge-Cloud Generative AI in Mobile Net- works: A Survey of AIGC Services,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.833892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.833892Z digest=sha256:e227880da7167742e3918f2259d51daf11cf038fea0fc4ab8c27168aa5f21916

Observation 3fed0e64-dd77-42c9-a911-0c71f61ca0e6 · outbound

This paper cites Machine and Deep Learning for Resource Allocation in Multi-Access Edge Computing: A Survey,.

Deploying Foundation Model Powered Agent Services: A Survey Machine and Deep Learning for Resource Allocation in Multi-Access Edge Computing: A Survey,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.839870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.839870Z digest=sha256:d6129c1887bc5a5c48a694a08912a73b5b8f908d4247632d1fd1513dc8363d45

Observation 793af5de-f3cc-4c38-b0ed-9d1eeab33b4e · outbound

This paper cites A Survey of Large Language Models,.

Deploying Foundation Model Powered Agent Services: A Survey A Survey of Large Language Models,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.846821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.846821Z digest=sha256:fd6de492cd9f8e3a91ae74b4638acf1186ff69f2821deb083cddac2d55ae007e

Observation badc3e58-19fe-452c-93a1-e05524f3d910 · outbound

This paper cites A Comprehensive Overview of Large Language Models,.

Deploying Foundation Model Powered Agent Services: A Survey A Comprehensive Overview of Large Language Models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.854139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.854139Z digest=sha256:f2a5ec26a1d8b5752bf07b7c60997f79906d2f4a0efcc659160f2e7f76c1af74

Observation 9df22d56-6622-40bc-8c33-4066fd3767dc · outbound

This paper cites A Survey on Model Compression for Large Language Models,.

Deploying Foundation Model Powered Agent Services: A Survey A Survey on Model Compression for Large Language Models,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.863625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.863625Z digest=sha256:4a8978166bb41af08c116e71292873583f60a6b6d00df251e08991602c198905

Observation effbb557-e21f-4d1d-adb0-efc5b6fadc44 · outbound

This paper cites Model Compression and Efficient Inference for Large Language Models: A Survey,.

Deploying Foundation Model Powered Agent Services: A Survey Model Compression and Efficient Inference for Large Language Models: A Survey,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.869985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.869985Z digest=sha256:e5cb68f122aa3608f8e0dcd07e5567593d41dd0579512bab37b7864e6130f5fd

Observation 78f134e6-b2b7-460c-b33e-400b8e0aa1d9 · outbound

This paper cites A Survey on Knowledge Distillation of Large Language Models.

Deploying Foundation Model Powered Agent Services: A Survey A Survey on Knowledge Distillation of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.875640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.875640Z digest=sha256:220802fbef62eb7d0383d1e08940e7c1472667e5732abd9b1ffccd6c52336053

Observation b40e5f34-c181-4e85-b872-a94fb6508daf · outbound

This paper cites A survey on large language model based autonomous agents,.

Deploying Foundation Model Powered Agent Services: A Survey A survey on large language model based autonomous agents,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.881689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.881689Z digest=sha256:b1de74583b3c8c8f4fe9d89babc5dd31bbf2e195bcf11ac4d58087437ca2eb8c

Observation e73b7eed-4e37-4425-a734-0fdf64dbceb1 · outbound

This paper cites Large Language Model based Multi-Agents: A Survey of Progress and Challenges.

Deploying Foundation Model Powered Agent Services: A Survey Large Language Model based Multi-Agents: A Survey of Progress and Challenges

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.887811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.887811Z digest=sha256:73fa1c89cd386bb92c3b77547ea337b8bfbeea66a39789797a0ce72cfb57306a

Observation 10afd8dd-fe79-4ae4-bf9f-f99442365c87 · outbound

This paper cites FedDSE: Distribution-aware Sub-model Extraction for Fed- erated Learning over Resource-constrained Devices,.

Deploying Foundation Model Powered Agent Services: A Survey FedDSE: Distribution-aware Sub-model Extraction for Fed- erated Learning over Resource-constrained Devices,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.894927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.894927Z digest=sha256:58907dab9f73bb5cebdc85ae7982d8246c4e3d20a53b54bf89bc1b77a65331f0

Observation 36e0b7aa-c032-4dc6-b092-48b6d5dc29b7 · outbound

This paper cites Hardware accelerator for multi-head attention and position-wise feed-forward in the trans- former,.

Deploying Foundation Model Powered Agent Services: A Survey Hardware accelerator for multi-head attention and position-wise feed-forward in the trans- former,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.902027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.902027Z digest=sha256:c4c52022915b4ff4bc37a011848a65d2b7200144ef988c751f10294ba795d0bd

Observation 00dc1268-87da-4ba2-b2d0-8b143623b051 · outbound

This paper cites Mnnfast: A fast and scalable system architecture for memory-augmented neural networks,.

Deploying Foundation Model Powered Agent Services: A Survey Mnnfast: A fast and scalable system architecture for memory-augmented neural networks,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.909274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.909274Z digest=sha256:4b57632851e7a1b607627afb341b8792f51078dd0bd63647b8fd8ceeaa1c2b17

Observation 6e2db958-2a28-479b-b2ac-7d27222e241f · outbound

This paper cites NPE: An FPGA-based Overlay Processor for Natural Language Processing.

Deploying Foundation Model Powered Agent Services: A Survey NPE: An FPGA-based Overlay Processor for Natural Language Processing

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.916243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.916243Z digest=sha256:e29274bcdd1c301dbe84645c348ec0adfa092eb3b28dfc353bdaee3085a6ca3e

Observation 9c3b5ec6-566f-4bf4-9f14-e5cd3356b8c2 · outbound

This paper cites Dfx: A low-latency multi-fpga appliance for accelerating transformer- based text generation,.

Deploying Foundation Model Powered Agent Services: A Survey Dfx: A low-latency multi-fpga appliance for accelerating transformer- based text generation,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.923086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.923086Z digest=sha256:14d0ad1861aad6a760d556261759fa6edcdbfae07cb764e70518f57d0d7b52b5

Observation 958d7f88-e6d5-4c14-8d05-db571355d9c0 · outbound

This paper cites Transformer- opu: An fpga-based overlay processor for transformer networks,.

Deploying Foundation Model Powered Agent Services: A Survey Transformer- opu: An fpga-based overlay processor for transformer networks,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.931785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.931785Z digest=sha256:13f3c2979a603b429f14e0c86d53780ab636d780f72a1dac5a6eb9ccd1c9993d

Observation 25e56444-1341-4d47-8ccb-c5297de9bf2f · outbound

This paper cites A Cost-Efficient FPGA Implementation of Tiny Transformer Model using Neural ODE.

Deploying Foundation Model Powered Agent Services: A Survey A Cost-Efficient FPGA Implementation of Tiny Transformer Model using Neural ODE

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.938941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.938941Z digest=sha256:d2ed349622451bb8aadb3d445b6f5bd4268a4b436ab9ff40e3b14a8938495cb0

Observation 8d850878-d709-4225-905f-2e22e90ae9f6 · outbound

This paper cites FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs.

Deploying Foundation Model Powered Agent Services: A Survey FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.947506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.947506Z digest=sha256:7f85f468f897a53a4a925a1f6a1d9152914c80d94e10e48bfc346dc9deec77cf

Observation 9a83747f-a5dd-475a-b5c3-16531a638ed5 · outbound

This paper cites Aˆ 3: Accelerating attention mechanisms in neural networks with approximation,.

Deploying Foundation Model Powered Agent Services: A Survey Aˆ 3: Accelerating attention mechanisms in neural networks with approximation,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.958912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.958912Z digest=sha256:9c7deb861ece198e2fee2e8ffd6d488b0938d7e26a82b3eb89dd00fce361b0dc

Observation 92be0b6b-b2b9-43a5-801f-ef4f8f2a3cad · outbound

This paper cites Elsa: Hardware-software co-design for efficient, lightweight self- attention mechanism in neural networks,.

Deploying Foundation Model Powered Agent Services: A Survey Elsa: Hardware-software co-design for efficient, lightweight self- attention mechanism in neural networks,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.971634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.971634Z digest=sha256:25e87aa13242ee9a3318b698790e81fd70044bd71d1553a3fe17abeace9d4c6c

Observation c1c09a46-5180-4830-a000-4d930dc38e50 · outbound

This paper cites Spatten: Efficient sparse attention architecture with cascade token and head pruning,.

Deploying Foundation Model Powered Agent Services: A Survey Spatten: Efficient sparse attention architecture with cascade token and head pruning,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.978791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.978791Z digest=sha256:97e178982b365dd2c6fea154e15d181fb3a081224d656767a9ed5bd0e83a6415

Observation a1fc3833-4dbd-4d2b-b44a-06ffe9b0e933 · outbound

This paper cites Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture,.

Deploying Foundation Model Powered Agent Services: A Survey Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:44.996917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:44.996917Z digest=sha256:98a4a3c2c6bb031393665fbf9d7f52133a04d73d31ace6ec1cef0206fb142e5a

Observation 5551b2b3-0629-40e1-be89-8a4260699275 · outbound

This paper cites Energon: Toward efficient acceleration of transformers using dynamic sparse attention,.

Deploying Foundation Model Powered Agent Services: A Survey Energon: Toward efficient acceleration of transformers using dynamic sparse attention,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.003951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.003951Z digest=sha256:3b1131fb685f589774f41f5affe4162bda501547ba544fea30d64b19c927ab44

Observation a7f72601-fa9a-48ec-9c8b-5477c2f56578 · outbound

This paper cites Att: A fault-tolerant reram accelerator for attention-based neural networks,.

Deploying Foundation Model Powered Agent Services: A Survey Att: A fault-tolerant reram accelerator for attention-based neural networks,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.017639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.017639Z digest=sha256:095981c1fc3d5fd85b0ca2324ea76ea085891cfbe75b74b351da4732860f6ffd

Observation c9bf2ea7-a315-4215-ab99-37163a26b14e · outbound

This paper cites In-memory com- puting based accelerator for transformer networks for long sequences,.

Deploying Foundation Model Powered Agent Services: A Survey In-memory com- puting based accelerator for transformer networks for long sequences,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.025799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.025799Z digest=sha256:59cd19da12838eb40bf2b10c428b537c07dc5b8e3c3bf971cb532fc84967c081

Observation d268528a-6776-44ed-bacb-755be315ab66 · outbound

This paper cites Work in progress: Real-time transformer inference on edge ai accelerators,.

Deploying Foundation Model Powered Agent Services: A Survey Work in progress: Real-time transformer inference on edge ai accelerators,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.032215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.032215Z digest=sha256:a024f590bc4f310545b5d6394eb773dac99e3eefd8f6739cae74ece75220e572

Observation 1165d6dd-7cfd-4073-a9aa-474fb5fe93d7 · outbound

This paper cites Simplifying Transformer Blocks.

Deploying Foundation Model Powered Agent Services: A Survey Simplifying Transformer Blocks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.044288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.044288Z digest=sha256:cb8afa8fc2152aa7604bc3fd2c8f0485009daa6d266e42499c21df96c10bddcb

Observation 3359070c-c84a-43bc-af38-20c5e8f4957b · outbound

This paper cites Accelerating transformer networks through recomposing softmax layers,.

Deploying Foundation Model Powered Agent Services: A Survey Accelerating transformer networks through recomposing softmax layers,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.053937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.053937Z digest=sha256:df723cb85c86f3a9fd0cb1af280ba3bd724c6bbb60a00686aefd5963ef877154

Observation 1d4ce839-6ebc-417f-8c23-3babb8e89872 · outbound

This paper cites Inference with Reference: Lossless Acceleration of Large Language Models.

Deploying Foundation Model Powered Agent Services: A Survey Inference with Reference: Lossless Acceleration of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.062703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.062703Z digest=sha256:2b7cf9cb9e21ab850ee175df3233c63a7e2363939a0fe48a737714059d07d650

Observation 4489ab2f-f63f-459b-a8c0-2dc647e6af75 · outbound

This paper cites Exponentially Faster Language Modelling.

Deploying Foundation Model Powered Agent Services: A Survey Exponentially Faster Language Modelling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.072049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.072049Z digest=sha256:50280236761aaf6ca8477448786224cc30d458576699b24c2ff5773450c19671

Observation 70d4e60c-576b-4e68-b3e0-ccfa1aa6eb3d · outbound

This paper cites Efficient LLM Inference on CPUs.

Deploying Foundation Model Powered Agent Services: A Survey Efficient LLM Inference on CPUs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.080567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.080567Z digest=sha256:bb265acea964c9668420d180685a4824e06b8862f98f7daeea68100af17e28e5

Observation 937ce440-38e5-4671-ba2b-0b4e032edc80 · outbound

This paper cites PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU.

Deploying Foundation Model Powered Agent Services: A Survey PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.091178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.091178Z digest=sha256:e4bdcde728f86b90db87f859f579cc3f6fe0ab3ed142d6eed050a5655777b59a

Observation 78edecc1-f912-4d40-81f1-8d8254a6173a · outbound

This paper cites HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices.

Deploying Foundation Model Powered Agent Services: A Survey HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.100032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.100032Z digest=sha256:cbd7568dc79003c255c924bc6097919953ba0e2b87c427c4c29f8855da44d751

Observation 8b9fe4ed-e832-4f3e-9ae6-38af59082e21 · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu,.

Deploying Foundation Model Powered Agent Services: A Survey Flexgen: High-throughput generative inference of large language models with a single gpu,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.111120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.111120Z digest=sha256:7b685291898c6b685032a60986ab3dd588f302e56cc076e036e3b28c2a5af27d

Observation 0832d36a-a6b7-4033-b45b-c30a5519f097 · outbound

This paper cites Deja vu: Contextual sparsity MANUSCRIPT 35 for efficient llms at inference time,.

Deploying Foundation Model Powered Agent Services: A Survey Deja vu: Contextual sparsity MANUSCRIPT 35 for efficient llms at inference time,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.118529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.118529Z digest=sha256:6d15955e8d7b7c47da9f77cd835bbe74769c654bfe92689efec782946a99f4ac

Observation 2a5fadc7-b381-441b-b00e-acea44634311 · outbound

This paper cites LLM in a flash: Efficient Large Language Model Inference with Limited Memory.

Deploying Foundation Model Powered Agent Services: A Survey LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.128373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.128373Z digest=sha256:b415e833b5e20f19763705e528c2d045e4f39a7dd5d94fc34824f5d533a3cec5

Observation a79fbb82-f456-40d1-b78c-a8626196ad55 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Deploying Foundation Model Powered Agent Services: A Survey Fast Transformer Decoding: One Write-Head is All You Need

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.137019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.137019Z digest=sha256:a77c64239f44d1a1522c0d5561e3892c37f8dd53159e04c7e38f130730cfa79b

Observation 3747b98c-656e-4eef-a35c-82d65972b90e · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints,.

Deploying Foundation Model Powered Agent Services: A Survey Gqa: Training generalized multi-query transformer models from multi-head checkpoints,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.143695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.143695Z digest=sha256:da1313b83624a0598c9b8ae777fe61a98912c5fe19bd8d7c96ccef56ee99d91e

Observation 72700bc9-a544-44a5-8d29-d84ad4447e8f · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Deploying Foundation Model Powered Agent Services: A Survey Efficient memory management for large language model serving with pagedattention,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.151153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.151153Z digest=sha256:66a87b7cf52a897822f3f067517787e7a652d6946bbf660460a562ed0d02813a

Observation d7ca7edd-d7ea-4fd8-8611-14655cb66182 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness,.

Deploying Foundation Model Powered Agent Services: A Survey Flashattention: Fast and memory-efficient exact attention with io-awareness,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.157005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.157005Z digest=sha256:13d039457a12e03b0f152969b2a2fb26a0bfe6eb3cba4e7df07b2b9f4b9762ed

Observation 81063395-0423-41e5-a02b-197af16a1802 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Deploying Foundation Model Powered Agent Services: A Survey FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.164356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.164356Z digest=sha256:65011ad470a0078c2e7b5cec7687bf5de8aefa2e913a1c3dda9c03b4b5820b86

Observation b2d271f7-4f40-4f20-acc7-8e835381542a · outbound

This paper cites FlashDecoding++: Faster Large Language Model Inference on GPUs.

Deploying Foundation Model Powered Agent Services: A Survey FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.171490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.171490Z digest=sha256:42b14dcae9b5122930ee50caa25502a9e4acef81c7ba41560ef5539249098cf1

Observation 89d0c1cc-4114-44ce-963c-016e6aa6f9fa · outbound

This paper cites Bminf: An efficient toolkit for big model inference and tuning,.

Deploying Foundation Model Powered Agent Services: A Survey Bminf: An efficient toolkit for big model inference and tuning,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.182893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.182893Z digest=sha256:a85d85bc85dad8bb9b1a7523d8b4f6b65bca1ceb77968cb4315b8e993dbd2e54

Observation a3b902e7-1f47-49c9-bd53-b728007f5130 · outbound

This paper cites Splitwise: Efficient generative LLM inference using phase splitting.

Deploying Foundation Model Powered Agent Services: A Survey Splitwise: Efficient generative LLM inference using phase splitting

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.192998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.192998Z digest=sha256:bb49acf54c7e27a9d4476a3f101d4a57cc0be41a85c1334035a859f62ab5a958

Observation d86ed139-674d-41c4-8fcf-2198a5f43e87 · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

Deploying Foundation Model Powered Agent Services: A Survey Fast Distributed Inference Serving for Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.199950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.199950Z digest=sha256:1c07de89b4b17fbbe1f7a9b8b93e383b7d704181f038ff1340a89257fb391b4f

Observation 5571c984-08c6-4e1f-9b6a-89dce54e8ef9 · outbound

This paper cites Specinfer: Accelerating generative large language model serving with speculative inference and token tree verification,.

Deploying Foundation Model Powered Agent Services: A Survey Specinfer: Accelerating generative large language model serving with speculative inference and token tree verification,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.206519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.206519Z digest=sha256:6f7cbb47e117781d577e194f867b67ef5e7314fe2b2a69c03f8e664f9cd73d97

Observation 31ee100c-7934-4068-9c2a-4b4055b25329 · outbound

This paper cites LLMCad: Fast and Scalable On-device Large Language Model Inference.

Deploying Foundation Model Powered Agent Services: A Survey LLMCad: Fast and Scalable On-device Large Language Model Inference

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.214235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.214235Z digest=sha256:d75258746e8a9c3c76467141ad2776d74b78e3e868540fc99e018dc56fd59ff4

Observation 8694023e-a2d2-4efa-9872-f6cbd7e881db · outbound

This paper cites Deepspeed-moe: Advancing mixture- of-experts inference and training to power next-generation ai scale,.

Deploying Foundation Model Powered Agent Services: A Survey Deepspeed-moe: Advancing mixture- of-experts inference and training to power next-generation ai scale,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.224317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.224317Z digest=sha256:4f794eb85b7925e0d006a1e356dec8a0fc7799f572a892d2849693900a9a8dfc

Observation 07db7bd8-db2d-4a4b-a1ab-19d0b9b6def2 · outbound

This paper cites EdgeMoE: Empowering Sparse Large Language Models on Mobile Devices.

Deploying Foundation Model Powered Agent Services: A Survey EdgeMoE: Empowering Sparse Large Language Models on Mobile Devices

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.233641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.233641Z digest=sha256:4728a161e7d8427fff37017af19616ee1f3a7d76264f30176590c29ba2150f4c

Observation 05763e97-7237-497d-8947-ffbd0783a2d1 · outbound

This paper cites Fast Inference of Mixture-of-Experts Language Models with Offloading.

Deploying Foundation Model Powered Agent Services: A Survey Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.244176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.244176Z digest=sha256:3fefc85824dec310f1ae6107f03d652df34d2c300d1292c7d5f1289e5445c126

Observation b8ffe746-a7fa-4456-b87f-0f53b7471532 · outbound

This paper cites MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache.

Deploying Foundation Model Powered Agent Services: A Survey MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.250980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.250980Z digest=sha256:328a4da743bf0cabc3354191514446889f76467e54d5e520d44aa63c42ac57c4

Observation ffd3fa2d-16e1-4bf8-b077-17f94695450d · outbound

This paper cites Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models.

Deploying Foundation Model Powered Agent Services: A Survey Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.263239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.263239Z digest=sha256:4ea8ec71c9bed5be97b932888a2c362d5041b51fe93a7c7d331ea2006377a1b1

Observation 14284966-53a6-48d0-bcf5-ea3fbd966e5c · outbound

This paper cites Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference.

Deploying Foundation Model Powered Agent Services: A Survey Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.279321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.279321Z digest=sha256:fd741ef12eeba7f3635a2b7a3ca43f2a04d10e2e349457b4073f43455b71a6e0

Observation 861938d7-6021-4c73-925d-0e00fa6017a8 · outbound

This paper cites Intelligence-endogenous management platform for computing and network convergence,.

Deploying Foundation Model Powered Agent Services: A Survey Intelligence-endogenous management platform for computing and network convergence,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.291684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.291684Z digest=sha256:4ab96dade974da1fd328e889a54f3502f68dac1e9e054b8f25f03cfbd6266fa1

Observation 01cd9dab-518c-438f-8e77-076f2480bfad · outbound

This paper cites Resource Allocation in Large Language Model Integrated 6G Vehicular Networks.

Deploying Foundation Model Powered Agent Services: A Survey Resource Allocation in Large Language Model Integrated 6G Vehicular Networks

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.300538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.300538Z digest=sha256:96fe8a6acda5ef8a06386fa0ac169cc8b45cb84c9ca5769e70f6473391b942f2

Observation 732159fa-0d3a-4387-926b-34cf931f7aaa · outbound

This paper cites LinguaLinked: A Distributed Large Language Model Inference System for Mobile Devices.

Deploying Foundation Model Powered Agent Services: A Survey LinguaLinked: A Distributed Large Language Model Inference System for Mobile Devices

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.307664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.307664Z digest=sha256:72b0d52560d44139db82a4e3b099f25d4be2f3210538ac060cab70bf3c350674

Observation f77681b2-d7a4-4b67-afc8-0caeb7c5479d · outbound

This paper cites {MegaScale}: Scaling large language model training to more than 10,000 {GPUs},.

Deploying Foundation Model Powered Agent Services: A Survey {MegaScale}: Scaling large language model training to more than 10,000 {GPUs},

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.323472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.323472Z digest=sha256:ec78ee94ce5c9e486826fdd808ef92088313560a474eb88349344c239c12e938

Observation efb56433-4b92-4de0-a47c-80068e1bf697 · outbound

This paper cites LOSP: Overlap synchronization parallel with local compensation for fast dis- tributed training,.

Deploying Foundation Model Powered Agent Services: A Survey LOSP: Overlap synchronization parallel with local compensation for fast dis- tributed training,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.330947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.330947Z digest=sha256:04a509be89e234c7bd391aae4558b70eb6f4ec3f08e30cb2eb665a90ee3cb441

Observation 8867df43-c025-4684-909d-8fdcb7c72fc3 · outbound

This paper cites Heterogeneous semantic and bit communications: A semi-noma scheme,.

Deploying Foundation Model Powered Agent Services: A Survey Heterogeneous semantic and bit communications: A semi-noma scheme,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.342869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.342869Z digest=sha256:d70893509ecd9fda5fb8d5f2d8e90ae98aa8f3d741f99eca731b644a182c41dc

Observation 5b693c08-7272-44e3-801e-307799372b04 · outbound

This paper cites Computing networks enabled semantic communications,.

Deploying Foundation Model Powered Agent Services: A Survey Computing networks enabled semantic communications,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.358217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.358217Z digest=sha256:504312ee2f0616f5cb77f806e22427567286c8f2cf4cf3ae564749f0c51b4b67

Observation ca6f87d5-ed48-48c3-b37d-5f7c7f2e949f · outbound

This paper cites ggerganov/llama.cpp: Port of facebook’s llama model in c/c++.

Deploying Foundation Model Powered Agent Services: A Survey ggerganov/llama.cpp: Port of facebook’s llama model in c/c++

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.366842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.366842Z digest=sha256:f43d30b45a7698a8dacadc2699e9c022a59ac19efa56d4cde7fb1162d0ead16a

Observation 355bf72e-8e52-436a-957a-85c990b00ab7 · outbound

This paper cites MLC-LLM,.

Deploying Foundation Model Powered Agent Services: A Survey MLC-LLM,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.378727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.378727Z digest=sha256:69c82e67d0bb688699ced522bca7a76e851d06f686c1eb5b95559f5c988aab48

Observation 98ccf142-2d37-402e-839a-b6cda2190e1f · outbound

This paper cites mnn-llm: llm deploy project based mnn.

Deploying Foundation Model Powered Agent Services: A Survey mnn-llm: llm deploy project based mnn

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.386862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.386862Z digest=sha256:97beab1bde19c5874d9ad01d8bce9716df66c9dd938c7fdd1172f2ac0aa18233

Observation 3c0dadb3-33c4-4dc4-af41-ded901776204 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

Deploying Foundation Model Powered Agent Services: A Survey Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.399188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.399188Z digest=sha256:a0b0263056a2c880c3f1a0cebd6ef8ec118781a9970d5ae72ea981cc0f9b02a5

Observation c4b2c57f-13c2-48b3-904f-b38640a1a19a · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,.

Deploying Foundation Model Powered Agent Services: A Survey Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.407081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.407081Z digest=sha256:4cb7447a8008d76e78c20adfc10c520a3aa41013694db0e0c1a9eb7eb54341ad

Observation 76dd67d8-c208-490e-aa92-810664184221 · outbound

This paper cites Openvino deep learning workbench: Comprehensive analysis and tuning of neural networks inference,.

Deploying Foundation Model Powered Agent Services: A Survey Openvino deep learning workbench: Comprehensive analysis and tuning of neural networks inference,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.417285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.417285Z digest=sha256:34032aa2a6b7bf69396b76790cee51b57544da62a89309010e71bd3cd65b639a

Observation b6b6abe1-10ed-42f3-a287-6dd4b7d766ba · outbound

This paper cites mllm is a fast and lightweight multimodal llm inference engine for mobile and edge devices.

Deploying Foundation Model Powered Agent Services: A Survey mllm is a fast and lightweight multimodal llm inference engine for mobile and edge devices

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.428842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.428842Z digest=sha256:ca8dbb5433f9c1a59aec9201f3902b8a6778d91d6492e5e3d059ceccf2bcd2fd

Observation a60b3c2a-b243-4907-8aaa-b2139f5b397e · outbound

This paper cites FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design.

Deploying Foundation Model Powered Agent Services: A Survey FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.435290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.435290Z digest=sha256:bcc0dd059ce0276145ad8bd6f6dc8e88009d2a4f3f8e50c9f2bb74147a93c3ce

Observation dad5ea92-7201-457d-948a-d4a6b5fbf9e8 · outbound

This paper cites Colossal-ai: A unified deep learning system for large-scale parallel training,.

Deploying Foundation Model Powered Agent Services: A Survey Colossal-ai: A unified deep learning system for large-scale parallel training,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.442991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.442991Z digest=sha256:0e4144276786f4f5d7aafdf957ffe5c80122ed7d2a990bd089bd83301e2e4db4

Observation bcbcdd4f-44a3-46f7-8ca4-ededf8da6f03 · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm,.

Deploying Foundation Model Powered Agent Services: A Survey Efficient large-scale language model training on gpu clusters using megatron-lm,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.451469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.451469Z digest=sha256:da9cd4f11d4f721c7c8441d2f9a487ad77495c4a6dbe62fc1150553e9d69c034

Observation 96fc2a21-cc3a-4522-9d46-b0e91f7abd4d · outbound

This paper cites A tensorrt toolbox for optimized large language model inference,.

Deploying Foundation Model Powered Agent Services: A Survey A tensorrt toolbox for optimized large language model inference,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.459594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.459594Z digest=sha256:00311e061382073510d9b079a493816456cb3281bd7a15cf1bc71732c5c111ec

Observation b23bdc69-310f-4295-a139-8d59289bf32c · outbound

This paper cites Langchain,.

Deploying Foundation Model Powered Agent Services: A Survey Langchain,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.466233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.466233Z digest=sha256:5b5f2bfd82546aedc2ab915f6ddacca4224df68b7420849dfa425d0fe5a9a71e

Observation a3d30d45-072f-4778-8262-918ba436cc0d · outbound

This paper cites Parrot: Efficient Serving of LLM-based Applications with Semantic Variable,.

Deploying Foundation Model Powered Agent Services: A Survey Parrot: Efficient Serving of LLM-based Applications with Semantic Variable,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.471857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.471857Z digest=sha256:cd45c5363aae4686e50766f6fe86902ff43542aa4da71e17641fd4eb076d3779

Observation 8b942c15-3a46-48dc-9f52-a0090250d19e · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Pro- grams,.

Deploying Foundation Model Powered Agent Services: A Survey SGLang: Efficient Execution of Structured Language Model Pro- grams,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.478780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.478780Z digest=sha256:51c8a77d9ba90a64a4c0d01d0b348ecff74c2396ac6c00e3d1a1449d2476b76a

Observation ac887d7e-c96a-413e-9f9a-e6ee15b42898 · outbound

This paper cites Clipper: A {Low-Latency} online prediction serving system,.

Deploying Foundation Model Powered Agent Services: A Survey Clipper: A {Low-Latency} online prediction serving system,

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.484686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.484686Z digest=sha256:e0c199e73a723fc2465aecc552960ca384fad320a8d68c81cb97608086a0b85c

Observation 8c75d58a-cda4-4079-844a-492d2e914a6c · outbound

This paper cites {MArk}: Exploiting cloud services for {Cost-Effective},{SLO-Aware} machine learning infer- ence serving,.

Deploying Foundation Model Powered Agent Services: A Survey {MArk}: Exploiting cloud services for {Cost-Effective},{SLO-Aware} machine learning infer- ence serving,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.491891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.491891Z digest=sha256:eaa4260a324399f849af6a5396409d51ff13327b88faaee88dcb12f1ae782115

Observation 030d2edc-7d65-408f-9de6-48f87d81d57a · outbound

This paper cites Nexus: A GPU cluster engine for accelerating DNN-based video analysis,.

Deploying Foundation Model Powered Agent Services: A Survey Nexus: A GPU cluster engine for accelerating DNN-based video analysis,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.503290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.503290Z digest=sha256:df15918eae04751d4be94aec671c722ea242ac76b0c168d6d1394f51d5b22522

Observation ebeae4ce-07a8-4d6e-9af6-25ef18f0038c · outbound

This paper cites InferLine: latency-aware provisioning and scaling for prediction serving pipelines,.

Deploying Foundation Model Powered Agent Services: A Survey InferLine: latency-aware provisioning and scaling for prediction serving pipelines,

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.512310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.512310Z digest=sha256:bd44d9652e44b68f81f99c734d50b6faa0022c6f915b4f76586bf514898f0fb9

Observation 7ec38905-e53b-4d14-93f1-d35f6ff26610 · outbound

This paper cites Serving {DNNs} like clockwork: Performance predictability from the bottom up,.

Deploying Foundation Model Powered Agent Services: A Survey Serving {DNNs} like clockwork: Performance predictability from the bottom up,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.523724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.523724Z digest=sha256:d6c212f34e550d9ee3998feaf1befc9d09fb29ce3f5620968152ac11c8b2ca6f

Observation 12c97abf-1e7e-4adf-b179-3bb11cb4d26a · outbound

This paper cites {INFaaS}: Automated model-less inference serving,.

Deploying Foundation Model Powered Agent Services: A Survey {INFaaS}: Automated model-less inference serving,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.535944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.535944Z digest=sha256:3fb85bb6574987f88b9a2fff6a685637eab52dce2b0f19efa7d0614143d7ff81

Observation 79c828ba-4997-41a6-9011-2abe536cc5ef · outbound

This paper cites Morphling: Fast, near-optimal auto-configuration for cloud- native model serving,.

Deploying Foundation Model Powered Agent Services: A Survey Morphling: Fast, near-optimal auto-configuration for cloud- native model serving,

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.548707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.548707Z digest=sha256:e921f527f04e46cc23747beb5079ce609de9e37a83b2dc5663b78e0619a08ed5

Observation f59425d5-cc09-42a4-b12e-e8e0c3df4370 · outbound

This paper cites Cocktail: A multidimensional optimization for model serving in cloud,.

Deploying Foundation Model Powered Agent Services: A Survey Cocktail: A multidimensional optimization for model serving in cloud,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.561212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.561212Z digest=sha256:4ce44e99a5d43f0bcc98fe673ce848c4b888b4b8257c02dd7840d71003cf84c2

Observation 29ff4f30-31c1-442b-9fa4-ac8997312935 · outbound

This paper cites Kairos: Building cost- efficient machine learning inference systems with heterogeneous cloud resources,.

Deploying Foundation Model Powered Agent Services: A Survey Kairos: Building cost- efficient machine learning inference systems with heterogeneous cloud resources,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.571495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.571495Z digest=sha256:3ae6afe80201ab635754bdc170886b6432d03a27843d14cbd1b7bd8983d48cc0

Observation 7b50d3ce-3ab0-4234-a896-ecbe3508f2a7 · outbound

This paper cites {SHEPHERD}: Serving {DNNs} in the wild,.

Deploying Foundation Model Powered Agent Services: A Survey {SHEPHERD}: Serving {DNNs} in the wild,

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.577939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.577939Z digest=sha256:101657e3be3c7bb0e1b57c9c3ec5e1943d229dec892bd05a60e17d7d6f6897fc

Observation 07bab909-0504-446c-9e35-980890ebb3d3 · outbound

This paper cites SpotServe: Serving Generative Large Language Models on Preemptible Instances.

Deploying Foundation Model Powered Agent Services: A Survey SpotServe: Serving Generative Large Language Models on Preemptible Instances

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.586601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.586601Z digest=sha256:3b977eee68298fbeedd11d24f38e47fb4f9884b97a819ae6f0d66b71e2b773b5

Observation 6f08276c-940a-4d38-85f2-0977bf46de00 · outbound

This paper cites Frequency resource allocation and interference management in mobile edge com- puting for an Internet of Things system,.

Deploying Foundation Model Powered Agent Services: A Survey Frequency resource allocation and interference management in mobile edge com- puting for an Internet of Things system,

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.596352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.596352Z digest=sha256:9170b358c27750bdb0447cf3f7b4a3a9d8daeac7cf21517abff19c3640e8917a

Observation 7b68c0dc-42db-45fb-a4aa-edffeab85ec5 · outbound

This paper cites Decentralized resource auctioning for latency-sensitive edge computing,.

Deploying Foundation Model Powered Agent Services: A Survey Decentralized resource auctioning for latency-sensitive edge computing,

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.606093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.606093Z digest=sha256:8f6c0d00710e2387613d3ca354bf337f98d1f09472a99350cf90292b4d6227f2

Observation 84201086-57d4-4728-b300-9db4c5e69dcd · outbound

This paper cites Joint computation partitioning and resource allocation for latency sensitive applications in mobile edge clouds,.

Deploying Foundation Model Powered Agent Services: A Survey Joint computation partitioning and resource allocation for latency sensitive applications in mobile edge clouds,

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.618571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.618571Z digest=sha256:bb075b528c4f6e209705efac8412832552cbe7ed55f1d168189d65002df25d3f

Observation 4724aa46-9daf-4365-9e04-376fe80dc311 · outbound

This paper cites Adaptive computation offloading and resource allocation strategy in a mobile edge computing environment,.

Deploying Foundation Model Powered Agent Services: A Survey Adaptive computation offloading and resource allocation strategy in a mobile edge computing environment,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.629199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.629199Z digest=sha256:b4f057f33ff49427e7fb8d3d4214350ddffadd19cb2f57a6578c0306b15dd2dc

Observation 7c911632-5049-4b4d-9ba5-b8a73f4032fc · outbound

This paper cites Resource allocation based on deep reinforcement learning in IoT edge computing,.

Deploying Foundation Model Powered Agent Services: A Survey Resource allocation based on deep reinforcement learning in IoT edge computing,

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.639635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.639635Z digest=sha256:0f67b3e53629267d990982feea722dc5b239246490bd7dd69b36aa29f009b279

Observation dd23b577-6374-47c8-88ce-3c09a1d6d2d5 · outbound

This paper cites CE-IoT: Cost-effective cloud- edge resource provisioning for heterogeneous IoT applications,.

Deploying Foundation Model Powered Agent Services: A Survey CE-IoT: Cost-effective cloud- edge resource provisioning for heterogeneous IoT applications,

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.655729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.655729Z digest=sha256:e95d64248a9ef876f085af99064fbc92519167346843d4a7106df04f215f164f

Observation 5f2b4c56-efc9-4244-9b74-fa68f59b551f · outbound

This paper cites Dynamic resource allocation and computation offloading for IoT fog computing system,.

Deploying Foundation Model Powered Agent Services: A Survey Dynamic resource allocation and computation offloading for IoT fog computing system,

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.667356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.667356Z digest=sha256:fd49361cd7088eb8d1219c90f8ff1b14c6481a7048ee089844dc360633372eaa

Observation 22f70170-8dd8-47b5-95ed-e6cd26485d8a · outbound

This paper cites Lass: Running latency sensitive serverless computations at the edge,.

Deploying Foundation Model Powered Agent Services: A Survey Lass: Running latency sensitive serverless computations at the edge,

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.679429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.679429Z digest=sha256:773fc2a1923d757e652fcbd854ec6edd2a6924a063396432874947045547e768

Observation d0cabc15-1ee0-429e-89e6-962312b923ca · outbound

This paper cites Resource provisioning and allocation in function- as-a-service edge-clouds,.

Deploying Foundation Model Powered Agent Services: A Survey Resource provisioning and allocation in function- as-a-service edge-clouds,

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.686643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.686643Z digest=sha256:15953c8c7eae1aecd3b8d85331437ab7d184dab71947522106ce4fc2eec83f39

Observation cb3aa6b7-dbae-48d6-ba56-9f0d6eb62aa8 · outbound

This paper cites KneeScale: Efficient resource scaling for serverless computing at the edge,.

Deploying Foundation Model Powered Agent Services: A Survey KneeScale: Efficient resource scaling for serverless computing at the edge,

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.699402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.699402Z digest=sha256:8eaf7a68ae60f90dc1b552d51ec7eaee5b084739662902243da763158ca063fc

Pith citing papers

Observation 767d2de8-bf3f-4ff9-b7a4-41c62355cd52 · inbound

Enhancing RAG with Active Learning on Conversation Records: Reject Incapables and Answer Capables cites this paper.

Enhancing RAG with Active Learning on Conversation Records: Reject Incapables and Answer Capables Deploying Foundation Model Powered Agent Services: A Survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T22:47:39.121487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:47:39.121487Z digest=sha256:47bc7ed9038238396bb22336d2138d184af43f1ab0fdf355c293424c24fa5e59

Observation a9a7292d-627d-4ca3-b01f-acf9078e605a · inbound

LLM-Powered AI Agent Systems and Their Applications in Industry cites this paper.

LLM-Powered AI Agent Systems and Their Applications in Industry Deploying Foundation Model Powered Agent Services: A Survey

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:06:38.134799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T14:05:54.535411Z digest=sha256:7e612c8eaec58dcf3b3fb89c6679ef0b35f06fb91d27786cc5a8641b41357c14