Pith. sign in

Paper Citation Record · LEDGER

EfficientLLM: Efficiency in Large Language Models

As of 18 August 2026, this Paper Citation Record lists 100 of 299 outbound references and 3 inbound Pith citation observations for arXiv:2505.13840.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13840 v1

Coverage vector

measured 100 of 299 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:13:35.213663Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:00:53.896063Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T23:40:51.690711Z

Reference resolution

100 of 299 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5f948e72-2454-4c13-abdf-ce21234e286f · outbound

This paper cites Language models are few-shot learners.

EfficientLLM: Efficiency in Large Language Models Language models are few-shot learners

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:32.775291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:32.775291Z digest=sha256:1c89e8df7f558d20a3272bcde65fdf189e66149b8e3e027fb898330ab221bd68

Observation 951f3895-dc47-4662-a830-0c2f960bb8c2 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

EfficientLLM: Efficiency in Large Language Models PaLM: Scaling Language Modeling with Pathways

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:32.843859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:32.843859Z digest=sha256:42a2f3457c733164ab124fe17dcb712013523192604b370e3680c35ad8f1e597

Observation 02bca011-5edc-4131-a34e-2dea500c45a8 · outbound

This paper cites Scaling Laws for Neural Language Models.

EfficientLLM: Efficiency in Large Language Models Scaling Laws for Neural Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:32.910913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:32.910913Z digest=sha256:cfdea577107d675ea86464a8601946c780317e8156714f1bdcc458b209291f3b

Observation a23588cd-a102-4764-b917-f66df29ce839 · outbound

This paper cites Training Compute-Optimal Large Language Models.

EfficientLLM: Efficiency in Large Language Models Training Compute-Optimal Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:32.915706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:32.915706Z digest=sha256:b6279b7d581d3c6988db3c13dfdd401ceb65f761ee561e9c92ea7b64592c8209

Observation 4cc69ba7-d955-4e55-ad42-fe9943bf5698 · outbound

This paper cites Automating Customer Service using LangChain: Building custom open-source GPT Chatbot for organizations.

EfficientLLM: Efficiency in Large Language Models Automating Customer Service using LangChain: Building custom open-source GPT Chatbot for organizations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:32.920235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:32.920235Z digest=sha256:39bac8f02616f5046f41c6dbbb4265e43538ff8b58d6c30151abe6fe8eb4b5f2

Observation 5191e7f7-57e8-470a-9d65-1fb5189b0623 · outbound

This paper cites LitLLM: A Toolkit for Scientific Literature Review.

EfficientLLM: Efficiency in Large Language Models LitLLM: A Toolkit for Scientific Literature Review

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:32.925981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:32.925981Z digest=sha256:329fbc6b475a09ee72801b0e4f9db5562f9d6e92d74d8ef27dabdf06a1ab3af3

Observation b06304d5-da3b-4db4-9180-6dbade87b3ed · outbound

This paper cites A general purpose device for interaction with llms.

EfficientLLM: Efficiency in Large Language Models A general purpose device for interaction with llms

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:32.931131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:32.931131Z digest=sha256:16189bae335091a1325cdc6c0a878043e31a410cd7355f59e3718c99cf904b0b

Observation 74260f2c-8fe1-4863-b61f-1e373c9e69d5 · outbound

This paper cites Large language models are zero-shot reasoners.

EfficientLLM: Efficiency in Large Language Models Large language models are zero-shot reasoners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:32.934958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:32.934958Z digest=sha256:ffe0b5071c15b7591be0df7171d0eab6aae2d85a2420bc12bd8dcf07a3f9ec40

Observation 408748da-a0d3-47e4-b650-78ba41069bf1 · outbound

This paper cites Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment.

EfficientLLM: Efficiency in Large Language Models Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:32.981763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:32.981763Z digest=sha256:386a1ca6a9e2986bc8c8a0f300bbed296dea500a23d0549095a9f11a0dd376b0

Observation 247275bc-a738-4160-b06e-2e7f7d53bb81 · outbound

This paper cites an unresolved cited work.

EfficientLLM: Efficiency in Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.043865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.043865Z digest=sha256:8e1dd083a13b63f2a144d5cd2389dcca854d4fde49f373552922d4a80137ef4a

Observation b6ebfa13-ac11-4a88-a8eb-03c59e1fc13b · outbound

This paper cites On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective.

EfficientLLM: Efficiency in Large Language Models On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.088459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.088459Z digest=sha256:32b093f1c0186d6c5d0769c33e072d7452301d35f4ab03fbff2b7f3e2a36d837

Observation 15eeec11-219c-4574-82bc-fd8a7ed95118 · outbound

This paper cites Energy and policy considerations for deep learning in NLP.

EfficientLLM: Efficiency in Large Language Models Energy and policy considerations for deep learning in NLP

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.093273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.093273Z digest=sha256:3fa15eda852c8e449ccba5a9e5bf4747936cb574cc2c7911fe03207b3119e534

Observation f0633638-b793-4dd2-aee3-0c481a5c102d · outbound

This paper cites Efficient Large Language Models: A Survey.

EfficientLLM: Efficiency in Large Language Models Efficient Large Language Models: A Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.097404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.097404Z digest=sha256:765a08e3eeb87e2ccb6bc6f3e474b4f2eb03704fef2a4c4fd5a58acce6318d31

Observation 37be1ec0-ffc4-4329-ae96-7dcbefb0419f · outbound

This paper cites Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models.

EfficientLLM: Efficiency in Large Language Models Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.166455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.166455Z digest=sha256:dfe8da5139660489146ffd76d24a6f97a3708229214491f9d5352680caa6d44d

Observation 1c109bf2-a736-417c-be82-669126288436 · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

EfficientLLM: Efficiency in Large Language Models A Survey on Efficient Inference for Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.206951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.206951Z digest=sha256:f80fbd58891375fb4d74fdb8d7b7996c136a987f82189b422f52f4c309bb41f4

Observation d10b0f08-0a2a-4872-bf71-a9e2d99396a6 · outbound

This paper cites Compressing context to enhance inference efficiency of large language models.

EfficientLLM: Efficiency in Large Language Models Compressing context to enhance inference efficiency of large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.211852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.211852Z digest=sha256:a29bc865367d37e8fb8c113ac8b0d5d69cd9c2747d5decb5ce1abb014a498bc4

Observation 4dc62e1a-1d94-472f-b7eb-d485e0246dd6 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

EfficientLLM: Efficiency in Large Language Models Generating Long Sequences with Sparse Transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.361427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.361427Z digest=sha256:67c605f3ac581219427ae517d17eed0ba1cd10381443dfae7c86e97e52946623

Observation e6f25923-976f-4631-985f-4a2a2020d564 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

EfficientLLM: Efficiency in Large Language Models GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.365768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.365768Z digest=sha256:18629db0c6f854e6918543815d253a9e61886f2cff9cee611894f6d2c30a8765

Observation 2390607b-fed6-464f-b818-b95805d48321 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

EfficientLLM: Efficiency in Large Language Models Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.441903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.441903Z digest=sha256:7b401c0058303b46ac75c5b859b8245549c758ca90c6f9ff8756e42864a41c0a

Observation e951f492-9ba3-4c93-b26f-6a2b2bc62811 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

EfficientLLM: Efficiency in Large Language Models Fast Transformer Decoding: One Write-Head is All You Need

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.447216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.447216Z digest=sha256:1f908750b88a6a4644fa37dc629bb4b65ebdf31642e0b799640dc167eacfa17e

Observation fa138352-6cfd-4fc5-89c2-bd8b7c6ab0b9 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

EfficientLLM: Efficiency in Large Language Models GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.452412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.452412Z digest=sha256:5931c6f1a1792b5c0bb4bd2d519cb25dd9fc8f0179e340d25a78ffd2c3aa66ed

Observation 21977f1d-bd59-4d33-b735-23abe8c52859 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

EfficientLLM: Efficiency in Large Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.455821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.455821Z digest=sha256:13ca217113b79b9573205ad3a2736b96d393fbf7de7a25d61975f23185a9c792

Observation ec3b582f-70ec-4398-b36b-27601724498f · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

EfficientLLM: Efficiency in Large Language Models Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.533836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.533836Z digest=sha256:201b382678c858d57d1c3b0ef28bc2f8537eb0a57783d7c38de6def26718859d

Observation b0c35916-e686-4cbf-a2d5-df5727359c6a · outbound

This paper cites Mixed precision training.

EfficientLLM: Efficiency in Large Language Models Mixed precision training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.545753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.545753Z digest=sha256:96590db59db2e351afd1c818b3e41ac1e1bc3d89eeaa973c8785883a32c928e2

Observation 9904abd2-ea3b-4b79-a686-a42188455d68 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

EfficientLLM: Efficiency in Large Language Models Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.550464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.550464Z digest=sha256:5a2cf0347977fdc6c561afa2c206b6290ee417dfb8c2ed01b4bd1b1a5bb2ab01

Observation e7df93f8-7814-47f1-82cb-3ba198c422ba · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

EfficientLLM: Efficiency in Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.553949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.553949Z digest=sha256:1d483f9853e20e8b8584dcb580873a068b38e11b69d8319b2640e6419cc3fd4e

Observation 840d8869-d794-428a-994e-6fad304411c3 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

EfficientLLM: Efficiency in Large Language Models Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.558857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.558857Z digest=sha256:b4c3407ff0434fe14d0564b49b29487b45d72f5e38da0a17f353533178a4dd00

Observation 5dea72ef-8714-42f3-9235-168769055cfb · outbound

This paper cites FedPara: Low-Rank Hadamard Product for Communication-Efficient Federated Learning.

EfficientLLM: Efficiency in Large Language Models FedPara: Low-Rank Hadamard Product for Communication-Efficient Federated Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.658558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.658558Z digest=sha256:17adb6ea7d02de5f88085ce29ccf538f052f22ae0c137fad891d3d29ac4034c4

Observation f6ac40c6-22ff-4af7-97dc-34c23c1d8637 · outbound

This paper cites KronA: Parameter Efficient Tuning with Kronecker Adapter.

EfficientLLM: Efficiency in Large Language Models KronA: Parameter Efficient Tuning with Kronecker Adapter

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.758621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.758621Z digest=sha256:df55a7b064b3548e6f3f8c8526cd9925a72e62ff4dece16b4b1893f022dcbdf3

Observation c941cb12-03bc-40fd-80f9-abb6e1ab8426 · outbound

This paper cites One-for-All: Generalized LoRA for Parameter-Efficient Fine-tuning.

EfficientLLM: Efficiency in Large Language Models One-for-All: Generalized LoRA for Parameter-Efficient Fine-tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.762904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.762904Z digest=sha256:373d3db4fc4bb2033d37c773af16f48f513f73fc95322087e62c46255a691953

Observation 96279bc5-09a3-455b-ac71-acd20eebb4fe · outbound

This paper cites LLM.int8() : 8-bit matrix multiplication for transformers at scale.

EfficientLLM: Efficiency in Large Language Models LLM.int8() : 8-bit matrix multiplication for transformers at scale

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.766675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.766675Z digest=sha256:8ab7b889b7f0023d4f42f7183581b899dd7c743bb1466807b874a50b2ef7a6c1

Observation 225c4292-cdb3-48d4-be87-ea37c0ad57dd · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

EfficientLLM: Efficiency in Large Language Models GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.770083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.770083Z digest=sha256:5532301f3d4ca94b30b07406539d17a55ee835d5632c3be13db7307650f1d36c

Observation c2f8f926-35f7-47dc-bf34-f4c19ea0df25 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms, 2023.

EfficientLLM: Efficiency in Large Language Models Qlora: Efficient finetuning of quantized llms, 2023

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.774000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.774000Z digest=sha256:c1644f0d263a224ae27bcb15251e15164278d97f1a4d423d8f101e95012e175d

Observation d6bceeee-f504-4867-b899-29df2411ef45 · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot, 2023.

EfficientLLM: Efficiency in Large Language Models Sparsegpt: Massive language models can be accurately pruned in one-shot, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.777703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.777703Z digest=sha256:a255abdb17f1457a43695dab306d8ba592699bc891f69e814d31971c41e8f5ea

Observation c55fd4dc-efab-4214-a4f3-a703fe60e17a · outbound

This paper cites an unresolved cited work.

EfficientLLM: Efficiency in Large Language Models Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.782240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.782240Z digest=sha256:39d4110e480e0a18ed711be3294ac74962cfbf65eea2d60ef023dbeed5e0e5c9

Observation d8c2a781-2262-4634-944b-5c89422d5524 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

EfficientLLM: Efficiency in Large Language Models Distilling the Knowledge in a Neural Network

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.785522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.785522Z digest=sha256:1f0b53558346f3fc956f2e13a921bdb03e9cf425c03116cd22fd10bb30fb42f4

Observation df0086fc-f07b-467a-866c-d1770cdd195a · outbound

This paper cites Fast inference from transformers via speculative decoding.

EfficientLLM: Efficiency in Large Language Models Fast inference from transformers via speculative decoding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.790094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.790094Z digest=sha256:330e52237d6199f108fc657dfdaf9af0f924c085f589e491704f3c3421369d13

Observation a640565f-a4fc-432e-92c8-b0834ef23d95 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

EfficientLLM: Efficiency in Large Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.849882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.849882Z digest=sha256:cbaaaa27bbc9ef8610122012ae3d23b6bdab6e5c69bc5358e5a848b0140abff8

Observation c577c9e2-9d47-4016-9059-d96e05ad15ec · outbound

This paper cites Mixtral of Experts.

EfficientLLM: Efficiency in Large Language Models Mixtral of Experts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:33.962623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:33.962623Z digest=sha256:9a094facc3191c0bec37cf41de61c3ca9846ee8df94e41ba20eae3aa380a031a

Observation 79c369e2-102c-459a-8fb1-89d77455e4b1 · outbound

This paper cites Understanding int4 quantization for language models: latency speedup, composability, and failure cases.

EfficientLLM: Efficiency in Large Language Models Understanding int4 quantization for language models: latency speedup, composability, and failure cases

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.062116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.062116Z digest=sha256:038ff7f9ffb7aac950942af7d75402debddd3d7b3bf52812e579243a4583e115

Observation 03bae4aa-ae5c-4d84-a61b-845ff155eb06 · outbound

This paper cites No free lunch theorems for optimization.

EfficientLLM: Efficiency in Large Language Models No free lunch theorems for optimization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.066106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.066106Z digest=sha256:6452e2d377a299e38534021985b2d70a6ee62896329940a91851f95f33139b5d

Observation a90cfc53-507f-4d48-8026-c824f393effc · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen.

EfficientLLM: Efficiency in Large Language Models Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.070720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.070720Z digest=sha256:0d700fd4a24ca10db4214b5b8a7c74a203e8818b34429e6e85cd88e703a7bbcc

Observation 9026baac-3e40-4ea8-939e-c79763fc8e2e · outbound

This paper cites Dora: Weight-decomposed low-rank adaptation.

EfficientLLM: Efficiency in Large Language Models Dora: Weight-decomposed low-rank adaptation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.075297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.075297Z digest=sha256:32ecb40778536522293115a14ca8c0983ad42c6a623a8143afeacf91dbb1e4ff

Observation 1f45e39c-4c2a-4e45-a147-63087728887f · outbound

This paper cites A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA.

EfficientLLM: Efficiency in Large Language Models A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.078911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.078911Z digest=sha256:de9d1fd2554591559e375e2d3399ee39e08c42bc80e8321382fbac70937be624

Observation 2e05c53b-d997-4503-8c40-9bf323b6d2c6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

EfficientLLM: Efficiency in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.206439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.206439Z digest=sha256:586f61ede2b97cfb84bfeedaddcd0fcd0c6bbba0182de70ec9e3f27d391a7442

Observation 65417263-fabd-470b-96dd-d7ce36ce13ec · outbound

This paper cites Improving language understanding by generative pre-training.

EfficientLLM: Efficiency in Large Language Models Improving language understanding by generative pre-training

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.210360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.210360Z digest=sha256:c371e87bcbef81fc873b6990d0863d663c8a239c99ad9acfb7d07f7040bd47ac

Observation 1d810ba8-f918-4182-944c-015f5303d448 · outbound

This paper cites Language models are unsupervised multitask learners.

EfficientLLM: Efficiency in Large Language Models Language models are unsupervised multitask learners

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.214268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.214268Z digest=sha256:791570c1a6a038945187e1015acf6e4ade475ad46b354f49232b5380ecec8772

Observation 3c7335ed-1e4d-46b1-8ef6-90ed16735ebd · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

EfficientLLM: Efficiency in Large Language Models Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.218165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.218165Z digest=sha256:11b8a1b0d93eb79c2db5dcfeb821172d8e6097161b5ca651f79693ce70305bf7

Observation 381bc3d4-d614-4826-aff7-360610655f70 · outbound

This paper cites Roberta: A robustly optimized bert pretraining approach, 2019.

EfficientLLM: Efficiency in Large Language Models Roberta: A robustly optimized bert pretraining approach, 2019

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.221887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.221887Z digest=sha256:f8d89a336e217fa3c57705cf0d0514b90c30de4bf9542e8976dc63f6452d1a22

Observation 0c332393-9890-4716-b4fb-b1bda9cde52e · outbound

This paper cites ALBERT: A Lite BERT for Self-supervised Learning of Language Representations.

EfficientLLM: Efficiency in Large Language Models ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.225437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.225437Z digest=sha256:c379f7ab4c05c5ba779003ef53af5042426f5f35168450d2bb3459953a1aaf5b

Observation 073f8829-1df3-4323-b9a5-a162af22d78f · outbound

This paper cites Leveraging large language models to enhance personalized recommendations in e-commerce.

EfficientLLM: Efficiency in Large Language Models Leveraging large language models to enhance personalized recommendations in e-commerce

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.229182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.229182Z digest=sha256:5ecac1084a4fe2372ad2631d6a67fb540ca117e6b9358bed2a4492bbc1ccaefc

Observation d280520a-2404-4ef7-b61f-abe840441ce1 · outbound

This paper cites Text Understanding and Generation Using Transformer Models for Intelligent E-commerce Recommendations.

EfficientLLM: Efficiency in Large Language Models Text Understanding and Generation Using Transformer Models for Intelligent E-commerce Recommendations

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.232657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.232657Z digest=sha256:4d5de2a55fb1d60287259055f06bbab4058fb0eb060290a4a335b8fc112ea31d

Observation c20d9a92-a261-4b69-b3cf-00b0ff7f4528 · outbound

This paper cites Recommendation systems in the era of llms.

EfficientLLM: Efficiency in Large Language Models Recommendation systems in the era of llms

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.237407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.237407Z digest=sha256:64b043a5c8fa2eebe06fdce5ffd1031e2d14ffd09bc1b881721f10336ef6c994

Observation a7796d54-75c9-41c3-814d-5872e8654b02 · outbound

This paper cites Recommender systems in the era of large language models (llms).

EfficientLLM: Efficiency in Large Language Models Recommender systems in the era of large language models (llms)

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.240934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.240934Z digest=sha256:ade59419aceb27fa6b7f6bf2d60af0ef05a74246592bb89445883c23c58f91fd

Observation 8ef89c45-6811-4ae3-a8d5-e18aa57f4dc3 · outbound

This paper cites Adapting large language models for education: Foundational capabilities, potentials, and challenges, 2023.

EfficientLLM: Efficiency in Large Language Models Adapting large language models for education: Foundational capabilities, potentials, and challenges, 2023

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.349703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.349703Z digest=sha256:d6562cfe6e0098240fedc90304ca0b3b06380870102dd61f35a69ae69f04b2b6

Observation 4ff14712-aa08-4466-a10a-21da4390b35a · outbound

This paper cites Yu, and Qingsong Wen.

EfficientLLM: Efficiency in Large Language Models Yu, and Qingsong Wen

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.464995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.464995Z digest=sha256:c6a5d7e01992085e0f0dc7c9954aee2d194333198d2c4b3cbcce8c5a77fe1fbd

Observation 58e5de48-de15-48e6-8352-79cf666473ac · outbound

This paper cites Simulating Classroom Education with LLM-Empowered Agents.

EfficientLLM: Efficiency in Large Language Models Simulating Classroom Education with LLM-Empowered Agents

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.468926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.468926Z digest=sha256:e51cd0363f19d96d66fe70520c0b63c358535b234268026275f68dd886402084

Observation e3708e14-7f72-43da-9f57-36a182acf613 · outbound

This paper cites A Survey on Large Language Models for Code Generation.

EfficientLLM: Efficiency in Large Language Models A Survey on Large Language Models for Code Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.472114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.472114Z digest=sha256:c4a5268f53eb49b692dae00f0ed6746019ef965b17d66d2076fd98e4b2e6a5e9

Observation b33b1f37-544e-4701-b782-0c31e4fbf690 · outbound

This paper cites Make Every Move Count: LLM-based High-Quality RTL Code Generation Using MCTS.

EfficientLLM: Efficiency in Large Language Models Make Every Move Count: LLM-based High-Quality RTL Code Generation Using MCTS

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.475984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.475984Z digest=sha256:5c818c4c7ad846e2acf4946e6795de1f652725988d6c3ddd48732f56eed25222

Observation e3982a78-3b50-4d29-9675-97a774b16514 · outbound

This paper cites Mapping the Increasing Use of LLMs in Scientific Papers.

EfficientLLM: Efficiency in Large Language Models Mapping the Increasing Use of LLMs in Scientific Papers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.479364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.479364Z digest=sha256:883e24e086091a9d1702ef5d160cfe0bd4f9d74f270a62ca1ca29db2d30ec6bc

Observation 1a3b47ca-d499-4605-8191-ea0989f16cce · outbound

This paper cites u bler, Jiaji Huang, Matth \.

EfficientLLM: Efficiency in Large Language Models u bler, Jiaji Huang, Matth \

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.483404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.483404Z digest=sha256:44d2eb42f59caef295cfac672bd09ba59f9b1accece4dcbb8c95c27d7ff22583

Observation a31e0ed4-755a-448e-b724-06596ba10dfd · outbound

This paper cites Leveraging Compute-in-Memory for Efficient Generative Model Inference in TPUs.

EfficientLLM: Efficiency in Large Language Models Leveraging Compute-in-Memory for Efficient Generative Model Inference in TPUs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.503630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.503630Z digest=sha256:16d767bebc7a87f7f161a387ed4b74b3f57dc601048885e09acb059170b84f9d

Observation e9680936-20dc-496e-96ca-a5ff1eee8afd · outbound

This paper cites In-datacenter performance analysis of a tensor processing unit.

EfficientLLM: Efficiency in Large Language Models In-datacenter performance analysis of a tensor processing unit

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.613816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.613816Z digest=sha256:994b6ca6b4c3ce24d26b594ba147fd8b02739c80607786f617ae98c4bd4f644d

Observation 3b350739-a135-4548-8fcc-aa8b6a3e3bd4 · outbound

This paper cites A novel neuromorphic processors realization of spiking deep reinforcement learning for portfolio management.

EfficientLLM: Efficiency in Large Language Models A novel neuromorphic processors realization of spiking deep reinforcement learning for portfolio management

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.739702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.739702Z digest=sha256:208384dc672e07eafff5854fd148433e1626ec5b4edb59ca2f285880594297cb

Observation 3bfaf2ef-2dcb-4d92-9b24-39845350dc3a · outbound

This paper cites Machine learning with neuromorphic photonics.

EfficientLLM: Efficiency in Large Language Models Machine learning with neuromorphic photonics

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.769968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.769968Z digest=sha256:e8f117512a796a213a4b96faa6eaceaf245d7229852bb1edf28ff9a158adf35b

Observation 3d5e7229-2610-4f2b-abed-b6b0bcffde8e · outbound

This paper cites Efficient Training of Large Language Models on Distributed Infrastructures: A Survey.

EfficientLLM: Efficiency in Large Language Models Efficient Training of Large Language Models on Distributed Infrastructures: A Survey

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.773701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.773701Z digest=sha256:b50e61a8e75e94afd548199b1bdc8f20852419b5fbafb337e553c2fb4ec85cf8

Observation 965c4839-83cc-45d6-8ba3-d112f4b053bc · outbound

This paper cites A survey on distributed machine learning.

EfficientLLM: Efficiency in Large Language Models A survey on distributed machine learning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.777002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.777002Z digest=sha256:1bb969608f76247dac7e1f34432192cf4a86fa0a7e8b97b55a04d1f1e9cf1812

Observation c1310b56-6ac4-4264-8559-7fe6a711bf47 · outbound

This paper cites A Survey of Model Compression and Acceleration for Deep Neural Networks.

EfficientLLM: Efficiency in Large Language Models A Survey of Model Compression and Acceleration for Deep Neural Networks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.780807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.780807Z digest=sha256:c3d9cc56eef4a3ab517c275441564a56d329f325dd89bcd7143f0a7b376fcd64

Observation 4c455c56-7938-4db1-bc97-74d720506e05 · outbound

This paper cites Model compression via distillation and quantization.

EfficientLLM: Efficiency in Large Language Models Model compression via distillation and quantization

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.784223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.784223Z digest=sha256:f8cfdce680a108444b7ead0ed78ced6dcfd07536dc5ce08eb9da3ce0bb2119be

Observation 3688f486-281f-4257-b9a0-45c78884bac8 · outbound

This paper cites A survey on model compression for large language models.

EfficientLLM: Efficiency in Large Language Models A survey on model compression for large language models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.787230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.787230Z digest=sha256:d62788591bb501d9dc85ab35d83ed181c65a620983d025a425a7498ab8bb9ce0

Observation fbef8810-5e93-467a-b3ba-e40a4bf8d30d · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

EfficientLLM: Efficiency in Large Language Models Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.790739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.790739Z digest=sha256:bafc28fe3d0d5c23a762ce4455eb1619d0febda4fad89985faab8777c2f2ffe9

Observation 331b132d-00f0-46ae-a309-87e22eb43198 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

EfficientLLM: Efficiency in Large Language Models FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.793711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.793711Z digest=sha256:46d5a764175fc1379ac996787eef587c0d6bdc1f4a832bb0b83cd0546b5ac288

Observation 42836e9b-95a0-4c09-9d0a-9822a6c58313 · outbound

This paper cites Operator Fusion in XLA: Analysis and Evaluation.

EfficientLLM: Efficiency in Large Language Models Operator Fusion in XLA: Analysis and Evaluation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.797401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.797401Z digest=sha256:d643250a714bc1009468ea7bf91d07315f8efd5addccae984951012867ba90d8

Observation 1709dafd-a6a3-4bd5-a229-0c5b3609cd8d · outbound

This paper cites u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \.

EfficientLLM: Efficiency in Large Language Models u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.800970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.800970Z digest=sha256:d55a3e4203b0d0005fdd3ab5544dac8569b2377dd62b6aaa979b98c66cdf5709

Observation b09721a9-b129-4996-ab44-3bfb4c2ceac1 · outbound

This paper cites Knowledge distillation: A survey.

EfficientLLM: Efficiency in Large Language Models Knowledge distillation: A survey

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.844723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.844723Z digest=sha256:2bffc0d74f23f754450dc9fb8419dfa66282546ad698192aa5d31b74839d92ec

Observation f6d5b9e6-5bde-4fc1-bf19-e6260ca364ca · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

EfficientLLM: Efficiency in Large Language Models DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.917392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.917392Z digest=sha256:b71ab03e084d9f0302ca28ca63bcdbe14248b5cbefd23fbfe1384d2b5f8effef

Observation 7c2a6931-accf-4465-bee4-212f264bca15 · outbound

This paper cites Improving language models by retrieving from trillions of tokens.

EfficientLLM: Efficiency in Large Language Models Improving language models by retrieving from trillions of tokens

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.926329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.926329Z digest=sha256:43864c7494a20925277d3c1b2112b445e57b92897c2cef4e38ba0aa4eaa5a98f

Observation 7b85382c-5bdb-41be-968e-3e3dbd2e15ec · outbound

This paper cites Jurassic-1: Technical details and evaluation.

EfficientLLM: Efficiency in Large Language Models Jurassic-1: Technical details and evaluation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.932030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.932030Z digest=sha256:04841667502fff74b92a5bb449fe5962b06ecf844264f8774e070987d1d541a9

Observation 080ea051-be5f-444c-8aab-c4752ea943f0 · outbound

This paper cites Enhancing LLM Factual Accuracy with RAG to Counter Hallucinations: A Case Study on Domain-Specific Queries in Private Knowledge-Bases.

EfficientLLM: Efficiency in Large Language Models Enhancing LLM Factual Accuracy with RAG to Counter Hallucinations: A Case Study on Domain-Specific Queries in Private Knowledge-Bases

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.935549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.935549Z digest=sha256:fd39628b5d93b7d0e5172e5f84a98a624b9b7cfe8af29b467c4ec2b2f6abd016

Observation fd5fe7d0-4b68-424e-94bb-e113dc26c3f6 · outbound

This paper cites Benchmarking retrieval-augmented generation for medicine.

EfficientLLM: Efficiency in Large Language Models Benchmarking retrieval-augmented generation for medicine

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.939409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.939409Z digest=sha256:5e54f2ce2607b6740f7dbaa18b2a420d0eee7318c3bd6c83a6bc0abb2d2db152

Observation ea25360c-4175-4db2-be08-75862683f310 · outbound

This paper cites Longformer: The Long-Document Transformer.

EfficientLLM: Efficiency in Large Language Models Longformer: The Long-Document Transformer

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.942813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.942813Z digest=sha256:4f5eff3423e6608ffc5ef3c5780071ca30fd689056c8549b547dc20eae4dbd04

Observation 8263bf1b-6e86-4db6-9682-91cc3a86a857 · outbound

This paper cites Big bird: Transformers for longer sequences.

EfficientLLM: Efficiency in Large Language Models Big bird: Transformers for longer sequences

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.947235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.947235Z digest=sha256:5d995c622fe30f44ebe767716586b4f9c3b57da0fffe8b38fa9212b50ae882ac

Observation 5a5e21c2-cde9-47c8-b7bc-bb0707cd8ebf · outbound

This paper cites Reformer: The efficient transformer.

EfficientLLM: Efficiency in Large Language Models Reformer: The efficient transformer

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.950254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.950254Z digest=sha256:39b00144a215aa2b3e3c557341170efde4372c6a378d0bb05ad77fa55aecde2f

Observation 5936166b-955d-43ac-ac4f-78c4551a39b7 · outbound

This paper cites Mixture of experts: a literature survey.

EfficientLLM: Efficiency in Large Language Models Mixture of experts: a literature survey

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.953746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.953746Z digest=sha256:94cd01fe41390441c8e050a8e1bb8eae647bf61a09e80114c1d0864f67583adb

Observation a4122517-25d8-43be-afd8-7323b0230180 · outbound

This paper cites Exploring the Benefit of Activation Sparsity in Pre-training.

EfficientLLM: Efficiency in Large Language Models Exploring the Benefit of Activation Sparsity in Pre-training

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.957015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.957015Z digest=sha256:fca04ed023803e9467b48a47fe267b4c08862ed336ccbccd5413d0938a72c1e1

Observation 6f611781-0158-4b29-8180-5d86a9a9b824 · outbound

This paper cites MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts.

EfficientLLM: Efficiency in Large Language Models MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.960454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.960454Z digest=sha256:7d07c4c210ecd18bdbf04ccbd3b90eb08ccbcc1659e367302c0d60bbded56e36

Observation 039dbc16-ec97-4204-819c-ed3e1c854d19 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

EfficientLLM: Efficiency in Large Language Models Adam: A Method for Stochastic Optimization

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.963706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.963706Z digest=sha256:2e59639dbcce16a30d727108fe96a605493586bccb83380be1e998b883ef33fa

Observation a7143b88-4dfd-43a5-8f76-687bd85d0f8f · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic optimization.

EfficientLLM: Efficiency in Large Language Models Adaptive subgradient methods for online learning and stochastic optimization

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.966652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.966652Z digest=sha256:4d53634a1f0a9180ab8093e56a5691ed54d7ace7722bb3fc5f54d68a1889c03b

Observation 2758ef09-1d30-4dc4-ae9a-2de74c88b1d0 · outbound

This paper cites Large Batch Optimization for Deep Learning: Training BERT in 76 minutes.

EfficientLLM: Efficiency in Large Language Models Large Batch Optimization for Deep Learning: Training BERT in 76 minutes

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:34.970197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:34.970197Z digest=sha256:9ab2c82e85b2a5267d52bab7cc1dd086857b49690af29b5b74ec4c556f4be60c

Observation 9653765b-2ae6-4e30-bf02-869eda0ba19c · outbound

This paper cites Hyp-RL : Hyperparameter Optimization by Reinforcement Learning.

EfficientLLM: Efficiency in Large Language Models Hyp-RL : Hyperparameter Optimization by Reinforcement Learning

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:35.031035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:35.031035Z digest=sha256:a5545a0bb02dc1a08dc4c5da30cfb03df02b07487d7da7398a8098716367b3d1

Observation 6911b3e2-0906-4ff1-a6b8-8efced9511f0 · outbound

This paper cites Zeus: Understanding and optimizing GPU energy consumption of DNN training.

EfficientLLM: Efficiency in Large Language Models Zeus: Understanding and optimizing GPU energy consumption of DNN training

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:35.100401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:35.100401Z digest=sha256:eb1d9f32154a3e20f20e1fc6209a479b9cc9244f8ce53d56b8dd6d3dccc40966

Observation fd4f1c46-25e9-429d-b73a-ffdc391c3e1a · outbound

This paper cites A self-tuning actor-critic algorithm.

EfficientLLM: Efficiency in Large Language Models A self-tuning actor-critic algorithm

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:35.103637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:35.103637Z digest=sha256:480f71a8142706e4b2971a553b8248993f24d779b9d25b77990364b86bead7a7

Observation 86f330c2-c1d8-46fd-b051-62092ac2356c · outbound

This paper cites A comprehensive survey of neural architecture search: Challenges and solutions.

EfficientLLM: Efficiency in Large Language Models A comprehensive survey of neural architecture search: Challenges and solutions

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:35.107890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:35.107890Z digest=sha256:7bcddc881108ebbed79afcde41405146f28f13936996507597d66a9fa24706b6

Observation 5158d4af-e173-4cda-b10a-82548f7b7d33 · outbound

This paper cites Efficientnet: Rethinking model scaling for convolutional neural networks.

EfficientLLM: Efficiency in Large Language Models Efficientnet: Rethinking model scaling for convolutional neural networks

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:35.112121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:35.112121Z digest=sha256:1943f3bc4cb5daa0047a6073f36af8e4a068b44d711ccd0ebc8d8f5e7a52af78

Observation c3b46dae-82da-4cce-ab08-518e97aea87d · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

EfficientLLM: Efficiency in Large Language Models Teaching Large Language Models to Reason with Reinforcement Learning

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:35.115307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:35.115307Z digest=sha256:32a50f021b7500e79682d46b8b749d79feefc2042330f9d9f16ea9a79c412d18

Observation 961f0705-dff2-40ce-b452-174b38088b9e · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

EfficientLLM: Efficiency in Large Language Models DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:35.120113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:35.120113Z digest=sha256:274ac9d1e0a3dd13c88305cef728a621df411f15c84e8cb6250724993b1281df

Observation d85b7d3b-7abd-403b-b6bb-e5e2c83b44f6 · outbound

This paper cites Fine-Tuning Language Models with Advantage-Induced Policy Alignment.

EfficientLLM: Efficiency in Large Language Models Fine-Tuning Language Models with Advantage-Induced Policy Alignment

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:35.124526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:35.124526Z digest=sha256:b0fcaa81e235593252a5f434ec79828d7f1e5b1869f4f6b8b811526a7c716962

Observation dc2c1813-3b42-41c7-8d5d-789d5f42ab0d · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

EfficientLLM: Efficiency in Large Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:35.128987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:35.128987Z digest=sha256:cbe4747c7379574a07b1ebf69e09163ddaa89a4d7b2270d98990a5a000499870

Observation a41aaf69-82db-4343-97e5-8d3cda9a0b66 · outbound

This paper cites The efficiency spectrum of large language models: An algorithmic survey.

EfficientLLM: Efficiency in Large Language Models The efficiency spectrum of large language models: An algorithmic survey

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:35.133271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:35.133271Z digest=sha256:b9902564f61ea245afd90f57a5382aa392ecdfe7276b2066eaaf7be100498f90

Observation 7dca577a-6daa-422e-8760-60b3e58ceedd · outbound

This paper cites Energy and policy considerations for deep learning in nlp.

EfficientLLM: Efficiency in Large Language Models Energy and policy considerations for deep learning in nlp

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:35.213663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:35.213663Z digest=sha256:fcec22c5cdbc7f1021af26ff909f5c6f64955db7913f08a4fee8486a34dcd97a

Pith citing papers

Observation de19ad69-08bb-40da-a065-06d671a7492a · inbound

SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation cites this paper.

SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation EfficientLLM: Efficiency in Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:48.384737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:48.384737Z digest=sha256:6f1bd943b8585fedab92d730b380026bc06aa5fd708f4b5ffab36788d5e97a21

Observation cdd5c10d-32a5-4e3a-8d9f-14c7eb85c04d · inbound

MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU cites this paper.

MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU EfficientLLM: Efficiency in Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:51.693768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:57:25.256574Z digest=sha256:f6386357b3843d699d14e4ddbf325b9c58a3c22786876432cd72a508d28d300e

Observation 99a9a95c-ad7b-4a1b-ae3d-a7f8c2c974eb · inbound

APQF: Agentic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-Tuning cites this paper.

APQF: Agentic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-Tuning EfficientLLM: Efficiency in Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T12:00:53.896063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:00:53.896063Z digest=sha256:0cb7fbacd3503f320424c60da56a6313de8b63a4c0f6cc70838c539d8865f578