Pith. sign in

Paper Citation Record · LEDGER

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications

As of 19 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 4 inbound Pith citation observations for arXiv:2505.07203.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.07203 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:27:53.745929Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T12:55:10.063756Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T08:51:08.950391Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0243f9e1-63e6-4859-ac0a-d6beb68c31c6 · outbound

This paper cites https: //character.ai/.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications https: //character.ai/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.771323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.459444Z digest=sha256:7c6e8353d29c7239f280a3c4a7c8a2ecf932ff05e886219bff4f9a3440c1b46f

Observation 53d40eb0-3707-4757-b408-28f2da28b813 · outbound

This paper cites [Online; accessed 2025-04-17].

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications [Online; accessed 2025-04-17]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.752863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.465242Z digest=sha256:539b71903357d62863745af5600153f361b57b7c06b3b933a35a52f2aecd9ba3

Observation 3bf8fc66-045e-467d-ad8f-671410d9738d · outbound

This paper cites Taming{Throughput-Latency} tradeoff in{LLM} inference with {Sarathi-Serve}.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Taming{Throughput-Latency} tradeoff in{LLM} inference with {Sarathi-Serve}

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.735576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.470802Z digest=sha256:5e08b4c05f6e5aef0442b838fc8e59f0f1c8bf46aac654081b01dda413fd1009

Observation bf5283b4-44a0-4f54-b508-f081fe947733 · outbound

This paper cites Pytorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Pytorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.717255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.475751Z digest=sha256:99b0239dd9458c117ede3f8d0ae32e3c4605053af03fdcf862594ee6828c61a3

Observation ea0f3eef-bd27-4fbe-99c8-d8ab76bd2752 · outbound

This paper cites The ai code editor.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications The ai code editor

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.698347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.480802Z digest=sha256:7e0238508a47d1f9b13b77e75b4ea679a551fcf62dcc390257eca456e1390115

Observation b9b9b9d6-3fd9-41da-ae41-a92fbcb18c4f · outbound

This paper cites Glimpse: Continuous, real-time object recog- nition on mobile devices.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Glimpse: Continuous, real-time object recog- nition on mobile devices

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.681002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.485859Z digest=sha256:310e35c0cfb33140c495d5f7cd7b95a9b5e81cd7070b7c86d3bc210a7b1e36f8

Observation dd654ef5-101c-418e-a198-43a3279fafe2 · outbound

This paper cites Feature engineering for machine learning and data analytics.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Feature engineering for machine learning and data analytics

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.663159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.491208Z digest=sha256:ac307e5e88f0576935e30b8479059067c794a9dd47e39061431cf0f394be57c1

Observation e9d28194-8d4e-46e4-8d4e-5690219f1d89 · outbound

This paper cites XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.495832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.495832Z digest=sha256:e04d76195298a6fcebc5e0bf1cd0e70c61bc329d250c355387956bfd1598ab23

Observation 01c3e016-8daa-4660-b225-d83c3c8a69e6 · outbound

This paper cites Oneadapt: Fast adaptation for deep learning applications via back- propagation.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Oneadapt: Fast adaptation for deep learning applications via back- propagation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.646204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.500782Z digest=sha256:fe9f988b3a4221801f7f04c9175b624152d179f374a71019d4cebc1206e2d37e

Observation e9c30c07-e086-4ae1-afaa-1bb06b2273dd · outbound

This paper cites Accmpeg: Optimizing video encoding for accurate video analytics.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Accmpeg: Optimizing video encoding for accurate video analytics

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.629176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.505752Z digest=sha256:d249e28506ceeca2797b8fe0d7b059476f7672ccbf6d5d481e3c5494a89fc4d7

Observation fd00cd96-e99b-45a2-86cf-6162d77e606f · outbound

This paper cites Empowering Many, Biasing a Few: Generalist Credit Scoring through Large Language Models.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Empowering Many, Biasing a Few: Generalist Credit Scoring through Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.510969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.510969Z digest=sha256:cbed7a0cc9f4a7a0db35ee2566ba146a39270ff9056fa3038889e5c927097e7d

Observation 4919d590-6d62-4440-ba4f-6481f1d2c410 · outbound

This paper cites 360brew: A decoder-only foundation model for personalized ranking and recommendation.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications 360brew: A decoder-only foundation model for personalized ranking and recommendation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.516094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.516094Z digest=sha256:67087e38a096030890422535f0328d4bc31c923448221e6568a70a083ec84286

Observation a6044832-a151-4b28-ab8c-a25cf61898ab · outbound

This paper cites Gpt-3: Its nature, scope, limits, and consequences.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Gpt-3: Its nature, scope, limits, and consequences

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.520895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.520895Z digest=sha256:eb628f04439ba5eae5f0f551818896f06e158ffd8548195d0a437cf42b5a69fd

Observation 5b5c677f-e06e-4c82-a30f-12d19dbcc643 · outbound

This paper cites Github copilot - write code faster.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Github copilot - write code faster

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.592014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.525803Z digest=sha256:abec04a8faa1367fd22227fbe287ac5782160029bcbe4ec6b057861a98b2a25c

Observation 87bf785e-f30f-46f5-bde2-8443c4eac836 · outbound

This paper cites Tiresias: A {GPU} cluster manager for distributed deep learning.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Tiresias: A {GPU} cluster manager for distributed deep learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.575925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.530420Z digest=sha256:8003146686965169fe416ae8e729690f3b694b5ac7d24477d3f6767b24f34b36

Observation b9c6bfc9-2db1-4ff4-bbd7-c9a99cb3d920 · outbound

This paper cites AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.535008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.535008Z digest=sha256:3896ca880a713c948f80658400edc7162917d76436d9072c44224c2855761543

Observation 070ec353-da31-46bf-9097-ec604a14edf7 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.540334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.540334Z digest=sha256:181aa299c3a38c67b398dc5ef30961679d6bbffce32ce9401ac30d85d674564d

Observation 87b1e661-6b83-44e0-b97f-63c38887672c · outbound

This paper cites Epic: Efficient position-independent context caching for serving large language models, 2024.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Epic: Efficient position-independent context caching for serving large language models, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.561055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.545556Z digest=sha256:9edd15d763f4ea24924c779d4c7bca813ac40eaafc93f223c76be4d4860d4612

Observation afae9e35-653b-42c7-86ab-913980a668e6 · outbound

This paper cites RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.550541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.550541Z digest=sha256:9114d75d56cbf8fd41d09752f0c6c1e04a222abf2859716289aa268013dd274b

Observation d26bc59e-c4cb-4528-b497-48d5f1fe3b59 · outbound

This paper cites Gear: An efficient kv cache com- pression recipe for near-lossless generative inference of llm, 2024.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Gear: An efficient kv cache com- pression recipe for near-lossless generative inference of llm, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.546569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.555355Z digest=sha256:cc20b1385b6554ac63eeda3b69ea030a2f94791cbfebdf5b92e87a74b13d0f3d

Observation 24550dde-8603-49d5-8f9d-0c92bdc5ac57 · outbound

This paper cites Over-fitting and model tuning.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Over-fitting and model tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.532326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.559851Z digest=sha256:b3297aa3c571d910bd37c2a1ba225bfe45baf0771e4b7f4453ce3af150fafb24

Observation 7a2840c0-8f67-4532-a4f5-63002c5bcb22 · outbound

This paper cites Efficient memory management for large language model serving with 13 pagedattention.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Efficient memory management for large language model serving with 13 pagedattention

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.517795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.564979Z digest=sha256:6e7bb39f9956515cba209d55409216834be936cf94936e7b137511e488dcd65f

Observation a8a825b4-1a37-424a-91a2-8606daec7c64 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Efficient memory management for large language model serving with pagedattention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.569776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.569776Z digest=sha256:8a4fdd8ef98cc0e92baf3d51a227cf2524a27440dfc3728b3594114dcf967336

Observation 3af2354d-a1b5-4add-80e9-ae4b4bc9219a · outbound

This paper cites Spam-T5: Benchmarking Large Language Models for Few-Shot Email Spam Detection.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Spam-T5: Benchmarking Large Language Models for Few-Shot Email Spam Detection

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.574575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.574575Z digest=sha256:d439e415fd58f171a0e78357bfbcc2e4d2787f8cdf47679f6733b2d1abd89f89

Observation c4d1eb7f-1ea4-4c94-87a7-cf06ed827ba1 · outbound

This paper cites Depression detection on social media with large language models.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Depression detection on social media with large language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.579596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.579596Z digest=sha256:b553fd348a06f989e891ff2ba03c23f38dd8ff0fd2554cbb53eb89df388edc05

Observation 40c1f9e7-a35a-4767-85d1-9e452da647be · outbound

This paper cites Reducto: On-camera filtering for resource-efficient real-time video analytics.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Reducto: On-camera filtering for resource-efficient real-time video analytics

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.495183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.584164Z digest=sha256:3c63dc1fafcee1fc12224557fe4f10e3c3e1bfd2bbee03879a27586bb78de96d

Observation 3217ccf5-0339-4387-a970-07e14193a430 · outbound

This paper cites Terapipe: Token-level pipeline parallelism for training large-scale language models.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Terapipe: Token-level pipeline parallelism for training large-scale language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.588957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.588957Z digest=sha256:4a68976f679f3405bedf30085141e9a8f02071a63a7e171f2d817248876eaecc

Observation 6496282b-7c4e-4f0c-8a51-4d1587ba24f1 · outbound

This paper cites Edge assisted real-time object detection for mobile augmented reality.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Edge assisted real-time object detection for mobile augmented reality

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.471905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.594621Z digest=sha256:9638ca66b936b0fbb7eb5e98f6f7aab251f121faeb3215291c9c62e0aae2e82c

Observation 4f462cbd-2303-4fe8-a5dd-47995a8a342c · outbound

This paper cites Gonzalez, Ion Stoica, and Matei Zaharia.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Gonzalez, Ion Stoica, and Matei Zaharia

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.599044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.599044Z digest=sha256:86f289beffe4c0418b2b2288c095324b094701e369ac7193e5222e8dd0363e33

Observation f536499a-cf67-4118-b5c5-3fd2b85d6881 · outbound

This paper cites FinGPT: Democratizing Internet-scale Data for Financial Large Language Models.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications FinGPT: Democratizing Internet-scale Data for Financial Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.604822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.604822Z digest=sha256:7457fa59e2c33e499b64cd974b69e43e99c1f87b12f4b325dd98d427331d6f1b

Observation cc014c3c-8e22-4e48-902d-c1ad9d1fc420 · outbound

This paper cites Cachegen: Kv cache compression and streaming for fast large language model serving.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Cachegen: Kv cache compression and streaming for fast large language model serving

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.447997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.609858Z digest=sha256:5d5c8916d67b6b696b55d22cd227287ef2c3dc84114b9c0ae38e07fa53ea04f8

Observation 3e442f03-9673-4d2a-b0de-4f72bf8e1b4d · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.620412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.620412Z digest=sha256:3f245013e14db170e877d18d96d8754ac267fc2de3c16e29a1f0d3065da723e9

Observation 047d1cb1-4bb2-4c18-8fed-76d619473651 · outbound

This paper cites Github - lmcache/lmcache: Redis for llms.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Github - lmcache/lmcache: Redis for llms

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.423570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.625061Z digest=sha256:79e0d03fd1a28799708e70bc9f3747d3bb3dc9fe66705d5216ea3823b7d167a9

Observation 07b981e4-390d-4878-af50-bed6259c4cd0 · outbound

This paper cites Chatgpt: Conversational language model.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Chatgpt: Conversational language model

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.409538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.630135Z digest=sha256:a80a0e79a646cbee0dee8696a211979449d6e3190c7acc5376acfa4000a98aca

Observation 4dc108da-b91e-4536-91b3-14426a1005dd · outbound

This paper cites Generative agents: Interac- tive simulacra of human behavior.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Generative agents: Interac- tive simulacra of human behavior

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.394551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.635605Z digest=sha256:92ddaf631d1d6b8e4e66c969c38981dfd9211e313a42fd8c387169d12b80b344

Observation f0185cbd-b785-46d2-9b5c-d417a446525c · outbound

This paper cites Optimus: an efficient dynamic resource scheduler for deep learn- ing clusters.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Optimus: an efficient dynamic resource scheduler for deep learn- ing clusters

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.378675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.640003Z digest=sha256:806a9d83bdf1e402cad371afa157977eddc054579857b507356994e433710657

Observation 8c9461ee-d736-4803-9a83-a4178d128321 · outbound

This paper cites Perplexity is a free ai search engine.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Perplexity is a free ai search engine

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.363796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.644768Z digest=sha256:4f10af573e4685eb3815b9e8e1bac49b532baeaf2dedd9f787133160e924e3ce

Observation 9b1acce1-2265-4fe2-bd50-d0ab5d9a0bcd · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.650313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.650313Z digest=sha256:9b5bbeb3d5f25ccd831b8e06d86cba7ca16541efb751d33a19b891292d0a71e3

Observation b3927c8b-6136-4ba6-b91a-264d15d65b3a · outbound

This paper cites Rah! recsys–assistant–human: A human-centered recommendation framework with llm agents.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Rah! recsys–assistant–human: A human-centered recommendation framework with llm agents

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.348979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.655137Z digest=sha256:158ee7827ea36fa6525fb2558e3dc7eb4f22cb61872961f24e7ecaffe27133f0

Observation 637b7295-4500-4d2f-98b3-98f68cc2d8c1 · outbound

This paper cites Playing games with ais: the limits of gpt-3 and similar large language models.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Playing games with ais: the limits of gpt-3 and similar large language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.334300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.660664Z digest=sha256:50b725ab977a5c869eab2cd21340893e921a6306a3d95f10361dd7ad1f44f146

Observation 2ef96d4e-921a-48b0-a532-543f457614fc · outbound

This paper cites Beyond Classification: Financial Reasoning in State-of-the-Art Language Models.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Beyond Classification: Financial Reasoning in State-of-the-Art Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.665157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.665157Z digest=sha256:6594593e25629b50b3790d73c326e1a28dc8b28e7502d473091cc3b609c0be37

Observation 3925683a-e76d-4d6b-ae8e-78799ab38ecd · outbound

This paper cites Enhancing Recommender Systems with Large Language Model Reasoning Graphs.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Enhancing Recommender Systems with Large Language Model Reasoning Graphs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.671082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.671082Z digest=sha256:a271f14df74ed1f3419bc62743dc29048bec1e60f6e5fbe4a72034828a628e98

Observation becd3254-6802-4aa7-bf8c-d64054083e98 · outbound

This paper cites Chain-of-thought prompt- ing elicits reasoning in large language models.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Chain-of-thought prompt- ing elicits reasoning in large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.319407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.677053Z digest=sha256:14157f203f0b1f9b120da3b31287546505423aa3c18cb63c4c41504f288269a6

Observation d6ff7978-52d9-447f-a655-bb4f442ad64e · outbound

This paper cites A survey on large language models for recommendation.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications A survey on large language models for recommendation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.681837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.681837Z digest=sha256:ff7254716c99b03d304f646f902902f73b5b772523d4b684d57e77eec58ce3b8

Observation d5d1acd8-b8a7-4907-9529-8b63e542db75 · outbound

This paper cites Predict- ing loan default in peer-to-peer lending using narrative data.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Predict- ing loan default in peer-to-peer lending using narrative data

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.293937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.687345Z digest=sha256:4a179a1e231b24fafe26287f3c757a61ee33280d5e597951a35a1b28d226d186

Observation 30cf5885-a546-4145-8433-9491ff7869fe · outbound

This paper cites Auto-GPT for Online Decision Making: Benchmarks and Additional Opinions.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Auto-GPT for Online Decision Making: Benchmarks and Additional Opinions

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.691773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.691773Z digest=sha256:5a61594aa4ddc381281756a47e1ef31a4e68dcb6efbc9d3f210cf139aecef3f6

Observation a9da14eb-f966-4f74-9245-ba3710c9315f · outbound

This paper cites Mini- mizing hallucinations and communication costs: Adversarial debate and voting mechanisms in llm-based multi-agents.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Mini- mizing hallucinations and communication costs: Adversarial debate and voting mechanisms in llm-based multi-agents

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.278564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.696361Z digest=sha256:09ba07e2c368543d642203ab9b83afafa1357673ccb1e27e25ddeb03a8250523

Observation 0eeba333-fa42-4744-aaf8-95a48fcbdd05 · outbound

This paper cites Cacheblend: Fast large language model serving for rag with cached knowledge fusion.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Cacheblend: Fast large language model serving for rag with cached knowledge fusion

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.264743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.700859Z digest=sha256:a5a29c63a4217f3e12eb1dd7b462017cd149db713e2f3d680059fc7e14ca3ca8

Observation 899ff7b9-bb18-44ad-bec8-6fb6014d425a · outbound

This paper cites FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.706664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.706664Z digest=sha256:ca741b295ec9bc451ac15bb87c6970dfeade2cd80dfbc7f2bebf8828cd646a96

Observation e751b87d-52b3-4af0-b225-c9c631042b71 · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based} generative models.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Orca: A distributed serving system for {Transformer-Based} generative models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.249901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.711790Z digest=sha256:e72979067efd923fa8da6e78fd4adb4ed0225c2c9ce4632db0517f8d40ebc73c

Observation 8e54bc2c-54a8-4dbf-be4f-4961db2a084d · outbound

This paper cites Towards explaining the effects of data preprocessing on machine learning.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Towards explaining the effects of data preprocessing on machine learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.232773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.717119Z digest=sha256:713da37f0607fe337af4b7603644b9ba1a4743d82e1e66b236aa7887957f62ee

Observation 51bae741-9221-4c33-a038-53bcb0c90103 · outbound

This paper cites Caravan: practical online learning of in-network ml models with labeling agents.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Caravan: practical online learning of in-network ml models with labeling agents

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.217874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.721623Z digest=sha256:c49830ccec49a2cb69de3ef505e60bbd43aa68efa7834eae27e25c2baa069bc7

Observation af44b6bd-290f-4c2e-a910-3c0c899563da · outbound

This paper cites Ll- maaa: Making large language models as active annotators.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Ll- maaa: Making large language models as active annotators

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.202824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.727069Z digest=sha256:27e12500e8771d298496e3bd8d2e2b4bbdf8cb0819e2441e793daaa06253289e

Observation 46778d6b-fc90-4c16-93a6-ce1a3260d7e2 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative in- ference of large language models.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications H2o: Heavy-hitter oracle for efficient generative in- ference of large language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.188221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.731317Z digest=sha256:e8f46812ea0d4b16c8dfebe325b869fdeffe6f8958ae5b1cbe69089d589e95e2

Observation ed04e13f-d7a1-46ca-a052-8cef1fad50b4 · outbound

This paper cites Mpic: Position-independent multimodal context caching system for efficient mllm serving, 2025.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Mpic: Position-independent multimodal context caching system for efficient mllm serving, 2025

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.173986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.735827Z digest=sha256:d800018694b2bc49ae923f107e70028994535dda4a4919b3cd0f9d8260dfcbd2

Observation 8488794d-3968-4225-9121-505c1348aa6b · outbound

This paper cites Gonzalez, Clark Barrett, and Ying Sheng.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Gonzalez, Clark Barrett, and Ying Sheng

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.741450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.741450Z digest=sha256:f5b7eae49abc9a2f35d0bd41d9eb885d8648167e822f971fadace19d7972c527

Observation 82f531d2-ea1b-43f4-931e-94e4302d6649 · outbound

This paper cites {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serv- ing.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serv- ing

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:27:54.150339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:27:53.745929Z digest=sha256:010dae6fe8a93251ac175b8f96a496e41be9e9e42ec6f407caffc84f7cabd72c

Observation 7929bbf0-80a4-4cab-94d4-14e929b0c38c · outbound

This paper cites an unresolved cited work.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.614510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.614510Z digest=sha256:b5072eeb3fca4f05d8978292f52e4692933fc6ddc77b11a8d3f95fae138bf591

Pith citing papers

Observation eb2e10b5-71de-4c19-8415-8f9c15abf136 · inbound

TetriServe: Efficiently Serving Mixed DiT Workloads cites this paper.

TetriServe: Efficiently Serving Mixed DiT Workloads PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:10.063756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:10.063756Z digest=sha256:6b2a2d6c25a77bae41d4996c84c73b44827b1ec97d09efedb298b2c60a165c5b

Observation 80f2c909-7b31-4f36-b684-7462641c5c4f · inbound

From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill cites this paper.

From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:51:08.952905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T08:47:29.759674Z digest=sha256:b8ea057968f6964fa4455ba78d52f725394d6e32482e9a7a94dce959902127a8

Observation f8d7267a-7304-4f84-a113-e6452afaff3c · inbound

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing cites this paper.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:48:27.073874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:3aaf16980ef362a1d459fa3a45b30cee77c5f18084f44c3f4d0345912c08b10b

Observation 5179647a-9af1-4c01-919e-362ad2d2382b · inbound

Beyond Storage: State as a Runtime Control Problem in Parallel and Distributed Systems cites this paper.

Beyond Storage: State as a Runtime Control Problem in Parallel and Distributed Systems PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T19:51:03.178136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:51:03.178136Z digest=sha256:15b21808ac9e866b9b75f48d18f6c65f5ad6742137811c395c1b36351ba3738b