Pith. sign in

Paper Citation Record · LEDGER

Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models

As of 21 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 6 inbound Pith citation observations for arXiv:2412.12687.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12687 v3

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:55:45.235032Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:53:26.821910Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T17:05:02.945302Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8982ab12-7561-4c0e-b348-5c320e3df8a3 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models Training Compute-Optimal Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:55:45.163262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:55:45.163262Z digest=sha256:3f3c2d43ad1b953cd66c84aee9dcb40e3e37f48e1268454cb8a53d4dc2172036

Observation dbcb050e-499a-4e27-896e-3bdf9460e414 · outbound

This paper cites Harnessing the power of llms in practice: A survey on chatgpt and beyond,.

Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models Harnessing the power of llms in practice: A survey on chatgpt and beyond,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:55:45.471463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:55:45.169883Z digest=sha256:eb6875731088b3c6b3ca05365199363098fb3487c302db381cbbdff074717f94

Observation 7f57e927-76ce-4e92-afd8-0cfaa7c3dae9 · outbound

This paper cites Llm-pruner: On the structural pruning of large language models,.

Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models Llm-pruner: On the structural pruning of large language models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:55:45.454338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:55:45.174399Z digest=sha256:de3fbef014d4c240d82868acabf90c882a02c44eff391e456222ea539a177dc6

Observation 09eb085e-17ca-4abc-ac74-87654ef31a8f · outbound

This paper cites LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models.

Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:55:45.180184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:55:45.180184Z digest=sha256:4deda6a0cb6f044b69504eca8359e9a49724349581fa9dc843e042f62cfa469e

Observation dcd60f6c-a7d5-47c5-80ca-f0d8753e08a4 · outbound

This paper cites DistillSpec: Improving Speculative Decoding via Knowledge Distillation.

Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models DistillSpec: Improving Speculative Decoding via Knowledge Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T13:55:45.185076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:55:45.185076Z digest=sha256:cbc5808dc498b71adb4278215850a6b3ecb85702e1af7a4e0c57a36655959311

Observation ee27de9a-7b8e-40d6-a3a6-8b1cb174733d · outbound

This paper cites Hybrid slm and llm for edge-cloud collaborative inference,.

Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models Hybrid slm and llm for edge-cloud collaborative inference,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:55:45.439985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:55:45.192117Z digest=sha256:802d3735b0ff27d3c14e33f58ee2ef646bf9f8d4cc33856030da1636be98bbb3

Observation c396383c-43c3-4c48-ade8-bc941dcefdbf · outbound

This paper cites Fast inference from transform- ers via speculative decoding,.

Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models Fast inference from transform- ers via speculative decoding,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:55:45.425572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:55:45.197135Z digest=sha256:64b5879cb580b499bfe11db8dc36280fa176545bca68768a51e8c26d71d40b77

Observation f219e540-4030-46ea-afaa-e526d3f5f1f6 · outbound

This paper cites Understanding the metropolis-hastings algorithm,.

Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models Understanding the metropolis-hastings algorithm,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:55:45.202129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:55:45.202129Z digest=sha256:4a2a14e40f50064655a9e7cfaa7777a4bfb4a28000e2caaec37cb07557dd60e6

Observation e9f3a79e-7c5c-460c-957d-995813ec170a · outbound

This paper cites Dropout as a bayesian approximation: Representing model uncertainty in deep learning,.

Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models Dropout as a bayesian approximation: Representing model uncertainty in deep learning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:55:45.206967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:55:45.206967Z digest=sha256:aaaabd5bb210c91b9e99f9b4c10cf9bf002aa81f4ef5ac9f540749730a0320a1

Observation 3f6f6ccc-3a78-4c15-a236-945adbca0013 · outbound

This paper cites Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models.

Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T13:55:45.211650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:55:45.211650Z digest=sha256:dec04f4bd938ede28f6e25a70b95be386c225ed60b1be5b54b5dc12d312cb193

Observation 523e6b7f-88e4-4ba8-84b4-3e7d68e4accc · outbound

This paper cites SPUQ: Perturbation-Based Uncertainty Quantification for Large Language Models.

Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models SPUQ: Perturbation-Based Uncertainty Quantification for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:55:45.217207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:55:45.217207Z digest=sha256:9ab5899fb2b2753d915998e2f05e0d66374241a6108a46169e3398e636418b11

Observation 69cb4e8e-17d4-4317-947e-13ae09a95eee · outbound

This paper cites Wordnet: a lexical database for english,.

Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models Wordnet: a lexical database for english,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T13:55:45.222000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:55:45.222000Z digest=sha256:9c855eb05d701663ca7dc3247657d8a542b408a4ee306ca3d425b6985f07d7c2

Observation 6ec2920f-1329-4c20-889a-7ab334710ae2 · outbound

This paper cites Stanford alpaca: An instruction-following llama model,.

Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models Stanford alpaca: An instruction-following llama model,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:55:45.381778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:55:45.226468Z digest=sha256:f653cd4cd41da1c8fe66cd35f83e21cc078c5b831efb49ef53f8957ed61c942a

Observation c08093f0-8d7e-4a69-899d-37120135de1a · outbound

This paper cites The flan collection: Designing data and methods for effective instruction tuning,.

Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models The flan collection: Designing data and methods for effective instruction tuning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:55:45.367543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:55:45.231224Z digest=sha256:e399fb284e4c92dd9088675b85d4499bff5bce0bb4770c1830ced7aac4e4b7fd

Observation adda9605-51ae-48ae-b5d0-868fb525a047 · outbound

This paper cites Universal Sentence Encoder.

Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models Universal Sentence Encoder

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T13:55:45.235032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:55:45.235032Z digest=sha256:2954a02c7a81d10a9a1f44ca383563478a65067b68ea86a1ca511895f830e0e2

Pith citing papers

Observation ac5c0747-3d9d-4a73-b154-a2819754f540 · inbound

Semantic Packet Aggregation for Token Communication via Genetic Beam Search cites this paper.

Semantic Packet Aggregation for Token Communication via Genetic Beam Search Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:26.821910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:26.821910Z digest=sha256:df31ff9cb00ff293de107e84b1ec5d89cabcfe9202344db0cdbfbc54d1dc5559

Observation 3619cabc-515e-4763-a9a8-7821d0ffc3e9 · inbound

The Larger the Merrier? Efficient Large AI Model Inference in Wireless Edge Networks cites this paper.

The Larger the Merrier? Efficient Large AI Model Inference in Wireless Edge Networks Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:43:56.705132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:43:56.705132Z digest=sha256:1dc81dd6c2f9825f49f76fc8e0bf599d81db6f0b4ad44f3fea0ed70420b87775

Observation bacd9d4a-a979-40c7-bde9-06f33b506bd8 · inbound

Communication-Efficient Hybrid Language Model via Uncertainty-Aware Opportunistic and Compressed Transmission cites this paper.

Communication-Efficient Hybrid Language Model via Uncertainty-Aware Opportunistic and Compressed Transmission Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:55:06.669092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:55:06.669092Z digest=sha256:c181578e5c7a4b5b949be260f6cba36ca1206a8408431e39a3a098f5a393ec99

Observation ef0da03f-33c7-4127-b10c-6dd7f4416723 · inbound

Prompting Wireless Networks: Reinforced In-Context Learning for Power Control cites this paper.

Prompting Wireless Networks: Reinforced In-Context Learning for Power Control Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:02.769207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:59:02.769207Z digest=sha256:46158f7229bf3b9e4b95978d366ba88bd520fcca3b8a9e3b42059ea4632980e2

Observation 18aed4b0-f397-4acc-9500-50d23ce2a4f0 · inbound

Low-Complexity Semantic Packet Aggregation for Token Communication via Lookahead Search cites this paper.

Low-Complexity Semantic Packet Aggregation for Token Communication via Lookahead Search Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:02.245722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:02.245722Z digest=sha256:d2b78764e4b1b6a101c1fed9cda468360c7eae65ddfeee198da0b7dae78689a7

Observation 87813549-5074-4d4a-9217-32a086a73519 · inbound

DSSD: Efficient Edge-Device LLM Deployment and Collaborative Inference via Distributed Split Speculative Decoding cites this paper.

DSSD: Efficient Edge-Device LLM Deployment and Collaborative Inference via Distributed Split Speculative Decoding Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:05:03.040669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T17:05:02.637601Z digest=sha256:0e38acc6f7d82cc142361ae39e80c413879c39c46e7af25dc154c72efe6d2bd6