REVIEW 3 cited by
TOPLOC: A Locality Sensitive Hashing Scheme for Trustless Verifiable Inference
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Large language models (LLMs) have proven to be very capable, but access to frontier models currently relies on inference providers. This introduces trust challenges: how can we be sure that the provider is using the model configuration they claim? We propose TOPLOC, a novel method for verifiable inference that addresses this problem. TOPLOC leverages a compact locality-sensitive hashing mechanism for intermediate activations, which can detect unauthorized modifications to models, prompts, or precision with 100% accuracy, achieving no false positives or negatives in our empirical evaluations. Our approach is robust across diverse hardware configurations, GPU types, and algebraic reorderings, which allows for validation speeds significantly faster than the original inference. By introducing a polynomial encoding scheme, TOPLOC minimizes the memory overhead of the generated proofs by $1000\times$, requiring only 258 bytes of storage per 32 new tokens, compared to the 262 KB requirement of storing the token embeddings directly for Llama 3.1-8B-Instruct. Our method empowers users to verify LLM inference computations efficiently, fostering greater trust and transparency in open ecosystems and laying a foundation for decentralized, verifiable and trustless AI services.
Forward citations
Cited by 3 Pith papers
-
Which Model Is Actually Serving You? IRIS: Budgeted Black-Box Auditing of Model Substitution and Routing Dilution in LLM Gateways
Random-generation probes plus a pilot-fitted budget let a text-only auditor detect model substitution, estimate the routing dilution fraction, and attribute the served backend across LLM gateways.
-
Integrity of peer-to-peer distributed LLM inference under malicious nodes
Under a simulated isotropic noise model, a canary-trap activation-drift detector achieves perfect AUROC separation of one malicious shard in multi-hop LLM inference.
-
Towards Anonymous Neural Network Inference
funion applies Echomix's Pigeonhole storage and BACAP capabilities to run neural network inference through a store-compute-store pipeline, claiming sender-receiver unlinkability inherited from the mixnet.
Discussion (0). Continue with ORCID to comment.