Pith. sign in

REVIEW 1 cited by

Evaluating the Correctness of Inference Patterns Used by LLMs for Judgment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.09083 v2 pith:OG32KAPF submitted 2024-10-06 cs.AI cs.CLcs.CVcs.LG

Evaluating the Correctness of Inference Patterns Used by LLMs for Judgment

classification cs.AI cs.CLcs.CVcs.LG
keywords inferencepatternsllmsusedjudgmentlanguagecorrectcorrectness
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

This paper presents a method to analyze the inference patterns used by Large Language Models (LLMs) for judgment in a case study on legal LLMs, so as to identify potential incorrect representations of the LLM, according to human domain knowledge. Unlike traditional evaluations on language generation results, we propose to evaluate the correctness of the detailed inference patterns of an LLM behind its seemingly correct outputs. To this end, we quantify the interactions between input phrases used by the LLM as primitive inference patterns, because recent theoretical achievements have proven several mathematical guarantees of the faithfulness of the interaction-based explanation. We design a set of metrics to evaluate the detailed inference patterns of LLMs. Experiments show that even when the language generation results appear correct, a significant portion of the inference patterns used by the LLM for the legal judgment may represent misleading or irrelevant logic.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Auditing Provenance Sensitivity in LLM Agent Action Selection

    cs.AI 2026-07 conditional novelty 7.0

    LLM agents respond to source-authority cues yet show measurable sensitivity to unauthorized context under controlled tests: 5.4% action discordance under competition and a 2.4% retained-invalid error pattern.