Pith. sign in

REVIEW 1 cited by

Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.08440 v4 pith:3GH6QNZV submitted 2024-07-11 cs.CL cs.AI

classification cs.CLcs.AI
keywords inferentialrule-followingllmscapabilityfollowingresultsscenariosabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although Large Language Models (LLMs) have demonstrated strong ability, they are further supposed to be controlled and guided by in real-world scenarios to be safe, accurate, and intelligent. This demands the possession of capability of LLMs. However, no prior work has made a clear evaluation of the inferential rule-following capability of LLMs. Previous studies that try to evaluate the inferential rule-following capability of LLMs fail to distinguish the inferential rule-following scenarios from the instruction-following scenarios. Therefore, this paper first clarifies the concept of inferential rule-following and proposes a comprehensive benchmark, RuleBench, to evaluate a diversified range of inferential rule-following abilities. Our experimental results on a variety of LLMs show that they are still limited in following rules. Our analysis based on the evaluation results provides insights into the improvements for LLMs toward a better inferential rule-following intelligent agent. We further propose Inferential Rule-Following Tuning (IRFT). The experimental results show that through IRFT, LLMs can learn abstract rule-following abilities from purely synthetic data and then generalize to RuleBench. The data and code can be found at: https://anonymous.4open.science/r/llm-rule-following-B3E3/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Shuttle Between the Instructions and the Parameters of Large Language Models

    cs.LG 2025-02 conditional novelty 6.0 of 10

    SHIP jointly trains an encoder and decoder so a language model's soft-prompt parameters can be reconstructed from instructions and vice versa, improving instruction induction and inductive reasoning.

Pith tools