Pith. sign in

Paper Citation Record · LEDGER

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios

As of 14 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2412.08972.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08972 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:25:31.762978Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:58:02.335060Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T12:58:02.541932Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fe74d52f-9e9a-4981-9381-23d17cb9ce5a · outbound

This paper cites online" 'onlinestring :=.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.466590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.466590Z digest=sha256:c99ad9817cd6514bf185a14492c6568448a9af9d0e014c8c708942aec4eba766

Observation c9c95e2b-5f9e-46d6-bf7c-2a2d276398b7 · outbound

This paper cites write newline.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.472716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.472716Z digest=sha256:a3cd86f92727f398c9f2ae17ae04a71dc59109e1193b3490936fb41fd76f6d8f

Observation 599ee7b7-00f5-4437-a3b6-b38fe3bb4ac1 · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.478406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.478406Z digest=sha256:0d34d3e7fc2389340896ec1bf31478f05724072f881f9eb505ee34f740e9ea24

Observation 472af47e-a485-4b67-bd82-325ffc628e56 · outbound

This paper cites an unresolved cited work.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:25:32.634419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T17:25:31.484261Z digest=sha256:83988095b9f0c99e1198be93aa2f2b5ce52204a1308ed592c33de478e9386dbc

Observation 8c645659-9cea-455c-896a-efcc98ce5329 · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.490688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.490688Z digest=sha256:a2eb9dda9c8ed349096611658755a120a0245849fe181a81a6947f93c6280d0c

Observation c9fcdded-cafc-41bd-aeb4-2870f989629b · outbound

This paper cites Benchmarking Large Language Models on Controllable Generation under Diversified Instructions.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Benchmarking Large Language Models on Controllable Generation under Diversified Instructions

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:25:32.521547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T17:25:31.502542Z digest=sha256:44cf29e2647bfe53919314fd9a0ecc98a493cbae907d56af1b45144c9c87c90f

Observation b3ab5312-a0a8-4a93-8215-efe0af9a4d2f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.507934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.507934Z digest=sha256:e5ecfefa386db6dfc0488ef68a0296a3537e8f8d94a6dd2d9a566eef7ff8fd9d

Observation 75cfa14a-b8f9-4487-896f-7ae1df5c893d · outbound

This paper cites A Survey on In-context Learning.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios A Survey on In-context Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.514193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.514193Z digest=sha256:9321797cceaeffcdf6d0d4c1f7ad5c0b227ce73b7d21825d3f531f3631379f7f

Observation ab56d529-30ee-4481-861e-93a4c9952e9d · outbound

This paper cites The Llama 3 Herd of Models.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.522554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.522554Z digest=sha256:48044b46d4072c1132dd648d4103d9d7f8641eaa34fd9d4f27f734846d3e4324

Observation cef593c2-1dbc-4197-adc3-f9923440cff7 · outbound

This paper cites NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.528740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.528740Z digest=sha256:340745fdf90420bc32441624195370a0da277b0ff7ae5e274b23bba99e01f52d

Observation dcd5b8b4-e197-468e-b98f-c474912133a6 · outbound

This paper cites Specializing Smaller Language Models towards Multi-Step Reasoning.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Specializing Smaller Language Models towards Multi-Step Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.534205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.534205Z digest=sha256:033fd642d7c525427d6b69e3dbe0bda79cb534ec5f2b03b062372e724a14967f

Observation faa6572b-b8ad-4e9f-8cc2-f61ace86641f · outbound

This paper cites Neural Module Networks for Reasoning over Text.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Neural Module Networks for Reasoning over Text

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.539635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.539635Z digest=sha256:806a1a9b5bf63358665eedac96cc64ce8a8d9841d2e0359742b68e91b8c0c815

Observation 80815c01-31e8-47f4-862e-111eaa9ec44e · outbound

This paper cites FOLIO: Natural Language Reasoning with First-Order Logic.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios FOLIO: Natural Language Reasoning with First-Order Logic

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.545315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.545315Z digest=sha256:1bc29529246d9b6f98c3799bfa0d63320231bb6cd7ac370ea79edd2afe876c22

Observation a36dd340-b6a0-4515-8470-c043cb3bfde8 · outbound

This paper cites Can Large Language Models Understand Real-World Complex Instructions?.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Can Large Language Models Understand Real-World Complex Instructions?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.551256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.551256Z digest=sha256:58e1a764f413b14f20e830f60c48b77a2e321a38b69ce0dd2f9b651fd5c5b236

Observation 3b3b2d29-82b6-4140-8d85-e1d892dbcc71 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.556838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.556838Z digest=sha256:105d1f3ef5798f171dac22491b0fb529d1ef00057ecf355790148273732abc7d

Observation dad0c5e5-51b0-4dea-a78e-47e6e0f51821 · outbound

This paper cites Disentangling Logic: The Role of Context in Large Language Model Reasoning Capabilities.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Disentangling Logic: The Role of Context in Large Language Model Reasoning Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.562964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.562964Z digest=sha256:9596f424bfeb29e3f1b4d2aaa0a79d2be9d4f6b95f7d05b6640c24d5a733a2e7

Observation 1bcfb457-4e42-403d-b33e-24e07603a894 · outbound

This paper cites Fine-tuning and Utilization Methods of Domain-specific LLMs.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Fine-tuning and Utilization Methods of Domain-specific LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.568737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.568737Z digest=sha256:874a0def6c5f52045128d1313a0d642d3b141953ced236531c64e77931eac4e5

Observation 84047906-9d58-4a8b-b19f-970db3f7ca96 · outbound

This paper cites FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.573876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.573876Z digest=sha256:96340bb31fb0298596a42f3ef13c5ef14a2fc9e3bd7a68f777eae5c5ee990298

Observation 554c006a-5a8d-431e-90e0-d7d6e714a2b1 · outbound

This paper cites Large Language Models are Zero-Shot Reasoners.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Large Language Models are Zero-Shot Reasoners

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.579584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.579584Z digest=sha256:abc5d3dee084ace3c6456fa776f324ce1fa3a0056a4ca5ba2a794bd73cc1d730

Observation 8c110deb-f1f6-4d3d-b68f-216350ef4e65 · outbound

This paper cites an unresolved cited work.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:25:32.618511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T17:25:31.590606Z digest=sha256:a92124f7a58d96861a51acd0f9ab5029c2118601da54ca6e6377d98681b7a098

Observation e8113bb3-d5cd-4f4a-aaa8-2ac2ecb38777 · outbound

This paper cites Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.596604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.596604Z digest=sha256:ccf52efa9eab03c8aa56f987cad42c69cd38ea4cfed63c396b3a1a92af3b28a1

Observation 66f26733-5eae-44f1-b1d6-7627e89aeeab · outbound

This paper cites AlignBench: Benchmarking Chinese Alignment of Large Language Models.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.603431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.603431Z digest=sha256:07b064b48b067244476cdb316bb80027f580d74e6ec76afbc857c663781711f3

Observation d3cafca9-863a-4946-9151-f0ca0cfdaa72 · outbound

This paper cites The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.609161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.609161Z digest=sha256:357d3019082a688e633690d9aa3b80fb8daead8c44414dcb73ba5ca762e27f60

Observation 8e8f1c96-c593-444a-ba49-79d753438df8 · outbound

This paper cites GPT-4 Technical Report.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios GPT-4 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.620525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.620525Z digest=sha256:c6cef52c67fb202a620bd2473d5a3b75b1f7a0fb0c588c64ef426b1c92d722b3

Observation a7fe685c-4e24-41a8-9943-9522fe5cded6 · outbound

This paper cites an unresolved cited work.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:25:32.601585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T17:25:31.626065Z digest=sha256:163d3eea685a4c5ec44dc6684233ca52bb672adfba5a48f2a82fc47b3fecd040

Observation a2997379-5792-4bc2-ba04-86a4b9da5cdb · outbound

This paper cites Can LLMs Follow Simple Rules?.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Can LLMs Follow Simple Rules?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.631000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.631000Z digest=sha256:1f5f3be2763c2c6df9d2486a8c5845e28db21693ff016360a70833a3403a97e3

Observation d07f299e-107c-41cc-a9d8-1f3fb7fbe234 · outbound

This paper cites Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.635627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.635627Z digest=sha256:4399dbb75844fec4ae1c091d49390a18c4408da997784169ddaee46fccacc4b9

Observation 54d7ec07-72ff-4892-8fef-63be4a50f400 · outbound

This paper cites an unresolved cited work.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.640961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.640961Z digest=sha256:b861584cc24a91a5244bac6ca56b7f26282123d6655002526aecce6b8058a8f6

Observation 29a4a05b-dc45-474b-8d11-7feff824a42e · outbound

This paper cites Evaluating Large Language Models on Controlled Generation Tasks.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Evaluating Large Language Models on Controlled Generation Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.645790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.645790Z digest=sha256:000b275f06e8520c6d6df82fb8977e25c09ca31d4aaee5b9f365223e948ab12f

Observation ff534545-f620-4513-9392-6df483194e48 · outbound

This paper cites Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.650877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.650877Z digest=sha256:d0fa011459e0ca2dbd3e752fbd361c70e38eebd52dfcc712847ba377d82993df

Observation 93c81ae0-192f-4a1e-a76b-4770122d047f · outbound

This paper cites ProofWriter: Generating Implications, Proofs, and Abductive Statements over Natural Language.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios ProofWriter: Generating Implications, Proofs, and Abductive Statements over Natural Language

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.657742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.657742Z digest=sha256:f27be9c6c3d73d1511d8ce7380e89107d2ac76128d24e944045be8e09a4febcb

Observation 1cfbde95-e41a-4638-a9fc-2b37c86c2a4e · outbound

This paper cites Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data?.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.663743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.663743Z digest=sha256:a0e5334e11da3f5e1299214c240dcfa26eb6ca5ecb013a5b4aa5f3dcac45da71

Observation ecbe2bb7-1901-4741-8d75-912bb2e577b5 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Gemini: A Family of Highly Capable Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.669363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.669363Z digest=sha256:c924052f4b1215e40378ada0a338b2171dfb634190968e7f4fbaa91d4547151f

Observation e0baaebf-292d-4d06-a7f2-7c8db95538c2 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.674596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.674596Z digest=sha256:822b8c3307f0aa177744c1f76a3e8122f32c18c25732bbd5d9432e2bb5be14e9

Observation d1b92c73-8178-4984-90f9-730491d49f0c · outbound

This paper cites Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.679713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.679713Z digest=sha256:476964f6c5d7d8e1070013c657f90e2adc828d25527f2628c1a1404fc9974f5f

Observation c88a85ba-37a6-4566-a606-1ce458d5726c · outbound

This paper cites Symbolic Working Memory Enhances Language Models for Complex Rule Application.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Symbolic Working Memory Enhances Language Models for Complex Rule Application

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.685026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.685026Z digest=sha256:73175273a326d885b63f176eb1ebd8cac244569ad35c8c1b704b98ae1b34f7d5

Observation c782645d-ffea-4ab0-b09b-eed7f272ade5 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.690218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.690218Z digest=sha256:5301ad380fdb39f6acc5bb694be4139e3b71e642888a416b97b4ab6681c99bf3

Observation 1a85ad85-4774-4a3b-9fa8-dffe70884ab0 · outbound

This paper cites Larger language models do in-context learning differently.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Larger language models do in-context learning differently

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.695578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.695578Z digest=sha256:18ffde563b1dc1bf1786c371832014fc834cee883d756ed0e8c175a54e551343

Observation 4cdb2072-3836-4bf3-93a8-6fbb1f18f0f0 · outbound

This paper cites Benchmarking Complex Instruction-Following with Multiple Constraints Composition.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Benchmarking Complex Instruction-Following with Multiple Constraints Composition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.700799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.700799Z digest=sha256:b6cf25fb75890a68332a17a3301be99cb83dc2391aad80542c2b6ca654457894

Observation 89f1b186-2986-43fc-b091-9e364538064d · outbound

This paper cites an unresolved cited work.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:25:32.573102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T17:25:31.706510Z digest=sha256:ac70500428a61e86e2aab4b66211e8d079b4e82abc3dcd8009344032ddff077a

Observation 5c57647e-2132-4c7d-9d04-3f09215c7e4b · outbound

This paper cites AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.711425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.711425Z digest=sha256:7e98a7f33f68fb07b94b6bc1947c3fd916574777851525bf05ecf298aa75f5e8

Observation 3ffddade-a12f-4d0c-822a-75a995be22b8 · outbound

This paper cites FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.716582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.716582Z digest=sha256:2f4768007015a20847e1e19c7f31eac9b857af1c2e6226dd649015106fa49aa3

Observation f9aa2e2f-1cc9-4ae8-9b06-958f5effb3e4 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.721684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.721684Z digest=sha256:addff4116bbe7a84eb0ff3683c5fab1c450aaf27bee1efa61ee6960e2e57fe1e

Observation a8f55908-4f7a-4154-a59b-33998ccec510 · outbound

This paper cites Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.727116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.727116Z digest=sha256:c71b3aeb75086b48b13f115da3ee3667e0c55ee6b096d0cae902d67dc901d376

Observation 844f10de-a47f-4635-bcba-4a3297884042 · outbound

This paper cites IDEAL: Influence-Driven Selective Annotations Empower In-Context Learners in Large Language Models.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios IDEAL: Influence-Driven Selective Annotations Empower In-Context Learners in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.732283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.732283Z digest=sha256:82c3ae4b73ef0655f1fa9a16fc6ae4bb875fb7f84dd5b7c4fc8e58ada60444fa

Observation 384d583b-315a-4a8a-adf5-48678e0a6465 · outbound

This paper cites Supervised Chain of Thought.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Supervised Chain of Thought

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.737463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.737463Z digest=sha256:09b469d4f5b527822ca2de59dab4fc847d75549b709623f194f6e3e1087f4289

Observation 73562581-8cd3-4c46-b5f0-be198b988482 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.742515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.742515Z digest=sha256:becdf2088d437606e9ffcadfad5a9a2513233d72f390b159ebccd2c5e2ffca63

Observation 85d4fd0c-8c37-4866-b2a1-3e3105cc69c4 · outbound

This paper cites AR-LSAT: Investigating Analytical Reasoning of Text.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios AR-LSAT: Investigating Analytical Reasoning of Text

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.747587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.747587Z digest=sha256:1a20f8bd751c604068eef407ddc21d1cfea798143e26412b34c9d8b49cd43421

Observation 74a6d8bb-5218-4579-9c27-2f087c4131ca · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Instruction-Following Evaluation for Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.752696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.752696Z digest=sha256:fe165655b8e029cecfab730b54d0e0c73d4af2b43c2301f91fa9a1b722615f74

Observation ab0a5149-1d39-483d-bf07-d5ecd39f7b4f · outbound

This paper cites TRAD: Enhancing LLM Agents with Step-Wise Thought Retrieval and Aligned Decision.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios TRAD: Enhancing LLM Agents with Step-Wise Thought Retrieval and Aligned Decision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.757823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.757823Z digest=sha256:fddb721951fd414c9fa81b16d829f6298b7337ac5b92fa998cf6d8e2357219c8

Observation 7b10c5b1-eb13-437b-a2d3-e8c4a2892122 · outbound

This paper cites DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.762978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.762978Z digest=sha256:e77886e19deff2c4238e99928e6d2cbb9054611b3747154f23470476bae0b619

Pith citing papers

Observation dd2a529d-24b1-45f7-8885-b74ad43c7ffc · inbound

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge cites this paper.

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:58:02.549983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:58:02.335060Z digest=sha256:04b6c6e1f7dcd0182a0f6f3238fd67eccd93ecae5f1c43720ae81117165bdcf6