REVIEW 1 cited by
STREET: A Multi-Task Structured Reasoning and Explanation Benchmark
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce STREET, a unified multi-task and multi-domain natural language reasoning and explanation benchmark. Unlike most existing question-answering (QA) datasets, we expect models to not only answer questions, but also produce step-by-step structured explanations describing how premises in the question are used to produce intermediate conclusions that can prove the correctness of a certain answer. We perform extensive evaluation with popular language models such as few-shot prompting GPT-3 and fine-tuned T5. We find that these models still lag behind human performance when producing such structured reasoning steps. We believe this work will provide a way for the community to better train and test systems on multi-step reasoning and explanations in natural language.
Forward citations
Cited by 1 Pith paper
-
Evaluating LLM Reasoning in the Operations Research Domain with ORQA
ORQA is a new 1,513-question multiple-choice benchmark showing that open-source LLMs score up to 77% on operations research modeling questions, well below a 93% expert baseline.
Discussion (0). Continue with ORCID to comment.