Pith. sign in

REVIEW 2 cited by

RiddleSense: Reasoning about Riddle Questions Featuring Linguistic Creativity and Commonsense Knowledge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2101.00376 v2 pith:PMXRDAHZ submitted 2021-01-02 cs.CL cs.AI

classification cs.CLcs.AI
keywords commonsensereasoningabilitiesansweringquestionadvancedcreativitylanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Question: I have five fingers but I am not alive. What am I? Answer: a glove. Answering such a riddle-style question is a challenging cognitive process, in that it requires complex commonsense reasoning abilities, an understanding of figurative language, and counterfactual reasoning skills, which are all important abilities for advanced natural language understanding (NLU). However, there are currently no dedicated datasets aiming to test these abilities. Herein, we present RiddleSense, a new multiple-choice question answering task, which comes with the first large dataset (5.7k examples) for answering riddle-style commonsense questions. We systematically evaluate a wide range of models over the challenge, and point out that there is a large gap between the best-supervised model and human performance -- suggesting intriguing future research in the direction of higher-order commonsense reasoning and linguistic creativity towards building advanced NLU systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges

    cs.AI 2025-02 conditional novelty 6.0 of 10

    EnigmaEval is a private benchmark of 1,184 puzzle-hunt problems on which state-of-the-art vision-language models score 7.0% on normal and 0% on hard items.

  2. A Comprehensive Graph Framework for Question Answering with Mode-Seeking Preference Alignment

    cs.CL 2025-06 conditional novelty 4.0 of 10

    GraphMPA combines an embedding-similarity hierarchical graph with mode-seeking preference optimization to improve RAG question answering on six datasets.

Pith tools