Pith. sign in

REVIEW 4 cited by

CommonGen: A Constrained Text Generation Challenge for Generative Commonsense Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.03705 v4 pith:OXSXQRV6 submitted 2019-11-09 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords commonsensereasoningcommongengenerationgenerativetasktextability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, large-scale pre-trained language models have demonstrated impressive performance on several commonsense-reasoning benchmark datasets. However, building machines with commonsense to compose realistically plausible sentences remains challenging. In this paper, we present a constrained text generation task, CommonGen associated with a benchmark dataset, to explicitly test machines for the ability of generative commonsense reasoning. Given a set of common concepts (e.g., {dog, frisbee, catch, throw}); the task is to generate a coherent sentence describing an everyday scenario using these concepts (e.g., "a man throws a frisbee and his dog catches it"). The CommonGen task is challenging because it inherently requires 1) relational reasoning with background commonsense knowledge, and 2) compositional generalization ability to work on unseen concept combinations. Our dataset, constructed through a combination of crowdsourced and existing caption corpora, consists of 79k commonsense descriptions over 35k unique concept-sets. Experiments show that there is a large gap between state-of-the-art text generation models (e.g., T5) and human performance. Furthermore, we demonstrate that the learned generative commonsense reasoning capability can be transferred to improve downstream tasks such as CommonsenseQA by generating additional context.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning to Select In-Context Demonstration Preferred by Large Language Model

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A generative preference-learning method trains a latent demonstration selector from LLM feedback and improves few-shot in-context learning performance on most of 19 benchmark datasets.

  2. Randomly Sampled Language Reasoning Problems Elucidate Limitations of In-Context Learning

    cs.LG 2025-01 conditional novelty 6.0 of 10

    On randomly sampled 3-state DFA language tasks, foundation LLMs underperform n-gram baselines under pure in-context-learning prompts.

  3. Generative Language Models Potential for Requirement Engineering Applications: Insights into Current Strengths and Limitations

    cs.SE 2024-12 conditional novelty 5.0 of 10

    ChatGPT and Gemini generally underperform task-specific models on requirements engineering benchmarks, except ChatGPT achieves a new top F1 score on the REQuestA question answering dataset.

  4. Generative Adversarial Networks Bridging Art and Machine Intelligence

    cs.LG 2025-02 unverdicted novelty 1.0 of 10

    This paper is a textbook-style review of generative adversarial networks, covering theory, classic variants, training methods, and applications; no new architecture, theorem, or experimental result is introduced.

Pith tools