REVIEW 8 cited by
Coupling Large Language Models with Logic Programming for Robust and General Reasoning from Text
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
While large language models (LLMs), such as GPT-3, appear to be robust and general, their reasoning ability is not at a level to compete with the best models trained for specific natural language reasoning problems. In this study, we observe that a large language model can serve as a highly effective few-shot semantic parser. It can convert natural language sentences into a logical form that serves as input for answer set programs, a logic-based declarative knowledge representation formalism. The combination results in a robust and general system that can handle multiple question-answering tasks without requiring retraining for each new task. It only needs a few examples to guide the LLM's adaptation to a specific task, along with reusable ASP knowledge modules that can be applied to multiple tasks. We demonstrate that this method achieves state-of-the-art performance on several NLP benchmarks, including bAbI, StepGame, CLUTRR, and gSCAN. Additionally, it successfully tackles robot planning tasks that an LLM alone fails to solve.
Forward citations
Cited by 8 Pith papers
-
Beyond the Surface: A Solution-Aware Retrieval Model for Competition-level Code Generation
SolveRank trains a contrastive retriever on LLM-generated logically equivalent problem variants and reports improved retrieval and code generation, though the retrieval evaluation is circular.
-
SOP-Agent: Empower General Purpose AI Agent with Domain-Specific SOPs
A decision-graph SOP navigator guides LLM agents through branching and looping workflows, with reported gains on household tasks, code generation, data cleaning, and a new customer-service benchmark.
-
Generative Agents for Multi-Agent Autoformalization of Interaction Scenarios
GAMA uses LLM agents to turn natural language game descriptions into validated executable logic programs, reaching about 77% semantic correctness on 110 scenarios from five 2x2 games.
-
Building a Stable Planner: An Extended Finite State Machine Based Planning Module for Mobile GUI Agent
A hand-authored EFSM planning module boosts Qwen2.5-VL-72B on AndroidWorld from 35.0% to 63.8% task success.
-
A Comparative Study of Neurosymbolic AI Approaches to Interpretable Logical Reasoning
A comparison of two neurosymbolic designs concludes that the hybrid design, pairing an LLM with a separate symbolic solver, is the more promising path to general logical reasoning.
-
Synergizing Logical Reasoning, Knowledge Management and Collaboration in Multi-Agent LLM System
SynergyMAS combines a graph database with a Clingo logic solver, corrective RAG, and Theory of Mind prompts in a hierarchical multi-agent team, demonstrated on a Smart Home Energy Management case study.
-
LLMSR@XLLM25: An Empirical Study of LLM for Structural Reasoning
A few-shot, untuned Meta-Llama-3-8B-Instruct system with regex post-processing ranks 5th on the LLMSR@XLLM25 structural reasoning shared task.
-
Dspy-based Neural-Symbolic Pipeline to Enhance Spatial Reasoning in LLMs
A DSPy-orchestrated LLM plus Answer Set Programming pipeline reports 82% average accuracy on StepGame and 69% on SparQA, well above direct prompting baselines.
Discussion (0). Continue with ORCID to comment.