REVIEW 4 cited by
Are large language models superhuman chemists?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) have gained widespread interest due to their ability to process human language and perform tasks on which they have not been explicitly trained. However, we possess only a limited systematic understanding of the chemical capabilities of LLMs, which would be required to improve models and mitigate potential harm. Here, we introduce "ChemBench," an automated framework for evaluating the chemical knowledge and reasoning abilities of state-of-the-art LLMs against the expertise of chemists. We curated more than 2,700 question-answer pairs, evaluated leading open- and closed-source LLMs, and found that the best models outperformed the best human chemists in our study on average. However, the models struggle with some basic tasks and provide overconfident predictions. These findings reveal LLMs' impressive chemical capabilities while emphasizing the need for further research to improve their safety and usefulness. They also suggest adapting chemistry education and show the value of benchmarking frameworks for evaluating LLMs in specific domains.
Forward citations
Cited by 4 Pith papers
-
Divergence Decoding: Training-Free Capability Fusion
Divergence Decoding routes each token to either a domain specialist or a general reasoning LLM based on Jensen-Shannon divergence, outperforming either model alone on most tested scientific tasks.
-
Scientific-Intention Driven Embodied Intelligent Solar Telescope: Conceptual Design
A three-layer AI-agent design for intention-driven autonomous solar telescopes is proposed; only the precision temperature-control prototype was tested.
-
Domain Specific Benchmarks for Evaluating Multimodal Large Language Models
A review paper that organizes domain-specific MLLM benchmarks into an eight-discipline taxonomy, with summary tables and performance highlights.
-
Large Language Models Transform Organic Synthesis From Reaction Prediction to Automation
This survey claims to be the first systematic review of LLMs for organic synthesis, but its central 'evaluation' is never actually performed.
Discussion (0). Sign in to comment.