Pith. sign in

REVIEW 2 cited by

Holistic chemical evaluation reveals pitfalls in reaction prediction models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.09004 v1 pith:ZLN7HS42 submitted 2023-12-14 physics.chem-ph cs.LG

classification physics.chem-phcs.LG
keywords chemicalevaluationholisticmodelmodelspredictionlimitationsmetrics
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The prediction of chemical reactions has gained significant interest within the machine learning community in recent years, owing to its complexity and crucial applications in chemistry. However, model evaluation for this task has been mostly limited to simple metrics like top-k accuracy, which obfuscates fine details of a model's limitations. Inspired by progress in other fields, we propose a new assessment scheme that builds on top of current approaches, steering towards a more holistic evaluation. We introduce the following key components for this goal: CHORISO, a curated dataset along with multiple tailored splits to recreate chemically relevant scenarios, and a collection of metrics that provide a holistic view of a model's advantages and limitations. Application of this method to state-of-the-art models reveals important differences on sensitive fronts, especially stereoselectivity and chemical out-of-distribution generalization. Our work paves the way towards robust prediction models that can ultimately accelerate chemical discovery.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Agentic generation of verifiable rules for deterministic, self-expanding reaction classification

    cs.AI 2026-07 unverdicted novelty 7.0 of 10

    Multi-agent LLMs generate and verify 14,073 deterministic reaction rules from 665,901 patents, enabling 97.7% classification of unseen reactions with finer resolution than fixed proprietary systems.

  2. Challenging reaction prediction models to generalize to novel chemistry

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A Transformer reaction predictor evaluated on document-, author-, time-, and reaction-class splits reveals that in-distribution benchmarks overstate real-world accuracy and that models only partially extrapolate to un...

Pith tools