Pith. sign in

REVIEW 1 cited by

GraphBPE: Molecular Graphs Meet Byte-Pair Encoding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.19039 v1 pith:WPODIVM7 submitted 2024-07-26 cs.LG cs.AIphysics.chem-phq-bio.BM

classification cs.LGcs.AIphysics.chem-phq-bio.BM
keywords moleculardifferentgraphbpegraphsmodelpreprocessingarchitecturesboost
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the increasing attention to molecular machine learning, various innovations have been made in designing better models or proposing more comprehensive benchmarks. However, less is studied on the data preprocessing schedule for molecular graphs, where a different view of the molecular graph could potentially boost the model's performance. Inspired by the Byte-Pair Encoding (BPE) algorithm, a subword tokenization method popularly adopted in Natural Language Processing, we propose GraphBPE, which tokenizes a molecular graph into different substructures and acts as a preprocessing schedule independent of the model architectures. Our experiments on 3 graph-level classification and 3 graph-level regression datasets show that data preprocessing could boost the performance of models for molecular graphs, and GraphBPE is effective for small classification datasets and it performs on par with other tokenization methods across different model architectures.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ECG-Byte: A Tokenizer for End-to-End Generative Electrocardiogram Language Modeling

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A BPE-based tokenizer lets an LLM generate clinical text directly from quantized ECG signals, matching two-stage encoder methods with roughly 3x faster training and 48% of the data.

Pith tools