Pith. sign in

REVIEW 5 cited by

xVal: A Continuous Numerical Tokenization for Scientific Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.02989 v2 pith:EN7KHWA7 submitted 2023-10-04 stat.ML cs.AIcs.CLcs.LG

classification stat.MLcs.AIcs.CLcs.LG
keywords modelsscientificlanguagedatasetsxvalnumbersnumericaltext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Due in part to their discontinuous and discrete default encodings for numbers, Large Language Models (LLMs) have not yet been commonly used to process numerically-dense scientific datasets. Rendering datasets as text, however, could help aggregate diverse and multi-modal scientific data into a single training corpus, thereby potentially facilitating the development of foundation models for science. In this work, we introduce xVal, a strategy for continuously tokenizing numbers within language models that results in a more appropriate inductive bias for scientific applications. By training specially-modified language models from scratch on a variety of scientific datasets formatted as text, we find that xVal generally outperforms other common numerical tokenization strategies on metrics including out-of-distribution generalization and computational efficiency.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 7 citations worldwide. Full citation record

  1. multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data

    cs.LG 2025-05 conditional novelty 6.0 of 10

    multivariateGPT extends next-token prediction to jointly predict the class and continuous value of mixed categorical and numeric time series, with Gaussian uncertainty, and outperforms discrete-token baselines on clin...

  2. POI-Enhancer: An LLM-based Semantic Enhancement Framework for POI Representation Learning

    cs.AI 2025-02 conditional novelty 6.0 of 10

    POI-Enhancer uses LLM-generated text features and attention-based fusion to improve POI embeddings from six classic models, gaining consistent accuracy on three real-world mobility datasets.

  3. A Multimodal PDE Foundation Model for Prediction and Scientific Text Descriptions

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A multimodal transformer predicts ODE/PDE solutions and generates correct scientific text descriptions from numerical and symbolic inputs, with low error on in-distribution and out-of-distribution tests.

  4. 3DMolFormer: A Dual-channel Framework for Structure-based Drug Discovery

    cs.CE 2025-02 conditional novelty 6.0 of 10

    A dual-channel transformer that reads and writes 3D coordinates as continuous numbers alongside chemical tokens achieves state-of-the-art docking and pocket-aware molecule generation.

  5. FinTeam: A Multi-Agent Collaborative Intelligence System for Comprehensive Financial Scenarios

    cs.CE 2025-07 conditional novelty 5.0 of 10

    A four-agent LLM pipeline trained with role-specific data improves human preference on comprehensive Chinese financial analysis tasks.

Pith tools