Pith. sign in

REVIEW 4 cited by

Chronologically Consistent Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.21206 v3 pith:IMJYCWRC submitted 2025-02-28 q-fin.GN q-fin.TR

classification q-fin.GNq-fin.TR
keywords languagemodelmodelstrainingchronologicallyconsistentdatabias
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models are increasingly used in social sciences, but their training data can introduce lookahead bias and training leakage. A good chronologically consistent language model requires efficient use of training data to maintain accuracy despite time-restricted data. Here, we overcome this challenge by training a suite of chronologically consistent large language models, ChronoBERT and ChronoGPT, which incorporate only the text data that would have been available at each point in time. Despite this strict temporal constraint, our models achieve strong performance on natural language processing benchmarks, outperforming or matching widely used models (e.g., BERT), and remain competitive with larger open-weight models. Lookahead bias is model and application-specific because even if a chronologically consistent language model has poorer language comprehension, a regression or prediction model applied on top of the language model can compensate. In an asset pricing application predicting next-day stock returns from financial news, we find that ChronoBERT and ChronoGPT's real-time outputs achieve Sharpe ratios comparable to a much larger Llama model, indicating that lookahead bias is modest. Our results demonstrate a scalable, practical framework to mitigate training leakage, ensuring more credible backtests and predictions across finance and other social science domains.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Talking to Digital Twins: Selective Disclosure and Belief Measurement in Financial Social Media

    econ.GN 2026-08 conditional novelty 7.0 of 10

    Daily real-time LLM digital-twin interviews of finfluencer accounts predict cross-sectional large-cap returns over the next ten trading days, mainly in the silent region with no concurrent public post.

  2. HindsightBench: A Black-Box Behavioral Audit Protocol for Parametric Hindsight in Time-Indexed LLM Decision Tasks

    cs.LG 2026-07 conditional novelty 6.0 of 10

    An integrated black-box audit protocol profiles parametric hindsight in LLMs and shows the date-trigger reflex tracks training generation, not scale, while effective knowledge cutoffs span 22 months.

  3. Scaling Point-in-Time Language Models

    cs.CL 2026-04 conditional novelty 5.5 of 10

    Scaling point-in-time LLMs to 4B parameters and 1T temporally filtered tokens narrows the gap to unrestricted models to about 8–11 average points and yields positive out-of-sample Sharpe ratios from news embeddings.

  4. DatedGPT: Preventing Lookahead Bias in Large Language Models with Time-Aware Pretraining

    cs.CL 2026-03 reject novelty 5.0 of 10

    Twelve yearly-cutoff language models are trained to prevent lookahead bias, but the headline finance result in the abstract does not appear in the paper.

Pith tools