Pith. sign in

REVIEW 4 cited by

CAREER: A Foundation Model for Labor Sequence Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.08370 v4 pith:IPACBRHJ submitted 2022-02-16 cs.LG econ.EM

classification cs.LGecon.EM
keywords careerdatasetsdatasurveyeconometricmodelmodelspredictions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Labor economists regularly analyze employment data by fitting predictive models to small, carefully constructed longitudinal survey datasets. Although machine learning methods offer promise for such problems, these survey datasets are too small to take advantage of them. In recent years large datasets of online resumes have also become available, providing data about the career trajectories of millions of individuals. However, standard econometric models cannot take advantage of their scale or incorporate them into the analysis of survey data. To this end we develop CAREER, a foundation model for job sequences. CAREER is first fit to large, passively-collected resume data and then fine-tuned to smaller, better-curated datasets for economic inferences. We fit CAREER to a dataset of 24 million job sequences from resumes, and adjust it on small longitudinal survey datasets. We find that CAREER forms accurate predictions of job sequences, outperforming econometric baselines on three widely-used economics datasets. We further find that CAREER can be used to form good predictions of other downstream variables. For example, incorporating CAREER into a wage model provides better predictions than the econometric models currently in use.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Book of Life approach: Enabling richness and scale for life course research

    cs.CL 2025-07 conditional novelty 6.0 of 10

    The authors build and release an open-source toolkit that turns Dutch population registry data into millions of plain-text personal life histories for LLM-based analysis.

  2. Employee Turnover Prediction: A Cross-component Attention Transformer with Consideration of Competitor Influence and Contagious Effect

    cs.LG 2025-01 reject novelty 6.0 of 10

    A cross-component attention transformer predicts individual employee turnover across firms by combining employee, company, competitor, and departed-colleague signals, reporting large top-K precision gains over baselines.

  3. Predictions as Surrogates: Revisiting Surrogate Outcomes in the Age of AI

    stat.ML 2025-01 accept novelty 6.0 of 10

    Recalibrated prediction-powered inference estimates the optimal imputed loss by cross-fitted machine learning, always improving on label-only inference and matching the best possible variance among PPI estimators when...

  4. KARRIEREWEGE: A Large Scale Career Path Prediction Dataset

    cs.CL 2024-12 conditional novelty 6.0 of 10

    KARRIEREWEGE is a large public career path dataset with ESCO mapping and LLM-synthesized free-text titles for improving career trajectory prediction.

Pith tools