Pith. sign in

REVIEW 2 major objections 2 minor 93 references

LifeSentence: Language models can encode human life course trajectories from longitudinal panel data

T0 review · 2 major / 2 minor · reviewed 2026-06-30 · grok-4.3

Pith's one-line read A language model trained on life-event sequences from panel data predicts future outcomes three times better than prior methods and recovers social stratification patterns without supervision.

desk verdict LifeSentence shows an LLM can be tuned on converted SOEP life events to beat baselines on prediction and order tasks while surfacing known stratification patterns, but missing method details leave the source of the gains unclear. read the letter →

arxiv 2606.11220 v1 pith:ADZV65QU submitted 2026-05-14 cs.CL

classification cs.CL
keywords lifecoursetrajectorieslanguagemodelslongitudinalpaneldataeventpredictionsocialstratificationinstructiontuningchronologicalreconstruction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper presents LifeSentence, which converts discrete life events from longitudinal panel studies into structured natural-language descriptions and instruction-tunes a pretrained 24-billion-parameter language model on them. This approach supplements the limited panel data with knowledge from pretraining, yielding substantially higher accuracy on prediction, ordering, and reasoning tasks than classical statistics or other deep-learning baselines. The model also identifies known patterns such as the education premium, gender wage gap, and motherhood penalty directly from event sequences. Readers might care because accurate life-course modeling could support better forecasting of individual trajectories and new ways to explore how early events shape later outcomes.

What carries the argument

LifeSentence, the instruction-tuned language model that ingests life events as natural-language records and performs prediction, robustness, and reasoning tasks over them.

What would settle it

A controlled experiment that trains an otherwise identical model on the same panel data but without the natural-language event representation and then checks whether it still achieves comparable accuracy on joint prediction and still surfaces the education premium, gender wage gap, and motherhood penalty at similar rates.

Watch

Extended reading notes

Core claim

LifeSentence represents each life event as a structured natural-language record and instruction-tunes a 24B-parameter language model on roughly 65,000 individuals from the German Socio-Economic Panel. The resulting model outperforms baselines across an 18-task taxonomy, delivers a threefold improvement in joint event-and-timing prediction, achieves 91.2 percent Kendall's tau on chronological reconstruction from timestamp-free event sets, and recovers documented social stratification patterns from discrete sequences alone.

Load-bearing premise

Turning discrete life events into natural-language records lets the language model usefully add pretraining knowledge without the representation or tuning process creating distortions that would invalidate the performance gains or the recovered stratification patterns.

Editorial extensions

If this is right

  • Joint event-and-timing prediction improves threefold over the best prior baselines.
  • Chronological order can be reconstructed at 91.2 percent Kendall's tau from event sets that lack timestamps.
  • Documented stratification patterns emerge from event sequences without any explicit supervision on those patterns.
  • Natural-language queries allow direct exploration of connections between early-life histories and specified late-life endpoints.
  • The same model handles prediction, robustness checks, and reasoning tasks within a single trained system.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The natural-language interface could let researchers test counterfactual biographies that are difficult to express in traditional statistical models.
  • If the approach generalizes, other panel studies with similar event structures might yield comparable gains without needing massive new training sets.
  • The unsupervised recovery of stratification patterns suggests the model is capturing distributional regularities in life sequences that align with external sociological findings.
  • Privacy considerations would need examination if such models were applied to individual-level forecasting outside research settings.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The manuscript introduces LifeSentence, which converts longitudinal panel data from the German Socio-Economic Panel (~65k individuals) into structured natural-language event records and instruction-tunes a pretrained 24B-parameter LLM across an 18-task taxonomy covering prediction, robustness, and reasoning. It claims this approach outperforms classical and deep-learning baselines, with a threefold improvement in joint event-and-timing prediction and 91.2% Kendall's tau on chronological-order reconstruction from timestamp-stripped sequences. The work further asserts that, without explicit supervision, the model recovers documented social-stratification patterns (education premium, gender wage gap, motherhood penalty) from discrete event sequences alone and enables natural-language counterfactual queries.

Significance. If the reported gains and unsupervised pattern recoveries are shown to arise from the panel trajectories rather than pretraining distributional knowledge, the approach would demonstrate that LLMs can usefully augment small-scale longitudinal data for life-course modeling, potentially enabling new forms of predictive and counterfactual analysis that classical methods cannot support.

major comments (2)
  1. [Abstract] Abstract: the claim that stratification patterns (education premium, gender wage gap, motherhood penalty) are recovered 'from discrete event sequences alone' without explicit supervision lacks any reported control comparing outputs of the base 24B model (prior to SOEP instruction-tuning) against the fine-tuned LifeSentence model on identical prompts. Without this ablation, it remains possible that the patterns reflect pretraining co-occurrences rather than inferences drawn from the ~65k trajectories.
  2. [Abstract] Abstract and evaluation sections: performance numbers (threefold joint-prediction improvement, 91.2% Kendall's tau) are stated without any description of train/test splits on the 65k individuals, exact baseline implementations, hyperparameter matching, or statistical significance tests, preventing evaluation of whether the central empirical claims hold.
minor comments (2)
  1. The 18-task taxonomy is referenced but not enumerated; a table listing task definitions, input formats, and metrics would improve reproducibility.
  2. [Abstract] The statement that the approach uses 'roughly 45 times fewer' individuals than prior transformer methods should include the specific prior data sizes and citations for direct comparison.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive comments, which help clarify the presentation of our empirical claims. We address each major point below and commit to revisions that directly respond to the concerns raised.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the claim that stratification patterns (education premium, gender wage gap, motherhood penalty) are recovered 'from discrete event sequences alone' without explicit supervision lacks any reported control comparing outputs of the base 24B model (prior to SOEP instruction-tuning) against the fine-tuned LifeSentence model on identical prompts. Without this ablation, it remains possible that the patterns reflect pretraining co-occurrences rather than inferences drawn from the ~65k trajectories.

    Authors: We agree this ablation is necessary to isolate the contribution of the SOEP trajectories. In the revised manuscript we will add a direct comparison: the identical prompts used for LifeSentence will be run on the untuned 24B base model, with quantitative and qualitative differences in recovered stratification patterns reported. This will be placed in a new subsection of the results. revision: yes

  2. Referee: [Abstract] Abstract and evaluation sections: performance numbers (threefold joint-prediction improvement, 91.2% Kendall's tau) are stated without any description of train/test splits on the 65k individuals, exact baseline implementations, hyperparameter matching, or statistical significance tests, preventing evaluation of whether the central empirical claims hold.

    Authors: We acknowledge that the abstract and main evaluation sections currently omit these details. The methods section of the full manuscript already specifies an 80/20 individual-level split and lists the classical and deep-learning baselines, but we will expand both the abstract (within length limits) and the evaluation sections to include: (i) explicit train/test split ratios and leakage controls, (ii) precise baseline re-implementations and hyperparameter grids, (iii) matching criteria across models, and (iv) statistical significance results (bootstrap confidence intervals and paired tests). These additions will make the reported gains fully verifiable. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical results on held-out data against external baselines

full rationale

The paper presents an empirical ML study: instruction-tuning a pretrained LLM on ~65k SOEP trajectories, then reporting accuracy, Kendall tau, and unsupervised pattern recovery on held-out test sets against classical and deep-learning baselines. No equations, derivations, fitted parameters renamed as predictions, or self-citation chains appear in the provided text. All central claims (3x joint prediction improvement, 91.2% tau, recovery of stratification patterns) are measured quantities on external data, not quantities defined by the authors' own inputs. The pretraining-knowledge concern raised in the skeptic note is an interpretability question, not a circularity reduction.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The central claims rest on the untested transfer of distributional knowledge from general LLM pretraining to life-event sequences and on the assumption that the chosen text representation preserves the necessary sequential and causal structure.

assumptions (1)
  • domain assumption Pretrained large language models encode transferable distributional knowledge about human life courses that instruction tuning on panel data can usefully activate.
    This is the explicit bridging assumption stated in the abstract.

how reviews work

0 comments
Cite this review

Pith. "Pith review of LifeSentence: Language models can encode human life course trajectories from longitudinal panel data." pith.science (2026). https://pith.science/paper/ADZV65QU

@misc{pith2026260611220,
  author       = {Pith},
  title        = {Pith review of: LifeSentence: Language models can encode human life course trajectories from longitudinal panel data},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ADZV65QU}},
  note         = {Machine review of arXiv:2606.11220}
}
read the original abstract

Forecasting human life outcomes is important to gain insights into how individuals attain long and healthy lives. Conventional statistical approaches yield limited accuracy, potentially due to discarding the sequential structure of the life course. Modern methods such as transformer architectures require large scale training data that most longitudinal panel studies lack. Here we introduce LifeSentence, a model for life-course reasoning that bridges large language models with longitudinal panel data. By representing each life event as a structured natural-language record and instruction-tuning a pretrained 24-billion-parameter language model across an 18-task evaluation taxonomy spanning prediction, robustness and reasoning, LifeSentence supplements panel data with distributional knowledge already encoded during pretraining. Trained on approximately 65,000 individuals from the German Socio-Economic Panel - roughly 45 times fewer than prior transformer-based approaches - LifeSentence outperforms classical and deep learning baselines across all task families, achieving a threefold improvement in joint event-and-timing prediction from best baselines and 91.2% Kendall's tau when reconstructing chronological order from timestamp-stripped event sets. Without explicit supervision, the model recovers documented patterns of social stratification, including the education premium, the gender wage gap and the motherhood penalty, from discrete event sequences alone. A natural-language interface further enables qualitatively new research queries, such as connecting an early-life history to a specified late-life endpoint, establishing LifeSentence as both a predictive tool and a probe for counterfactual exploration of human biographies.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

93 extracted references · 93 canonical work pages

  1. [1]

    Vaillant, G. E. & Mukamal, K. Successful aging. Am. J. Psychiatry 158, 839–847 (2001)

  2. [2]

    & Seligman, M

    Diener, E. & Seligman, M. E. P. Very happy people. Psychol. Sci. 13, 81–84 (2002)

  3. [3]

    Moffitt, T. E. et al. A gradient of childhood self-control predicts health, wealth, and public safety. Proc. Natl Acad. Sci. USA 108, 2693–2698 (2011)

  4. [4]

    Salganik, M. J. et al. Measuring the predictability of life outcomes with a scientific mass collaboration. Proc. Natl Acad. Sci. USA 117, 8398–8403 (2020)

  5. [5]

    Grossmann, I. et al. Insights into the accuracy of social scientists'forecasts of societal change. Nat. Hum. Behav. 7, 484–501 (2023)

  6. [6]

    Oparina, E. et al. Machine learning in the prediction of human wellbeing. Sci. Rep. 15, 1632 (2025)

  7. [7]

    R., Raymo, J

    Halpern-Manners, A., Warren, J. R., Raymo, J. M. & Nicholson, D. A. The impact of work and family life histories on economic well-being at older ages. Social Forces 93, 1369–1396 (2015)

  8. [8]

    Wills, A. K. et al. Life course trajectories of systolic blood pressure using longitudinal data from eight UK cohorts. PLoS Medicine 8, e1000440 (2011)

Show all 93 references
  1. [9]

    Hayward, M. D. & Gorman, B. K. The long arm of childhood: the influence of early-life social conditions on men's mortality. Demography 41, 87–107 (2004)

  2. [10]

    F., Schafer, M

    Ferraro, K. F., Schafer, M. H. & Wilkinson, L. R. Childhood disadvantage and health problems in middle and later life: early imprints on physical health? American Sociological Review 81, 107–133 (2016)

  3. [11]

    E., Shuey, K

    Willson, A. E., Shuey, K. M. & Elder, G. H. Jr. Cumulative advantage processes as mechanisms of inequality in life course health. American Journal of Sociology 112, 1886–1924 (2007). LifeSentence: language models for life-course trajectories21

  4. [12]

    S., Catalano, R

    Rook, K. S., Catalano, R. & Dooley, D. The timing of major life events: effects of departing from the social clock. American Journal of Community Psychology 17, 233–258 (1989)

  5. [13]

    Elder, G. H., Jr. The life course as developmental theory. Child Dev. 69, 1–12 (1998)

  6. [14]

    Elder, G. H., Jr. & Shanahan, M. J. The life course and human development. In Handbook of Child Psychology Vol. 1 (ed. Lerner, R. M.) 665–715 (Wiley, 2006)

  7. [15]

    Topel, R. H. & Ward, M. P. Job mobility and the careers of young men. Q. J. Econ. 107, 439–479 (1992)

  8. [16]

    W., Kaplan, G

    Lynch, J. W., Kaplan, G. A. & Shema, S. J. Cumulative impact of sustained economic hardship on physical, cognitive, psychological, and social functioning. N. Engl. J. Med. 337, 1889–1895 (1997)

  9. [17]

    Vaswani, A. et al. Attention is all you need. In Advances in Neural Information Processing Systems 30 (NeurIPS) (2017)

  10. [18]

    & Toutanova, K

    Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proc. NAACL-HLT 2019, 4171–4186 (2019)

  11. [19]

    Shmatko, A. et al. Learning the natural history of human disease with generative transformers. Nature 647, 248–256 (2025)

  12. [20]

    Renc, P. et al. Zero shot health trajectory prediction using transformer. npj Digit. Med. 7, 256 (2024)

  13. [21]

    & Blei, D

    Vafa, K., Palikot, E., Du, T., Kanodia, A., Athey, S. & Blei, D. M. CAREER: a foundation model for labor sequence data. Preprint at https://arxiv.org/abs/2202.08370 (2024)

  14. [22]

    Savcisens, G. et al. Using sequences of life-events to predict human lives. Nat. Comput. Sci. 4, 43–56 (2024)

  15. [23]

    D., Raghavan, S., Deng, F., Lang, M., Succi, M

    Yang, E., Li, M. D., Raghavan, S., Deng, F., Lang, M., Succi, M. D., Huang, A. J. & Kalpathy-Cramer, J. Transformer versus traditional natural language processing: how much data is enough for automated radiology report classification? Br. J. Radiol. 96, 20220769 (2023)

  16. [24]

    & Varoquaux, G

    Grinsztajn, L., Oyallon, E. & Varoquaux, G. Why do tree-based models still outperform deep learning on typical tabular data? in Advances in Neural Information Processing Systems 35 (NeurIPS, 2022)

  17. [25]

    Raffel, C. et al. Exploring the limits of transfer learning with a unified text-to-text transfer transformer. J. Mach. Learn. Res. 21, 1–67 (2020)

  18. [26]

    Wei, J. et al. Finetuned language models are zero-shot learners. In Proc. International Conference on Learning Representations (2022)

  19. [27]

    & Athey, S

    Du, T., Kanodia, A., Brunborg, H., Vafa, K. & Athey, S. LABOR-LLM: language-based occupational representations with large language models. Preprint at https://arxiv.org/abs/2406.17972 (2024)

  20. [28]

    Aribandi, V. et al. ExT5: towards extreme multi-task scaling for transfer learning. In Proc. International Conference on Learning Representations (2022)

  21. [29]

    J., Yang, V., Edin, K

    Lundberg, I., Brown-Weinstock, R., Clampet-Lundquist, S., Pachman, S., Nelson, T. J., Yang, V., Edin, K. & Salganik, M. J. The origins of unpredictability in life outcome prediction tasks. Proc. Natl Acad. Sci. USA 121, e2322973121 (2024)

  22. [30]

    R., Kim, C

    Tamborini, C. R., Kim, C. & Sakamoto, A. Education and lifetime earnings in the United States. Demography 52, 1383–1407 (2015)

  23. [31]

    & Salvanes, K

    Bhuller, M., Mogstad, M. & Salvanes, K. G. Life-cycle earnings, education premiums, and internal rates of return. J. Labor Econ. 35, 993–1030 (2017)

  24. [32]

    & Katz, L

    Bertrand, M., Goldin, C. & Katz, L. F. Dynamics of the gender gap for young professionals in the financial and corporate sectors. Am. Econ. J. Appl. Econ. 2, 228–255 (2010)

  25. [33]

    Barth, E., Kerr, S. P. & Olivetti, C. The dynamics of gender earnings differentials: evidence from establishment data. Eur. Econ. Rev. 134, 103713 (2021)

  26. [34]

    & Søgaard, J

    Kleven, H., Landais, C. & Søgaard, J. E. Children and gender inequality: evidence from Denmark. Am. Econ. J. Appl. Econ. 11, 181–209 (2019)

  27. [35]

    & Machado, C

    Almond, D., Cheng, Y. & Machado, C. Large motherhood penalties in US administrative microdata. Proc. Natl Acad. Sci. USA 120, e2209740120 (2023)

  28. [36]

    Goebel, J. et al. The German Socio-Economic Panel (SOEP). Jahrb. Natl.ökon. Stat. 239, 345–360 (2019)

  29. [37]

    G., Frick, J

    Wagner, G. G., Frick, J. R. & Schupp, J. The German Socio-Economic Panel Study (SOEP) – scope, evolution and enhancements. Schmollers Jahrb. 127, 139–169 (2007)

  30. [38]

    & Tegmark, M

    Gurnee, W. & Tegmark, M. Language models represent space and time. In Proc. International Conference on Learning Representations (2024)

  31. [39]

    & Zettlemoyer, L

    Aghajanyan, A., Gupta, S. & Zettlemoyer, L. Intrinsic dimensionality explains the effectiveness of language model fine-tuning. In Proc. 59th Annual Meeting of the Association for Computational Linguistics 7319–7328 (ACL, 2021). LifeSentence: language models for life-course tra...

  32. [40]

    Hu, E. J. et al. LoRA: low-rank adaptation of large language models. In Proc. 10th International Conference on Learning Representations https://openreview.net/forum?id=nZeVKeeFYf9 (2022)

  33. [41]

    Shaw, C. et al. Comparison of imputation strategies for incomplete longitudinal data in life-course epidemiology. Am. J. Epidemiol. 192, 2075–2084 (2023)

  34. [42]

    A., Wu, K., Sim, A

    Lazar, A., Jin, L., Spurlock, C. A., Wu, K., Sim, A. & Todd, A. Evaluating the effects of missing values and mixed data types on social sequence clustering using t-SNE visualization. J. Data Inf. Qual. 11, 7:1–7:22 (2019)

  35. [43]

    Rindfuss, R. R. The young adult years: diversity, structural change, and fertility. Demography 28, 493–512 (1991)

  36. [44]

    Manning, W. D. Young adulthood relationships in an era of uncertainty: a case for cohabitation. Demography 57, 799–819 (2020)

  37. [45]

    R., Mortimer, J

    Eliason, S. R., Mortimer, J. T. & Vuolo, M. The transition to adulthood: life course structures and subjective perceptions. Soc. Psychol. Q. 78, 205–227 (2015)

  38. [46]

    Elzinga, C. H. Complexity of categorical time series. Sociol. Methods Res. 38, 463–481 (2010)

  39. [47]

    Family trajectories across time and space: increasing complexity in family life courses in Europe? Demography 55, 135–164 (2018)

    Van Winkle, Z. Family trajectories across time and space: increasing complexity in family life courses in Europe? Demography 55, 135–164 (2018)

  40. [48]

    The distribution of the flora in the alpine zone

    Jaccard, P. The distribution of the flora in the alpine zone. New Phytologist 11, 37–50 (1912)

  41. [49]

    Levenshtein, V. I. Binary codes capable of correcting deletions, insertions, and reversals. Sov. Phys. Dokl. 10, 707–710 (1966)

  42. [50]

    Vaserstein, L. N. Markov processes over denumerable products of spaces, describing large systems of automata. Probl. Peredachi Inf. 5, 64–72 (1969)

  43. [51]

    & Winter, D

    Levandowsky, M. & Winter, D. Distance between sets. Nature 234, 34–35 (1971)

  44. [52]

    & Grunow, D

    Begall, K. & Grunow, D. Labour force transitions around first childbirth in the Netherlands. Eur. Sociol. Rev. 31, 697–712 (2015)

  45. [53]

    & Blane, D

    Hallqvist, J., Lynch, J., Bartley, M., Lang, T. & Blane, D. Can we disentangle life course processes of accumulation, critical period and social mobility? An analysis of disadvantaged socio-economic positions and myocardial infarction in the Stockholm Heart Epidemiology Progra...

  46. [54]

    T., Pavlick, E

    McCoy, R. T., Pavlick, E. & Linzen, T. Right for the wrong reasons: diagnosing syntactic heuristics in natural language inference. Proc. 57th Annu. Meet. Assoc. Comput. Linguist. 3428–3448 (2019)

  47. [55]

    & Kao, H.-Y

    Niven, T. & Kao, H.-Y. Probing neural network comprehension of natural language arguments. Proc. 57th Annu. Meet. Assoc. Comput. Linguist. 4658–4664 (2019)

  48. [56]

    & Mayer, K

    Brückner, H. & Mayer, K. U. De-standardization of the life course: what it might mean? And if it means anything, whether it actually took place? Adv. Life Course Res. 9, 27–53 (2005)

  49. [57]

    Kendall, M. G. A new measure of rank correlation. Biometrika 30, 81–93 (1938)

  50. [58]

    Hirschberg, D. S. A linear space algorithm for computing maximal common subsequences. Commun. ACM 18, 341–343 (1975)

  51. [59]

    Petroni, F. et al. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) 2463–2473 (Association for Computational Li...

  52. [60]

    & Liu, Q

    Wu, Z., Chen, Y., Kao, B. & Liu, Q. Perturbed masking: parameter-free probing for analyzing and interpreting BERT. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics 4166–4176 (Association for Computational Linguistics, 2020)

  53. [61]

    & Berant, J

    Talmor, A., Elazar, Y., Goldberg, Y. & Berant, J. oLMpics – on what language model pre-training captures. Trans. Assoc. Comput. Linguist. 8, 743–758 (2020)

  54. [62]

    Zeiler, M. D. & Fergus, R. Visualizing and understanding convolutional networks. In Computer Vision – ECCV 2014 (eds Fleet, D. et al.) 818–833 (Springer, 2014)

  55. [63]

    Elder, G. H. Jr. Military times and turning points in men's lives. Dev. Psychol. 22, 233–245 (1986)

  56. [64]

    Sampson, R. J. & Laub, J. H. Socioeconomic achievement in the life course of disadvantaged men: military service as a turning point, circa 1940–1965. Am. Sociol. Rev. 61, 347–367 (1996)

  57. [65]

    E., Diener, E., Georgellis, Y

    Clark, A. E., Diener, E., Georgellis, Y. & Lucas, R. E. Lags and leads in life satisfaction: a test of the baseline hypothesis. The Economic Journal 118, F222–F243 (2008)

  58. [66]

    S., LaLonde, R

    Jacobson, L. S., LaLonde, R. J. & Sullivan, D. G. Earnings losses of displaced workers. American Economic Review 83, 685–709 (1993)

  59. [67]

    & von Wachter, T

    Sullivan, D. & von Wachter, T. Job displacement and mortality: an analysis using administrative data. The Quarterly Journal of Economics 124, 1265–1306 (2009)

  60. [68]

    & Sauzet, O

    Stacherl, B. & Sauzet, O. Chronic disease onset and wellbeing development: longitudinal analysis and the role of healthcare access. Eur. J. Public Health 34, 29–34 (2024). LifeSentence: language models for life-course trajectories23

  61. [69]

    & Brouwer, W

    Stöckel, J., van Exel, J. & Brouwer, W. B. F. Adaptation in life satisfaction and self-assessed health to disability—evidence from the UK. Soc. Sci. Med. 328, 115996 (2023)

  62. [70]

    & Chen, D

    Gao, T., Fisch, A. & Chen, D. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP) 3816–38...

  63. [71]

    & Shazeer, N

    Roberts, A., Raffel, C. & Shazeer, N. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) 5418–5426 (Association for Computational Linguistics, 2020)

  64. [72]

    & Sontag, D

    Hegselmann, S., Buendia, A., Lang, H., Agrawal, M., Jiang, X. & Sontag, D. TabLLM: Few-shot Classification of Tabular Data with Large Language Models. Proc. 26th International Conference on Artificial Intelligence and Statistics (AISTATS) PMLR 206, 5549–5581 (2023)

  65. [73]

    & Lee, K

    Dinh, T., Zeng, Y., Zhang, R., Lin, Z., Gira, M., Rajput, S., Sohn, J., Papailiopoulos, D. & Lee, K. LIFT: Language-Interfaced Fine-Tuning for Non-Language Machine Learning Tasks. Advances in Neural Information Processing Systems 35, 11763–11784 (2022)

  66. [74]

    & Ruder, S

    Howard, J. & Ruder, S. Universal Language Model Fine-tuning for Text Classification. Proc. 56th Annual Meeting of the Association for Computational Linguistics 1, 328–339 (2018)

  67. [75]

    & Fasang, A

    Aisenbrey, S. & Fasang, A. E. The interplay of work and family trajectories over the life course: Germany and the United States in comparison. Am. J. Sociol. 122, 1448–1484 (2017)

  68. [76]

    Bratberg, E., Davis, J., Mazumder, B., Nybom, M., Schnitzlein, D. D. & Vaage, K. A comparison of intergenerational mobility curves in Germany, Norway, Sweden, and the US. Scand. J. Econ. 119, 72–101 (2017)

  69. [77]

    & Bolan, M

    Stovel, K. & Bolan, M. Residential trajectories: using optimal alignment to reveal the structure of residential mobility. Sociol. Methods Res. 32, 559–598 (2004)

  70. [78]

    Incarceration as exposure: the prison, infectious disease, and other stress-related illnesses

    Massoglia, M. Incarceration as exposure: the prison, infectious disease, and other stress-related illnesses. J. Health Soc. Behav. 49, 56–71 (2008)

  71. [79]

    & Dean, J

    Hinton, G., Vinyals, O. & Dean, J. Distilling the knowledge in a neural network. Preprint at https://arxiv.org/abs/1503.02531 (2015)

  72. [80]

    & Weinberger, K

    Guo, C., Pleiss, G., Sun, Y. & Weinberger, K. Q. On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning 1321–1330 (PMLR, 2017)

  73. [81]

    & Farquhar, S

    Kuhn, L., Gal, Y. & Farquhar, S. Semantic uncertainty: linguistic invariances for uncertainty estimation in natural language generation. In Proceedings of the 11th International Conference on Learning Representations (ICLR, 2023)

  74. [82]

    & Candès, E

    Romano, Y., Patterson, E. & Candès, E. J. Conformalized quantile regression. In Advances in Neural Information Processing Systems 32, 3538–3548 (NeurIPS, 2019)

  75. [83]

    Y., Saligrama, V

    Bolukbasi, T., Chang, K.-W., Zou, J. Y., Saligrama, V. & Kalai, A. T. Man is to computer programmer as woman is to homemaker? Debiasing word embeddings. In Advances in Neural Information Processing Systems 29, 4349–4357 (NeurIPS, 2016)

  76. [84]

    Caliskan, A., Bryson, J. J. & Narayanan, A. Semantics derived automatically from language corpora contain human-like biases. Science 356, 183–186 (2017)

  77. [85]

    & Chang, K.-W

    Zhao, J., Wang, T., Yatskar, M., Ordonez, V. & Chang, K.-W. Gender bias in coreference resolution: evaluation and debiasing methods. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologie...

  78. [86]

    & Zettlemoyer, L

    Dettmers, T., Pagnoni, A., Holtzman, A. & Zettlemoyer, L. QLoRA: Efficient Finetuning of Quantized LLMs. Adv. Neural Inf. Process. Syst. 36, 10088–10115 (2023)

  79. [87]

    & Hutter, F

    Loshchilov, I. & Hutter, F. Decoupled weight decay regularization. in International Conference on Learning Representations (ICLR, 2019)

  80. [88]

    & Yarats, D

    Ma, J. & Yarats, D. On the adequacy of untuned warmup for adaptive optimization. Proc. AAAI Conf. Artif. Intell. 35, 8828–8836 (2021)

  81. [89]

    L., Kindermans, P.-J., Ying, C

    Smith, S. L., Kindermans, P.-J., Ying, C. & Le, Q. V. Don't decay the learning rate, increase the batch size. in International Conference on Learning Representations (ICLR, 2018)

  82. [90]

    & Guestrin, C

    Chen, T. & Guestrin, C. XGBoost: a scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 785–794 (ACM, 2016)

  83. [91]

    & Schmidhuber, J

    Hochreiter, S. & Schmidhuber, J. Long short-term memory. Neural Comput. 9, 1735–1780 (1997)

  84. [92]

    Cho, K. et al. Learning phrase representations using RNN encoder–decoder for statistical machine translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) 1724–1734 (Association for Computational Linguistics, 2014). LifeSent...

  85. [93]

    employment_type

    Yao, Y., Rosasco, L. & Caponnetto, A. On early stopping in gradient descent learning. Constr. Approx. 26, 289–315 (2007). LifeSentence — Supplementary Information25 Supplementary Information Language models can encode human life course trajectories from longitudinal panel data...

Pith tools

Reviewed June 30, 2026 · model on record in the stance chip above.