REVIEW 5 cited by
Transformer models: an introduction and catalog
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In the past few years we have seen the meteoric appearance of dozens of foundation models of the Transformer family, all of which have memorable and sometimes funny, but not self-explanatory, names. The goal of this paper is to offer a somewhat comprehensive but simple catalog and classification of the most popular Transformer models. The paper also includes an introduction to the most important aspects and innovations in Transformer models. Our catalog will include models that are trained using self-supervised learning (e.g., BERT or GPT3) as well as those that are further trained using a human-in-the-loop (e.g. the InstructGPT model used by ChatGPT).
Forward citations
Cited by 5 Pith papers
-
Machine Learning for the Cluster Reconstruction in the CALIFA Calorimeter at R3B
An edge-detection neural network combined with agglomerative pre-clustering improves cluster reconstruction accuracy in simulated CALIFA calorimeter events by about 34% relative to the standard R3B algorithm.
-
Bridging the Evaluation Gap: Leveraging Large Language Models for Topic Model Evaluation
LLM-based metrics for coherence, repetitiveness, diversity, and topic-document alignment rate topic models, but scores shift substantially depending on which LLM does the judging.
-
IG-GAN: A Generative Adversarial Network for Aerodynamic Data Generation Based on Intrinsic Geometry
A Bézier-coefficient generator plus RBF discriminator yields 97% and 83% lower test MSE than SSL-Transformer on Burgers’ velocity and ONERA M6 aerodynamic coefficients.
-
A Hybrid Framework for Subject Analysis: Integrating Embedding-Based Regression Models with Large Language Models
Using an ML-predicted label count to constrain LLM generation and post-editing outputs to the LCSH vocabulary lifts subject-heading prediction F1 from 0.135 to 0.300 on a 2,100-book test set.
-
AttentionSmithy: A Modular Framework for Rapid Transformer Development and Customization
A modular transformer framework is introduced and validated by replicating the original transformer, running neural architecture search on WMT14 translation, and fine-tuning a BERT-style model for cell type classification.
Discussion (0). Continue with ORCID to comment.