REVIEW 3 cited by
We need to talk about random seeds
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Modern neural network libraries all take as a hyperparameter a random seed, typically used to determine the initial state of the model parameters. This opinion piece argues that there are some safe uses for random seeds: as part of the hyperparameter search to select a good model, creating an ensemble of several models, or measuring the sensitivity of the training algorithm to the random seed hyperparameter. It argues that some uses for random seeds are risky: using a fixed random seed for "replicability" and varying only the random seed to create score distributions for performance comparison. An analysis of 85 recent publications from the ACL Anthology finds that more than 50% contain risky uses of random seeds.
Forward citations
Cited by 3 Pith papers
-
Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence
An agentic two-stage optimizer (LLM-generated question scoring plus Bayesian hyperparameter search) improves image-to-video prompt adherence, winning human preference tests up to 69% over random search.
-
A neural network approach to learning solutions of a class of elliptic variational inequalities
A weak adversarial neural network method, based on regularized gap functions, solves elliptic obstacle problems, including nonsymmetric and biactive cases, with an a priori error analysis.
-
Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility
The paper formalizes test-time scaling into three regimes, introduces a discovery-stability profile for repeated-sampling evaluation, and releases nearly two million reasoning traces.
Discussion (0). Continue with ORCID to comment.