{"id":"9c9b2e87-ccfb-4707-a8ce-077e21cac071","arxiv_id":"2608.06048","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Name popularity in the US and France is extremely unequal and has stayed that way for over a century; the authors fit this pattern with a two-parameter Rayleigh-Jeans condensation model borrowed from wealth physics.","lead":"Using 100+ years of US and French baby name records, the authors show that name popularity is as unequal as global wealth: half of all names account for about 1% of births, and the pattern barely changes over a century. They argue this stability reflects a universal 'Rayleigh-Jeans condensation' mechanism previously applied to wealth, energy, and voting, but the argument rests on curve fitting with two tuned parameters.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central causal claim that RJ condensation is the driving mechanism for name-inequality rests on a two-parameter curve fit; after matching the Gini, only one shape parameter remains, and no name-choice dynamics with conserved integrals is provided.","rationale":"The reader's CONDITIONAL verdict is appropriate. The descriptive part of the paper is solid: the Lorenz and Pareto curves and the Gini stability are direct computations on public administrative data, and the time-correlation analysis is a reasonable auxiliary observation. The weak point is exactly the inference from RJE curve agreement to a thermalization mechanism. The fitting protocol fixes epsilon through the Gini and a through the Lorenz distance, leaving the Pareto check as a consistency test of the same distribution; it does not validate the RJ form against other flexible families. Moreover, the RJ distribution is derived from two conserved integrals, but no name-choice dynamics that conserve norm and energy is given, and the parameters are refit independently each year, so the conservation relations are not tested across time. Section VI's own admission that a two-parameter fit is not a sufficient argument makes the subsequent claim of fundamental confirmation an overreach. The concern is about evidential support rather than internal inconsistency: the model is not self-contradictory, and the authors may well be correct. The cleanest way to decide whether the specific RJ functional form is doing identifiable work is a null-model comparison. Since the reader already conditioned on this weakness, no change to the verdict is needed.","tokens_in":17569,"tokens_out":6188,"duration_ms":67925,"concrete_test":"Refit the same US-F and FR-F yearly datasets with a two-parameter null model that has no conserved quantities, using the same protocol: fix the scale by the observed mean frequency, choose the shape parameter to minimize the geometric distance to the data Lorenz curve, and report the Lorenz distance and the Pareto curve it produces. A suitable null is a truncated Zipf distribution or a lognormal distribution with a threshold. If the null model achieves an average Lorenz distance comparable to the reported RJE values below 5e-3 and reproduces the stable Gini range, then the RJ functional form is not identifiable from these curves and the causal conclusion is unsupported. If RJE is substantially closer on out-of-sample years (fit on year t, predict year t+1), the concern is weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The empirical description of the data is probably sound: the Lorenz and Pareto curves are direct statistics on public data, and their reported stability is a real observation. The load-bearing step is the inference in Section VI that agreement between the RJE model and these curves means RJ condensation is the driving mechanism. In the fitting protocol of Section IV, the parameter epsilon is fixed by forcing the model Gini to equal the data Gini, and a is chosen to minimize the geometric distance to the data Lorenz curve. The Pareto curves use the same fitted distribution, so they are a consistency check on the same fitted object, not an independent prediction. No equation of motion, conservation law, or relaxation timescale for the name-choice process is specified; the two integrals of motion are asserted by analogy in Section II, and the parameters are re-fit separately for each year. The paper itself states in Section VI that a two-parameter fit 'is not a sufficient argument,' yet the same paragraph concludes that universality across systems 'gives the fundamental confirmation of the validity of the RJ thermalization and condensation theory.' That is the central logical gap. The claim may be true, but the evidence as presented cannot distinguish RJE from any other flexible two-parameter distribution family that fits the same Lorenz shape.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes official US and French baby-name frequency data over more than a century. It constructs Lorenz curves, computes Gini coefficients, and reports Pareto curves, finding that inequality in name frequencies is stable, with Gini values in the range 0.85–0.95. The authors then fit an extended Rayleigh-Jeans (RJE) model: the parameter epsilon is fixed by matching the model Gini coefficient to the data Gini coefficient, and the parameter a is chosen to minimize the geometric distance between the model and data Lorenz curves. The same fitted distribution is used to produce model Pareto curves. The paper also studies time correlations of top names using Pearson coefficients. The authors conclude that RJ condensation is the driving mechanism for the observed inequality in name frequencies.","tokens_in":17748,"tokens_out":3495,"duration_ms":38763,"significance":"The descriptive part of the paper is a solid empirical contribution: the documented stability of name-frequency Lorenz and Pareto curves over 100+ years, using public administrative data, is a real and interesting observation. The time-correlation analysis adds a useful cultural-historical dimension. If the causal claim were established, the paper would significantly extend the wealth-thermalization hypothesis to a new class of social data. However, the current evidence for the central claim is a two-parameter curve fit to the same data used for fitting; no name-choice dynamics, conserved quantities, or out-of-sample predictions are provided. The manuscript is transparent about the fitting procedure, but the inference in Section VI goes beyond what the evidence can support.","major_comments":[{"comment":"The fitting protocol determines epsilon by the condition that the RJE Lorenz curve has the same Gini coefficient as the real data Lorenz curve, and it determines a by minimizing the geometric distance between the RJE and real Lorenz curves. Therefore the model Lorenz curve is guaranteed to match in one parameter and is optimized in the other, and the Pareto curves shown with the same parameters are a consistency check on the same fitted object rather than an independent prediction. The statement in Section VI that 'the RJ condensation also provides a very good description of Lorenz and Pareto curves' is consequently not independent evidence for the mechanism. To make this load-bearing inference, the paper needs at least one out-of-sample test, such as fitting parameters on one year and predicting another year's curves, or using the fitted parameters to predict a quantity not used in the fit.","section":"Section IV and Section VI"},{"comment":"The physical mapping for names is asserted by analogy: the paper sets f_m = w_m = E_m and invokes the Rayleigh-Jeans form rho_m = T/(E_m - mu) from conservation of total energy and probability norm. No dynamics of name choice is specified that would conserve these two integrals, and T and mu are re-fit separately for each year. As presented, the RJE model is a flexible two-parameter family used to fit empirical Lorenz curves rather than a thermal-equilibrium prediction derived for the name-formation process. A concrete test would be to derive the density of states from a plausible name-choice dynamics, or to show that the fitted parameters satisfy a conservation law across years; without this, the central claim reduces to curve fitting.","section":"Section II"},{"comment":"The paper itself acknowledges that 'the fact that the RJE model with two parameters fits the real Lorenz and Pareto curves is not a sufficient argument in the full favor of RJ thermalization theory.' The subsequent appeal to universality across many systems as the 'fundamental confirmation' of RJ condensation does not close this gap: the observed stability and similarity of Lorenz curves across systems is an empirical regularity, but it does not by itself identify RJ condensation as the mechanism. The argument would need a mechanism that generates the specific RJE density of states for the name system, or a falsifiable prediction that distinguishes RJE from other two-parameter distribution families with similar Lorenz shapes.","section":"Section VI"}],"minor_comments":[{"comment":"The caption and text are inconsistent: the text says 'US M 1880, FR F 1900' while the figure is for FR male names, and the caption lists 'FR M 1900'; this should be corrected.","section":"Appendix Fig. A7"},{"comment":"There is a typo in the first sentence: 'deacribed' should be 'described'; in Section VI, 'mecanism' should be 'mechanism'.","section":"Section V"},{"comment":"The bottom-panel year list in Fig. 3 begins '1990' but the text and figure cover 1900 to 2023; the same issue appears in Fig. A3, where '1900 to 2022' should likely be '1900 to 2023'.","section":"Fig. 3 and Fig. A3 captions"},{"comment":"The procedure of assigning frequency zero to names absent in a given year may artificially inflate Pearson correlations; the paper notes this but does not quantify the effect. A sensitivity check using only names present in both years would strengthen the correlation analysis.","section":"Section V"}],"recommendation":"major_revision","confidential_remarks":"The paper relies heavily on the authors' own prior and preprint works [13-15,28,29] for the WTH and RJE framework. The descriptive statistics are likely acceptable for the journal's scope, but the causal conclusion needs independent validation or at least a clear out-of-sample test before the central claim can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper is worth reading for its empirical core, not for its causal conclusion. The data work is straightforward and credible. From public SSA and INSEE files they compute Lorenz and Pareto curves for US and French given names over 145/124 years, report Gini coefficients staying in 0.85–0.95, and document a clear regime change in correlation of top names after roughly 1960. Those are real observations, reproducible, and new as far as I know. The rank trajectories of individual names, like Louise dropping to near-forgotten and returning to the top, are a nice concrete touch.\n\nWhat is less solid is the move from curve fitting to thermodynamics. The RJE model has two parameters. Epsilon is fixed by matching the data Gini; a is chosen by minimizing distance to the data Lorenz curve. So the Lorenz agreement is partly built in, and the Pareto curves use the same fitted distribution, so they are a consistency check on the same fitted object, not an independent prediction. The paper itself says in Section VI that a two-parameter fit 'is not a sufficient argument,' then basically ignores that and concludes that universality across systems gives 'fundamental confirmation' of RJ thermalization. That is an overreach. No dynamics, no conserved quantities for name choice, no relaxation timescale are provided; the two integrals of motion are asserted by analogy, and the parameters are refit every year. Any flexible two-parameter family with the same Lorenz shape would likely do as well.\n\nI don't think the paper is worthless or dishonest. It is transparent about the fitting procedure, and the descriptive claims look solid. The problem is the interpretive leap: 'RJ condensation is the driving mechanism' is a hypothesis, not a demonstrated result. The existing program in their earlier papers may be right, but this paper does not add enough evidence to convert a skeptic. It also leans on self-citations, though that is not itself a flaw because the RJE machinery is their own and formally derived elsewhere.\n\nWho is this for? People working in sociophysics and inequality, and anyone interested in null models for name-frequency distributions. A serious referee should engage: the empirical section deserves publication as a statistical study, and the interpretation should be either softened or supported by out-of-sample predictions, such as fixed parameters across years or an explicit dynamical model. I would send it to review rather than desk-reject, with expectation of major revision on the interpretation.","headline":"Solid descriptive statistics on 100+ years of name data, but the thermodynamic mechanism is asserted, not derived; the RJE 'agreement' is mostly a two-parameter fit.","tokens_in":18347,"tokens_out":2081,"would_cite":false,"duration_ms":21958,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that given-name frequency distributions in the USA and France are thermodynamic equilibria described by the Rayleigh-Jeans law, stable for over a century.","keywords":["Rayleigh-Jeans condensation","given-name statistics","Lorenz curve","Pareto distribution","Gini coefficient","thermalization","inequality universality","name popularity"],"falsifier":"Truncate the official data by raising the minimum-frequency cutoff (for example, from 5 to 50 occurrences in the US data) and re-fit the RJE parameters $a$ and $\\varepsilon$; if the Rayleigh-Jeans description is real, the fitted parameters and the inferred Gini should stay constant up to sampling noise, whereas a systematic drift with the cutoff would show that the 'thermal' distribution is an artifact of the data's lower bound.","tokens_in":17180,"feed_emoji":"👶","tokens_out":13557,"duration_ms":123658,"temperature":0.7,"pith_summary":"The paper claims that the frequency distribution of given names in the USA (1880–2025) and France (1900–2023) is not a product of arbitrary fashion but a thermal equilibrium. Using government records, it builds Lorenz and Pareto curves of name popularity and finds them stable for more than a century, with the Gini coefficient confined to 0.85–0.95. It shows that these curves are reproduced by the Rayleigh-Jeans (RJ) thermal distribution, the same two-integral statistical-mechanics law used for wealth inequality: each name acts as an energy level, and the population of names occupies those levels with the RJ occupation formula. If the claim is right, name popularity belongs to a universal family of inequality phenomena that includes wealth, energy consumption, and voting.","feed_headline":"Name popularity follows a thermodynamic law","feed_subtitle":"US and French name-frequency curves stay stable for 100+ years, with Gini 0.85–0.95, fitting Rayleigh-Jeans condensation.","key_machinery":"The load-bearing object is the Rayleigh-Jeans occupation formula $\\rho_m = T/(E_m - \\mu)$ for a system with two conserved integrals of motion (energy and probability norm), together with the extended RJ density-of-states model $\\nu(E_k) = dk/dE_k = N(e^a - 1)/[a(1 + (e^a-1)E_k]$ used to build the Lorenz and Pareto curves. The formula produces the characteristic 'condensation' of many agents on low-energy states, which the paper identifies with the observed long tail of rare names; the spectrum parameter $a$ and the rescaled energy $\\varepsilon = E/B$ are the two fitting parameters that let the model match the real curves.","core_discovery":"The central discovery is that name popularity obeys the Rayleigh-Jeans distribution $\\rho_m = T/(E_m - \\mu)$, where $E_m = w_m = f_m$ is the (shifted) frequency of name $m$, while the temperature $T$ and chemical potential $\\mu$ are fixed by conservation of total energy $E = \\sum_m E_m \\rho_m$ and total norm $\\eta = \\sum_m \\rho_m = 1$. In this picture the most popular names are high-energy states and the long tail of rare names is the low-energy condensate; the two parameters are re-fitted each year, and the resulting Lorenz and Pareto curves agree with the observed curves to a typical geometric distance below $5\\times10^{-3}$. The Gini coefficient stays in the narrow band $G \\in [0.85, 0.95]$ for both countries for over a century, and the RJ description therefore captures the inequality structure of name choice as a steady-state thermal system.","pith_inferences":["A testable extension would be an agent-based model in which parents copy or exchange name preferences through pairwise 'collisions'; if it thermalizes to the same $\\rho_m \\propto 1/(E_m - \\mu)$ with the same spectrum, the equilibrium claim would be mechanistically supported rather than only phenomenologically.","The paper re-fits $T$ and $\\mu$ every year but does not report their time series; tracking the annual drift of the chemical potential could reveal what cultural force corresponds to changing scarcity of names.","Because the fit uses the same functional form for every year, a sharper test is whether the fitted parameters vary smoothly in time; a discontinuous jump would mark a cultural phase transition rather than adiabatic thermalization.","The mapping $E_m = f_m$ is chosen for analogy, not derived; checking whether an alternative monotone mapping (for example, log-frequency) destroys the parameter stability would isolate which quantity is genuinely conserved."],"forward_implications":["Name popularity is a predictable statistical-mechanical quantity: once $T$ and $\\mu$ (or $a$ and $\\varepsilon$) are fixed for a year, the entire Lorenz and Pareto curves follow, so the observed inequality is not a collection of arbitrary choices.","The Gini coefficient should remain near 0.85–0.95 as long as the system stays in the same thermal regime, meaning the extreme concentration of popularity among a few names is a stable equilibrium, not a transient fashion.","The paper's universality argument places name popularity in the same family as wealth, energy consumption, and voting, for which the same RJ condensation has been reported.","The stability of the correlation structure until the mid-20th century and its change afterwards indicates that the thermal analogy tolerates slow drift of the parameters while preserving the equilibrium form."],"supporting_citations":[{"why":"Supplies the US given-name frequency data from 1880 to 2025 that all empirical Lorenz and Pareto curves are built from.","marker":"[3]"},{"why":"Supplies the French given-name frequency data from 1900 to 2023 used for the same construction.","marker":"[4]"},{"why":"Introduces the Wealth Thermalization Hypothesis and the Rayleigh-Jeans form for inequality that this paper extends to names.","marker":"[13]"},{"why":"Provides the Rayleigh-Jeans extended (RJE) model and the analytic Lorenz/Pareto curve expressions used for the fits.","marker":"[14]"},{"why":"Gives the statistical-physics basis of the Rayleigh-Jeans distribution from energy and norm conservation.","marker":"[18]"},{"why":"Establishes RJ thermalization for classical wave turbulence with two integrals of motion, the physical mechanism invoked here.","marker":"[19]"},{"why":"Documents condensation of classical nonlinear waves, the phenomenon the paper identifies with the rare-name tail.","marker":"[21]"},{"why":"Shows the same RJ condensation describes energy and carbon-emission distributions, part of the universality argument.","marker":"[28]"},{"why":"Shows the same RJ description applies to EU election voting distributions, another instance of the universal pattern.","marker":"[29]"},{"why":"Supplies the world wealth-inequality benchmark (Gini 0.889) to which the name-inequality values are compared.","marker":"[7]"}],"fun_headline_variants":["Name popularity follows thermodynamic law","Why name trends obey Rayleigh-Jeans stats","Stat mech of baby names: Gini 0.85–0.95","US-France name data show 100-year thermal stability","Name inequality fits thermal condensation model"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a society's name-choice process conserves two global quantities—an 'energy' equal to name frequency and a probability norm—so that the Rayleigh-Jeans formula $\\rho_m = T/(E_m - \\mu)$ is the true distribution rather than just a flexible two-parameter curve.","fun_headline_variants_meta":{"raw":{"variants":["Name popularity follows thermodynamic law","Why name trends obey Rayleigh-Jeans stats","Stat mech of baby names: Gini 0.85–0.95","US-France name data show 100-year thermal stability","Name inequality fits thermal condensation model"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000193,"raw_usage":{"total_tokens":1328,"prompt_tokens":899,"completion_tokens":429,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":515,"completion_tokens_details":{"reasoning_tokens":355}},"tokens_in":515,"tokens_out":429,"duration_ms":5184,"temperature":1.0,"reasoning_tokens":355,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T18:44:21.372118+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Truncate the official data by raising the minimum-frequency cutoff (for example, from 5 to 50 occurrences in the US data) and re-fit the RJE parameters $a$ and $\\varepsilon$; if the Rayleigh-Jeans description is real, the fitted parameters and the inferred Gini should stay constant up to sampling noise, whereas a systematic drift with the cutoff would show that the 'thermal' distribution is an artifact of the data's lower bound.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the US given-name frequency data from 1880 to 2025 that all empirical Lorenz and Pareto curves are built from."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the French given-name frequency data from 1900 to 2023 used for the same construction."},{"cited_title":"Pareto, Manuale di Economia Politica (1906); EGEA – Universit` a Bocconi Editore: Milan, (2006) ISBN 88- 8350-084-9","cited_arxiv_id":null,"evidence_quote":"Gives the statistical-physics basis of the Rayleigh-Jeans distribution from energy and norm conservation."},{"cited_title":"Zakharov, V.S","cited_arxiv_id":null,"evidence_quote":"Establishes RJ thermalization for classical wave turbulence with two integrals of motion, the physical mechanism invoked here."},{"cited_title":"Connaughton, N","cited_arxiv_id":null,"evidence_quote":"Documents condensation of classical nonlinear waves, the phenomenon the paper identifies with the rare-name tail."},{"cited_title":"Picozzi, J","cited_arxiv_id":null,"evidence_quote":"Shows the same RJ condensation describes energy and carbon-emission distributions, part of the universality argument."},{"cited_title":"Baudin, A","cited_arxiv_id":null,"evidence_quote":"Shows the same RJ description applies to EU election voting distributions, another instance of the universal pattern."},{"cited_title":"Chancel, T","cited_arxiv_id":null,"evidence_quote":"Supplies the world wealth-inequality benchmark (Gini 0.889) to which the name-inequality values are compared."}],"review_version":1}