{"id":"efe42918-b289-422d-b243-ae9384fd7304","arxiv_id":"2411.14469","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across GPT-4, Gemini, and Claude, adding race or gender to a name changes predicted points of interest in ways that systematically disadvantage Black, Hispanic, and female individuals.","lead":"The authors asked three large language models to guess which places people visit based on their names, with and without race and gender labels. The models attached minority groups and women to less wealthy and less career-related destinations, and the paper argues this amplifies real-world disparities.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'amplify' claim rests on an unvalidated mapping from career POIs to NHTS work trips; without that mapping, the paper only shows LLM predictions are skewed, not that they exceed real-world disparities.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the career-related POI list is never validated against actual work-related travel, yet the amplification claim is built on comparing LLM career-POI shares to NHTS work-trip shares. I agree that this is the central weak point. The descriptive bias results—minority groups assigned to poverty-related POIs, women assigned to fewer career-related POIs—are robust and worth reporting, but they support a claim of stereotyped predictions, not necessarily amplification relative to measured human behavior. The paper's own limitations section acknowledges the limited POI set but does not address the construct-validity problem directly. The high refusal rates in Table 2 add a selection-bias concern, but they are secondary; the primary issue is the unvalidated mapping. Because the reader's CONDITIONAL verdict already captures this uncertainty, no verdict change is needed: the paper should be accepted only if the mapping is validated or the amplification claim is softened to a descriptive bias claim.","tokens_in":11401,"tokens_out":2761,"duration_ms":34820,"concrete_test":"Re-run Experiment I with a POI list validated against NHTS work destinations: have independent annotators classify each candidate POI as a workplace or not, keep only POIs unanimously rated as workplaces (e.g., office buildings, factories, worksites), and recompute the LLM career-related shares and the male/female and race gaps. If the recomputed LLM gaps no longer exceed the NHTS work-travel gaps, or if the ranking across racial groups changes materially, the amplification claim is an artifact of the unvalidated POI mapping. Additionally, as a robustness check, re-analyze the main comparisons by imputing all rejected responses as the worst-case category for the affected group to test whether exclusion of refusals changes the conclusion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that LLMs do not merely reflect but amplify race and gender disparities in human mobility. The support for 'amplify' comes from comparing LLM career-related POI choice shares (Appendix A: Industry conference center, Employment center, Professional training center, Career consultation center) to NHTS 2022 work-trip proportions (Table 1). This comparison assumes those four POIs measure the same construct as 'work-related travel' in NHTS. That assumption is unvalidated: two of the POIs (career consultation center, professional training center) are not workplaces, and the LLM task is a forced choice among four POIs, so the resulting shares are not directly comparable to trip-purpose proportions from a travel survey. The descriptive finding that LLMs associate minority groups with poverty-related POIs and women with fewer career-related POIs is credible and consistent across models. But the headline 'amplification' claim requires the NHTS comparison to be meaningful; if the POI list does not operationalize work trips, the paper has shown only that LLM outputs are stereotyped, not that they exaggerate actual measured disparities. A secondary aggravator is that high refusal rates (Table 2: Gemini rejects 87.5% of Black wealth-related queries) are excluded without re-weighting, so the non-refused responses may be a selected sample. Both issues are load-bearing for the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper asks whether commercial LLMs (GPT-4o, Gemini-1.5-pro, Claude-3.5-sonnet) encode race and gender bias when predicting human mobility. In Experiment I, the authors prompt the models to choose, from a small fixed set of points of interest (POIs), the place a named individual most likely visited, with or without explicit race/gender labels. They report that minority groups, especially Black and Hispanic individuals, are less often assigned wealth-related POIs and more often assigned poverty-related POIs, while women are less often assigned career-related POIs than men. They compare the career-related POI shares to work-trip shares from NHTS 2022 and conclude that the models do not merely reflect but amplify real-world disparities. Experiment II repeats the comparison in a paired-assignment design, and logistic regressions quantify name-based and label-based effects. The paper's descriptive claim—that LLM outputs are demographically skewed in these ways—is credible, but the stronger 'amplification' conclusion rests on an unvalidated mapping between the experiment's career POIs and NHTS work trips, and on analyses that exclude high, group-correlated refusal rates without reweighting.","tokens_in":11637,"tokens_out":4861,"duration_ms":54209,"significance":"If the descriptive findings hold, they are significant: LLMs are already proposed for travel planning, urban analytics, and mobility simulation, and the demonstration that three major commercial models produce strongly stereotyped predictions about where people go is a valuable fairness result. The paper has genuine strengths: it tests multiple independent model families through their APIs, anchors the bias question against an external survey (NHTS 2022), provides a name-level logistic regression that separates name cues from explicit labels, and reports refusal rates rather than hiding them. These features make the core observation reproducible and worth publishing. However, the headline interpretation—that LLMs 'amplify' rather than merely reflect or distort observed disparities—is currently supported only by a direct comparison of two non-equivalent quantities: a forced-choice share over four hand-picked career POIs and an NHTS trip-purpose share. The paper also lacks uncertainty quantification.","major_comments":[{"comment":"The central 'amplify' claim compares LLM shares of career-related POIs with NHTS work-trip shares, but this assumes that the four career POIs in Table S1 (Industry conference center, Employment center, Professional training center, Career consultation center) operationalize work-related travel. Professional training centers and career consultation centers are not workplaces, and the LLM task is a forced choice among four listed POIs, so the resulting category share is not a trip-purpose proportion. No validation of the POI list against actual workplace destinations is provided. As written, the comparison in Table 1 establishes only that the LLM outputs differ from the NHTS benchmark; it does not establish that the model amplifies the real-world disparity. Please either validate the mapping (for example, against workplace visit frequencies in a mobility or location-based services dataset) or reframe the 'amplification' statements in the Abstract, Results, and Discussion as descriptive bias relative to the survey, with the mapping limitation stated explicitly.","section":"Comparison with Survey Data (Table 1)"},{"comment":"Refusal rates are large and group-dependent—Gemini rejects 87.5% of Black wealth-related queries, and Claude rejects 37.5%—yet the Methods state that refusals were excluded from the reported analyses without any reweighting or sensitivity analysis. If refusal propensity is correlated with demographic group and POI category, the non-refused responses are a selected sample, and the extreme estimates in Figure 2 (such as Gemini's wealth-related coefficient below -6 for Black) could reflect differential refusal rather than bias in answered queries. The paper should report refusal-adjusted estimates or bounds (for example, worst-case reallocation of refusals), or explicitly state and defend the assumption that refusals are ignorable. As it stands, this is a load-bearing gap in the support for the quantitative disparity claims, not a minor caveat.","section":"Methods (last paragraph) and Table 2"},{"comment":"No uncertainty quantification is reported for any of the headline numbers. The paper reports exact point estimates for career-related and wealth-related shares (e.g., 12.2% vs. 2% for White males and females in Figure 1b; 48.0% vs. 0.5% for Black males in Figure 4b) without confidence intervals, bootstrap intervals, repeated API runs, or temperature settings. The logistic regressions in Figure 2 also lack coefficient standard errors despite the text describing some variables as statistically insignificant. Because the choice sets are randomized and the names are sampled, the stability of the cross-group rankings and the extreme coefficients cannot be assessed. At minimum, the authors should supply repeated-run or bootstrap intervals for the main tables and figures, or clearly state the single-run design and its implications for the strength of the claims.","section":"Figures 1, 2, and 4; Logistic Regression Analysis"}],"minor_comments":[{"comment":"The 'Black Male' list of representative names includes 'Gwendolyn', which is a female name; if this list was used to create male subgroup prompts in Experiment II, the gender comparison for the Black subgroup is contaminated. Please verify the name lists and correct this entry.","section":"Table S2"},{"comment":"The caption says the figure uses GPT-4o, but the surrounding text describes the results as if they apply to LLMs generally; please clarify which model and which experiment each panel reports, and state where the corresponding results for Gemini and Claude appear.","section":"Amplified Gender and Race Disparities (Figure 4)"},{"comment":"The logistic regression section does not specify the model equation, the handling of multiple observations per name, the standard error calculation, or the threshold used to mark coefficients statistically insignificant in Figure 2; please add this information or point to a supplement.","section":"Methods, Logistic Regression Analysis"},{"comment":"The description of output extraction as 'simple natural language processing techniques' is vague; please specify the JSON parsing and retry procedure, including how malformed or partial outputs were handled, since refusal and parse failures both affect the effective sample.","section":"Methods, large language model employment"},{"comment":"The NHTS citation is incompletely formatted ('Federal, H.: Administration'); please supply the full report title, year, and URL or DOI.","section":"Reference [20]"}],"recommendation":"major_revision","confidential_remarks":"This is a timely and potentially useful empirical contribution, and the descriptive bias results across three commercial models are likely to interest the cs.CL community. The main risk is overclaiming 'amplification' from a comparison that has not established construct equivalence between the experiment's career POIs and NHTS work trips. I believe the authors can address this either by validating the POI mapping or by reframing the central claim as descriptive bias with explicit caveats, so this is a major revision rather than a rejection. The refusal-rate issue is also fixable with sensitivity analyses. No concerns about novelty disclosure or citation patterns beyond the reference formatting; the paper's scope fits the journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper has a credible descriptive finding—commercial LLMs, given names or explicit demographics, systematically skew predicted POI visits in ways that track race and gender stereotypes—but the headline claim that they 'amplify' real-world disparities is not supported by the evidence as presented. The amplification conclusion depends on a comparison to NHTS work-trip shares that the paper never validates: two of the four career POIs (career consultation center, professional training center) are not workplaces, and the LLM task is a forced choice among four POIs, not a trip-purpose survey. So the model's 'career-related' share is not the same construct as NHTS work travel. Without that mapping, the paper shows the models produce stereotyped predictions, not that they exaggerate measured disparities.\n\nWhat is genuinely new is the application of name-based bias probing to POI visitation prediction with three current commercial models, and the demonstration that adding explicit race/gender labels sharply shifts predictions (e.g., Black male wealth-related share drops from 48% to 0.5%). That is reproducible, consistent across models, and policy-relevant for anyone deploying LLMs in travel planning or urban analytics. The logistic regression analysis is a reasonable addition, though it reports coefficients without error bars or significance thresholds in the figure, so the 'statistically insignificant' markers are not verifiable from the paper alone.\n\nThe soft spots are real. First, the NHTS comparison is load-bearing for the word 'amplify' in the title, and the mapping is unvalidated. Second, rejection rates are high for some model/subgroup pairs—Gemini rejects 87.5% of Black wealth-related queries—and the analysis excludes refusals without re-weighting. The paper acknowledges this but treats it as a side observation; it actually conditions the reported shares on a potentially non-random subset of responses. Third, there are no uncertainty estimates anywhere; the forced-choice percentages come from 2,675 names, so the patterns are probably stable, but the paper gives no sense of variability.\n\nWho is this for? AI-fairness researchers and urban-computing folks will get a useful, if methodologically imperfect, demonstration of bias in LLM mobility predictions. It deserves a serious referee—the descriptive result is worth publishing, but the framing needs to change. My recommendation: engage with it, but in review push hard on the NHTS comparison and the refusal handling. If the authors can validate the POI set against actual workplace visits or recast the claim as one about stereotyped predictions, it becomes a solid contribution.","headline":"Credible evidence of stereotyped POI predictions, but the 'amplify' claim is unsupported because the NHTS comparison relies on an unvalidated POI-to-work-trip mapping.","tokens_in":12089,"tokens_out":2988,"would_cite":false,"duration_ms":31331,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Popular LLMs predict women fewer career-related places and Black and Hispanic individuals far fewer wealth-related places than White individuals, and these predicted gaps exceed real survey gaps.","keywords":["large language models","societal bias","race and gender disparities","human mobility prediction","points of interest","bias amplification","prompt-based fairness evaluation","travel behavior"],"falsifier":"Collect real mobility traces with the same eight race-gender groups, classify visited places into the paper's four POI categories, and compare group visit shares with the LLM predictions: if the observed shares match the LLM predictions rather than the national survey, the amplification claim fails, while if they match the survey, it survives.","tokens_in":11215,"feed_emoji":"📍","tokens_out":9154,"duration_ms":85033,"temperature":0.7,"pith_summary":"This paper tries to establish that commercial large language models, when asked to predict where a named person is likely to go, reproduce and in some cases amplify racial and gender stereotypes about mobility. The authors asked three chatbots—GPT-4o, Gemini, and Claude—to choose points of interest for people in eight race-gender subgroups, using real first names with and without explicit demographic labels. Women were consistently assigned fewer career-related destinations than men of the same race, and when race was stated, Black and Hispanic individuals were assigned wealth-related destinations far less often than White individuals. Comparing the career-related predictions to a national travel survey showed that the models exaggerate the real gender gap in work travel by a wide margin. The paper concludes that LLM-based travel planning, urban analytics, and similar tools can systematically disadvantage women and racial minorities.","feed_headline":"LLMs give women fewer career places and minorities more poverty places","feed_subtitle":"The bias appears in all three tested chatbots: a warning for AI travel and urban tools.","key_machinery":"The machinery is a forced-choice POI task with paired categories. Each prompt names one person (or two people in Experiment II) and a randomized list of four points of interest; the model must select one (or two) places. The POI lists are deliberately paired: career-related versus everyday-needs, and wealth-related versus poverty-related. Name-level demographic ratios come from a name-race dataset and a name-gender dataset, feeding logistic regressions that separate name-inferred from explicitly stated demographics. The national travel survey provides the real-world baseline for work travel against which the career-related predictions are judged as reflection or amplification.","core_discovery":"The central discovery is that demographic cues in a prompt change an LLM's predicted destinations, and the changes are systematically unequal. In the wealth-versus-poverty comparisons, when race is not mentioned, GPT-4o assigns wealth-related places to all groups at similar high rates; once race is specified, Black and Hispanic individuals' wealth-related shares collapse to near zero in pairwise comparisons, while White individuals' shares remain high. In the career-versus-everyday comparisons, women of every racial group are less likely than men of that group to receive career-related POIs, and the gender gaps are much larger than the gaps found in work-related travel in the 2022 national travel survey. The authors call this amplification: the models do not simply mirror aggregate human behavior, they exaggerate its demographic structure, associating minority men with career places and poverty places in ways the survey does not show.","pith_inferences":["Not tested in the paper: the four career-related POIs include 'industry conference center' and 'career consultation center,' which are not routine workplaces; a human rating study of whether these POIs represent work trips could change the size of the reported amplification.","The same prompt design could be extended to other protected attributes such as age, disability, or sexual orientation, or to other behavior categories such as health-care visits and school locations; the paper's framing suggests these are likely to show similar disparities.","Because the race-specified probes use the same first name with different demographic labels, the models' responses appear label-driven rather than name-driven; this could enable a cheap mitigation test by prepending neutral statistics or counter-stereotypical context before the choice prompt.","A direct behavioral validation with GPS or cell-phone traces on the exact POI categories would either confirm or overturn the amplification claim; this is the natural next experiment."],"forward_implications":["Any downstream system using these models for travel recommendations or urban analytics would inherit a systematic skew: women get fewer career-related destinations, and Black and Hispanic users get fewer wealthy destinations, even when the model is not told their demographics but only their name.","Bias audits that only measure refusal rates are insufficient: models with high refusal rates still amplify disparities in the answers they do return.","The reflection-versus-amplification comparison against a national travel survey provides a quantitative test that can be rerun for other models and other POI categories to track fairness over time.","The paper's results imply that demographic labels are a strong control variable in LLM mobility tasks, so applications that intend to be neutral should either avoid such labels or otherwise correct for their effect."],"supporting_citations":[{"why":"Supplies the real-world baseline of work-related travel shares by gender and race against which the LLM career-related predictions are compared.","marker":"[20]"},{"why":"Supplies name-to-race proportions used to derive Black, Hispanic, and Asian name ratios for the prompts and regressions.","marker":"[24]"},{"why":"Supplies name-to-gender proportions used to build the Female ratio variable for each name.","marker":"[25]"},{"why":"Identifies GPT-4o as one of the three tested models whose API predictions are analyzed.","marker":"[21]"},{"why":"Identifies Gemini as one of the three tested models whose API predictions are analyzed.","marker":"[22]"},{"why":"Identifies Claude as one of the three tested models whose API predictions are analyzed.","marker":"[23]"}],"fun_headline_variants":["LLMs skew travel predictions: minorities less wealth, women less career","GPT-4, Gemini, Claude all show race and gender bias in place prediction","LLM travel predictions widen race and gender gaps","Race and gender cues make LLMs predict biased destinations","LLMs exaggerate real-world mobility biases by race and gender"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The amplification claim rests on treating the four career-related POI labels as measuring work-related travel; if those labels are not what people visit on work trips, the model-vs-survey gap measures category mismatch rather than bias amplification.","fun_headline_variants_meta":{"raw":{"variants":["LLMs skew travel predictions: minorities less wealth, women less career","GPT-4, Gemini, Claude all show race and gender bias in place prediction","LLM travel predictions widen race and gender gaps","Race and gender cues make LLMs predict biased destinations","LLMs exaggerate real-world mobility biases by race and gender"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000681,"raw_usage":{"total_tokens":3058,"prompt_tokens":878,"completion_tokens":2180,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":494,"completion_tokens_details":{"reasoning_tokens":2094}},"tokens_in":494,"tokens_out":2180,"duration_ms":14768,"temperature":1.0,"reasoning_tokens":2094,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:00:29.039150+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect real mobility traces with the same eight race-gender groups, classify visited places into the paper's four POI categories, and compare group visit shares with the LLM predictions: if the observed shares match the LLM predictions rather than the national survey, the amplification claim fails, while if they match the survey, it survives.","supporting_citations":[{"cited_title":"2022 nextgen national household travel survey core data","cited_arxiv_id":null,"evidence_quote":"Supplies the real-world baseline of work-related travel shares by gender and race against which the LLM career-related predictions are compared."},{"cited_title":"Scientific Data 10(1), 299 (2023)","cited_arxiv_id":null,"evidence_quote":"Supplies name-to-race proportions used to derive Black, Hispanic, and Asian name ratios for the prompts and regressions."},{"cited_title":"DOI: https://doi.org/10.24432/C55G7X (2020) 16 Appendix A Specification of experiments In the experiments, the POI sets for four categories are presented as follows","cited_arxiv_id":null,"evidence_quote":"Supplies name-to-gender proportions used to build the Female ratio variable for each name."},{"cited_title":"Version: GPT-4o-2024-08-06 (2024)","cited_arxiv_id":null,"evidence_quote":"Identifies GPT-4o as one of the three tested models whose API predictions are analyzed."},{"cited_title":"Version: Gemini-1.5-pro (2024)","cited_arxiv_id":null,"evidence_quote":"Identifies Gemini as one of the three tested models whose API predictions are analyzed."},{"cited_title":"Version: Claude-3-5-sonnet-20240620 (2024)","cited_arxiv_id":null,"evidence_quote":"Identifies Claude as one of the three tested models whose API predictions are analyzed."}],"review_version":1}