Pith. sign in

REVIEW 4 major objections 5 minor 6 references

Signal or Noise? Evaluating Large Language Models in Resume Screening Across Contextual Variations and Human Expert Benchmarks

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Large language models screen resumes with stable, context-dependent patterns that systematically diverge from human expert judgments in every condition tested.

desk verdict A well-designed question undermined by pseudo-replication: the headline p-values treat repeated scores on identical resumes as independent, so the central claims don't survive as analyzed. read the letter →

arxiv 2507.08019 v1 pith:Y37WSMAY submitted 2025-07-08 cs.CL econ.GNq-fin.EC

classification cs.CLecon.GNq-fin.EC
keywords largelanguagemodelsresumescreeningcontextsensitivityhuman-AIcomparisonstatisticalanalysisrecruitmentautomationmeta-cognitionsignaldetection
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Three large language models — Claude, GPT, and Gemini — were asked to score resumes for an Associate Product Manager role under four prompt conditions, and their scores were compared with those of three experienced recruiters. The paper's central claim is that LLM scoring is not random noise: it shows repeatable, context-dependent patterns, particularly in GPT's large score shifts when company context is added. At the same time, LLM scores differed significantly from human expert scores in every condition ($p < 0.01$), with human experts scoring more conservatively and using the full range of the 0–100 scale. Under a stripped-down job description, GPT inflated its average score from 64.3 to 82.9 while human experts stayed cautious, producing the widest human–AI gap. The paper concludes that LLMs can provide interpretable and efficient screening, but uncalibrated LLM scores would change which candidates are selected.

What carries the argument

The central mechanism is the controlled factorial comparison design. Three LLMs and three human experts each scored 10 resumes on a 0–100 scale under four prompt conditions — job description only, an MNC profile (Firm1), a startup profile (Firm2), and a reduced job description — using both ten identical copies of one resume and ten randomly selected resumes. One-way ANOVA and paired $t$-tests are the instruments that turn these score tables into claims about inter-model consistency, human–AI alignment, and contextual adaptation, with Cohen's $d$ used to gauge effect size. The meta-cognition prompt, which asked each model to state its percentage weights for resume components, is the additional mechanism that exposes how LLMs allocate emphasis and lets the paper contrast GPT's mechanical weight redistribution with human experts' qualitative rationales.

What would settle it

Re-analyze the Same Resume data with evaluator as a random effect and the ten identical copies treated as repeated measures rather than independent rows. If GPT's No-Company-to-Firm1 and No-Company-to-Firm2 paired $t$-tests no longer reach $p < 0.05$ once within-evaluator correlation is modeled, the study's headline contextual-adaptation result is an artifact of pseudo-replication.

Watch

Extended reading notes

Core claim

The paper's discovery is that LLM resume screening is neither uniformly signal nor pure noise: it is reliably patterned but systematically misaligned with human expert judgment. Across 420 evaluations of three LLMs and three human recruiters on identical and randomized resumes, ANOVA found significant mean differences among the LLMs in four of eight LLM-only conditions, and significant LLM-versus-human differences in all contexts at $p < 0.01$. Contextual sensitivity is model-specific: GPT's mean scores rose by roughly 25 points when an MNC or startup profile was added to the prompt (paired $t$ test, $p < 0.001$), Gemini rose only for the MNC context ($p = 0.038$), and Claude showed no reliable movement ($p > 0.1$). When the job description was reduced to essentials, GPT's mean rose from 64.3 to 82.9 while human experts became more conservative, producing the largest systematic disagreement in the study. The paper concludes that LLMs decompose evaluation into explicit weighted criteria but lack the qualitative, context-aware risk judgment behind human expert scoring.

Load-bearing premise

The main load-bearing premise is that the ten scores each evaluator gave to ten identical copies of the same resume count as ten independent measurements; if repeated judgments on identical material are correlated, the reported $p$-values are too optimistic.

Editorial extensions

If this is right

  • LLM-only screening would rank different candidates than a human panel: with LLM means 15–25 points higher, the same resume can pass an LLM cutoff and fail a human one.
  • Score interpretation is prompt-dependent: GPT's roughly 25-point swing when company context is added means adding one sentence about firm type can move a candidate across a hiring threshold.
  • Organizations cannot assume a single 'LLM behavior' in hiring: the three models differed significantly from each other in four of eight LLM-only conditions and reacted differently to the same context.
  • Short or incomplete job descriptions are a known failure mode: GPT's score inflation from 64.3 to 82.9 in the Reduced Context condition implies that sparse postings invite over-crediting.
  • The LLMs' explicit weighting schemes are not the same logic experts use, so using those schemes as explainability for automated decisions would inherit the measured human–AI divergence.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The study frames signal versus noise as a property of models and conditions, but it never varies sampling temperature; re-running each model multiple times under different temperatures would directly measure output randomness, which is the more literal test of the title's question.
  • The human benchmark effectively rests on two experts in the later conditions, since one expert began assigning zero scores; claims about 'human expert judgment' as a stable standard need a larger and more engaged panel before they generalize.
  • A natural extension is a label-swap experiment: exchange the Firm1 and Firm2 descriptions, or anonymize the company names, to test whether the observed contextual adaptation tracks the organizational profiles themselves or incidental wording.
  • If the pseudo-replication concern in the Same Resume condition is confirmed by a repeated-measures analysis, the practical takeaway that LLM and human scores differ survives, but the per-model adaptation claims, especially for GPT, would need to be re-benchmarked.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper reports a controlled experiment in which three LLMs (Claude, GPT, Gemini) and three human recruitment experts scored resumes under four context conditions (No Company, Firm1, Firm2, Reduced Context) and two resume sets (ten identical copies and ten different resumes). The authors claim that ANOVAs reveal significant LLM-human differences (p < 0.01), that paired t-tests show GPT adapts strongly to company context, Gemini partially, and Claude minimally, and that meta-cognition analyses reveal different weighting strategies. The central statistical claims rest on treating the ten scores produced by each evaluator as independent observations, which is invalid because the evaluator is the experimental unit; this issue affects virtually all reported p-values and undermines the main conclusions.

Significance. If the claims were valid, the study would be a useful empirical contribution to the debate on LLM reliability in high-stakes screening tasks, with practical implications for hybrid recruitment systems. The paper has some strengths: a controlled within-subject design, standardized prompts, a clearly described human expert benchmark, and a qualitative meta-cognition analysis that distinguishes LLM weighting schemes from human rationales. However, the central quantitative claims do not survive scrutiny because the analyses ignore the nested structure of the data. The reported p-values are anti-conservative, and with only three evaluators per group, an evaluator-level permutation test cannot yield two-sided p-values below 0.10 for any LLM-versus-human contrast. The paper's core message therefore lacks statistical support at the stated level of confidence. I also note that no data or code are provided, which prevents independent re-analysis of the problematic statistics.

major comments (4)
  1. [Statistical Analysis Plan / Study 1: Same Resume Analysis] The primary inferential tests treat 10 scores per evaluator as independent observations. In the Same Resume condition the ten resumes are identical copies of the same material, so the effective sample size is the number of evaluators (n = 3 per group), not 10 times that. All reported F and t statistics with df = 27, 54, and 9 (e.g., LLM-only F(2,27), combined F(5,54), GPT paired t(9)) therefore have far too many degrees of freedom. An evaluator-level permutation test with 3 LLMs versus 3 humans cannot produce a two-sided p-value below 0.10 for any contrast, yet the paper reports p = 0.0015, p = 0.0007, and p = 0.0002. The claim of 'consistently significant differences between LLM and human evaluations (p < 0.01)' is not supported by the analysis as run.
  2. [Study 1, Contextual Sensitivity Analysis] The paired t-tests for context adaptation (e.g., GPT t(9) = -6.07, p < 0.001; Gemini t(9) = -2.42, p = 0.038) pair 10 identical resume copies across conditions. This is not a valid paired design because the units are evaluators, not resumes; with 3 evaluators per model the test would have at most 2 degrees of freedom. Consequently, the model-specific adaptation hierarchy (GPT strong, Gemini partial, Claude minimal) is not statistically established.
  3. [Study 2: Randomized Resume Analysis] The LLM-only ANOVA for the Firm2 randomized condition reports F(2,27) = 3.83, p = 0.344. With df = (2,27), an F of 3.83 corresponds to a p-value of about 0.034, not 0.344. This internal inconsistency is a clear reporting error and raises questions about the reliability of the other ANOVA output. The authors should verify all reported F and p values and provide the raw data so that the analyses can be checked.
  4. [Study 3: Reduced Context Analysis / Limitations] The ANOVA comparing all evaluators in the Reduced Context condition is reported as F(5,54) = 18.23, p = 0.0002, implying six groups of ten scores. This is inconsistent with the Limitations section, which states that one human expert was lost in later experimental conditions, and with the earlier observation that Expert 3 provided zero scores in some conditions. If Expert 3's zero scores were excluded, the degrees of freedom would be different; if they were included, the narrative about expert loss is inaccurate. The treatment of missing or zero human scores must be clarified and reflected in the reported degrees of freedom.
minor comments (5)
  1. [Abstract and Results] The abstract states 'consistently significant differences between LLM and human evaluations (p < 0.01)', while the body uses α = 0.1 for several LLM-only comparisons; the summary should distinguish the α levels used for different families of tests.
  2. [Results, 'four of eight LLM-only conditions'] The phrase 'four of eight LLM-only conditions' is not defined in the text. The authors should list exactly which conditions produced significant effects and at which alpha level, ideally in a summary table.
  3. [Statistical Analysis Plan] The plan mentions false discovery rate corrections, but no FDR-adjusted p-values are reported anywhere in the Results. The authors should either apply the correction or explain why it was omitted for the primary analyses.
  4. [Method / Data Availability] No raw scores, analysis scripts, or software version are provided. Given the centrality of the statistical claims and the errors noted above, a data supplement with evaluator-level scores is essential for any revision.
  5. [Results, Descriptive Statistics] The treatment of Expert 3's zero scores is inconsistent: the descriptive text says they 'provided zero scores for all candidates in some later conditions,' but later analyses report means for Expert 3 in the Reduced Context condition. Please specify the exact exclusion rule applied to each analysis.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the empirical comparisons are self-contained; the evaluator-independence weakness is a statistical validity concern, not a circular derivation.

full rationale

I walked the paper's claimed derivation chain: the study collects raw LLM and human expert scores, applies one-way ANOVA and paired t-tests, and interprets the resulting p-values and effect sizes. No model parameter is fitted from a subset of data and then renamed as a prediction; no result is defined in terms of another result it is supposed to explain; and no load-bearing uniqueness theorem or prior-work ansatz is imported from the authors' own publications. The Signal Detection Theory framing appears only in the Introduction and Discussion as an interpretive vocabulary, not as a source of derived predictions: the paper does not compute d-prime or criterion measures, and the signal-versus-noise language is a descriptive label for the observed significant differences and variability, not a mechanism by which those differences are produced. The one substantive concern is statistical, not circular: the Same Resume condition uses ten identical copies per evaluator, and the reported ANOVA and paired t-tests treat the ten scores per evaluator as independent observations, which likely makes p-values anti-conservative because repeated judgments by the same evaluator on identical material are correlated. That issue affects whether the p<0.01 significance claims are valid as stated, but it is not a circularity: the claims are not forced by construction, by definition, or by self-citation. Accordingly, no circular step meets the evidentiary standard required for this analysis, and the appropriate score is 0.

Assumptions & free parameters 1 free parameters · 5 assumptions · 0 invented entities

The central comparisons rest on statistical independence and representativeness assumptions that are not satisfied, plus an unstated reliance on unversioned LLM outputs. There are no fitted physical parameters and no invented entities, so the ledger is short, but the domain assumptions are load-bearing and fragile.

free parameters (1)
  • alpha threshold for significance = 0.1
    Primary analyses report significance at both 0.1 and 0.01; several LLM-only context effects are only significant at 0.1, and no multiple-comparison correction is applied, inflating the number of significant findings.
assumptions (5)
  • domain assumption Scores from ten identical resume copies are independent observations.
    ANOVA and paired t-tests treat each same-resume score as an independent unit; repeated presentations from one evaluator are likely correlated, so this assumption is violated.
  • domain assumption Three human experts are a representative benchmark for human recruiter judgment.
    Human-AI difference claims are generalized from n=3 experts, one of whom zero-scored in later conditions.
  • domain assumption Unversioned LLM API outputs characterize the model's stable behavior.
    No model version, temperature, or sampling parameters are reported; a single run may capture random noise.
  • standard math ANOVA normality and homogeneity assumptions hold.
    The paper says assumption testing was performed, but no residual analyses or transformation results are shown.
  • domain assumption Signal Detection Theory can be inferred from score means alone.
    No hit/false-alarm matrix, d', or criterion is computed; SDT is used as a narrative frame.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Signal or Noise? Evaluating Large Language Models in Resume Screening Across Contextual Variations and Human Expert Benchmarks." pith.science (2026). https://pith.science/paper/Y37WSMAY

@misc{pith2026250708019,
  author       = {Pith},
  title        = {Pith review of: Signal or Noise? Evaluating Large Language Models in Resume Screening Across Contextual Variations and Human Expert Benchmarks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Y37WSMAY}},
  note         = {Machine review of arXiv:2507.08019}
}
read the original abstract

This study investigates whether large language models (LLMs) exhibit consistent behavior (signal) or random variation (noise) when screening resumes against job descriptions, and how their performance compares to human experts. Using controlled datasets, we tested three LLMs (Claude, GPT, and Gemini) across contexts (No Company, Firm1 [MNC], Firm2 [Startup], Reduced Context) with identical and randomized resumes, benchmarked against three human recruitment experts. Analysis of variance revealed significant mean differences in four of eight LLM-only conditions and consistently significant differences between LLM and human evaluations (p < 0.01). Paired t-tests showed GPT adapts strongly to company context (p < 0.001), Gemini partially (p = 0.038 for Firm1), and Claude minimally (p > 0.1), while all LLMs differed significantly from human experts across contexts. Meta-cognition analysis highlighted adaptive weighting patterns that differ markedly from human evaluation approaches. Findings suggest LLMs offer interpretable patterns with detailed prompts but diverge substantially from human judgment, informing their deployment in automated hiring systems.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

6 extracted references · 5 canonical work pages

  1. [1]

    stochastic parrots,

    abling advanced text understanding and generation. The introduction of the attention mechanism in transformers allows these models to weigh the importance of different parts of the input, leading to significant improvements in tasks such as text classification, summarization, and question answering (Fields et al., 2024; Wang et al., 2024). This architectu...

  2. [2]

    Your scoring pattern: How did you divide the total of 100 and what all parameters did you choose to score the resumes." For human expert evaluation, materials were distributed via email with comprehensive instructions explaining the evaluation task. Experts were asked to develop their own marking schemes based on their professional experience and apply th...

  3. [4]

    Claude remained largely non-responsive to contextual information across diverse candidates, with non-significant changes for both company contexts, reinforcing its characterization as having minimal contextual adaptability. Human experts demonstrated more nuanced adaptation patterns with diverse candidates that reflected their professional expertise and a...

  4. [18]

    Armstrong, L., Liu, A., MacNeil, S., & Metaxa, D. (2024). The silicon ceiling: Auditing GPT’s race and gender biases in hiring. Barocas, S., & Selbst, A. D. (2016). Big data’s disparate impact. California Law Review, 104(3), 671–732. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language m...

  5. [1995]

    All statistical analyses were conducted using appropriate software with significance levels and confidence intervals reported for all hypothesis tests

    to control for Type I error inflation across multiple hypothesis tests. All statistical analyses were conducted using appropriate software with significance levels and confidence intervals reported for all hypothesis tests. Results Descriptive Statistics The dataset comprised 420 total evaluations across all conditions and evaluators. LLM scores consisten...

  6. [2017]

    composite scoring based on professional experience

    and suggests that GPT may interpret information scarcity as grounds for optimistic evaluation rather than increased caution. Claude's relative stability (66.6 to 65.6) masked high individual variability, with a standard deviation of 15.1 indicating inconsistent application of evaluation criteria despite apparent mean stability. Gemini's moderate increase ...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.