Pith. sign in

REVIEW 2 major objections 4 minor 44 references

How Common Are Estimated Latent-Distribution Departures From Normality? Evidence From 504 Item-Response Data Sets

T0 review · 2 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read More than half of 504 real item-response datasets depart from the common normal-trait assumption by at least ten cumulative-probability points at some trait level, with a median peak difference near eleven points.

desk verdict First large-scale empirical map of estimated latent-trait departures from normality; carefully scoped and transparent, with a real but manageable identification caveat about short tests. read the letter →

arxiv 2608.06817 v1 pith:GNLXUSU5 submitted 2026-08-07 stat.AP

classification stat.AP
keywords itemresponsetheorylatenttraitdistributionnormalityassumptionempiricalhistogramgrid-boundaryKolmogorov-Smirnovdistancereliabilitysensitivitypersonscores
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks how often the routine assumption that an unmeasured trait is normally distributed among test-takers actually holds in real data. Fitting 504 item-response datasets twice, once with the normal assumption and once with a free-form latent distribution, the author finds that in more than half of the datasets the two fitted distributions differ by at least ten percentage points of cumulative probability at some point on the trait scale; the median maximum difference is about eleven points, and nearly one dataset in five reaches twenty points. The estimated shapes include skewness, heavy tails, flat regions, and occasional multimodality, with skewness and heavy tails the most common forms. Because the estimated densities are conditional on the working item models, the results are framed as sensitivity of the normality convention rather than as recovery of a true population density. The paper's recommended practice is that applied analyses should state the distributional assumption, examine a flexible alternative, and report sensitivity separately for each result used in interpretation or decision making.

What carries the argument

The central object is the grid-boundary Kolmogorov–Smirnov distance D_KS, the maximum absolute difference in cumulative probability between the standardized empirical-histogram estimate and the standard normal distribution, evaluated only at the midpoints between adjacent quadrature-grid cells. This quantity carries the comparison because it can be read directly as a discrepancy in cumulative mass: a value of .10 means that at one evaluated cell boundary the flexible estimate places ten percentage points more or less of its mass below that boundary than the normal distribution implies. The empirical-histogram method, which re-estimates masses on a 121-point quadrature grid inside the EM algorithm, supplies the flexible density, and the whole analysis is performed by fitting each dataset twice under identical item models, quadrature rules, and optimizer controls, differing only in the latent distribution, with post hoc standardization to a common metric.

What would settle it

Re-estimate the same 504 datasets using item response functions with free discriminations (2PL/GRM) and, where possible, a lower asymptote (3PL), and recompute the grid-boundary KS distances; if the median D_KS falls below .05 and the share above .10 drops well below 20%, the claim that departures from normality are common would be shown to depend on the restricted working models.

Watch

Extended reading notes

Core claim

The author establishes that substantial estimated departures from a normal latent trait distribution are common across a large, heterogeneous corpus of item-response data. On a standardized metric, the grid-boundary Kolmogorov–Smirnov distance between the empirical-histogram estimate and the matched normal distribution exceeds .10 in 56.0% of the 504 units, .15 in 30.8%, and .20 in 18.5%, with a median of .109 and a 90th percentile of .257. The paper then shows that the consequences of relaxing the normality restriction depend on the quantity reported: the fixed-item density component of reliability changes by a median absolute value of only .002, full-refit marginal reliability by .025, item threshold or location RMSD by .274, test-curve RMSD per item by .077, and standardized person scores by .093, with a median 4.8% of persons crossing a |z|=1 classification boundary. It also documents that alternative flexible estimators (extrapolated empirical histogram and Hannan–Quinn-selected Davidian curves) broadly agree on which datasets are most unusual but differ on the magnitude of the departure.

Load-bearing premise

The central claim assumes that the Rasch and GPCM working item models are close enough to the data that the flexible density estimate reflects latent-distribution shape rather than absorbing model misfit such as multidimensionality, local dependence, or unmodeled lower asymptotes.

Editorial extensions

If this is right

  • If departures of the measured size are as common as reported, simulation studies of IRT estimators should generate latent distributions with magnitudes anchored to real corpora—median D_KS around .11, with a substantial tail above .20—rather than arbitrary shapes.
  • Reliability alone is not a sufficient sensitivity check: the paper shows that the fixed-item density component of reliability changes little (median absolute .002) while item parameters, test curves, and person scores move noticeably, so a satisfactory reliability coefficient gives little assurance about other reported quantities.
  • Uses that attach consequences to absolute locations—cut scores, norm tables, growth percentiles, screening boundaries—are the ones most exposed to between-calibration differences of the size observed, and the paper's person-score shifts concentrate near tail thresholds.
  • The estimated departure is conditional on the working item model and density estimator, so any report of a latent-shape diagnostic should name the item model, density family, quadrature or smoothing controls, and order-selection rule.
  • The paper's descriptive exceedance shares are not formal test rejections; converting them into rejection rates would require a per-unit reference distribution for the fitted diagnostic under the same estimator and design, which the author identifies as a natural next step.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The numerical landmarks—median D_KS ~ .11, ~31% above .15, ~19% above .20—could serve as a benchmark scale of nonnormality magnitudes for future robustness studies, much as observed-score surveys did for raw score distributions.
  • If the working Rasch and GPCM models absorb multidimensionality or local dependence as shape, a larger-scale replication using 2PL or GRM fits might show how much of the measured departure is item-model artifact; the paper's small Rasch-to-2PL check (median absolute KS change about .015) hints the effect may be modest but not negligible.
  • The paper's finding that person-score shifts are largest in the tails suggests that screening and classification decisions using tail thresholds could be the most sensitive; a direct study linking D_KS magnitude to decision error under real stakes would be a testable extension.
  • The high median excess kurtosis (~3.6) implies realistic latent distributions are markedly heavier-tailed than the normal, which may matter for the accuracy of EAP shrinkage in the tails; one could examine whether flexible priors reduce bias for extreme scorers in the same corpus.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper carries out a large-scale empirical description of fitted latent trait distributions: it calibrates 504 real item-response data sets from the Item Response Warehouse under a normal latent density and under a 121-point empirical-histogram density, using Rasch for dichotomous and GPCM for polytomous units, with shared optimizer, item-model, and standardization settings. It summarizes the fitted difference with a grid-boundary Kolmogorov–Smirnov distance, skewness, kurtosis, and a Hartigan-dip descriptor, and it compares reliability, item parameters, test characteristic curves, and EAP scores between the two calibrations. The headline result is that more than half of the units (56.0%) show KS > .10, with a median KS of .109, and that full-refit quantities such as marginal reliability and person scores can move appreciably even when the fixed-item density component of reliability changes little. The authors repeatedly state that the estimated density is model-conditional and is used as a sensitivity measure rather than as a recovered true density, and they include robustness checks across quadrature/optimizer settings, alternative density estimators, and a known-truth simulation, with code and a replication manifest released.

Significance. If the headline finding is robust, this is a valuable empirical anchor for applied IRT: it gives the field a first large-scale description of how often the conventional normal latent-distribution assumption changes fitted results in real data, and it identifies which reported quantities are most sensitive. The design has genuine strengths that should be credited: both calibrations share the same item model and optimizer controls, all coordinates are standardized to a common metric before comparison, the unit-flow manifest prevents silent exclusions, multiple density estimators and a known-truth simulation are included, and the replication package is unusually complete. The principal reservation is that the headline exceedance rate is computed over the full corpus including very short, weakly identified tests; if the rate does not survive restricting to adequately identified units, the central claim would need substantial reframing. My major comments therefore ask for a test-length/identification-stratified reporting of the central frequencies and for a corpus-level check of item-model sensitivity.

major comments (2)
  1. [§3.1, Table 2; §3.2, Table 3; OSM C.2] The headline claim that 56.0% of data sets show D_KS > .10 is reported only for the full 504-unit frame, which includes units with as few as 3 items and 500 persons (Sec. 2.1). Table 3 reports a negative log(items) coefficient (β = -0.170, p = .001), showing that high-KS units are concentrated in shorter tests, and OSM C.2 states that the selected units' finite response patterns left degrees of freedom in the grid masses, so the 121-point empirical histogram is not point-identified in the weak regime. The known-truth simulation in Sec. 2.7 covers 10- and 20-item tests at N = 5,000, not the corpus minimum of 3 items / 500 persons, so recovery evidence does not cover the problematic regime. I request that the authors report the exceedance rates at KS > .10, .15, and .20 and the median KS for an adequately identified subset (e.g., units with at least 10 or 20 items, or with a larger number of observed response patterns), and that they state explicitly whether the 'more than half of data sets' assertion survives in that subset. Without this stratification, the central empirical assertion is not robust to the weakest-identified units in the corpus.
  2. [§2.2, §3.4, §4.5] The estimated departure from normality is conditional on the working Rasch or GPCM item model, as the authors acknowledge, and multidimensionality, local dependence, DIF, and unmodeled lower asymptotes can be absorbed into the flexible density estimate as shape. The item-response-function check in Sec. 3.4 covers only eight dichotomous units, so it is too narrow to establish corpus-wide whether large KS values are driven by item-model misspecification rather than latent shape. I ask the authors to add a corpus-level sensitivity analysis on a larger set of dichotomous units (e.g., refit high-KS cognitive units with 2PL/3PL, or apply a comparable polytomous robustness check) and to report how the exceedance frequencies change; at minimum, the abstract and the Section 3.1 summary should state more prominently that the frequencies are conditional on the working item models and the EH quadrature grid, so that readers do not read the headline as evidence about the true latent distribution.
minor comments (4)
  1. [§2.3, Eq. (1)] The indexing in the definition of D_KS is somewhat compressed: the reader must infer that k = 0 and k = K contribute zero differences (W_0 = 0, Φ(-∞) = 0; W_K = 1, Φ(∞) = 1). A sentence explicitly stating that the endpoints are included only for notational completeness would help.
  2. [§2.7, §4.4] The comparison between EH and EHW uses the subset of 429 complete pairs, while the Davidian comparisons use 384 GPCM pairs; the difference in denominators is reported in Table 5 but could be stated more explicitly in the main text to avoid the impression that all comparisons are on the same units.
  3. [§2.3, Table 2] The Hartigan-dip descriptor depends on the Monte Carlo draw size (4,000), the jitter SD (.05), and the deterministic per-unit seed; this is described in Sec. 2.3, but Table 2 does not remind the reader that the reported dip values are smoothed Monte Carlo descriptors rather than exact dip values of the unsmoothed masses. A footnote to the table would prevent misinterpretation.
  4. [§2.8] The replication materials are described as released at a GitHub URL, but no versioned DOI or commit hash is given; adding a persistent identifier would make the reproducibility claim more durable and citable.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the headline KS results are descriptive between-fit comparisons, explicitly model-conditional, and not predictions derived from fitted constants.

full rationale

The paper makes no predictive or first-principles claim; its central quantities are descriptive summaries of two fitted models applied to the same data. The grid-boundary KS distance is defined directly as max over cell boundaries of |W_k - Phi(B_k)| where W_k are the empirical-histogram grid masses and Phi is the standard normal CDF (Sec. 2.3, OSM A.2). The 'departures' reported are therefore, by construction, a comparison between the fitted flexible distribution and the normal distribution under the same item model, not an estimate of a true latent density. The paper explicitly states this: 'It does not identify an estimator-free population density, and differences between calibrations are accordingly interpreted as sensitivity measures rather than as bias with respect to a known truth' (Sec. 1.3), and 'The estimated distribution is conditional on this item model, density estimator, grid, and optimization rule' (Sec. 2.2). The exceedance shares are labeled descriptive: 'The exceedance shares reported below are therefore descriptive counts, not significance decisions' (Sec. 2.4). No fitted parameter is renamed as a prediction; no result is imported from a self-citation chain; the only citations to methods (Bock & Aitkin 1981; Woods 2007a; Woods & Lin 2009) are external prior work used as tools, not as authority for the empirical finding. The concern that short tests leave the flexible density weakly identified is a validity and robustness limitation that the paper itself acknowledges (Secs. 1.3, 2.2, 4.5; OSM C.2), and it affects interpretation of the descriptive statistic, but it is not circularity: the KS distance still measures exactly what the paper says it measures, namely the difference between the two fitted calibrations. The analysis is also self-contained against external checks: a known-truth simulation (OSM E), an alternative-estimator comparison (Table 5), and a public replication package. Verdict: no significant circularity; score 0.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central frequencies are descriptive summaries of a specific fitting pipeline. The most important unstated premises are that the chosen item response functions are adequate for these data and that the 121-point grid resolves the shapes. No invented entities are introduced.

free parameters (3)
  • Jitter SD for dip descriptor = 0.05
    Hand-chosen smoothing constant in the Monte Carlo Hartigan-dip descriptor (Section 2.3); affects the multimodality summary but not the primary KS frequency.
  • Davidian candidate order caps = 6 and 10
    Chosen ceilings for Hannan-Quinn order selection in the density-estimator comparison (Section 2.7); influences estimator-agreement magnitudes, not the primary KS claim.
  • EH quadrature grid size = 121 points
    Fixed grid used for empirical-histogram estimation; part of the estimator contract (Section 2.2).
assumptions (4)
  • domain assumption Rasch model is the correct item response function for dichotomous units
    Dichotomous data were fit with Rasch, not 2PL, to aid identification (Section 2.1); misspecification of the IRF could imprint shape on the density estimate.
  • domain assumption GPCM is the correct item response function for polytomous units
    Polytomous data were fit with GPCM (Section 2.1); misspecification could create spurious departures.
  • ad hoc to paper The 121-point empirical-histogram grid resolves latent shape sufficiently
    The grid width and spacing are fixed choices (Section 2.2); the paper acknowledges tail detail is limited by the grid (Section 4.5).
  • domain assumption The IRW corpus composition does not bias the descriptive frequencies for the analyzed units
    Frequencies describe the analyzed units only, and the paper explicitly disclaims generalization (Section 4.5).

how reviews work

0 comments
Cite this review

Pith. "Pith review of How Common Are Estimated Latent-Distribution Departures From Normality? Evidence From 504 Item-Response Data Sets." pith.science (2026). https://pith.science/paper/GNLXUSU5

@misc{pith2026260806817,
  author       = {Pith},
  title        = {Pith review of: How Common Are Estimated Latent-Distribution Departures From Normality? Evidence From 504 Item-Response Data Sets},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GNLXUSU5}},
  note         = {Machine review of arXiv:2608.06817}
}
read the original abstract

Item response models usually assume a normal trait distribution, yet little is known about how often fitted distributions in real studies differ substantially from normality or which reported results are most affected. We analyzed 504 itemresponse data sets from 273 studies in the Item Response Warehouse, fitting each with the normal assumption and with a flexible distribution estimated from the responses, and we compared reliability, item estimates, predicted test responses, and person scores between the two calibrations. In more than half of the data sets the two fitted distributions differed by at least 10 percentage points of cumulative probability at some point on the trait scale, with a median maximum difference of about 11 points; about one-third reached 15 points, and nearly one-fifth reached 20 points. The estimated shapes included skewness, heavy tails, flat regions, and occasional multimodality. Differences were larger in several attitudinal, affective, and behavioral domains, although these contrasts weakened when data sets were compared within item-model families. Reliability usually changed little when item estimates were held fixed, whereas refitting the full model produced larger changes in some data sets, and item estimates, predicted test responses, and person scores did not change in parallel. Alternative flexible methods generally identified the same data sets as most unusual but disagreed somewhat about magnitude. Applied analyses should state the distributional assumption, examine a flexible alternative, and report sensitivity separately for each result used in interpretation or decision making.

Figures

Figures reproduced from arXiv: 2608.06817 by the authors.

Figure 1
Figure 1. Selected Estimated Latent Distributions KS = 0.017 KS = 0.151 KS = 0.073 KS = 0.175 KS = 0.128 KS = 0.394 (d) right−skewed · Cognitive (e) left−skewed · Opinion (f) multimodal (a) near−normal · Cognitive (b) flat (c) heavy−tailed · Affective −2 0 2 −2 0 2 −2 0 2 Standardized latent trait θ Density Estimated latent density Matched normal Note. The solid curves are smoothed standardized EH estimates, and the dashed cu… view at source ↗
Figure 2
Figure 2. Magnitude and Frequency of Estimated Departures median 0.109 0 10 20 30 40 50 0.0 0.1 0.2 0.3 0.4 Grid−boundary KS distance Units (a) Distribution across 504 units KS > .05 87.3% (440/504) KS > .10 56.0% (282/504) KS > .15 30.8% (155/504) KS > .20 18.5% (93/504) 0% 20% 40% 60% 80% 100% 0.0 0.1 0.2 0.3 0.4 KS threshold Share of units (b) Share above each KS threshold Note. The left panel shows the grid-boundary KS di… view at source ↗
Figure 3
Figure 3. KS Distance by Construct and Item-Model Family Personality (n=28) Cognitive/educational (n=102) Other (n=9) Developmental (n=8) Unclassified (n=109) Physical health/functioning (n=20) Behavioral (n=30) Affective/mental health (n=113) Opinion/attitude (n=85) 0.0 0.1 0.2 0.3 0.4 Grid−boundary KS distance (a) Pooled unit distributions n=6 n=107 n=61 n=41 n=4 n=81 n=38 n=71 Cognitive/educational Unclassified Affective/m… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Between-Calibration Sensitivity and KS Distance 1 value off scale (6.72×10^20) Full−refit marginal reliability median 0.025 · n=496 Test−curve RMSD/item median 0.077 · n=496 Person−score MAD median 0.093 · n=496 0.0 0.1 0.2 0.3 0.4 0.0 0.1 0.2 0.3 0.4 0.0 0.1 0.2 0.3 0…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

44 extracted references · 41 canonical work pages

  1. [1]

    and Arnau, Jaume and López-Montiel, Dolores and Bono, Roser and Bendayan, Rebecca , title =

    Blanca, María J. and Arnau, Jaume and López-Montiel, Dolores and Bono, Roser and Bendayan, Rebecca , title =. Methodology , year =. doi:10.1027/1614-2241/a000057 , url =

  2. [2]

    and Zhang, Zhiyong and Yuan, Ke-Hai , title =

    Cain, Meghan K. and Zhang, Zhiyong and Yuan, Ke-Hai , title =. Behavior Research Methods , year =. doi:10.3758/s13428-016-0814-1 , url =

  3. [3]

    and Yu, Carol C

    Ho, Andrew D. and Yu, Carol C. , title =. Educational and Psychological Measurement , year =. doi:10.1177/0013164414548576 , url =

  4. [4]

    Darrell and Aitkin, Murray , title =

    Bock, R. Darrell and Aitkin, Murray , title =. Psychometrika , year =. doi:10.1007/BF02293801 , url =

  5. [5]

    and Lin, Nan , title =

    Woods, Carol M. and Lin, Nan , title =. Applied Psychological Measurement , year =. doi:10.1177/0146621608319512 , url =

  6. [6]

    and Thissen, David , title =

    Woods, Carol M. and Thissen, David , title =. Psychometrika , year =. doi:10.1007/s11336-004-1175-8 , url =

  7. [7]

    , title =

    Woods, Carol M. , title =. Educational and Psychological Measurement , year =. doi:10.1177/0013164406288163 , url =

  8. [8]

    Educational and Psychological Measurement , year =

    Li, Zhen and Cai, Li , title =. Educational and Psychological Measurement , year =. doi:10.1177/0013164417717024 , url =

Show all 44 references
  1. [9]

    British Journal of Mathematical and Statistical Psychology , year =

    Guastadisegni, Lucia and Cagnone, Silvia and Moustaki, Irini and Vasdekis, Vassilis , title =. British Journal of Mathematical and Statistical Psychology , year =. doi:10.1111/bmsp.12379 , url =

  2. [10]

    , title =

    Mislevy, Robert J. , title =. Psychometrika , year =. doi:10.1007/BF02306026 , url =

  3. [11]

    Sass, D. A. and Schmitt, T. A. and Walker, C. M. , title =. Applied Measurement in Education , year =. doi:10.1080/08957340701796415 , url =

  4. [12]

    Psychometrika , year =

    Samejima, Fumiko , title =. Psychometrika , year =. doi:10.1007/BF02294639 , url =

  5. [13]

    van den Oord, Edwin J. C. G. , title =. Applied Psychological Measurement , year =. doi:10.1177/0146621604269791 , url =

  6. [14]

    Journal of Educational and Behavioral Statistics , year =

    Monroe, Scott , title =. Journal of Educational and Behavioral Statistics , year =. doi:10.3102/1076998620953764 , url =

  7. [15]

    PLOS Mental Health , year =

    Ferr, Heury , title =. PLOS Mental Health , year =. doi:10.1371/journal.pmen.0000403 , url =

  8. [16]

    and Braginsky, Mika and Caffrey-Maffei, Lucy and Gilbert, Joshua B

    Domingue, Benjamin W. and Braginsky, Mika and Caffrey-Maffei, Lucy and Gilbert, Joshua B. and Kanopka, Klint and Kapoor, Radhika and Lee, Hansol and Liu, Yiqing and Nadela, Savira and Pan, Guanzhong and Zhang, Lijin and Zhang, Susu and Frank, Michael C. , title =. Behavior Res...

  9. [17]

    Psychological Bulletin , year =

    Micceri, Theodore , title =. Psychological Bulletin , year =

  10. [18]

    and Bock, R

    Green, Bert F. and Bock, R. Darrell and Humphreys, Lloyd G. and Linn, Robert L. and Reckase, Mark D. , title =. Journal of Educational Measurement , year =

  11. [19]

    , title =

    Woods, Carol M. , title =. Applied Psychological Measurement , year =

  12. [20]

    , title =

    Woods, Carol M. , title =. Psychological Methods , year =

  13. [21]

    Educational and Psychological Measurement , year =

    Monroe, Scott and Cai, Li , title =. Educational and Psychological Measurement , year =

  14. [22]

    and MacEachern, Steven N

    Duncan, Kristin A. and MacEachern, Steven N. , title =. Statistical Modelling , year =

  15. [23]

    San Mart. On the. Psychometrika , year =

  16. [24]

    , title =

    Finch, Holmes and Edwards, Julianne M. , title =. Educational and Psychological Measurement , year =

  17. [25]

    and Tao, Jian , title =

    Zhang, Xue and Wang, Chun and Weiss, David J. and Tao, Jian , title =. Multivariate Behavioral Research , year =

  18. [26]

    , title =

    Wang, Chun and Su, Shiyang and Weiss, David J. , title =. Multivariate Behavioral Research , year =

  19. [27]

    and Edwards, Michael C

    Manapat, Percival D. and Edwards, Michael C. , title =. Educational and Psychological Measurement , year =

  20. [28]

    Philip , title =

    Chalmers, R. Philip , title =. Journal of Statistical Software , year =

  21. [29]

    2025 , howpublished =

    Robitzsch, Alexander , title =. 2025 , howpublished =. doi:10.32614/CRAN.package.sirt , verify =

  22. [30]

    2025 , note =

    Pritikin, Joshua , title =. 2025 , note =. doi:10.32614/CRAN.package.rpf , url =

  23. [31]

    Hartigan, J. A. and Hartigan, P. M. , title =. The Annals of Statistics , year =. doi:10.1214/aos/1176346577 , verify =

  24. [32]

    2025 , doi =

    Maechler, Martin and Ringach, Dario , title =. 2025 , doi =

  25. [33]

    , title =

    Pustejovsky, James E. , title =. 2026 , doi =

  26. [34]

    2026 , doi =

    Barrett, Tyson and Dowle, Matt and Srinivasan, Arun and Gorecki, Jan and Chirico, Michael and Hocking, Toby and Schwendinger, Benjamin and Krylov, Ivan , title =. 2026 , doi =

  27. [35]

    2016 , isbn =

    Wickham, Hadley , title =. 2016 , isbn =

  28. [36]

    and Mayo-Wilson, Evan and Nezu, Arthur M

    Appelbaum, Mark and Cooper, Harris and Kline, Rex B. and Mayo-Wilson, Evan and Nezu, Arthur M. and Rao, Stephen M. , title =. American Psychologist , year =. doi:10.1037/amp0000191 , verify =

  29. [37]

    , title =

    Lord, Frederic M. , title =. Educational and Psychological Measurement , year =

  30. [38]

    , title =

    Cook, Desmond L. , title =. Educational and Psychological Measurement , year =

  31. [39]

    Simulation studies for methodological research in psychology:

    Siepe, Bj. Simulation studies for methodological research in psychology:. Psychological Methods , year =

  32. [40]

    Psychological Methods , year =

    Soland, James and Kuhfeld, Megan and Edwards, Kelly , title =. Psychological Methods , year =

  33. [41]

    Rasch, Georg , title =

  34. [42]

    Applied Psychological Measurement , year =

    Muraki, Eiji , title =. Applied Psychological Measurement , year =

  35. [43]

    and Beaton, Albert E

    Mislevy, Robert J. and Beaton, Albert E. and Kaplan, Bruce and Sheehan, Kathleen M. , title =. Journal of Educational Measurement , year =

  36. [44]

    and Cohen, Allan S

    Bolt, Daniel M. and Cohen, Allan S. and Wollack, James A. , title =. Journal of Educational and Behavioral Statistics , year =

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.