REVIEW 3 major objections 8 minor 35 references
Towards Measurement Theory for Artificial Intelligence
T0 review · 3 major / 8 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper argues that a formal measurement theory for AI—synthesizing representational measurement, metrology, and measure theory—would make AI evaluation commensurable and scientifically grounded.
desk verdict A well-motivated programmatic proposal for measurement theory in AI, but its commensurability promise rests on an unproved RTM precondition that needs a worked example. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the 'AI observable', defined as a measurable function $\varphi : (\Omega,\mathcal{F}) \to (\mathbb{R},\mathcal{B}(\mathbb{R}))$, where $\Omega$ is the space of AI system states and $\mathcal{F}$ is a $\sigma$-algebra of distinguishable events. This formal definition is coupled with a layered measurement stack spanning physical, systems, algorithm, task/behaviour, and contextual/emergent layers, each with its own observables, instrumentation requirements, and validity conditions. The argument also relies on representational measurement's homomorphism from an empirical relational structure to a numerical structure, and on a proposed meta-axiom requiring every measurement proposition to declare its empirical system, numerical system, representation function, and uniqueness group of scale transformations.
What would settle it
A demonstration that pairwise 'at least as capable as' judgments across a broad task suite are systematically intransitive (A beats B, B beats C, but C beats A), or that no stable ordering survives context shifts, would falsify the central promise of commensurable AI measurement.
Extended reading notes
Core claim
The paper's central claim is that a principled measurement theory for artificial intelligence would (and ought to) enable commensurable evaluations across models, tasks, and research groups, and would connect frontier AI evaluations with established quantitative risk-analysis techniques. The discovery is a diagnosis: AI evaluation currently lacks the scope, rigour, and depth of measurement practice in other sciences, and the remedy is a formal framework that defines AI observables, assigns scale types, and makes the choice of measurement operations explicit. The paper adopts a position of methodological realism, hypothesizing that stable latent attributes of AI systems exist and can be validated through consistent, coherent, predictive measurement models.
Load-bearing premise
The framework rests on the premise that AI attributes such as capability or reliability can be compared consistently enough—if A beats B and B beats C, then A beats C—for a meaningful numerical scale to exist.
Editorial extensions
If this is right
- If adopted, MTAI would make evaluation results commensurable across models, tasks, and research groups by assigning each AI attribute a scale type and an invariance class.
- It would allow frontier AI evaluations to plug into established quantitative risk-analysis and reliability-engineering techniques, which require variables with known scale properties.
- It would make the definition of AI capability explicitly contingent on the measurement operations and scales chosen, surfacing choices that current benchmarking hides.
- It would give regulators and auditors standardised, calibratable measurement protocols for compliance, reliability, and risk levels.
- It would expose ill-defined constructs: a proposed measure that cannot be cast as a measurable function on the system state space lacks mathematical foundation.
Reading between the lines
- Editorial extension: if this programme succeeds, benchmark leaderboards could be replaced or supplemented by explicit scale-type declarations, making it impossible to report a 'capability score' without stating its invariance class.
- Editorial extension: the framework implies a concrete research agenda: test whether pairwise 'at least as capable as' judgments on real model pairs are transitive across diverse task suites; current benchmarks do not establish this.
- Editorial extension: the AI-observable definition suggests a testable criterion for whether a proposed metric is well-founded—if it cannot be expressed as a measurable function with a stated uniqueness group, it is not yet an AI observable.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues for the development of a formal measurement theory for artificial intelligence (MTAI), synthesizing representational theory of measurement (RTM), measure theory, metrology, and psychometrics. It motivates the program by citing the fragmentation of current AI evaluation practice, the need for commensurable comparisons, standardization for risk analysis, and the contingency of AI capability on measurement choices. The paper sketches key components: a five-layer AI measurement stack, the distinction between direct and indirect observables, a formal definition of an AI observable as a measurable function, a meta-axiom quadruple for measurement propositions, and initial examples of how the framework might apply. The authors explicitly frame the work as an extended abstract that outlines and motivates a programme rather than a fully formalized theory.
Significance. If the proposed MTAI program is carried through, it could provide a much-needed metrological and psychometric foundation for AI evaluation, enabling cumulative comparison across models, tasks, and research groups and connecting AI risk assessment with established quantitative methods. The paper is unusually self-aware: it names its own obstacles (non-transitive orderings, context-dependence, high-dimensional evolution, black-box systems) and honestly distinguishes between motivation and achievement. Its historical synthesis of measurement theory is informative, and the layered measurement stack offers a useful organizing framework. At this stage, the value lies primarily in framing and agenda-setting rather than in a demonstrated theory; the central promises are conditional on empirical and axiomatic premises that are not yet established.
major comments (3)
- [Sections 2.4.1, 2.6, 3.1] The paper's central claim that MTAI 'would (and ought to) enable commensurable evaluations across models, tasks, and research groups' rests on the existence of an empirical relational structure over AI constructs that satisfies the RTM axioms. Section 2.4.1 explicitly conditions the construction of an ordinal or interval scale on properties such as transitivity, and Section 3.1's meta-axiom requires a homomorphism f : E → N that preserves all empirical relations. However, Section 2.6 (item 2) acknowledges that 'a single linear scale for capability might not exist' and that partial or multi-dimensional orderings may be necessary. For a natural Pareto-dominance ordering over tasks, no homomorphism into (R, ≥) can preserve the partial order. The paper never provides a worked example of an AI construct satisfying the required axioms, nor does it state conditions under which the representation theorem would apply. Without such an example or a clear treatment of partial orderings, the promised commensurability is not a consequence of the framework. Please supply at least one concrete worked example and show how the framework accommodates partial or intransitive orderings, or explicitly restrict the domain of applicability.
- [Section 2.5] The formal condition 's1 ⪰ s2 =⇒ φT (s1) ≤ φT (s2)' is reversed relative to the intended meaning of ⪰ as 'is at least as capable as.' If s1 is at least as capable as s2, the numerical representation should be non-decreasing, not non-increasing. This also contradicts Section 3.1, where 'a ⪯E b ⇐⇒ f(a) ≥ f(b)'. The inconsistency affects the formal sketch of representation functions and should be corrected.
- [Section 3.1] In the definition of the meta-axiom quadruple, condition (iv) calls GN a 'uniqueness group consisting of every scale transformation τ : N → N that leaves the empirical information content of f invariant.' For ordinal scales, the admissible transformations are all strictly increasing functions, which do not generally form a group (they are not all bijections and do not all have inverses within the set). The terminology is therefore inaccurate for ordinal scales; either restrict the meta-axiom to interval or ratio scales, or replace 'group' with a more general 'uniqueness class' or 'invariance structure.'
minor comments (8)
- [Section 1.3] Typo: 'reqiure' should be 'require' in the second bullet point.
- [Section 2.3] The sentence 'An examples includesadditive conjoint measurement' contains a typo and missing spacing; it should read 'An example includes additive conjoint measurement.'
- [Section 2.4] Typo: 'suuch' should be 'such' in the opening paragraph.
- [Section 2.5] Typo: 'approahces' should be 'approaches' in the introductory sentence.
- [Section 2.7] In the description of the Contextual/Emergent Layer, 'measureable' should be 'measurable.'
- [Section 2.4.1] The phrase 'underyling' appears twice and should be 'underlying.'
- [Section 3.2] Typo: 'offereing' should be 'offering.'
- [General] The paper repeatedly refers to forthcoming work for formalization; consider stating explicitly in the main text which parts are established here versus deferred, since the extended abstract relies heavily on that promised work.
Circularity Check
No circularity: the paper is a programmatic proposal with no derived predictions, no fitted parameters, and no self-citation chain carrying the central claim.
full rationale
The paper explicitly frames itself as an extended abstract that 'motivates and outlines a programme' for a measurement theory of AI, and it contains no derivation that reduces to its own inputs. There are no fitted parameters, no empirical predictions, and no equations that are asserted to follow from prior results by construction. The meta-axiom in Section 3.1 is presented as a proposed format for declaring measurement propositions (a quadruple consisting of empirical system, numerical system, homomorphism, and uniqueness group), not as a theorem or a derived consequence. Section 2.4.1's conditional statement that 'if we ascertain that this relation is transitive, consistent across repeated tests, and so forth, we might prove an interval or ordinal scale exists' is explicitly conditional, and Section 2.6 openly acknowledges that 'a single linear scale for capability might not exist. Partial or multi-dimensional orderings might be necessary.' Thus the central promise of commensurability is not presented as an already-derived result but as a goal contingent on unverified empirical axioms. The self-citations to the author's adjacent work ([33] for the AI stack and [16] for AI-agents infrastructure) are contextual references and are not load-bearing for the paper's central argument: the proposed MTAI framework does not depend on the correctness of either cited paper. The claim that stable latent attributes exist is explicitly labelled a 'hypothesis' in the Conclusion, not a conclusion derived from the framework. Under the standard that circularity requires quoting a specific reduction, fitted parameter renamed as prediction, or a load-bearing self-citation chain, no such step exists here.
Assumptions & free parameters
assumptions (5)
- domain assumption Measurement practice and methodology determines what we can say about AI systems, how we compare them, and which quantitative risk tools we can soundly invoke.
- domain assumption AI systems admit a measurable state space (Omega, F) with a sigma-algebra of distinguishable events.
- domain assumption Stable latent attributes of AI systems exist (methodological realism).
- domain assumption Constructs such as capability form an empirical relational structure satisfying RTM axioms (e.g., transitivity, monotonicity, completeness).
- standard math Stevens' nominal/ordinal/interval/ratio hierarchy is the right taxonomy for AI measurement.
invented entities (3)
-
AI observable, a measurable function phi: (Omega, F) to (R, B(R))
-
Five-layer AI measurement stack (physical, systems, algorithm/model, task/behaviour, contextual/emergent)
-
Meta-axiom quadruple (E, N, f, G_N) for measurement propositions
Cite this review
Pith. "Pith review of Towards Measurement Theory for Artificial Intelligence." pith.science (2026). https://pith.science/paper/OHNUBLQF
@misc{pith2026250705587,
author = {Pith},
title = {Pith review of: Towards Measurement Theory for Artificial Intelligence},
year = {2026},
howpublished = {\url{https://pith.science/paper/OHNUBLQF}},
note = {Machine review of arXiv:2507.05587}
}
read the original abstract
We motivate and outline a programme for a formal theory of measurement of artificial intelligence. We argue that formalising measurement for AI will allow researchers, practitioners, and regulators to: (i) make comparisons between systems and the evaluation methods applied to them; (ii) connect frontier AI evaluations with established quantitative risk analysis techniques drawn from engineering and safety science; and (iii) foreground how what counts as AI capability is contingent upon the measurement operations and scales we elect to use. We sketch a layered measurement stack, distinguish direct from indirect observables, and signpost how these ingredients provide a pathway toward a unified, calibratable taxonomy of AI phenomena.
Reference graph
Works this paper leans on
-
[1]
Eran Tal. Measurement in Science. In Edward N. Zalta, editor, The Stanford Encyclopedia of Philosophy . Metaphysics Research Lab, Stanford University, Fall 2020 edition, 2020
work page 2020
-
[2]
Aristotle. Categories. In Jonathan Barnes, editor, The Complete Works of Aristotle, Volume I. Princeton University Press, Princeton, 1984
work page 1984
-
[3]
The Thirteen Books of Euclid’s Elements
Euclid. The Thirteen Books of Euclid’s Elements. Cambridge University Press, Cambridge, 1908. Translated by T.L. Heath
work page 1908
-
[4]
Intension and Remission of Forms
Elzbieta Jung. Intension and Remission of Forms. In Henrik Lagerlund, editor, Encyclopedia of Medieval Philosophy, pages 551–555. Springer, Netherlands, 2011
work page 2011
-
[5]
Nicole Oresme and the medieval geometry of qualities and motions
Marshall Clagett. Nicole Oresme and the medieval geometry of qualities and motions. University of Wisconsin Press, Madison, 1968
work page 1968
-
[6]
Medieval quantifications of qualities: The ‘Merton School’
Edith Sylla. Medieval quantifications of qualities: The ‘Merton School’. Archive for history of exact sciences , 8(1):9–39, 1971
work page 1971
-
[7]
The foundations of modern science in the middle ages
Edward Grant. The foundations of modern science in the middle ages. Cambridge University Press, Cambridge, 1996
work page 1996
-
[8]
Immanuel Kant. Critique of Pure Reason. Cambridge University Press, Cambridge, 1787. Translated by Paul Guyer and Allen W. Wood, 1998
work page 1998
Show all 35 references
-
[9]
Jorgensen
Larry M. Jorgensen. The Principle of Continuity and Leibniz’s Theory of Consciousness. Journal of the History of Philosophy, 47(2):223–248, 2009
2009
-
[10]
Christopher E. Diehl. The Theory of Intensive Magnitudes in Leibniz and Kant . Ph.d. dissertation, Princeton University, 2012. Available online
2012
-
[11]
Number and measure: Hermann von Helmholtz at the crossroads of mathematics, physics, and psychology
Olivier Darrigol. Number and measure: Hermann von Helmholtz at the crossroads of mathematics, physics, and psychology. Studies in History and Philosophy of Science Part A, 34(3):515–573, 2003
2003
-
[12]
The origins of the representational theory of measurement: Helmholtz, Hölder, and Russell.Studies in History and Philosophy of Science (Part A), 24(2):185–206, 1993
Joel Michell. The origins of the representational theory of measurement: Helmholtz, Hölder, and Russell.Studies in History and Philosophy of Science (Part A), 24(2):185–206, 1993
1993
-
[13]
Scheming AIs: Will AIs fake alignment during training in order to get power? arXiv preprint arXiv:2311.08379, 2023
Joe Carlsmith. Scheming AIs: Will AIs fake alignment during training in order to get power? arXiv preprint arXiv:2311.08379, 2023
2023 arXiv
-
[14]
AI control: Improving safety despite intentional subversion
Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger. AI control: Improving safety despite intentional subversion. arXiv preprint arXiv:2312.06942 (Published in Proceedings of the 41st International Conference on Machine Learning), 2023
2023 arXiv
-
[15]
A sketch of an AI control safety case
Tomek Korbak, Joshua Clymer, Benjamin Hilton, Buck Shlegeris, and Geoffrey Irving. A sketch of an AI control safety case. arXiv preprint arXiv:2501.17315, 2025
2025 arXiv
-
[16]
Infrastructure for ai agents
Alan Chan, Kevin Wei, Sihao Huang, Nitarshan Rajkumar, Elija Perrier, Seth Lazar, Gillian K Hadfield, and Markus Anderljung. Infrastructure for ai agents. arXiv preprint arXiv:2501.10114, 2025
2025 arXiv
-
[17]
Krantz, R
David H. Krantz, R. Duncan Luce, Patrick Suppes, and Amos Tversky. Foundations of Measurement Vol 1: Additive and Polynomial Representations. Academic Press, San Diego and London, 1971. 13
1971
-
[18]
Krantz, R
Patrick Suppes, David H. Krantz, R. Duncan Luce, and Amos Tversky. Foundations of Measurement Vol 2: Geometrical, Threshold and Probabilistic Representations. Academic Press, San Diego and London, 1989
1989
-
[19]
An introduction to measure theory, volume 126
Terence Tao. An introduction to measure theory, volume 126. American Mathematical Soc., 2011
2011
-
[20]
International Vocabulary of Metrology—Basic and general concepts and associated terms (VIM)
Joint Committee for Guides in Metrology (JCGM). International Vocabulary of Metrology—Basic and general concepts and associated terms (VIM). JCGM, Sèvres, 3rd edition, 2012. Available online
2012
-
[21]
Nunnally and Ira H
Jum C. Nunnally and Ira H. Bernstein. Psychometric Theory. McGraw-Hill, New York, 3rd edition, 1994
1994
-
[22]
Measuring the Mind: Conceptual Issues in Contemporary Psychometrics
Denny Borsboom. Measuring the Mind: Conceptual Issues in Contemporary Psychometrics . Cambridge Uni- versity Press, Cambridge, 2005
2005
-
[23]
Campbell
Norman R. Campbell. Physics: the Elements. Cambridge University Press, London, 1920
1920
-
[24]
Duncan Luce and John W
R. Duncan Luce and John W. Tukey. Simultaneous conjoint measurement: A new type of fundamental measure- ment. Journal of Mathematical Psychology, 1(1):1–27, 1964
1964
-
[25]
What is an ontology? In Steffen Staab and Rudi Studer, editors, Handbook on Ontologies, page 1–17
Nicola Guarino, Daniel Oberle, and Steffen Staab. What is an ontology? In Steffen Staab and Rudi Studer, editors, Handbook on Ontologies, page 1–17. Springer, Berlin, Heidelberg, 2009
2009
-
[26]
Foundations of measurement, Vol
David Krantz, Duncan Luce, Patrick Suppes, and Amos Tversky. Foundations of measurement, Vol. I: Additive and polynomial representations. Dover Publications, 1971
1971
-
[27]
Foundations of measurement: representation, axiom- atization, and invariance, volume 3
Robert Duncan Luce, Patrick Suppes, and David H Krantz. Foundations of measurement: representation, axiom- atization, and invariance, volume 3. Courier Corporation, 2007
2007
-
[28]
S. S. Stevens. On the theory of scales of measurement. Science, 103:677–680, 1946
1946
-
[29]
Outline of a general model of measurement
Alberto Frigerio, Andrea Giordani, and Luca Mari. Outline of a general model of measurement. Synthese, 175(2):123–149, 2010
2010
-
[30]
Science outside the laboratory: Measurement in field science and economics
Marcel Boumans. Science outside the laboratory: Measurement in field science and economics. Oxford Univer- sity Press, 2015
2015
-
[31]
International Vocabulary of Metrology—Basic and general concepts and associated terms (VIM)
Joint Committee for Guides in Metrology (JCGM). International Vocabulary of Metrology—Basic and general concepts and associated terms (VIM). JCGM, Sèvres, 3rd edition, 2012
2012
-
[32]
Centaur: a foundation model of human cognition
Marcel Binz, Elif Akata, Matthias Bethge, Franziska Brändle, Fred Callaway, Julian Coda-Forno, Peter Dayan, Can Demircan, Maria K Eckstein, Noémi Éltet ˝o, et al. Centaur: a foundation model of human cognition. arXiv preprint arXiv:2410.20268, 2024
-
[33]
Out of control–why alignment needs formal control theory (and an alignment control stack)
Elija Perrier. Out of control–why alignment needs formal control theory (and an alignment control stack). arXiv preprint arXiv:2506.17846, 2025
2025 arXiv
-
[34]
The Metaphysics of Measurement
Chris Swoyer. The Metaphysics of Measurement. In John Forge, editor, Measurement, Realism and Objectivity, pages 235–290. Reidel, Dordrecht, 1987
1987
-
[35]
Bridgman
Percy W. Bridgman. The Logic of Modern Physics. Macmillan, New York, 1927. 14
1927
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.