Pith. sign in

REVIEW 3 major objections 8 minor 35 references

Towards Measurement Theory for Artificial Intelligence

T0 review · 3 major / 8 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper argues that a formal measurement theory for AI—synthesizing representational measurement, metrology, and measure theory—would make AI evaluation commensurable and scientifically grounded.

desk verdict A well-motivated programmatic proposal for measurement theory in AI, but its commensurability promise rests on an unproved RTM precondition that needs a worked example. read the letter →

arxiv 2507.05587 v1 pith:OHNUBLQF submitted 2025-07-08 cs.AI

classification cs.AI
keywords measurementtheoryAIevaluationmetrologyrepresentationalofmeasurescaletypesobservablesbenchmarkcomparability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that current AI evaluation is a collection of incommensurable practices: benchmarks and metrics proliferate without a shared theoretical foundation, so results cannot be systematically compared across models, tasks, or research groups. It proposes a programme for a measurement theory of AI (MTAI) built from three pillars: representational theory of measurement, metrology, and measure theory, with psychometrics for latent constructs. If this programme succeeds, researchers and regulators would gain AI observables with defined scale types and invariance conditions, enabling cumulative science, standardized regulation, and sound use of quantitative risk-analysis tools. The paper is explicit that this is a motivation and outline, not a completed theory, and it identifies key obstacles such as non-transitive orderings and context dependence that could block the promise.

What carries the argument

The load-bearing machinery is the 'AI observable', defined as a measurable function $\varphi : (\Omega,\mathcal{F}) \to (\mathbb{R},\mathcal{B}(\mathbb{R}))$, where $\Omega$ is the space of AI system states and $\mathcal{F}$ is a $\sigma$-algebra of distinguishable events. This formal definition is coupled with a layered measurement stack spanning physical, systems, algorithm, task/behaviour, and contextual/emergent layers, each with its own observables, instrumentation requirements, and validity conditions. The argument also relies on representational measurement's homomorphism from an empirical relational structure to a numerical structure, and on a proposed meta-axiom requiring every measurement proposition to declare its empirical system, numerical system, representation function, and uniqueness group of scale transformations.

What would settle it

A demonstration that pairwise 'at least as capable as' judgments across a broad task suite are systematically intransitive (A beats B, B beats C, but C beats A), or that no stable ordering survives context shifts, would falsify the central promise of commensurable AI measurement.

Watch

Extended reading notes

Core claim

The paper's central claim is that a principled measurement theory for artificial intelligence would (and ought to) enable commensurable evaluations across models, tasks, and research groups, and would connect frontier AI evaluations with established quantitative risk-analysis techniques. The discovery is a diagnosis: AI evaluation currently lacks the scope, rigour, and depth of measurement practice in other sciences, and the remedy is a formal framework that defines AI observables, assigns scale types, and makes the choice of measurement operations explicit. The paper adopts a position of methodological realism, hypothesizing that stable latent attributes of AI systems exist and can be validated through consistent, coherent, predictive measurement models.

Load-bearing premise

The framework rests on the premise that AI attributes such as capability or reliability can be compared consistently enough—if A beats B and B beats C, then A beats C—for a meaningful numerical scale to exist.

Editorial extensions

If this is right

  • If adopted, MTAI would make evaluation results commensurable across models, tasks, and research groups by assigning each AI attribute a scale type and an invariance class.
  • It would allow frontier AI evaluations to plug into established quantitative risk-analysis and reliability-engineering techniques, which require variables with known scale properties.
  • It would make the definition of AI capability explicitly contingent on the measurement operations and scales chosen, surfacing choices that current benchmarking hides.
  • It would give regulators and auditors standardised, calibratable measurement protocols for compliance, reliability, and risk levels.
  • It would expose ill-defined constructs: a proposed measure that cannot be cast as a measurable function on the system state space lacks mathematical foundation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: if this programme succeeds, benchmark leaderboards could be replaced or supplemented by explicit scale-type declarations, making it impossible to report a 'capability score' without stating its invariance class.
  • Editorial extension: the framework implies a concrete research agenda: test whether pairwise 'at least as capable as' judgments on real model pairs are transitive across diverse task suites; current benchmarks do not establish this.
  • Editorial extension: the AI-observable definition suggests a testable criterion for whether a proposed metric is well-founded—if it cannot be expressed as a measurable function with a stated uniqueness group, it is not yet an AI observable.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 8 minor

Summary. The paper argues for the development of a formal measurement theory for artificial intelligence (MTAI), synthesizing representational theory of measurement (RTM), measure theory, metrology, and psychometrics. It motivates the program by citing the fragmentation of current AI evaluation practice, the need for commensurable comparisons, standardization for risk analysis, and the contingency of AI capability on measurement choices. The paper sketches key components: a five-layer AI measurement stack, the distinction between direct and indirect observables, a formal definition of an AI observable as a measurable function, a meta-axiom quadruple for measurement propositions, and initial examples of how the framework might apply. The authors explicitly frame the work as an extended abstract that outlines and motivates a programme rather than a fully formalized theory.

Significance. If the proposed MTAI program is carried through, it could provide a much-needed metrological and psychometric foundation for AI evaluation, enabling cumulative comparison across models, tasks, and research groups and connecting AI risk assessment with established quantitative methods. The paper is unusually self-aware: it names its own obstacles (non-transitive orderings, context-dependence, high-dimensional evolution, black-box systems) and honestly distinguishes between motivation and achievement. Its historical synthesis of measurement theory is informative, and the layered measurement stack offers a useful organizing framework. At this stage, the value lies primarily in framing and agenda-setting rather than in a demonstrated theory; the central promises are conditional on empirical and axiomatic premises that are not yet established.

major comments (3)
  1. [Sections 2.4.1, 2.6, 3.1] The paper's central claim that MTAI 'would (and ought to) enable commensurable evaluations across models, tasks, and research groups' rests on the existence of an empirical relational structure over AI constructs that satisfies the RTM axioms. Section 2.4.1 explicitly conditions the construction of an ordinal or interval scale on properties such as transitivity, and Section 3.1's meta-axiom requires a homomorphism f : E → N that preserves all empirical relations. However, Section 2.6 (item 2) acknowledges that 'a single linear scale for capability might not exist' and that partial or multi-dimensional orderings may be necessary. For a natural Pareto-dominance ordering over tasks, no homomorphism into (R, ≥) can preserve the partial order. The paper never provides a worked example of an AI construct satisfying the required axioms, nor does it state conditions under which the representation theorem would apply. Without such an example or a clear treatment of partial orderings, the promised commensurability is not a consequence of the framework. Please supply at least one concrete worked example and show how the framework accommodates partial or intransitive orderings, or explicitly restrict the domain of applicability.
  2. [Section 2.5] The formal condition 's1 ⪰ s2 =⇒ φT (s1) ≤ φT (s2)' is reversed relative to the intended meaning of ⪰ as 'is at least as capable as.' If s1 is at least as capable as s2, the numerical representation should be non-decreasing, not non-increasing. This also contradicts Section 3.1, where 'a ⪯E b ⇐⇒ f(a) ≥ f(b)'. The inconsistency affects the formal sketch of representation functions and should be corrected.
  3. [Section 3.1] In the definition of the meta-axiom quadruple, condition (iv) calls GN a 'uniqueness group consisting of every scale transformation τ : N → N that leaves the empirical information content of f invariant.' For ordinal scales, the admissible transformations are all strictly increasing functions, which do not generally form a group (they are not all bijections and do not all have inverses within the set). The terminology is therefore inaccurate for ordinal scales; either restrict the meta-axiom to interval or ratio scales, or replace 'group' with a more general 'uniqueness class' or 'invariance structure.'
minor comments (8)
  1. [Section 1.3] Typo: 'reqiure' should be 'require' in the second bullet point.
  2. [Section 2.3] The sentence 'An examples includesadditive conjoint measurement' contains a typo and missing spacing; it should read 'An example includes additive conjoint measurement.'
  3. [Section 2.4] Typo: 'suuch' should be 'such' in the opening paragraph.
  4. [Section 2.5] Typo: 'approahces' should be 'approaches' in the introductory sentence.
  5. [Section 2.7] In the description of the Contextual/Emergent Layer, 'measureable' should be 'measurable.'
  6. [Section 2.4.1] The phrase 'underyling' appears twice and should be 'underlying.'
  7. [Section 3.2] Typo: 'offereing' should be 'offering.'
  8. [General] The paper repeatedly refers to forthcoming work for formalization; consider stating explicitly in the main text which parts are established here versus deferred, since the extended abstract relies heavily on that promised work.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is a programmatic proposal with no derived predictions, no fitted parameters, and no self-citation chain carrying the central claim.

full rationale

The paper explicitly frames itself as an extended abstract that 'motivates and outlines a programme' for a measurement theory of AI, and it contains no derivation that reduces to its own inputs. There are no fitted parameters, no empirical predictions, and no equations that are asserted to follow from prior results by construction. The meta-axiom in Section 3.1 is presented as a proposed format for declaring measurement propositions (a quadruple consisting of empirical system, numerical system, homomorphism, and uniqueness group), not as a theorem or a derived consequence. Section 2.4.1's conditional statement that 'if we ascertain that this relation is transitive, consistent across repeated tests, and so forth, we might prove an interval or ordinal scale exists' is explicitly conditional, and Section 2.6 openly acknowledges that 'a single linear scale for capability might not exist. Partial or multi-dimensional orderings might be necessary.' Thus the central promise of commensurability is not presented as an already-derived result but as a goal contingent on unverified empirical axioms. The self-citations to the author's adjacent work ([33] for the AI stack and [16] for AI-agents infrastructure) are contextual references and are not load-bearing for the paper's central argument: the proposed MTAI framework does not depend on the correctness of either cited paper. The claim that stable latent attributes exist is explicitly labelled a 'hypothesis' in the Conclusion, not a conclusion derived from the framework. Under the standard that circularity requires quoting a specific reduction, fitted parameter renamed as prediction, or a load-bearing self-citation chain, no such step exists here.

Assumptions & free parameters 0 free parameters · 5 assumptions · 3 invented entities

The ledger is small because the paper is a proposal rather than a derivation. There are zero free parameters: nothing is fitted, and no numeric constants are introduced beyond standard measurement-theoretic objects. The assumptions are domain claims about AI: that measurement practice fixes what can be asserted about AI systems, that AI states form a measurable space with a sigma-algebra, that stable latent attributes exist (methodological realism), that capability-like constructs satisfy RTM axioms, and that Stevens' scale taxonomy applies. The invented entities are definitional proposals, the AI observable, the five-layer measurement stack, and the meta-axiom quadruple, none of which carries an independent empirical handle. The smallness of the ledger is honest, but it also means the paper asserts almost nothing concrete that could be tested.

assumptions (5)
  • domain assumption Measurement practice and methodology determines what we can say about AI systems, how we compare them, and which quantitative risk tools we can soundly invoke.
    Stated in Section 1 as 'Our central assumption'. It is the premise of the whole programme; the paper flags it as an assumption but does not argue for it or test it.
  • domain assumption AI systems admit a measurable state space (Omega, F) with a sigma-algebra of distinguishable events.
    Invoked in Section 3.2 to define AI observables. The paper calls it 'an important assumption about the nature of AI systems and their measurable properties', but gives no construction for real frontier systems, which are proprietary, evolving, and partially black-box.
  • domain assumption Stable latent attributes of AI systems exist (methodological realism).
    Section 4: 'We hypothesise that stable, latent attributes of AI systems exist'. This realism commitment justifies latent-trait models such as IRT; it is acknowledged as a hypothesis and is load-bearing for the psychometric pillar.
  • domain assumption Constructs such as capability form an empirical relational structure satisfying RTM axioms (e.g., transitivity, monotonicity, completeness).
    Required for the meta-axiom proposal (Section 3.1) and for ordinal or interval scale claims (Section 2.4.1). The paper itself lists 'Non-Transitive or Incomplete Orderings' as an open challenge (Section 2.6), so this premise is load-bearing and unverified.
  • standard math Stevens' nominal/ordinal/interval/ratio hierarchy is the right taxonomy for AI measurement.
    Table 1 adopts Stevens' scale-type hierarchy without caveat. It is the textbook classification within measurement theory, though it is contested in the psychometric literature the paper otherwise draws on.
invented entities (3)
  • AI observable, a measurable function phi: (Omega, F) to (R, B(R))
    purpose: Proposed formal object for AI measurement, intended to make constructs like capability measurable and to expose ill-defined proposals that cannot be expressed as measurable functions.
    Defined in Section 3.2; it is a standard measurability condition applied to an AI state space, a definitional proposal with no falsifiable handle outside the paper.
  • Five-layer AI measurement stack (physical, systems, algorithm/model, task/behaviour, contextual/emergent)
    purpose: Modular decomposition of AI systems so layer-specific measurement paradigms, from metrology to psychometrics, can be applied and composed into an end-to-end framework.
    Proposed in Section 2.7 as 'by no means complete or mutually exclusive'; a taxonomy proposal with no independent empirical content.
  • Meta-axiom quadruple (E, N, f, G_N) for measurement propositions
    purpose: A reporting requirement that every measurement proposition declare its empirical system, numerical system, representation homomorphism, and uniqueness group.
    Section 3.1; this is the standard RTM representation setup restated as a requirement on AI evaluation, with no evidence that real AI evaluations can satisfy it.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Measurement Theory for Artificial Intelligence." pith.science (2026). https://pith.science/paper/OHNUBLQF

@misc{pith2026250705587,
  author       = {Pith},
  title        = {Pith review of: Towards Measurement Theory for Artificial Intelligence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OHNUBLQF}},
  note         = {Machine review of arXiv:2507.05587}
}
read the original abstract

We motivate and outline a programme for a formal theory of measurement of artificial intelligence. We argue that formalising measurement for AI will allow researchers, practitioners, and regulators to: (i) make comparisons between systems and the evaluation methods applied to them; (ii) connect frontier AI evaluations with established quantitative risk analysis techniques drawn from engineering and safety science; and (iii) foreground how what counts as AI capability is contingent upon the measurement operations and scales we elect to use. We sketch a layered measurement stack, distinguish direct from indirect observables, and signpost how these ingredients provide a pathway toward a unified, calibratable taxonomy of AI phenomena.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

35 extracted references · 30 canonical work pages

  1. [1]

    Measurement in Science

    Eran Tal. Measurement in Science. In Edward N. Zalta, editor, The Stanford Encyclopedia of Philosophy . Metaphysics Research Lab, Stanford University, Fall 2020 edition, 2020

  2. [2]

    Categories

    Aristotle. Categories. In Jonathan Barnes, editor, The Complete Works of Aristotle, Volume I. Princeton University Press, Princeton, 1984

  3. [3]

    The Thirteen Books of Euclid’s Elements

    Euclid. The Thirteen Books of Euclid’s Elements. Cambridge University Press, Cambridge, 1908. Translated by T.L. Heath

  4. [4]

    Intension and Remission of Forms

    Elzbieta Jung. Intension and Remission of Forms. In Henrik Lagerlund, editor, Encyclopedia of Medieval Philosophy, pages 551–555. Springer, Netherlands, 2011

  5. [5]

    Nicole Oresme and the medieval geometry of qualities and motions

    Marshall Clagett. Nicole Oresme and the medieval geometry of qualities and motions. University of Wisconsin Press, Madison, 1968

  6. [6]

    Medieval quantifications of qualities: The ‘Merton School’

    Edith Sylla. Medieval quantifications of qualities: The ‘Merton School’. Archive for history of exact sciences , 8(1):9–39, 1971

  7. [7]

    The foundations of modern science in the middle ages

    Edward Grant. The foundations of modern science in the middle ages. Cambridge University Press, Cambridge, 1996

  8. [8]

    Critique of Pure Reason

    Immanuel Kant. Critique of Pure Reason. Cambridge University Press, Cambridge, 1787. Translated by Paul Guyer and Allen W. Wood, 1998

Show all 35 references
  1. [9]

    Jorgensen

    Larry M. Jorgensen. The Principle of Continuity and Leibniz’s Theory of Consciousness. Journal of the History of Philosophy, 47(2):223–248, 2009

  2. [10]

    Christopher E. Diehl. The Theory of Intensive Magnitudes in Leibniz and Kant . Ph.d. dissertation, Princeton University, 2012. Available online

  3. [11]

    Number and measure: Hermann von Helmholtz at the crossroads of mathematics, physics, and psychology

    Olivier Darrigol. Number and measure: Hermann von Helmholtz at the crossroads of mathematics, physics, and psychology. Studies in History and Philosophy of Science Part A, 34(3):515–573, 2003

  4. [12]

    The origins of the representational theory of measurement: Helmholtz, Hölder, and Russell.Studies in History and Philosophy of Science (Part A), 24(2):185–206, 1993

    Joel Michell. The origins of the representational theory of measurement: Helmholtz, Hölder, and Russell.Studies in History and Philosophy of Science (Part A), 24(2):185–206, 1993

  5. [13]

    Scheming AIs: Will AIs fake alignment during training in order to get power? arXiv preprint arXiv:2311.08379, 2023

    Joe Carlsmith. Scheming AIs: Will AIs fake alignment during training in order to get power? arXiv preprint arXiv:2311.08379, 2023

  6. [14]

    AI control: Improving safety despite intentional subversion

    Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger. AI control: Improving safety despite intentional subversion. arXiv preprint arXiv:2312.06942 (Published in Proceedings of the 41st International Conference on Machine Learning), 2023

  7. [15]

    A sketch of an AI control safety case

    Tomek Korbak, Joshua Clymer, Benjamin Hilton, Buck Shlegeris, and Geoffrey Irving. A sketch of an AI control safety case. arXiv preprint arXiv:2501.17315, 2025

  8. [16]

    Infrastructure for ai agents

    Alan Chan, Kevin Wei, Sihao Huang, Nitarshan Rajkumar, Elija Perrier, Seth Lazar, Gillian K Hadfield, and Markus Anderljung. Infrastructure for ai agents. arXiv preprint arXiv:2501.10114, 2025

  9. [17]

    Krantz, R

    David H. Krantz, R. Duncan Luce, Patrick Suppes, and Amos Tversky. Foundations of Measurement Vol 1: Additive and Polynomial Representations. Academic Press, San Diego and London, 1971. 13

  10. [18]

    Krantz, R

    Patrick Suppes, David H. Krantz, R. Duncan Luce, and Amos Tversky. Foundations of Measurement Vol 2: Geometrical, Threshold and Probabilistic Representations. Academic Press, San Diego and London, 1989

  11. [19]

    An introduction to measure theory, volume 126

    Terence Tao. An introduction to measure theory, volume 126. American Mathematical Soc., 2011

  12. [20]

    International Vocabulary of Metrology—Basic and general concepts and associated terms (VIM)

    Joint Committee for Guides in Metrology (JCGM). International Vocabulary of Metrology—Basic and general concepts and associated terms (VIM). JCGM, Sèvres, 3rd edition, 2012. Available online

  13. [21]

    Nunnally and Ira H

    Jum C. Nunnally and Ira H. Bernstein. Psychometric Theory. McGraw-Hill, New York, 3rd edition, 1994

  14. [22]

    Measuring the Mind: Conceptual Issues in Contemporary Psychometrics

    Denny Borsboom. Measuring the Mind: Conceptual Issues in Contemporary Psychometrics . Cambridge Uni- versity Press, Cambridge, 2005

  15. [23]

    Campbell

    Norman R. Campbell. Physics: the Elements. Cambridge University Press, London, 1920

  16. [24]

    Duncan Luce and John W

    R. Duncan Luce and John W. Tukey. Simultaneous conjoint measurement: A new type of fundamental measure- ment. Journal of Mathematical Psychology, 1(1):1–27, 1964

  17. [25]

    What is an ontology? In Steffen Staab and Rudi Studer, editors, Handbook on Ontologies, page 1–17

    Nicola Guarino, Daniel Oberle, and Steffen Staab. What is an ontology? In Steffen Staab and Rudi Studer, editors, Handbook on Ontologies, page 1–17. Springer, Berlin, Heidelberg, 2009

  18. [26]

    Foundations of measurement, Vol

    David Krantz, Duncan Luce, Patrick Suppes, and Amos Tversky. Foundations of measurement, Vol. I: Additive and polynomial representations. Dover Publications, 1971

  19. [27]

    Foundations of measurement: representation, axiom- atization, and invariance, volume 3

    Robert Duncan Luce, Patrick Suppes, and David H Krantz. Foundations of measurement: representation, axiom- atization, and invariance, volume 3. Courier Corporation, 2007

  20. [28]

    S. S. Stevens. On the theory of scales of measurement. Science, 103:677–680, 1946

  21. [29]

    Outline of a general model of measurement

    Alberto Frigerio, Andrea Giordani, and Luca Mari. Outline of a general model of measurement. Synthese, 175(2):123–149, 2010

  22. [30]

    Science outside the laboratory: Measurement in field science and economics

    Marcel Boumans. Science outside the laboratory: Measurement in field science and economics. Oxford Univer- sity Press, 2015

  23. [31]

    International Vocabulary of Metrology—Basic and general concepts and associated terms (VIM)

    Joint Committee for Guides in Metrology (JCGM). International Vocabulary of Metrology—Basic and general concepts and associated terms (VIM). JCGM, Sèvres, 3rd edition, 2012

  24. [32]

    Centaur: a foundation model of human cognition

    Marcel Binz, Elif Akata, Matthias Bethge, Franziska Brändle, Fred Callaway, Julian Coda-Forno, Peter Dayan, Can Demircan, Maria K Eckstein, Noémi Éltet ˝o, et al. Centaur: a foundation model of human cognition. arXiv preprint arXiv:2410.20268, 2024

  25. [33]

    Out of control–why alignment needs formal control theory (and an alignment control stack)

    Elija Perrier. Out of control–why alignment needs formal control theory (and an alignment control stack). arXiv preprint arXiv:2506.17846, 2025

  26. [34]

    The Metaphysics of Measurement

    Chris Swoyer. The Metaphysics of Measurement. In John Forge, editor, Measurement, Realism and Objectivity, pages 235–290. Reidel, Dordrecht, 1987

  27. [35]

    Bridgman

    Percy W. Bridgman. The Logic of Modern Physics. Macmillan, New York, 1927. 14

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.