REVIEW 3 major objections 4 minor 39 references
The geometry of moral decision making
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Moral and legal decision making, the paper claims, is one variational problem whose two terms are utilitarianism, $E_p[U]$, and deontology, the penalty $\frac{1}{\beta} R(p,q)$.
desk verdict A clean recapitulation of bounded rationality dressed as a moral theory, with a legal example that has a sign error making the deontic term a reward rather than a penalty. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the variational objective of Eq. (78), $\max_p \left( E_p[U] - \frac{1}{\beta} R(p,q) \right)$, in which $q$ is a prior over actions, $R$ is a divergence-based regularizer, and $\beta$ is the inverse temperature. That objective carries the argument because its support condition makes the prior the encoding of the legal or moral code — only actions in the support of $q$ are eligible — and the divergence term makes rule-following a graded pull rather than a hard constraint. Solving it yields the Boltzmann-Gibbs distribution, an exponential family on the probability simplex, and varying $\beta$ moves the solution along the $e$-geodesic through $q$. The constraint-robust variant replaces the fixed prior with a source distribution and organizes the same trade-off through the rate-utility function, whose slope is $1/\beta$ and whose tangency points with constant-mutual-information surfaces define the utility expansion path.
What would settle it
The framework predicts that, for fixed utility assignments, observed choices follow the Boltzmann-Gibbs form $p(i) \propto q_i e^{\beta u_i}$ for some prior $q$ and weight $\beta \geq 0$. That prediction is falsifiable by behavioral data: fit $\beta$ and $q$ to the response frequencies of a large set of moral-dilemma decisions with known utilities, and the claim fails if the residuals are systematic — for example, if choices persistently violate the ratio property of Luce's choice axiom, or if no single fitted pair $(q, \beta)$ reproduces the frequencies across variations of the dilemma. In the legal domain, the applied model predicts discontinuous switches between permitted restrictions as $\beta$ crosses thresholds; observing smooth, history-dependent restriction patterns that no piecewise-constant $\beta$ can reproduce would falsify the application.
Extended reading notes
Core claim
The central claim is that a resource-bounded decision maker resolves the conflict between deontology and utilitarianism inside a single variational objective, Eq. (78): the optimal policy maximizes expected utility $E_p[U]$ minus the scaled penalty $\frac{1}{\beta} R(p,q)$, where the expected-utility term is the utilitarian component and the regularizer $R$, anchored to a prior $q$ whose support is the set of permitted actions, is the deontological component. The solution is the Boltzmann-Gibbs distribution, $p^*_\beta(i) \propto q_i e^{\beta u_i}$, an exponential-family weighting of actions by utility that interpolates between pure rule-following at $\beta \to 0$ and unrestricted utility maximization at $\beta \to \infty$, tracing an $e$-geodesic through the prior as the weight is swept. In the constraint-robust version with a source distribution over world states, the same trade-off is organized by a rate-utility function whose slope at every point is $1/\beta$, and the optimal kernels solve self-consistent equations of rate-distortion type. The author's stated point is that neither moral theory determines the coupling constant: it remains a free parameter for the legislative or judicial authority, and that free parameter is the formal locus of the margin of discretion.
Load-bearing premise
The argument stands on the interpretive identification of rule-following (deontology) with a divergence penalty against a prior over permitted actions — a modeling choice that the paper asserts rather than derives from moral theory, legal texts, or data — together with the external supply of the penalty weight $\beta$.
Editorial extensions
If this is right
- If the framework is right, every moral choice problem is specified by the prior $q$, the regularizer $R$, and the coupling $\beta$; no separate moral theory is needed beyond these ingredients.
- The optimal moral policy is never a pure rule or a pure maximizer in the interior regime: it is the Boltzmann-Gibbs distribution, and the two classical theories are recovered only in the limits $\beta \to 0$ and $\beta \to \infty$.
- In the legal application, a restriction of a fundamental right is justified exactly when the public utility gain exceeds the disutility of the restriction, and the model predicts that the chosen restriction switches discontinuously as the authority's weight $\beta$ crosses critical thresholds.
- Because the constraint-robust problem is formally a rate-distortion problem, coarse-graining the space of states is governed by the data-processing inequality: abstraction cannot increase the information a decision can carry, so hierarchical and simplified moral reasoning follows from the same objective.
- Autonomous agents implementing Eq. (78) inherit a tunable deontology: their rule-following behavior is controlled by a single externally supplied constant rather than by a hard-coded rule list.
Reading between the lines
- Editorial extension: with $\beta$ treated as a quantity to be fitted rather than legislated, the framework becomes behaviorally testable — one could estimate $\beta$ and $q$ from observed choice frequencies in moral dilemmas and ask whether a single pair transfers across situations, a check the paper does not perform.
- Editorial extension: because the constraint-robust problem is formally a rate-distortion problem, moral or legal vagueness can be read as a compression phenomenon — coarse representations of a situation cost less information, so an 'optimal vagueness' would trade decision accuracy against coding cost, a consequence the paper leaves implicit.
- Editorial consequence of the mapping: if deontology is a regularizer, then disagreements between rule-based and consequentialist moral theories are disagreements about the support of the prior and the value of the coupling constant — a philosophical dispute relocated onto two numbers that the paper hands to the legislative authority.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript proposes that bounded-rational moral and legal decision making is captured by a variational objective in which expected utility plays the role of utilitarianism and a divergence/regularization term plays the role of deontology. After introducing Markov kernels, co-sheaves, and information-geometric tools, the paper derives the Boltzmann-Gibbs solution for the multiplier-robust control problem, studies a rate-utility version of the constraint robust-control problem, and applies the framework to a model of constitutional-right restrictions. The central claim is Eq. (78): the optimal policy maximizes E_p[U] minus (1/beta)R(p,q), where R is interpreted as a deontic regularizer and beta is a free coupling constant to be fixed by a legislative or judicial authority.
Significance. If the moral identification in Eq. (78) were established, the paper would offer a compact mathematical bridge between information-theoretic bounded rationality, moral theory, and legal doctrine. The variational derivations in Section 5 are standard and mostly correct: the KKT conditions yield the Gibbs distribution, and the convexity/concavity statements about the rate-utility function in Section 6 follow familiar arguments. The paper is also honest that beta is not determined internally. However, the central 'deontology as regularization' claim is not derived but assumed by labelling R as deontic, the legal model in Section 7 contains a sign inconsistency, and the paper provides no falsifiable prediction or independent constraint on q, R, or beta. The mathematical framework is therefore not yet the moral theory the title promises.
major comments (3)
- [§7, Eq. (77) and the definition of D(d)] The legal application has a sign inconsistency with the main objective. With D(d(p,q)) = d_max - D_KL(p||q), the objective in (77) becomes E_p[U] - (1/beta)(d_max - D_KL(p||q)) = E_p[U] - d_max/beta + (1/beta)D_KL(p||q). Since the constant term does not affect the argmax, the optimization is equivalent to maximizing E_p[U] + (1/beta)D_KL(p||q), so the KL term rewards departure from the prior instead of penalizing it. This contradicts the sign of the deontic term in Eqs. (25) and (78), where R is subtracted as a cut-off protecting the support of q. In addition, D(d) = d_max - d is affine rather than strictly convex, contrary to the paper's own condition that the disutility function be a strictly decreasing convex function of d. As written, Section 7 does not instantiate 'deontology as regularization'; it instantiates the opposite of Eq. (78).
- [§7 and Eq. (78)] The identification of the divergence penalty with deontology is an interpretive assumption rather than a derived result. The label 'deontic' is attached to R(p,q) in Eq. (78), and the support of q is described as a deontological cut-off, but no argument from deontological ethics, legal texts, or behavioral data establishes that a moral system's content is represented by a prior q and a regularizer R. Section 7 selects the disutility function from three ad-hoc 'basic types (not to scale)' without doctrinal derivation. The paper's own statement that beta 'will remain a free parameter to be determined by the competent legislative or judicial authority' concedes that the framework has no independent predictive content for the central mapping. Unless the mapping is constrained by a separate theory or by testable predictions, Eq. (78) is a relabeling of a known bounded-rationality optimization problem, not a discovery about deontology.
- [§6, Eqs. (47)–(53)] The transition from the stated optimization problem (47) to the minimization in (48) and (52) is not an equivalence. Problem (47) maximizes over the prior kernel κ, while (48) fixes K and minimizes D_KL(P⋊K || P⋊κ) over κ; these are different variational problems. For a fixed K, the minimizer κ* = K, or q* = K_*P in the constant-kernel case, does not in general maximize the free-energy expression in (47), which is driven toward priors concentrated on high-utility actions. Consequently the rate-utility problem (53) and the subsequent concavity analysis are not consequences of the stated starting point. This does not invalidate the Section 5 derivation of the Gibbs solution, but it undermines the claimed generalization in Section 6 unless an additional equivalence argument is supplied.
minor comments (4)
- [§1, Eq. (1)] The displayed objective in Eq. (1) uses a minimization with a plus sign in front of (1/beta)D_KL(p||q), whereas Eqs. (25) and (78) use maximization with a minus sign; this inconsistency should be corrected.
- [§1, after Eq. (1)] The text says 'The expression is referred to as multiplier robust-control problem []' with an empty citation; the reference should be filled in.
- [§6.3, Definitions 6.1–6.2] The existence and uniqueness of the utility expansion path satisfying (74) and the contraction path satisfying (76) are not proved; the claimed disjointness and reflection symmetry of the two paths should be stated as a proposition with explicit hypotheses.
- [Figure 9 caption] The caption says the z-axis shows F_beta[p] as a function of temperature 1/beta, but also lists a prior q = (0.7,0.2,0.1) in Delta_2 and a utility vector U = (7,5); the dimensional mismatch and the precise role of the displayed curve should be clarified.
Circularity Check
The central “deontology is a regularizer” claim is stipulated by labeling R in Eq. (78) as “deontic”; the Section 7 legal example then fails to instantiate even that sign.
-
self definitional
[Abstract; Section 7, Eq. (78) and surrounding summary text]
"In essence, it can be succinctly summarised by analysing the components of the main object of our investigation, which is the following type of formula max_{p∈Δ(supp(q))} { Ep[U ]|utilitarian − 1/β R(p,q)|deontic } , (78) where R is a regularisation function and q is a Bayesian prior. … In contrast, R, with the help of the prior q and its support, serves as a cut-off and thus incorporates deontological considerations for the protection of individual rights, which expected utility cannot provide."
The identification of deontology with regularization is not derived; it is inserted as the label “deontic” attached to R in the very formula presented as the paper’s conclusion. Eq. (25) already defines the same bounded-rationality objective max_p E_p[U] − (1/β)D_KL(p||q) using standard results, and the only new step at the summary is to call R the deontic term. No theorem, legal source, or data fixes that mapping, and the paper says the coupling constant is “ultimately up to a third party, such as the legislator”. The claimed finding is therefore equivalent to its own definition: R is named “deontic”, and the paper then reads “deontology is a regularizer” back out of Eq. (78).
full rationale
Sections 3–6 are a largely self-contained exposition of standard bounded-rationality and rate-distortion mathematics: the Gibbs/Boltzmann solution, the Legendre/geodesic discussion, and the rate-utility function are derived from stated optimization problems, and where results are cited (Mattsson–Weibull, Ortega–Braun, Genewein et al., Berger) they are used as ordinary background, not as a uniqueness theorem that forces the moral interpretation. The circularity lies in the philosophical wrapper. The abstract’s “we interpret deontology as a regularisation function” and Eq. (78)’s explicit labels make the central claim true by stipulation: R is called deontic, and the summary then presents the labeled objective as the paper’s insight. Because β is explicitly left to an external authority and no moral/legal dataset is used, there is no independent check that would let the labeling fail. The one concrete legal instantiation in Section 7 further weakens rather than supports the claim: with D(d(p,q)) := dmax − D_KL(p||q), the objective (77) is algebraically equivalent to maximizing E_p[U] + (1/β)D_KL(p||q), the opposite sign of the deontic penalty in (78). That is an internal-consistency/correctness problem rather than a circularity, but it confirms that the legal example is not an independent derivation of the deontology-as-regularization thesis. Overall: the mathematics is not circular; the central interpretive claim reduces by construction, giving a score of 6 rather than a higher score.
Assumptions & free parameters
free parameters (5)
- Inverse temperature beta (coupling constant)
- Prior q over actions =
e.g., q=(0.7,0.2,0.1) in Figure 9
- State-dependent utility function U(x,y) =
e.g., U=(7,5) in Figure 9
- Disutility function D and divergence d =
e.g., D(d)=dmax-DKL(p||q)
- Source probability P on world states =
not specified, assumed given
assumptions (7)
- domain assumption Finite discrete measurable spaces and Markov kernels encode all relevant uncertainty and actions.
- domain assumption The decision maker chooses a distribution p maximizing Ep[U] minus a divergence penalty rather than choosing an action directly.
- ad hoc to paper Deontology corresponds to the divergence regularizer R(p,q) and to the support of the prior q.
- ad hoc to paper A restriction of a constitutional right is justified iff the restricted action leads to a legal state at least as desirable and the net public utility exceeds the disutility cost.
- domain assumption The inverse temperature beta is external and must be fixed by a competent authority.
- domain assumption The constraint-robust setting assumes a source probability P on world states and a fixed rate R.
- standard math Standard background theorems: disintegration theorem, convexity of KL divergence, Bauer maximum principle, information monotonicity.
invented entities (2)
-
Ought co-sheaf OX
-
Utility expansion path gamma+ and contraction path gamma-
Cite this review
Pith. "Pith review of The geometry of moral decision making." pith.science (2026). https://pith.science/paper/LPXGOAAP
@misc{pith2026250108865,
author = {Pith},
title = {Pith review of: The geometry of moral decision making},
year = {2026},
howpublished = {\url{https://pith.science/paper/LPXGOAAP}},
note = {Machine review of arXiv:2501.08865}
}
read the original abstract
We show how (resource) bounded rationality can be understood as the interplay of two fundamental moral principles: deontology and utilitarianism. In particular, we interpret deontology as a regularisation function in an optimal control problem, coupled with a free parameter, the inverse temperature, to shield the individual from expected utility. We discuss the information geometry of bounded rationality and aspects of its relation to rate distortion theory. A central role is played by Markov kernels and regular conditional probability, which are also studied geometrically. A gradient equation is used to determine the utility expansion path. Finally, the framework is applied to the analysis of a disutility model of the restriction of constitutional rights that we derive from legal doctrine. The methods discussed here are also relevant to the theory of autonomous agents.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Aliprantis, C. D., and Border, K. C. Infinite dimensional analysis: A Hitchhiker’s Guide, 3 ed. Springer, 2006
work page 2006
-
[2]
Information geometry and its applications , vol
Amari, S.-i. Information geometry and its applications , vol. 194. Springer, 2016
work page 2016
-
[3]
Ay, N., Jost, J., Vˆan Lˆe, H., and Schwachh¨ofer, L. Information geometry, vol. 64. Springer, 2017
work page 2017
-
[4]
Rate distortion theory : a mathematical basis for data compression
Berger, T. Rate distortion theory : a mathematical basis for data compression . Prentice- Hall series in information and system sciences. Prentice-Hall, Englewood Cliffs - N.J, 1971
work page 1971
-
[5]
Boissonnat, J.-D., Nielsen, F., and Nock, R. Bregman Voronoi diagrams. Discrete & Computational Geometry 44 (2010), 281–307
work page 2010
-
[6]
Boyd, S., and V andenberghe, L. Convex optimization . Cambridge university press, 2004
work page 2004
-
[7]
Bredon, G. E. Sheaf Theory, 2nd ed. 1997. ed. Graduate Texts in Mathematics, 170. Springer New York, New York, NY, 1997
work page 1997
-
[8]
Calafiore, G., and El Ghaoui, L. Optimization Models. Control systems and opti- mization series. Cambridge University Press, October 2014
work page 2014
Show all 39 references
-
[9]
Casebeer, W. D. Moral cognition and its neural constituents. Nature Reviews Neuro- science 4, 10 (2003), 840–846
2003
-
[10]
Springer, 2011
C ¸ inlar, E.Probability and stochastics. Springer, 2011. 29
2011
-
[11]
Deontological and utilitarian inclinations in moral decision making: a process dissociation approach
Conway, P., and Gawronski, B. Deontological and utilitarian inclinations in moral decision making: a process dissociation approach. Journal of personality and social psy- chology 104, 2 (2013), 216
2013
-
[12]
Justice and Mathematics: Two Simple Ideas
Cooter, R. Justice and Mathematics: Two Simple Ideas. In New directions in economic justice, R. Skurski, Ed. University of Notre Dame Press, 1983
1983
-
[13]
Cover, T. M. Elements of information theory , 1 ed. John Wiley & Sons, 1991
1991
-
[14]
Springer, 2015
Cox, D., Little, J., O’shea, D., and Sweedler, M.Ideals, varieties, and algorithms, 4 ed. Springer, 2015
2015
-
[15]
P., F aisal, A
Deisenroth, M. P., F aisal, A. A., and Ong, C. S. Mathematics for machine learning. Cambridge University Press, 2020
2020
-
[16]
A note on rational inattention and rate distortion theory
Denti, T., Marinacci, M., and Montrucchio, L. A note on rational inattention and rate distortion theory. Decisions in Economics and Finance 43 (2020), 75–89
2020
-
[17]
Evans, J. S. B. In two minds: dual-process accounts of reasoning. Trends in cognitive sciences 7, 10 (2003), 454–459
2003
-
[18]
Rights and moral cognition: an information-theoretic account
Friedrich, R. Rights and moral cognition: an information-theoretic account. Bachelor’s thesis (BLaw), University of Zurich, Law School, 2023
2023
-
[19]
Genewein, T., Leibfried, F., Grau-Moya, J., and Braun, D. A. Bounded ratio- nality, abstraction, and hierarchical decision-making: An information-theoretic optimality principle. Frontiers in Robotics and AI 2 (2015), 27
2015
-
[20]
Moral satisficing: Rethinking moral behavior as bounded rationality
Gigerenzer, G. Moral satisficing: Rethinking moral behavior as bounded rationality. Topics in cognitive science 2 , 3 (2010), 528–554
2010
-
[21]
D., Nystrom, L
Greene, J. D., Nystrom, L. E., Engell, A. D., Darley, J. M., and Cohen, J. D. The neural bases of cognitive conflict and control in moral judgment. Neuron 44, 2 (2004), 389–400
2004
-
[22]
P., and Sargent, T
Hansen, L. P., and Sargent, T. J. Robust control and model uncertainty. American Economic Review 91 , 2 (2001), 60–66
2001
-
[23]
Geometry of EM and related iterative algo- rithms
Hino, H., Akaho, S., and Murata, N. Geometry of EM and related iterative algo- rithms. Information Geometry 7 , Suppl 1 (2024), 39–77
2024
-
[24]
Introducing social norms in game theory
L´opez-P´erez, R. Introducing social norms in game theory. Working paper se- ries/Institute for Empirical Research in Economics , 292 (2006)
2006
-
[25]
Luce, R. D. Individual choice behavior , vol. 4. Wiley New York, 1959
1959
-
[26]
D., and Green, J
Mas-Colell, A., Whinston, M. D., and Green, J. R. Microeconomic theory, 18th print. ed. Oxford University Press, New York, 1995
1995
-
[27]
Mattsson, L.-G., and Weibull, J. W. Probabilistic choice and procedurally bounded rationality. Games and Economic Behavior 41 , 1 (2002), 61–78
2002
-
[28]
Quantal choice analysis: A survey
Mcfadden, D. Quantal choice analysis: A survey. NBER Book Chapters 5 (02 2012)
2012
-
[29]
The q-exponential family in statistical physics
Naudts, J. The q-exponential family in statistical physics. Central European Journal of Physics 7 (2009), 405–413
2009
-
[30]
Die Moral des Rechts
Neumann, U. Die Moral des Rechts. Deontologische und konsequentialistische Argumen- tationen in Recht und Moral. JRE 2 (1994), 81. 30
1994
-
[31]
A., and Braun, D
Ortega, P. A., and Braun, D. A. Thermodynamics as a theory of decision-making with information-processing costs. Proceedings of the Royal Society A: Mathematical, Phys- ical and Engineering Sciences 469 , 2153 (2013), 20120683
2013
-
[32]
A., and Stocker, A
Ortega, P. A., and Stocker, A. A. Human decision-making under limited time. Advances in Neural Information Processing Systems 29 (2016)
2016
-
[33]
Arbeitsbuch Mathematik f¨ ur Wirtschaftswissenschaftler
Pampel, T. Arbeitsbuch Mathematik f¨ ur Wirtschaftswissenschaftler. Springer, 2017
2017
-
[34]
Information geometry of the probability simplex: A short course
Pistone, G. Information geometry of the probability simplex: A short course. Nonlinear Phenomena in Complex Systems 23 , 2 (2020), 221–242
2020
-
[35]
Simon, H. A. Models of bounded rationality: Empirically grounded economic reason , vol. 3. MIT press, 1997
1997
-
[36]
Sims, C. A. Implications of rational inattention. Journal of Monetary Economics 50 , 3 (2003), 665–690. Swiss National Bank/Study Center Gerzensee Conference on Monetary Policy under Incomplete Information
2003
-
[37]
C., and Bialek, W
Tishby, N., Pereira, F. C., and Bialek, W. The information bottleneck method. arXiv preprint physics/0004057 (2000)
2000 arXiv
-
[38]
A new proof of a result concerning computation of the capacity for a discrete channel
Topsøe, F. A new proof of a result concerning computation of the capacity for a discrete channel. Zeitschrift f¨ ur Wahrscheinlichkeitstheorie und Verwandte Gebiete 22, 2 (1972), 166–168
1972
-
[39]
Zum Notstandsproblem
Welzel, H. Zum Notstandsproblem. Zeitschrift f¨ ur die gesamte Strafrechtswissenschaft 63, Jahresband (1951), 47–56. 31
1951
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.