REVIEW 3 major objections 5 minor 13 references
How brains build higher order representations of uncertainty
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper proposes that reported confidence is not a direct readout of first-order uncertainty but the output of a hierarchical Bayesian inference that combines a current noise estimate with learned expectations about typical noise.
desk verdict A clear conceptual reorganization of Bayesian confidence with a real measurement gap for the likelihood-like component; worth reading but not as an empirical claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the hierarchical Bayesian decomposition of uncertainty HORs into three named components: likelihood-like HORs, which estimate the momentary reliability of a current first-order representation; prior-like HORs, which encode learned expectations about typical noise along task-relevant dimensions; and posterior-like HORs, which integrate the two to form experienced uncertainty. The formal identity is $p(\mathrm{uncertainty_{FOR}} | \mathrm{estimate_{HO}}) \propto p(\mathrm{estimate_{HO}} | \mathrm{uncertainty_{FOR}})\, p(\mathrm{uncertainty_{FOR}})$, and the paper uses it to argue that confidence reports are products of two inferential stages rather than direct reads. The machinery also includes the analytical tools proposed for measuring each component: probabilistic population codes and TAFKAP-style decoding for the likelihood-like term, and the NERD diffusion model trained on decoded neurofeedback data for the prior-like term.
What would settle it
Run a perceptual task in which FOR uncertainty is decoded trial by trial from early visual cortex while the experimenter manipulates the distribution of noise the observer experiences over blocks. If confidence reports are unchanged whenever decoded FOR uncertainty is held constant, even when the experienced noise distribution changes, the claim that prior-like HORs enter confidence is falsified. The same experiment would also fail if neural patterns carrying the learned noise distribution cannot be recombined with patterns carrying current noise to predict confidence.
Extended reading notes
Core claim
The central claim is that metacognitive confidence is not a direct readout of first-order uncertainty. The brain is proposed to build a second-order Bayesian posterior over its own uncertainty: a current, noisy estimate of how unreliable the first-order representation is (a likelihood-like HOR) is combined with a learned distribution over how unreliable that kind of representation usually is (a prior-like HOR), and the resulting posterior-like HOR is what gets read out as confidence. The paper expresses this as $p(\mathrm{uncertainty_{FOR}} | \mathrm{estimate_{HO}}) \propto p(\mathrm{estimate_{HO}} | \mathrm{uncertainty_{FOR}})\, p(\mathrm{uncertainty_{FOR}})$. It argues that nearly all existing work studies only the posterior side of this chain, and that isolating the likelihood-like and prior-like components is required to explain dissociations between actual and reported uncertainty and to arbitrate between theories of metacognition and consciousness.
Load-bearing premise
The argument depends on the brain storing two distinguishable kinds of uncertainty information at once: how noisy the current signal feels and how noisy that kind of signal usually is, and on those being separable, measurable neural states. If they are not separable, the Bayesian decomposition is a mathematical metaphor with no neural target.
Editorial extensions
If this is right
- Confidence reports should be treated as posterior readouts, not direct measurements of first-order uncertainty, so experiments that use confidence to infer sensory noise need to control for prior-like expectations.
- Individual differences in metacognitive calibration can be explained by different learned noise priors rather than by different sensitivity to current uncertainty, which changes how metacognitive training would be designed.
- Confidence models that add a single noise term to a decision variable, such as meta-d'-type models, are incomplete by this account because they collapse the likelihood-like and posterior-like stages that the paper separates.
- Learning tasks without external feedback, such as decoded neurofeedback, can be reinterpreted as updating prior-like uncertainty distributions, making the NERD model a candidate mechanism for how such learning proceeds.
Reading between the lines
- If the decomposition is right, the established dissociation between objective and subjective uncertainty in the visual periphery is most naturally a mismatch between the prior-like HOR and the current likelihood-like HOR; this could be tested by retraining the prior while leaving the sensory representation untouched.
- The same Bayesian HOR structure could be extended to other properties of first-order representations, such as signal strength, source (external versus internal), or content, giving each higher-order theory of consciousness a measurable set of dimensions.
- A direct test would train subjects in two environments with different noise statistics but identical task stimuli, then measure confidence on probe trials where decoded FOR uncertainty is matched: if confidence tracks the training environment, prior-like HORs are causally implicated.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes that metacognitive estimates of uncertainty are not direct readouts of first-order uncertainty but rather the product of a hierarchical Bayesian inference process over the brain's own representations. It decomposes higher-order representations (HORs) of uncertainty into three components: likelihood-like HORs (momentary estimates of current FOR uncertainty), prior-like HORs (learned expectations about typical FOR noise), and posterior-like HORs (the integrated result that drives confidence reports). The authors survey existing methods (GLMsingle, GSN, TAFKAP/PRINCE, NERD) and argue that these can be adapted or extended to isolate the hypothesized components. The paper is explicitly a conceptual proposal and includes no new experimental data or quantitative model, but it offers a concrete framework and a research agenda for separating the contributions of current noise estimates and learned noise priors to metacognitive judgments.
Significance. If the proposed decomposition is correct, it would reframe the study of metacognition and confidence as a hierarchical inference problem, aligning uncertainty monitoring with the Bayesian brain framework and connecting metacognition to generative-model and reinforcement-learning approaches. The paper is clearly written and unusually transparent about its own limitations; for example, it explicitly notes that TAFKAP and PRINCE measure FOR uncertainty rather than the likelihood-like HOR about that uncertainty, and that GLMsingle/GSN measure voxel noise rather than FOR noise. That transparency is a strength. However, the framework's central empirical claim—that brains literally maintain separable likelihood-like and prior-like HORs and combine them multiplicatively—is currently supported mainly by analogical arguments and by the authors' own unpublished work (NERD). No method described in the paper independently measures the likelihood-like HOR, and no direct empirical test separating the three components is offered.
major comments (3)
- [Box 1; The 'likelihood' section] The central equation p(uncertainty_FOR | estimate_HO) ∝ p(estimate_HO | uncertainty_FOR) p(uncertainty_FOR) is not identifiable from the measurements the paper describes. The 'likelihood' section concedes that TAFKAP and PRINCE estimate FOR uncertainty 'rather than the (likelihood-like) HOR about that uncertainty,' and that GLMsingle and GSN measure voxel noise, not FOR noise. The paper does not propose any way to observe or estimate the likelihood-like HOR itself. Without an independent observable for this term, confidence reports alone are compatible with many non-Bayesian or non-hierarchical models (e.g., meta-d' or CASANDRE), and the Bayesian decomposition is underdetermined. Please specify a concrete measurement protocol or a set of falsifiable predictions that would distinguish the proposed hierarchical Bayesian process from a direct, non-hierarchical readout of first-order uncertainty.
- [The 'prior' section] The prior-like HOR rests almost entirely on the NERD model, which is supported by an unreviewed preprint (Azimi Asrari & Peters, 2025) and a conference abstract (Azimi Azrari & Peters, 2024). The analogy between denoising diffusion models and DecNef learning is interesting, but the paper does not demonstrate that the brain stores a separable, learnable distribution over FOR uncertainty, nor that NERD's learned noise distribution corresponds to a neural prior-like HOR. The claim that 'the lower-dimensional prior-like uncertainty HORs discovered by NERD could indeed capture individual variation' is presented without details of the analysis, sample size, or statistical results. Please clarify what evidence would confirm or refute the existence of a prior-like HOR and provide details of the NERD results so that readers can assess this load-bearing claim.
- [Box 1; Figure 1] The independence and multiplicative combination of the likelihood-like and prior-like components is assumed as an axiom, but no justification is given for why these two HOR types are independent or why they combine as a product. In standard Bayesian perception, the independence assumption is motivated by generative models of the environment; here, the 'estimate_HO' variable is not precisely defined, and it is not clear what physiological or computational constraint would enforce independence between a current noise estimate and a learned noise prior. Please state the conditions under which the multiplicative decomposition would fail (e.g., correlated noise estimates or context-dependent priors) and describe how a failure would be detected empirically.
minor comments (5)
- [The 'likelihood' section] There is a typo in the sentence 'the relationship between voxelwise noise and FOR noise is as complex a the relationship between voxelwise patterns and the mental structures they represent'; the word 'as' should be 'as the'.
- [The 'likelihood' section] The text says GLMsingle quantifies 'how much a given voxel's activity is predicted by a task-relevent variable'; 'relevent' should be 'relevant'.
- [Figure 2] The figure label 'reward calculatuion' contains a typo: 'calculatuion' should be 'calculation'.
- [References] The author name is spelled inconsistently: 'Azimi Asrari' in the text and reference list appears as 'Azimi Azrari' in the 2024 conference abstract reference; please standardize.
- [The 'prior' section] The phrase 'inferiortemporal cortex' should be 'inferior temporal cortex'.
Circularity Check
No significant circularity; the Bayesian HOR taxonomy is a proposed organizing framework, with self-citations serving as motivation rather than as a derivation of the conclusion.
full rationale
This is a perspective/proposal paper, not a derivation. The central claim—that metacognitive estimates of uncertainty reflect posterior-like HORs built from likelihood-like and prior-like components—is introduced as a hypothesis ('we propose'), and Box 1 explicitly labels the Bayesian equation as a formulation, not as an established result. No parameters are fitted and then re-labeled as predictions, and no equation is shown to equal its own input. The main self-citations (Winter & Peters 2022; Azimi Asrari & Peters 2025) provide empirical and modeling motivation for the existence of prior-like HORs, but the paper does not invoke them as a uniqueness theorem or as a proof that the Bayesian decomposition is forced. In fact, the footnote concedes that even a non-Bayesian comparison process would require some representation of FOR-uncertainty distributions. The acknowledged absence of a direct measure of the likelihood-like HOR (TAFKAP/PRINCE read FOR uncertainty rather than the HOR about it) is an identifiability and measurement gap, not a circular reduction. Therefore the reasoning chain is not circular; at most it leans on the authors' own prior work for motivation, which is a self-citation concern rather than a circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption The brain constructs FORs and HORs as internal representations about the world and about FORs.
- domain assumption Perception and metacognition can be described as Bayesian inference.
- domain assumption Behavioral confidence reports are read-outs of posterior-like HORs.
- ad hoc to paper The likelihood and prior components are independent and multiplicatively combined.
invented entities (3)
-
likelihood-like HOR
-
prior-like HOR
-
posterior-like HOR
Cite this review
Pith. "Pith review of How brains build higher order representations of uncertainty." pith.science (2026). https://pith.science/paper/CL4VX2R6
@misc{pith2026250619057,
author = {Pith},
title = {Pith review of: How brains build higher order representations of uncertainty},
year = {2026},
howpublished = {\url{https://pith.science/paper/CL4VX2R6}},
note = {Machine review of arXiv:2506.19057}
}
read the original abstract
Higher-order representations (HORs) are neural or computational states that are "about" first-order representations (FORs), encoding information not about the external world per se but about the agent's own representational processes -- such as the reliability, source, or structure of a FOR. These HORs appear critical to metacognition, learning, and even consciousness by some accounts, yet their dimensionality, construction, and neural substrates remain poorly understood. Here, we propose that metacognitive estimates of uncertainty or noise reflect a read-out of "posterior-like" HORs from a Bayesian perspective. We then discuss how these posterior-like HORs reflect a combination of "likelihood-like" estimates of current FOR uncertainty and "prior-like" learned distributions over expected FOR uncertainty, and how various emerging engineering and theory-based analytical approaches may be employed to examine the estimation processes and neural correlates associated with these highly under-explored components of our experienced uncertainty.
Figures
Reference graph
Works this paper leans on
-
[1]
Adams, W. J., Graf, E. W., & Ernst, M. O. (2004). Experience can change the ’light-from-above’ prior. Nature Neuroscience, 7(10), 1057–1058. https://doi.org/10.1038/nn1312 Azimi Asrari, H., & Peters, M. A. K. (2025). Revealing higher-order neural representations of uncertainty with the noise estimation through reinforcement-based diffusion (nerd) model. a...
work page Pith review arXiv 2004
-
[18]
Girshick, A. R., Landy, M. S., & Simoncelli, E. P. (2011). Cardinal rules: Visual orientation per- ception reflects knowledge of environmental statistics.Nature Neuroscience, 14(7), 926–932. https://doi.org/10.1038/nn.2831 Gold, J. I., & Shadlen, M. N. (2007). The neural basis of decision making. Annu. Rev. Neurosci., 30, 535–574. Goris, R. L., Movshon, J...
doi:10.1038/nn.2831 2011
-
[69]
https://doi.org/ 10.1038/S41597-021-00845-7 De Ridder, D., Vanneste, S., & Freeman, W. (2014). The bayesian brain: Phantom percepts resolve sensory uncertainty. Journal of Neuroscience , 34(46), 15094–15100. https://doi.org/10. 1523/JNEUROSCI.3360-14.2014 Dennett, D. C. (1991). Consciousness explained. Little, Brown; Company. Dunlosky, J., & Metcalfe, J. ...
arXiv 2014
-
[86]
Cleeremans, A., Achoui, D., Beauny, A., Keuninckx, L., Martin, J.-R., Mu˜ noz-Moldes, S., Vuil- laume, L., & de Heering, A. (2019). Learning to be conscious. Trends in Cognitive Sciences, 23(12), 921–934. https://doi.org/10.1016/j.tics.2019.09.011 Cleeremans, A., Timmermans, B., & Pasquali, A. (2007). Conscious access to first-order and higher- order repr...
-
[139]
Pospisil, D. A., & Pillow, J. W. (2024). Revisiting the high-dimensional geometry of population responses in visual cortex. bioRxiv. https://doi.org/10.1101/2024.02.16.580726 Prince, J. S., Charest, I., Kurzawski, J. W., Pyles, J. A., Tarr, M. J., & Kay, K. N. (2022). Improving the accuracy of single-trial fmri response estimates using glmsingle. Elife, 1...
-
[192]
Rosenthal, D. M. (2005). Consciousness and mind. Oxford University Press. https://www.amazon. com/Consciousness-Mind-David-Rosenthal/dp/0198236964 Rosenthal, D. M. (2012). Higher-order awareness, misrepresentation and function. Philosophical Transactions of the Royal Society B: Biological Sciences , 367(1594), 1424–1438. https:// doi.org/10.1098/rstb.2011...
-
[443]
Friston, K. (2010). The free-energy principle: A unified brain theory? Nat. Rev. Neurosci. , 11(2), 127–138. Friston, K., Lin, M., Frith, C. D., Pezzulo, G., Hobson, J. A., & Ondobaka, S. (2017). Active inference, curiosity and insight. Cognitive Neuroscience, 8(2), 144–153. https://doi.org/10. 1080/17588928.2016.1267040 Fr¨ omer, R., Nassar, M. R., Bruck...
-
[620]
https://doi.org/10.1038/s41467-024-00620-0 van Bergen, R. S., & Jehee, J. F. M. (2021). Tafkap: An improved method for probabilistic decoding of cortical activity. bioRxiv. https://doi.org/10.1101/2021.03.04.433946 van Bergen, R. S., Ma, W. J., Pratte, M. S., & Jehee, J. F. M. (2015). Sensory uncertainty decoded from visual cortex predicts behavior. Natur...
arXiv 2021
Show all 13 references
-
[656]
Shibata, K., Watanabe, T., Sasaki, Y., & Kawato, M. (2011). Perceptual learning incepted by decoded fmri neurofeedback without stimulus presentation. Science, 334(6061), 1413–1415. https://doi.org/10.1126/SCIENCE.1212003 Smith, M. A., & Kohn, A. (2008). Spatial and temporal sc...
2011 doi
-
[668]
https://doi.org/10.3389/fnhum.2013.00668 Shekhar, M., & Rahnev, D. (2024). How do humans give confidence? a comprehensive comparison of process models of perceptual metacognition. Journal of Experimental Psychology: General , 153(3),
2024
-
[1458]
https://doi.org/10.3389/fpsyg.2017.01458 Von Eckardt, B. (2012). The representational theory of mind. The Cambridge handbook of cognitive science, 1(29-50). Walker, E. Y., Pohl, S., Denison, R. N., Barack, D. L., Lee, J., Block, N., Ma, W. J., & Meyniel, F. (2023). Studying th...
2012
-
[1867]
https://doi.org/10.1038/s41593-023-01444-y Watanabe, T., Sasaki, Y., Shibata, K., & Kawato, M. (2017). Advances in fmri real-time neuro- feedback. Trends in cognitive sciences, 21(12), 997–1010. Weisberg, M. (2013). Simulation and similarity: Using models to understand the wor...
2017 doi
-
[5602]
https://doi.org/10.1038/s41598-018-23936-9 Haynes, J.-D., & Rees, G. (2005). Predicting the orientation of invisible stimuli from activity in human primary visual cortex. Nature Neuroscience, 8(5), 686–691. https://doi.org/10. 1038/nn1445 13 Jehee, J. F. M., Ling, S., Swisher,...
2005
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.