REVIEW 3 major objections 7 minor 22 references
Deep learning the holographic black hole with charge
T0 review · 3 major / 7 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A ten-layer neural network whose weights encode the scalar field equation recovers the charged anti-de Sitter black hole metric from boundary data.
desk verdict The charged black hole extension rests on a wrong change-of-variables, so the network learns an input function rather than the metric. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a 10-layer feedforward network in which the linear transformation between layers is the discretized scalar equation of motion in the coordinate $\eta$; the weight matrix is $W^{(n)} = \begin{pmatrix} 1 & \Delta\eta \\ \Delta\eta\,m^2 & 1-\Delta\eta\,R(\eta^{(n)}) \end{pmatrix}$, and the activation function encodes the potential $V(\varphi)$. A coordinate transformation $d\eta = dr/\sqrt{f}$ removes the metric function $f$ from the weights and activation, so the only unknown to be learned is the metric function $R(\eta)$ at each layer. The final activation is a smoothed step function of the horizon boundary condition, and the loss is $L^1$ distance to the labels plus a per-case regularization term.
What would settle it
Fix a single regularization coefficient and $\eta$-exponent, say the RN1 values, and train all six cases; if the recovered metrics separate into good and bad fits with errors far outside the reported range, the per-case tuning is doing essential work and the metric is not being learned from data alone.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that the expected metric can be obtained by training this network: for each of six cases (charge $Q=0.5$ or $0.25$; topology $k=0,1,-1$), the trained ten-point emergent metric matches the analytic metric with mean-square errors between 0.00155 and 0.27903. The architecture encodes the discretized equation of motion $\partial_\eta \pi + R(\eta)\pi - m^2\varphi - \delta V/\delta\varphi = 0$ in its weights, and the data are boundary values of $\varphi$ and $\pi$ labeled by whether they satisfy the horizon condition at $\eta_{\rm fin}$. The authors take this as evidence that by learning boundary CFT data, the network reproduces the AdS metric even for spacetimes with charge and different topology.
Load-bearing premise
The reconstruction stands on the assumption that the per-case regularization term can be chosen without already knowing the answer; the paper gives no rule for selecting the coefficient or the $\eta$-exponent, and the two largest reported errors suggest the choice is fragile.
Editorial extensions
If this is right
- For any of the tested charges $Q=0.5,0.25$ and topologies $k=0,1,-1$, a trained network yields a metric whose ten sampled values approximate the analytic Reissner-Nordström–AdS metric.
- Training with a suitable learning rate and batch size makes the loss converge to the same value given enough epochs, while a learning rate near 0.1 prevents convergence; this gives practical guidance for similar reconstructions.
- Because the coordinate transformation $d\eta = dr/\sqrt{f}$ moves the metric out of the activation function, the same architecture can in principle be applied to other static black holes once the transformation is computed numerically.
- The recovered metric is discrete by construction, so the method directly produces the metric only at the ten $\eta$ layers used in the network.
Reading between the lines
- A sharper version of the paper's test would hold out the analytic metric: fix one regularization rule, train on all six cases, and compare errors; the reported table does not show that this succeeds.
- If the regularization hurdle can be removed, the same discretized-equation network could be aimed at unknown bulk metrics, with the horizon labeling condition as the only input beyond boundary data.
- The ten-layer discretization means only ten points of $R(\eta)$ are learned; scaling to more layers and studying interpolation error would show whether the method converges to the full metric.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript adapts the deep-learning approach of Hashimoto et al. (Ref. [16]) to the Reissner-Nordström-AdS black hole. The authors construct a ten-layer neural network in which the weights are proposed to be the discretized scalar-field equation of motion in a new radial coordinate η defined by dη = dr/√f. They generate a training set by propagating boundary data through the analytic RN-AdS metric, label points with a horizon regularity condition, and train the network to reproduce the labels. The trained weights are interpreted as the RN-AdS metric function R(η), and the results are compared with the analytic metric for two charges (Q=0.5, 0.25) and three horizon topologies (k=0, ±1). The paper reports MSE values between 0.00155 and 0.279 and discusses the influence of learning rate, batch size, and initialization on the loss.
Significance. If the central identification of the network weight with the metric were correct, the paper would extend the holographic deep-learning program to charged black holes with nontrivial topology. The authors are transparent in reporting per-case numerical errors and in showing hyperparameter sensitivity, and they provide a quantitative comparison table. However, the main mathematical step connecting the coordinate-transformed equation to the network weight is incorrect, and the reported agreement is between the network output and the same quantity inserted into the data-generation procedure. As a result, the central claim that the neural network learns the RN-AdS metric from boundary data is not supported by the current derivation and experiments.
major comments (3)
- [Section II, Eqs. (13)-(16)] Under the coordinate transformation dη = dr/√f in Eq. (13), the scalar equation (8) does not become Eq. (14). If π in (14) is taken to be the new momentum ∂_η φ, the correct equation reads ∂_η(∂_η φ) + [f'(r(η))/(2√f) + 2√f/r(η)] ∂_η φ - m²φ - V' = 0 in the four-dimensional case used in the numerical section; if π is still the old φ'(r), then ∂_ηπ acquires a 1/√f factor on the source term and the coefficient is f'/√f + 2√f/r. In neither case is the coefficient equal to f(r(η)). Therefore the weight W_{22}^{(n)} = 1 - Δη R(η^{(n)}) in Eq. (16) is not the metric at the grid point, and the comparison in Appendix A between R and the analytic f is a comparison between two different quantities. The central claim in Section V that 'the expected metric can be obtained by training this network' is consequently unsupported.
- [Section III.A and Section IV] The training set is produced by propagating Eq. (14) through the analytic RN-AdS metric and then labeling the data with the horizon condition (18), which is derived from the same metric. The network architecture in Eq. (16) is exactly the discretized version of the same equation, with R(η) as the trainable weights. Training therefore recovers the R(η) that was used to assign the labels; the agreement in Table II is a check that the optimization finds the generating function of the forward model, not an independent inference of the physical metric from boundary data. To support the holographic claim, the authors should break this circularity—for example, by generating data with the r-coordinate equation (8) and testing whether the η-coordinate network still recovers f—or explicitly restrict the claim to inverting the network's own forward model.
- [Table I and Fig. 7] The regularizer terms in Table I have a different coefficient and a different power of η for each of the six RN cases, and Section IV states only that these values were found 'according to experience.' No rule is given for selecting them, and the two cases with the largest errors (RN4: MSE 0.176; RN6: MSE 0.279) show visible mismatches in Fig. 7(d) and 7(f) that are not discussed. This is load-bearing because it means the claimed successes are not reproducible from the described procedure and the method is fragile with respect to the auxiliary terms.
minor comments (7)
- [Throughout] Typos: 'INTRODUCTON' (Section title), 'T raining' (Section III.B heading), 'TABEL' (Appendix A), 'wether' (Section I), 'TyTorch' (Section I) versus 'PyTorch' (Section III.B), and 'epoches' (Section IV) should be corrected.
- [Eq. (12) and Eq. (17)] The second activation function is written as x2 + Δr δV(x1)/(f δφ), which is not a well-formed expression; after the coordinate transformation, Eq. (17) uses δV/δx1. The notation should be made consistent and dimensionally correct.
- [Eq. (18)] The horizon condition uses 2/η π without derivation. Since this condition is used to assign every label, the paper should derive it from the near-horizon limit of Eq. (14) or from the regularity condition of the scalar field.
- [Section IV] The statement that 'the optimal learning rate is around 0.001' should be supported by a quantitative criterion; Fig. 6 shows only that lr=0.1 fails to converge, which does not establish optimality.
- [Appendix A, Table II] The r values differ from case to case; please state how the radial grids are chosen and whether all cases use the same η-range and layer count.
- [Reproducibility] The paper does not provide the code, the initialization scheme, or the exact training schedule (number of epochs, optimizer parameters), so the numerical results are not independently reproducible from the text alone.
- [Fig. 7] Panels (e) and (f) are hard to read because the two curves have similar line styles; use distinct markers or colors for the reappeared and emergent metrics.
Circularity Check
Section III.A generates labels with the metric R(η), and Section II makes R(η) the network weight; the learned 'emergent metric' is the data-generation input, making the central claim a supervised consistency check.
-
fitted input called prediction
[Section II, Eqs. (14)-(16); Section III.A, Eq. (18); Appendix A]
"our data are generated by propagating the e.o.m for a given metric ... We then propagate data from ηini to ηfin(i.e. black hole horizon) by using the e.o.m(14) ... we use the boundary condition of the black hole horizon 0 =F≡[2/η π−m2φ− δV(φ)/δφ]_{η=ηf in} (18) to divide the data into two categories labeled by 0 and 1 ... Mapping the discrete e.o.m into deep neural network we have weights W(n)= ... 1−ΔηR(η(n)) (16)."
R(η) is both the coefficient in the e.o.m used to generate the training labels and the network weight to be fitted. Eq. (14) propagates (φ,π) through the known metric R; Eq. (18) labels the result using the near-horizon behavior 2/η of that same R; Eq. (16) places the same R(η(n)) in the weight matrix. Training therefore recovers the parameter that generated the data, and the Appendix A agreement is a closure test rather than an independent prediction. Separately, the coordinate change dη=dr/√f turns Eq. (8) into ∂ηπ+[fη/(2f)+2√f/r]π−m²φ−δV/δφ=0, so the coefficient that (14) calls R(η) is not f(r(η)) even if the circularity is set aside.
full rationale
The paper's central claim in Section V that 'the expected metric can be obtained by training this network' is built on training data that the authors themselves generate by propagating the scalar e.o.m through the known metric (Section III.A). The same R(η) appears as the adjustable weight in Eq. (16), so the learned 'emergent metric' is, by construction, the input to the label-generation rule. This is the pattern of a fitted parameter being presented as a prediction: the agreement in Fig. 7 and Appendix A confirms that gradient descent can recover the generating parameter, not that the metric was independently derived from boundary data. The per-case regularizer coefficients in Table I are selected 'according to experience' with no stated rule, adding a further tuning channel. There is no load-bearing self-citation chain; the debt to [16] is methodological. Score 6 reflects that the central claimed prediction reduces by construction to the data-generation input, but the network training itself is a nontrivial numerical test.
Assumptions & free parameters
free parameters (6)
- Regularizer coefficient =
0.039, 0.05, 0.0225, 0.05, 0.028, 0.0115 (RN1-RN6)
- Regularizer exponent on η =
3.6, 2.6, 2.4, 2.5, 3, 2 (RN1-RN6)
- Label threshold =
0.1
- Activation steepness =
100
- Learned metric values R(η) =
Ten values per case, e.g., RN1 ranges 3.045 to 7.249
- Learning rate and batch size =
0.001 and 10
assumptions (6)
- domain assumption AdS/CFT duality relates boundary fields to bulk geometry.
- domain assumption The metric has the RN-AdS form of Eq. (2) with known M, Q, L.
- standard math Eq. (8) is the correct equation of motion for the probe scalar.
- ad hoc to paper The horizon regularity condition (18) is the correct labeling rule.
- domain assumption Ten layers with Δη = -0.1 faithfully represent the continuum equations.
- ad hoc to paper The hand-chosen regularizers of Table I are necessary and sufficient.
Cite this review
Pith. "Pith review of Deep learning the holographic black hole with charge." pith.science (2026). https://pith.science/paper/5ZNGVV7F
@misc{pith2026190801470,
author = {Pith},
title = {Pith review of: Deep learning the holographic black hole with charge},
year = {2026},
howpublished = {\url{https://pith.science/paper/5ZNGVV7F}},
note = {Machine review of arXiv:1908.01470}
}
read the original abstract
We use the deep learning algorithm to learn the Reissner-Nordstr\"om(RN) black hole metric by building a deep neural network. Plenty of data is made in boundary of AdS and we propagate it to the black hole horizon through AdS metric and equation of motion(e.o.m). We label this data according to the values near the horizon, and together with initial data constitute a data set. Then we construct corresponding deep neural network and train it with the data set to obtain the Reissner-Nordstrom(RN) black hole metric. Finally, we discuss the effects of learning rate, batch-size and initialization on the training process.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[16]
Deep learning and the AdS/CFT correspon- 12 dence,
K. Hashimoto, S. Sugishita, A. Tanaka and A. Tomiya, “Deep learning and the AdS/CFT correspon- 12 dence,” Phys. Rev. D 98, no. 4, 046019 (2018) [arXiv:1802.08313 [hep-th]]
arXiv 2018
-
[1]
Gauge theory correlators from noncritical string theory,
S. S. Gubser, I. R. Klebanov and A. M. Polyakov, “Gauge theory correlators from noncritical string theory,” Phys. Lett. B 428, 105 (1998) [hep-th/9802109]
arXiv 1998
-
[2]
Anti-de Sitter space and holography,
E. Witten, “Anti-de Sitter space and holography,” Adv. Theor. Math. Phys. 2, 253 (1998) [hep- th/9802150]
arXiv 1998
-
[3]
The Large N limit of superconformal field theories and supergravity,
J. M. Maldacena, “The Large N limit of superconformal field theories and supergravity,” Int. J. Theor. Phys. 38, 1113 (1999) [Adv. Theor. Math. Phys. 2, 231 (1998)] [hep-th/9711200]
arXiv 1999
-
[4]
Dimensional Reduction in Quantum Gravity,
G. ’t Hooft,, “Dimensional Reduction in Quantum Gravity,” [gr-qc/9310026]
-
[5]
L. Susskind, “The World as a hologram,” J. Math. Phys. 36, 6377 (1995) [hep-th/9409089]
arXiv 1995
-
[6]
Reducing the dimensionality of data with neural networks.,
G. E. Hinton and R. R. Salakhutdinov, “Reducing the dimensionality of data with neural networks.,” Science 313, 504 (2006)
work page 2006
-
[7]
Scaling learning algorithms towards AI,
Y. Bengio, Y. LeCun, “Scaling learning algorithms towards AI,” Large-scale kernel machines 34 (2007)
work page 2007
Show all 22 references
-
[8]
Deep learning,
Y. LeCun, Y. Bengio, G. Hinton, “Deep learning,” Nature 521, 436 (2015)
2015
-
[9]
Deep Learning: Methods and Applications,
L. Deng, D. Yu, “ Deep Learning: Methods and Applications,” Foundations and Trends in Signal Pro- cessing. 7 (3C4) (2014)
2014
-
[10]
Learning representations by back-propagating errors,
Rumelhart, E. David ; G. E. Hinton, E. Geoffrey ; Williams, J. Ronald, “Learning representations by back-propagating errors,” Nature. 323 (6088): 533C536 (1986)
1986
-
[11]
The Perceptron–a perceiving and recognizing automaton,
R. Frank, “ The Perceptron–a perceiving and recognizing automaton,” Report 85-460-1, Cornell Aero- nautical Laboratory (1957)
1957
-
[12]
Deep learning and the renormalization group,
C. B´ eny, “Deep learning and the renormalization group,” [arXiv:1301.3124 [quant-ph]]
-
[13]
An exact mapping between the Variational Renormalization Group and Deep Learning,
P. Mehta, D. J. Schwab, “An exact mapping between the Variational Renormalization Group and Deep Learning,” [arXiv:1410.3831 [stat]]
-
[14]
Entanglement Renormalization and Holography,
B. Swingle, “Entanglement Renormalization and Holography,” Phys. Rev. D 86, 065007 (2012) [arXiv:0905.1317 [cond-mat.str-el]]
2012 arXiv
-
[15]
Holography as deep learning,
W. C. Gan and F. W. Shu, “Holography as deep learning,” Int. J. Mod. Phys. D 26, no. 12, 1743020 (2017) [arXiv:1705.05750 [gr-qc]]
2017 arXiv
-
[17]
Reissner, (1916), ” ¨Uber die Eigengravitation des elektrischen Feldes nach der Einsteinschen Theorie,” Annalen der Physik (in German)
H. Reissner, (1916), ” ¨Uber die Eigengravitation des elektrischen Feldes nach der Einsteinschen Theorie,” Annalen der Physik (in German). 50: 106C120
1916
-
[18]
Weyl, ”Zur Gravitationstheorie,” Annalen der Physik (in German), 54: 117C145 (1917)
H. Weyl, ”Zur Gravitationstheorie,” Annalen der Physik (in German), 54: 117C145 (1917)
1917
-
[19]
Nordstr¨ om, ”On the Energy of the Gravitational Field in Einstein’s Theory,” Verhandl
G. Nordstr¨ om, ”On the Energy of the Gravitational Field in Einstein’s Theory,” Verhandl. Koninkl. Ned. Akad. Wetenschap., Afdel. Natuurk., Amsterdam. 26: 1201C1208 (1918)
1918
-
[20]
G. B. Jeffery, ”The field of an electron on Einstein’s theory of gravitation,” Proc. Roy. Soc. Lond. A. 99: 123C134 (1921)
1921
-
[21]
Ketkar, ”Introduction to pytorch,” Deep learning with python
N. Ketkar, ”Introduction to pytorch,” Deep learning with python. Apress, Berkeley, CA, 2017: 195-208
2017
-
[22]
Paszke, S
A. Paszke, S. Gross, S. Chintala, et al, ”Automatic differentiation in pytorch,” Lerer, A. (2017). 13 Appendix A: Comparison and MSE As a direct comparison of result, TABEL II contrasts between the discrete metric which is from the neural network and the standard case. In the t...
2017
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.