Pith. sign in

REVIEW 5 major objections 4 minor 15 references

Mobile Image Analysis Application for Mantoux Skin Test

T0 review · 5 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper claims that a smartphone app using ARCore depth sensing, a scaling-sticker reference, and DeepLabv3 segmentation can measure Mantoux skin-test indurations at millimeter accuracy, replacing the manual ballpoint-pen method.

desk verdict The app is a plausible engineering prototype, but the reported millimeter accuracy is an artifact of calibrating and then evaluating on the same clay mocks, so the paper isn't yet a validated scientific contribution. read the letter →

arxiv 2506.17954 v1 pith:WKFII3NM submitted 2025-06-22 eess.IV cs.CV

classification eess.IVcs.CV
keywords MantouxtesttuberculinskinindurationmeasurementARCoredepthestimationDeepLabv3segmentationmobilehealthlatenttuberculosisscalingstickercalibration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that a smartphone can replace the manual ballpoint-pen/ruler reading of the Mantoux tuberculin skin test. The proposed Android app uses ARCore depth sensing, a scaling sticker placed next to the induration as a real-world reference, and a DeepLabv3 segmentation model to isolate the raised area and measure its diameter in millimeters. The authors argue this removes subjective interpretation and improves accessibility in low-resource settings, and they report that at the optimal capture distance of 219-220 mm a 10 mm clay mock induration was measured as 9.91 mm. The broader claim is that automated measurement of this standard TB-screening test can be done entirely on a mobile device.

What carries the argument

The load-bearing mechanism is the scaling-sticker calibration: a sticker with known physical size appears in the same image plane as the induration, and ARCore's Depth API supplies the millimeter depth of the skin surface. The app then measures the largest diameter of the segmented induration in pixels and converts it with a depth-calibrated factor chosen by pixel diameter (0.1197 below 50 pixels, 0.1523 between 50 and 80, 0.1499 between 80 and 200). This piecewise factor is what turns a dimensionless pixel measurement into a physical millimeter reading, and the paper's reported accuracy is essentially an evaluation of that calibration on the same clay objects used to derive it.

What would settle it

Capture a set of real TST indurations across varied skin tones and lighting, measure each with two blinded clinicians using the ballpoint-pen method and with the app, then compare both to an independent reference such as high-frequency ultrasound or calipers; if the app's errors exceed the 1-2 mm clinical tolerance or track systematically with skin tone, the central accuracy claim does not transfer to patients.

Watch

Extended reading notes

Core claim

The central discovery is a measurement pipeline in which ARCore's depth estimation provides the distance from camera to skin, and a scaling sticker supplies a known physical reference, so pixel distances can be converted to millimeters without photogrammetric 3D reconstruction. The paper reports piecewise conversion factors: 0.1197 for pixel diameters below 50, 0.1523 for 50-80, and 0.1499 for 80-200, calibrated on 5, 10, and 15 mm clay mocks. On the decisive test, a 10 mm mock at 219-220 mm capture depth was measured as 9.91 mm. The authors state that the evaluated system shows significant improvements in accuracy and reliability over standard clinical practice and conclude that on-device segmentation plus depth-based scaling is a viable path to standardized TST evaluation.

Load-bearing premise

The accuracy result rests on the assumption that clay mocks faithfully mimic the optical, depth, and edge properties of real human TST indurations, and that ARCore depth is equally accurate on skin as on the contrasting background used in the experiments.

Editorial extensions

If this is right

  • A clinician or trained worker can obtain an automated reading by holding the phone at the app-guided depth and orientation, removing the ballpoint-pen/ruler step and its inter-observer variability.
  • The on-device DeepLabv3 model computes the reading locally, so the app can work in settings without a network connection or specialized equipment.
  • Combining the measured diameter with the app's risk-factor questionnaire yields a positive or negative latent-TB classification, as demonstrated by the 10-15 mm induration threshold.
  • The reported optimal capture depth of 219-220 mm gives a concrete operating point for validation studies on real patients.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference beyond the paper: the same depth-plus-sticker pipeline could measure induration height or volume, not just diameter, because ARCore already provides per-pixel depth over the raised area.
  • Inference beyond the paper: the piecewise conversion factors are discontinuous at 50 and 80 pixels, so a smooth calibration curve fitted over the full range would likely reduce boundary artifacts and could be tested against the 5, 10, and 15 mm mocks.
  • Inference beyond the paper: if real-patient validation confirms the clay-mock accuracy, sticker-reference measurement could transfer to other raised skin lesions by retraining the segmentation model.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. The paper presents a mobile Android application (TBCheck) for automated measurement of Mantoux tuberculin skin test (TST) indurations. The system uses ARCore for depth estimation, DeepLabv3 for image segmentation, and a pixel-to-millimeter conversion calibrated through experiments with clay mock indurations of 5, 10, and 15 mm. The authors report an optimal capture depth of 219–220 mm, at which a 10 mm mock is measured as 9.91 mm, and claim that the application improves accuracy and reliability over standard clinical practice. The evaluation consists of four experiments: segmentation model selection (on the PAD-UFES-20 skin lesion dataset), depth optimization, calibration of conversion factors, and a usability study with ten participants.

Significance. If the measurement pipeline were shown to be accurate on real TST indurations, the application could offer a meaningful contribution to TB screening in resource-limited settings by standardizing and automating reading of the Mantoux test. The paper also demonstrates a systematic attempt to combine ARCore depth information, machine-learning segmentation, and a sticker-based scaling approach. However, the manuscript provides no validation on real patients, no comparison with the standard ballpoint pen and ruler method, no segmentation accuracy metrics on induration images, and no statistical analysis. The principal quantitative evidence is circular because the pixel-to-millimeter calibration factors are derived from and then evaluated on the same clay mock objects. The central claim of 'significant improvements in accuracy and reliability' is therefore unsupported by the evidence as presented.

major comments (5)
  1. [§4.5 and §4.6, Experiments 2 and 3] The load-bearing evidence for millimeter accuracy is circular. In §4.5, the pixel-to-millimeter conversion factors (0.1197, 0.1523, 0.1499) are presented as calibrated values, and §4.6 Experiment 3 states that these factors were obtained using mock indurations of 5, 10, and 15 mm. Experiment 2 then reports that the same type of clay mock (10 mm) is measured as 9.91 mm at the optimal depth. Because the evaluation and calibration sets are the same clay objects, the small reported error is a restatement of the fitted conversion factors rather than an independent test of accuracy. A valid evaluation would require held-out objects, a different ground-truth measurement modality, or real patient indurations with independent clinical measurement.
  2. [Abstract and §4.6 (conclusion of Experiments)] The abstract's claim that the application 'was evaluated against standard clinical practices, demonstrating significant improvements in accuracy and reliability' is not supported by the experiments in §4.6. No comparison to the ballpoint pen and ruler method or to any other clinical standard is reported; the only usability data come from ten participants completing a level-of-agreement questionnaire, with no quantitative accuracy comparison. The manuscript should either present such a comparison or explicitly restrict the claims to a technical demonstration on clay phantoms.
  3. [§4.4 and §4.6, Experiment 1] The segmentation model is trained exclusively on PAD-UFES-20, a dataset of dermatoscopic skin lesion images, and the manuscript provides no evidence that this model can segment TST indurations. No segmentation metrics (IoU, Dice, precision/recall) are reported for induration images, and Figure 3, labeled 'Successful and Failed Segmented of Skin Induration,' appears to include failures without quantitative analysis. The claim that DeepLabv3 provides 'robust image segmentation' for TST measurement is therefore not established, and the effect of segmentation errors on the reported diameter measurements is unknown.
  4. [§4.2 and §4.3] The ARCore depth estimation and plane detection are optimized using a 'contrasting background' (Figure 2), but the application is intended to operate on human skin, which exhibits varied pigmentation, hair, edema, and irregular surfaces. The transfer of depth accuracy from the laboratory setup with clay mocks on a contrasting cloth to real clinical conditions is a central assumption that is never tested. The authors themselves note that depth stability and accuracy under varied lighting conditions remain future work, which underscores that this assumption is unvalidated in the present manuscript.
  5. [§4.6, Experiments 2–4] No error bars, confidence intervals, or statistical tests are reported for any of the quantitative or usability measurements. For example, the finding that '60% felt confident' and 'only 50% found it easy to use' is based on ten participants, yet no uncertainty is attached to these percentages. More importantly, the repeated measurements from which the optimal depth of 219–220 mm and the conversion factors were derived are not reported with variance, making it impossible to assess the reliability of the calibration.
minor comments (4)
  1. [§4.5 (two sections numbered 4.5)] The manuscript contains two consecutive sections numbered 4.5: 'Automated Measurement of Induration' and 'Patient Information Collection and Reminder Function.' This numbering error should be corrected.
  2. [Reference list] The PAD-UFES-20 dataset is cited as [70] in §4.4, but the reference list contains only 24 entries; the reference should be added or the citation number corrected.
  3. [Throughout] There are several typographical and grammatical errors, for example 'pappers' in §2.1, 'di-agnostics' in the abstract, and the incomplete sentence in §2.1 starting with 'this app utilizes.' These should be fixed through a careful edit.
  4. [Figure 4] The three examples shown in Figure 4 (15.00 mm, 9.23 mm, and 4.97 mm) are presented as measurement results, but the figure does not identify which clay mock sizes these correspond to, making the figure uninformative without additional caption detail.

Circularity Check

2 steps flagged · score 7.0 of 10

Reported millimeter accuracy is a restatement of the calibration: conversion factors were fitted on the same 5/10/15 mm clay mocks later used as the 'validation' set.

  1. fitted input called prediction [§4.5 Automated Measurement of Induration and §4.6 Evaluation (Experiment 3, Experiment 2, Figure 4)]
    "This distance, initially measured in pixels, is then converted into millimeters using a calibrated conversion factor that varies dynamically with the measured pixel diameter: 0.1197 if the max diameter is less than 50 pixels, 0.1523 if the max diameter is between 50 and 80 pixels, and 0.1499 if the max diameter is between 80 and 200 pixels. ... Experiment 3 sought to find the optimal scalar factor for TST induration measurement using mock indurations of 5mm, 10mm, and 15mm. Optimal scalar factors were 0.1197 for 5mm, 0.1523 for 10mm, and 0.1399 for 15mm."

    The pixel-to-millimeter conversion factors in §4.5 were empirically fitted in Experiment 3 to make the 5mm, 10mm, and 15mm clay mocks read their known diameters. The evaluation then reports measurements of the same clay mocks — 9.91mm for the 10mm mock, 15.00mm, 4.97mm in Figure 4 — as evidence of accuracy. This is a restatement of the fit: the scale factors were chosen so those mock diameters produce those millimeter values, so the 'prediction' is forced by construction. The transfer from clay to real human TST indurations (different edge contrast, curvature, skin tone, hair, edema) is assumed without validation.

  2. fitted input called prediction [§4.6 Evaluation, Experiment 2, and §4.3 User Position and Image Capture Protocol]
    "Experiment 2 focused on determining the optimal depth value for accurate TST induration measurement. Images of a 10mm induration were captured at various depths, with the optimal range for accuracy being 219-220mm, where the predicted measurement was 9.91mm, demonstrating precise and reliable results in clinical settings."

    The 9.91mm value is not an out-of-sample prediction; it is the value at the depth that was selected because it produced the best accuracy for the same 10mm clay mock used to calibrate the conversion factor. Choosing the depth that gives 9.91mm and then reporting that depth as 'optimal' is parameter tuning, not validation. Both the depth and the scalar factor are fitted to the same 10mm mock, so the Experiment 2 result is a selected optimum of a two-parameter fit rather than an independent measurement of clinical accuracy.

full rationale

Most engineering components — ARCore plane detection, DeepLabV3 segmentation trained on PAD-UFES-20, the Android application — are not circular; they are evaluated on training/validation splits and are not derived from the target measurement claim. However, the central claim of millimeter-accurate TST induration measurement is circular in the narrow sense: §4.5 defines pixel-to-millimeter conversion factors (0.1197, 0.1523, 0.1499) that were fitted in Experiment 3 on exactly the 5mm, 10mm, and 15mm clay mocks, and the validation in Experiment 2 and Figure 4 reports measurements of the same clay mocks. The 9.91mm result for a 10mm mock is therefore a restatement of the calibration, not independent evidence. Experiment 2 adds a second fitted parameter, depth, by selecting 219–220mm because it produced the closest value for that same 10mm mock; reporting that selected optimum as 'precise and reliable' is fitted-input-called-prediction. No held-out real patient indurations, no comparison against the ballpoint/ruler standard on human skin, and no error bars are provided, so the abstract's claim of 'significant improvements in accuracy and reliability' is unsupported by independent evidence. There is no load-bearing self-citation: the cited University of Cape Town work is external. Score 7 rather than 10 because the segmentation model and app pipeline have independent content; the specific measurement-accuracy claim, however, reduces to its own calibration inputs.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The central measurement pipeline is loaded with unevidenced assumptions: calibration factors fitted to clay mocks, an untested assumption that ARCore works on skin, a segmentation model trained on unrelated images, and a simplified geometric model of induration diameter. These inputs are not independently verified or compared to clinical ground truth.

free parameters (3)
  • Pixel-to-millimeter conversion factor for diameters below 50 pixels = 0.1197
    Fitted in Experiment 3 to make a 5 mm clay mock measure correctly.
  • Pixel-to-millimeter conversion factor for diameters between 50 and 80 pixels = 0.1523
    Fitted in Experiment 3 to make a 10 mm clay mock measure correctly.
  • Pixel-to-millimeter conversion factor for diameters between 80 and 200 pixels = 0.1499 (algorithm) vs. 0.1399 (Experiment 3)
    The algorithm uses 0.1499, but Experiment 3 reports 0.1399 for the 15 mm mock, an inconsistency that suggests hand-tuning.
assumptions (5)
  • domain assumption Clay mock indurations of 5, 10, and 15 mm faithfully represent real TST indurations in appearance, depth, and edge properties.
    The app relies on this assumption to derive and validate the calibration factors, yet no comparison with real indurations is provided.
  • domain assumption ARCore depth estimation is accurate on human skin at phone-to-skin distances of 175-400 mm.
    ARCore depth APIs are typically designed for inanimate surfaces, and skin presents challenges such as texture, curvature, and occlusion. No validation on skin is presented.
  • domain assumption DeepLabv3 trained on the PAD-UFES-20 dermatoscopic skin lesion dataset generalizes to TST indurations captured with a smartphone camera.
    The training dataset is for skin lesions, not tuberculin reactions, and the paper reports no segmentation metrics on TST images.
  • domain assumption The longest line across the segmented edge corresponds to the clinically defined induration diameter measured perpendicular to the arm axis.
    The app uses the Euclidean distance between extreme edge points, but the clinical standard requires a perpendicular measurement. No geometry validation is reported.
  • standard math The Euclidean distance formula for pixel distance is correctly applied and the resulting conversion is linear in the specified ranges.
    The formula itself is standard, but the assumption that a single scalar factor per pixel range maps pixels to physical millimeters implies a linear camera model that is not justified for varying depth and skin curvature.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Mobile Image Analysis Application for Mantoux Skin Test." pith.science (2026). https://pith.science/paper/WKFII3NM

@misc{pith2026250617954,
  author       = {Pith},
  title        = {Pith review of: Mobile Image Analysis Application for Mantoux Skin Test},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WKFII3NM}},
  note         = {Machine review of arXiv:2506.17954}
}
read the original abstract

This paper presents a newly developed mobile application designed to diagnose Latent Tuberculosis Infection (LTBI) using the Mantoux Skin Test (TST). Traditional TST methods often suffer from low follow-up return rates, patient discomfort, and subjective manual interpretation, particularly with the ball-point pen method, leading to misdiagnosis and delayed treatment. Moreover, previous developed mobile applications that used 3D reconstruction, this app utilizes scaling stickers as reference objects for induration measurement. This mobile application integrates advanced image processing technologies, including ARCore, and machine learning algorithms such as DeepLabv3 for robust image segmentation and precise measurement of skin indurations indicative of LTBI. The system employs an edge detection algorithm to enhance accuracy. The application was evaluated against standard clinical practices, demonstrating significant improvements in accuracy and reliability. This innovation is crucial for effective tuberculosis management, especially in resource-limited regions. By automating and standardizing TST evaluations, the application enhances the accessibility and efficiency of TB di-agnostics. Future work will focus on refining machine learning models, optimizing measurement algorithms, expanding functionalities to include comprehensive patient data management, and enhancing ARCore's performance across various lighting conditions and operational settings.

Figures

Figures reproduced from arXiv: 2506.17954 by the authors.

Figure 1
Figure 1. Size Comparison of Modeled TB Indurations (5mm, 10mm, 15mm) Us [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Comparison of Plane Detection Without Cloth (Middle) and With Cloth [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Successful and Failed Segmented of Skin Induration [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Induration Measurements of 15.00 mm (left), 9.23 mm (middle), and [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

15 extracted references · 13 canonical work pages

  1. [5]

    Reliability of TB Skin-Test Measurement Using Ballpoint Pen,

    G. Pugliese and M. S. Favero, “Reliability of TB Skin-Test Measurement Using Ballpoint Pen,” Infection Control & Hospital Epidemiology, vol. 18, no. 10, pp. 730–730, Oct. 1997, doi: https://doi.org/10.1017/s0195941700000734

  2. [6]

    Trilateral overlap of tuberculosis, diabetes and HIV-1 in a high-burden African setting: implications for TB control,

    T. Oni, N. Berkowitz, M. Kubjane, R. Goliath, Naomi S. Levitt, and R. J. Wilkinson, “Trilateral overlap of tuberculosis, diabetes and HIV-1 in a high-burden African setting: implications for TB control,” European Respiratory Journal, vol. 50, no. 1, p. 1700004, Jul. 2017, doi: https://doi.org/10.1183/13993003.00004-2017

  3. [7]

    The Tuberculin Skin Test,

    R. E. Huebner, M. F. Schein, and J. B. Bass, “The Tuberculin Skin Test,” Clinical Infectious Diseases, vol. 17, no. 6, pp. 968–975, Dec. 1993, doi: https://doi.org/10.1093/clinids/17.6.968

  4. [8]

    Methodological approaches to shortening composite measurement scales,

    J. Coste, F. Guillemin, J. Pouchot, and J. Fermanian, “Methodological approaches to shortening composite measurement scales,” Journal of Clinical Epidemiology, vol. 50, no. 3, pp. 247–252, Mar. 1997, doi: https://doi.org/10.1016/s0895-4356(96)00363-0

  5. [9]

    Mobile phone-based evaluation of latent tuberculosis infection: Proof of concept for an integrated image capture and analysis system,

    S. Naraghi, T. Mutsvangwa, R. Goliath, M. X. Rangaka, and T. S. Douglas, “Mobile phone-based evaluation of latent tuberculosis infection: Proof of concept for an integrated image capture and analysis system,” Computers in Biology and Medicine, vol. 98, pp. 76– 84, Jul. 2018, doi: https://doi.org/10.1016/j.compbiomed.2018.05.009

  6. [10]

    Measurement of Skin Induration Size Using Smartphone Images and Photogrammetric Reconstruction: Pilot Study,

    R. Dendere, Tinashe Mutsvangwa, René Goliath, M. X. Rangaka, I. Abubakar, and T. S. Douglas, “Measurement of Skin Induration Size Using Smartphone Images and Photogrammetric Reconstruction: Pilot Study,” JMIR biomedical engineering, vol. 2, no. 1, pp. e3–e3, Dec. 2017, doi: https://doi.org/10.2196/biomedeng.8333

  7. [11]

    Image analysis for a mobile phone-based assessment of latent tuberculosis infection,

    S. Maclean, “Image analysis for a mobile phone-based assessment of latent tuberculosis infection,” open.uct.ac.za, 2020, Available: https://open.uct.ac.za/items/66b526c0-f39f- 41ad-ad54-034cba44c2d0

  8. [12]

    A mobile augmented reality application for supporting real-time skin lesion analysis based on deep learning,

    R. Francese, M. Frasca, M. Risi, and G. Tortora, “A mobile augmented reality application for supporting real-time skin lesion analysis based on deep learning,” Journal of Real-Time Image Processing, vol. 18, no. 4, pp. 1247–1259, May 2021, doi: https://doi.org/10.1007/s11554-021-01109-8

Show all 15 references
  1. [14]

    Mobile-Based Skin Disease Diagnosis System Using Convolutional Neural Networks (CNN),

    M. Maduranga and D. Nandasena, “Mobile-Based Skin Disease Diagnosis System Using Convolutional Neural Networks (CNN),” International Journal of Image, Graphics and Signal Processing, vol. 14, no. 3, pp. 47–57, Jun. 2022, doi: https://doi.org/10.5815/ijigsp.2022.03.05

  2. [15]

    https://arxiv.org/abs/1807.08891

  3. [16]

    Accuracy of a smartphone application using fractal image analysis of pigmented moles compared to clinical diagnosis and histological result,

    T. Maier et al., “Accuracy of a smartphone application using fractal image analysis of pigmented moles compared to clinical diagnosis and histological result,” Journal of the European Academy of Dermatology and Venereology, vol. 29, no. 4, pp. 663–667, Aug. 2014, doi: https://...

  4. [18]

    SkinSAM: Empowering Skin Cancer Segmentation with Segment Anything Model,

    M. Hu, Y. Li, and X. Yang, “SkinSAM: Empowering Skin Cancer Segmentation with Segment Anything Model,” arXiv (Cornell University), Apr. 2023, doi: https://doi.org/10.48550/arxiv.2304.13973

  5. [19]

    Acne Type Recognition for Mobile-Based Application Using YOLO,

    N. A. M. Isa and N. N. A. Mangshor, “Acne Type Recognition for Mobile-Based Application Using YOLO,” Journal of Physics: Conference Series, vol. 1962, no. 1, p. 012041, Jul. 2021, doi: https://doi.org/10.1088/1742-6596/1962/1/012041

  6. [21]

    Unconstrained automatic image matching using multiresolutional critical-point filters,

    Y. Shinagawa and T. L. Kunii, “Unconstrained automatic image matching using multiresolutional critical-point filters,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 20, no. 9, pp. 994–1010, 1998, doi: https://doi.org/10.1109/34.713364

  7. [2018]

    https://www.uppsatser.se/uppsats/a50b332f37/ 12 (accessed May 02, 2024)

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.