REVIEW 3 minor
Point Tracking in Surgery--The 2025 Surgical Tattoos in Infrared Challenge (STIRC2025)
T0 review · 0 major / 3 minor · reviewed 2026-07-15 · grok-4.5
Pith's one-line read STIRC2025 ranks seven surgical point trackers on infrared-tattoo accuracy and latency.
desk verdict Standard EndoVis challenge report: useful STIR bake-off with public data and dual accuracy/latency scoring, not a foundational result. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The STIR infrared-tattoo evaluation: ground-truth points marked by infrared tattoos on tissue, scored for tracking accuracy on in vivo and ex vivo video plus wall-clock inference latency.
What would settle it
Run the top-ranked STIRC2025 trackers on the same sequences without tattoos or with clinical landmarks, and check whether the accuracy ranking and latency numbers still hold under those conditions.
Extended reading notes
Core claim
STIRC2025 establishes a dual accuracy-and-efficiency ranking of seven submitted point-tracking algorithms on the STIR infrared-tattoo dataset, covering both in vivo and ex vivo surgical sequences, and releases the dataset and metrics so the ranking can be audited and extended.
Load-bearing premise
That infrared tattoo markers and the challenge’s accuracy-plus-latency scores are a good enough stand-in for the tracking quality that real surgical applications actually need.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript reports the 2025 Surgical Tattoos in Infrared Challenge (STIRC2025), held within MICCAI EndoVis 2025. Participants submit point-tracking algorithms that are scored on the public STIR (surgical tattoos in infrared) dataset along two quantitative axes: tracking accuracy on in vivo and ex vivo sequences, and inference latency. Seven teams participated. The paper summarizes challenge results and participant methods, and releases the challenge dataset and baseline/metrics code.
Significance. A public, dual-track (accuracy and efficiency) benchmark for surgical point tracking is a useful community resource, especially given the downstream tasks named in the abstract (segmentation, 3D reconstruction, virtual landmarking, autonomous probe scanning, subtask autonomy). Explicit release of the STIR challenge data and metrics/baseline code is a concrete reproducibility strength and supports follow-on comparison. The significance of any ranking conclusions depends on protocol, metric, and statistical details that are not assessable from the abstract alone.
minor comments (3)
- Abstract-only review: the full methods, evaluation protocol, result tables, and participant-method descriptions are not available, so ranking validity, metric definitions, and statistical treatment cannot be checked.
- The abstract frames STIR tattoos and the accuracy/latency metrics as enabling several clinical downstream tasks; a short limitations discussion of how well infrared tattoos and the chosen metrics transfer to those tasks would strengthen the framing (not required for the bake-off claim itself).
- Clarify whether this is a continuation of a prior STIR challenge iteration and, if so, what changed in 2025 (data splits, metrics, tracks) so readers can place the results.
Circularity Check
No significant circularity: challenge report evaluates external algorithms on a public labeled dataset with public metrics.
full rationale
This is an abstract-only challenge summary (STIRC2025 / EndoVis 2025). Its central claim is infrastructural and empirical: seven submitted algorithms were scored for tracking accuracy on in vivo and ex vivo STIR sequences and for inference latency, with the STIR dataset and metrics code released publicly. There is no derivation chain that reduces a claimed first-principles result or prediction to fitted parameters, self-defined quantities, or load-bearing self-citation uniqueness theorems. Evaluation is against an external labeled benchmark rather than a tautology constructed from the organizers' own model. Framing that point tracking enables downstream surgical tasks is motivational context, not a circular step. With only the abstract available, no equation-level self-definition, fitted-input-as-prediction, or ansatz-smuggling via self-citation is present. Score 0 is therefore the correct honest finding.
Assumptions & free parameters
assumptions (3)
- domain assumption Infrared tattoo markers in STIR provide reliable ground-truth correspondences for evaluating surgical point trackers on in vivo and ex vivo sequences.
- domain assumption Separate accuracy (in vivo/ex vivo) and inference-latency scores are an appropriate multi-objective ranking for algorithms intended for the listed downstream surgical tasks.
- standard math Standard supervised/self-supervised point-tracking evaluation methodology (public labeled sequences + fixed metrics) is a valid way to compare submitted algorithms.
Cite this review
Pith. "Pith review of Point Tracking in Surgery--The 2025 Surgical Tattoos in Infrared Challenge (STIRC2025)." pith.science (2026). https://pith.science/paper/NYKIQTLF
@misc{pith2026260712939,
author = {Pith},
title = {Pith review of: Point Tracking in Surgery--The 2025 Surgical Tattoos in Infrared Challenge (STIRC2025)},
year = {2026},
howpublished = {\url{https://pith.science/paper/NYKIQTLF}},
note = {Machine review of arXiv:2607.12939}
}
read the original abstract
Point tracking in surgery is crucial to enable applications in downstream tasks such as segmentation, 3D reconstruction, virtual tissue landmarking, autonomous probe-based scanning, and subtask autonomy. This paper introduces the 2025 iteration of a point tracking challenge to address this, wherein participants submit their algorithms for quantification. Their algorithms are evaluated using a dataset named surgical tattoos in infrared (STIR), with the challenge named the STIR Challenge 2025 (STIRC2025). The STIR Challenge 2025 comprises two quantitative components: accuracy and efficiency. The accuracy component tests the accuracy of algorithms on in vivo and ex vivo sequences. The efficiency component tests algorithm inference latency. The challenge was conducted as a part of MICCAI EndoVis 2025, and seven teams participated in this challenge. In this paper we summarize the challenge results and participant methods. The challenge dataset is available at: https://zenodo.org/records/20191078, and the code for baseline models and metrics calculation is available here: https://github.com/athaddius/STIRMetrics
Reviewed July 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.