Pith. sign in

REVIEW 3 minor

Point Tracking in Surgery--The 2025 Surgical Tattoos in Infrared Challenge (STIRC2025)

T0 review · 0 major / 3 minor · reviewed 2026-07-15 · grok-4.5

Pith's one-line read STIRC2025 ranks seven surgical point trackers on infrared-tattoo accuracy and latency.

desk verdict Standard EndoVis challenge report: useful STIR bake-off with public data and dual accuracy/latency scoring, not a foundational result. read the letter →

arxiv 2607.12939 v1 pith:NYKIQTLF submitted 2026-07-14 cs.CV

classification cs.CV
keywords pointtrackingsurgicalcomputervisioninfraredtattoosSTIRdatasetchallengebenchmarkMICCAIEndoVisinferencelatency
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reports the 2025 Surgical Tattoos in Infrared Challenge (STIRC2025), a public benchmark for point tracking in surgical video. Point tracking is framed as a building block for downstream surgical applications such as segmentation, 3D reconstruction, virtual tissue landmarking, autonomous probe scanning, and subtask autonomy. Participants submit tracking algorithms that are scored on the STIR infrared-tattoo dataset for both accuracy (on in vivo and ex vivo sequences) and efficiency (inference latency). Seven teams competed under MICCAI EndoVis 2025. The paper summarizes the rankings, the methods those teams used, and releases the dataset plus metric code so others can reproduce or extend the evaluation. A sympathetic reader cares because a shared, quantitative ranking on real surgical sequences makes it possible to compare trackers fairly rather than relying on private datasets or qualitative demos.

What carries the argument

The STIR infrared-tattoo evaluation: ground-truth points marked by infrared tattoos on tissue, scored for tracking accuracy on in vivo and ex vivo video plus wall-clock inference latency.

What would settle it

Run the top-ranked STIRC2025 trackers on the same sequences without tattoos or with clinical landmarks, and check whether the accuracy ranking and latency numbers still hold under those conditions.

Watch

Extended reading notes

Core claim

STIRC2025 establishes a dual accuracy-and-efficiency ranking of seven submitted point-tracking algorithms on the STIR infrared-tattoo dataset, covering both in vivo and ex vivo surgical sequences, and releases the dataset and metrics so the ranking can be audited and extended.

Load-bearing premise

That infrared tattoo markers and the challenge’s accuracy-plus-latency scores are a good enough stand-in for the tracking quality that real surgical applications actually need.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 3 minor

Summary. This manuscript reports the 2025 Surgical Tattoos in Infrared Challenge (STIRC2025), held within MICCAI EndoVis 2025. Participants submit point-tracking algorithms that are scored on the public STIR (surgical tattoos in infrared) dataset along two quantitative axes: tracking accuracy on in vivo and ex vivo sequences, and inference latency. Seven teams participated. The paper summarizes challenge results and participant methods, and releases the challenge dataset and baseline/metrics code.

Significance. A public, dual-track (accuracy and efficiency) benchmark for surgical point tracking is a useful community resource, especially given the downstream tasks named in the abstract (segmentation, 3D reconstruction, virtual landmarking, autonomous probe scanning, subtask autonomy). Explicit release of the STIR challenge data and metrics/baseline code is a concrete reproducibility strength and supports follow-on comparison. The significance of any ranking conclusions depends on protocol, metric, and statistical details that are not assessable from the abstract alone.

minor comments (3)
  1. Abstract-only review: the full methods, evaluation protocol, result tables, and participant-method descriptions are not available, so ranking validity, metric definitions, and statistical treatment cannot be checked.
  2. The abstract frames STIR tattoos and the accuracy/latency metrics as enabling several clinical downstream tasks; a short limitations discussion of how well infrared tattoos and the chosen metrics transfer to those tasks would strengthen the framing (not required for the bake-off claim itself).
  3. Clarify whether this is a continuation of a prior STIR challenge iteration and, if so, what changed in 2025 (data splits, metrics, tracks) so readers can place the results.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: challenge report evaluates external algorithms on a public labeled dataset with public metrics.

full rationale

This is an abstract-only challenge summary (STIRC2025 / EndoVis 2025). Its central claim is infrastructural and empirical: seven submitted algorithms were scored for tracking accuracy on in vivo and ex vivo STIR sequences and for inference latency, with the STIR dataset and metrics code released publicly. There is no derivation chain that reduces a claimed first-principles result or prediction to fitted parameters, self-defined quantities, or load-bearing self-citation uniqueness theorems. Evaluation is against an external labeled benchmark rather than a tautology constructed from the organizers' own model. Framing that point tracking enables downstream surgical tasks is motivational context, not a circular step. With only the abstract available, no equation-level self-definition, fitted-input-as-prediction, or ansatz-smuggling via self-citation is present. Score 0 is therefore the correct honest finding.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

As a challenge report abstract, the central claim rests on standard computer-vision evaluation practice and on domain assumptions that STIR’s infrared tattoos and the chosen accuracy/latency metrics represent surgical point tracking. No free parameters are fitted in the abstract’s claim, and no new physical entities are invented. The ledger is therefore thin on free parameters and invented entities and heavier on domain assumptions that only the full paper could fully justify.

assumptions (3)
  • domain assumption Infrared tattoo markers in STIR provide reliable ground-truth correspondences for evaluating surgical point trackers on in vivo and ex vivo sequences.
    The entire accuracy track depends on treating tattoo locations as the reference; if tattoos move relative to tissue of interest or are unrepresentative, rankings are misaligned with clinical tracking.
  • domain assumption Separate accuracy (in vivo/ex vivo) and inference-latency scores are an appropriate multi-objective ranking for algorithms intended for the listed downstream surgical tasks.
    The challenge design treats these two axes as the quantification of success; other factors (robustness to blood/smoke, long-horizon drift, stereo vs mono, calibration error) are not named as primary scores in the abstract.
  • standard math Standard supervised/self-supervised point-tracking evaluation methodology (public labeled sequences + fixed metrics) is a valid way to compare submitted algorithms.
    Challenge papers inherit the usual CV bake-off assumptions about held-out sequences and metric definitions; the abstract does not introduce a new formal theory of tracking error.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Point Tracking in Surgery--The 2025 Surgical Tattoos in Infrared Challenge (STIRC2025)." pith.science (2026). https://pith.science/paper/NYKIQTLF

@misc{pith2026260712939,
  author       = {Pith},
  title        = {Pith review of: Point Tracking in Surgery--The 2025 Surgical Tattoos in Infrared Challenge (STIRC2025)},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NYKIQTLF}},
  note         = {Machine review of arXiv:2607.12939}
}
read the original abstract

Point tracking in surgery is crucial to enable applications in downstream tasks such as segmentation, 3D reconstruction, virtual tissue landmarking, autonomous probe-based scanning, and subtask autonomy. This paper introduces the 2025 iteration of a point tracking challenge to address this, wherein participants submit their algorithms for quantification. Their algorithms are evaluated using a dataset named surgical tattoos in infrared (STIR), with the challenge named the STIR Challenge 2025 (STIRC2025). The STIR Challenge 2025 comprises two quantitative components: accuracy and efficiency. The accuracy component tests the accuracy of algorithms on in vivo and ex vivo sequences. The efficiency component tests algorithm inference latency. The challenge was conducted as a part of MICCAI EndoVis 2025, and seven teams participated in this challenge. In this paper we summarize the challenge results and participant methods. The challenge dataset is available at: https://zenodo.org/records/20191078, and the code for baseline models and metrics calculation is available here: https://github.com/athaddius/STIRMetrics

Discussion (0). Continue with ORCID to comment.

Pith tools

Reviewed July 15, 2026 · model on record in the stance chip above.