Pith. sign in

REVIEW 4 major objections 5 minor 15 references

Space-Creating versus Dead Possession: An Off-Ball Possession-Quality Index for Broadcast Football

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper claims that a junk-possession index—low-threat circulation in tied or losing states—predicts points where on-ball value (VAEP) does not, and that a spatial layer from broadcast video can tell sterile possession from…

desk verdict A genuinely novel two-layer possession-quality framework with an honest event-side core; the 'beyond VAEP' claim is real but under-supported by an unvalidated comparator. read the letter →

arxiv 2608.09887 v1 pith:7RHTY5V7 submitted 2026-08-10 cs.CV cs.CYcs.LG

classification cs.CVcs.CYcs.LG
keywords possessionqualityjunk-possessionindexexpectedthreatVAEPpitchcontrolspacecreationbroadcastvideoanalysisfootballanalytics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the most-cited football stat, possession percentage, hides a crucial off-ball distinction: whether holding the ball dragged the opponent's defensive block out of shape or simply circulated without threat. It builds a two-layer index that first flags "junk" possessions from event data—low-threat sequences in tied-or-losing game states—and then, for short flagged windows, uses broadcast video to measure whether the possession actually created space via a Space-Creation Index. The central quantitative claim is that the junk flag is not a repackaging of on-ball value: with team offensive VAEP and field tilt held fixed, the flag remains strongly negatively associated with points (p < $10^{-4}$) while VAEP is not significant (p = 0.34). The paper also reports that among 31 flagged windows from nine World Cup matches, 74% are spatially non-space-creating, 19% weak progression, and 6% space-creating—windows the event-only flag would score as failure. This matters because it offers a practical, broadcast-only route to measuring off-ball value at tournament scale.

What carries the argument

Two linked instruments carry the argument. First, the junk-possession index prices each possession sequence by value(s) = max(0, xT_max(s) − xT_start(s)) + 0.7·xG(s) + 0.1·1[box touch], normalized by a corpus-wide 90th percentile and clipped to [0,1]; sequences with q < 0.15 are flagged as junk, and the metric junk-open is the event share of such sequences in tied-or-losing states. Second, the Space-Creation Index (SCI) computes a net two-zone pitch-control change from broadcast video projected to pitch coordinates: SCI = Δ_own + Δ_opp, where Δ_own is the change in the possessing team's control of the attacking third and Δ_opp is the recession of the opponent's presence in the mirror zone. The event layer scans whole matches cheaply and plants flags; the spatial layer resolves why on short windows.

What would settle it

Run the same GSR pipeline on broadcast matches for which full commercial tracking data exist, compare the imputed off-screen player positions against true positions, and re-compute SCI verdicts on both; if the spatial classifications (non-creating vs weak vs space-creating) flip when imputed positions are perturbed or replaced with ground truth, the spatial layer's numbers are not trustworthy.

Watch

Extended reading notes

Core claim

The paper's core discovery is that the junk-possession flag—the on-ball-event share of low-threat possession sequences in tied-or-losing states—carries outcome-relevant variance that a standard on-ball action-value model (VAEP) does not capture on its own. In a same-match regression of 206 team-match rows from the 2026 FIFA World Cup, junk-open remains significantly negatively associated with points (p < $10^{-4}$, also under match-clustered errors) when team offensive VAEP and field tilt are held fixed, while VAEP is not significant (p = 0.34). The paper further shows that a spatial layer, computing a Space-Creation Index from pitch control on broadcast video, can adjudicate whether an event-flagged junk possession was spatially dead or space-creating, separating sterile domination (e.g., Germany's 73% possession in a penalty-shootout exit) from unlucky but genuinely space-creating play.

Load-bearing premise

The spatial layer's Space-Creation Index depends on the off-screen player imputation 'ghosting layer' from companion work, and the paper does not evaluate the accuracy of that imputation or filter windows on imputation sensitivity; if the ghosted positions are biased, the pitch-control estimates and the SCI verdicts (74% non-creating, etc.) could be unreliable.

Editorial extensions

If this is right

  • If the index is right, teams can diagnose sterile possession from event data alone and target the off-ball problem without full tracking data, which most leagues lack.
  • The junk flag behaves like a team trait: its value in a team's other matches predicts points and xG difference in a held-out match, so it could be used for pre-match assessment.
  • The two-layer design shows that event-only possession-value models miss a real axis of quality—space creation—and that the distinction is measurable from broadcast video.
  • The case evidence indicates that dominant ball share can coexist with elimination (Germany 73% possession, two non-creating windows), so possession-based narratives need spatial correction.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • One implication the author leaves implicit: a tracking-fed VAEP or EPV that includes defensive shape might subsume some of the junk flag's predictive content, since the non-reducibility test uses an event-fed VAEP; a natural extension would test whether the junk flag survives against tracking-based on-ball value on the same matches.
  • The 6% space-creating windows suggest event-only metrics systematically misclassify a small but consequential set of possessions—those that improve pitch control without producing a shot—so expected-threat and VAEP-style models could be augmented with a pitch-control term.
  • The macro/micro hybrid is directly transferable to other broadcast-only competitions (club leagues, national teams outside top tracking leagues), and the SCI thresholds (+12/+4) are operational values that a larger study could calibrate against actual scoring rates.
  • A testable extension would be to run the full pipeline on matches for which commercial tracking data exist, checking whether the imputed off-screen positions and SCI verdicts match the tracking-derived ground truth.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a two-layer possession-quality framework for broadcast football. The event-side junk-possession index segments each match into possession sequences, prices each by peak expected-threat gain plus a shot-xG term and a box-touch credit, normalizes by corpus-wide quantiles, and flags low-threat sequences in tied-or-losing game states. It reports that the resulting junk-open metric correlates negatively with points and xG difference, survives controls for field tilt and a self-implemented VAEP, and shows predictive associations in half-split and leave-one-match-out analyses. The spatial layer computes a Space-Creation Index (SCI) from broadcast video through a game-state-reconstruction pipeline with off-screen player imputation, and applies it to 31 flagged windows from nine World Cup matches, finding 74% non-space-creating, 19% weak progression, and 6% space-creating windows. The paper argues that the two layers together separate sterile from space-creating possession in a way that event-only on-ball value models cannot.

Significance. If the claims hold, the paper contributes a practical macro/micro hybrid for off-ball possession analysis from broadcast footage, an area where most event-based metrics are blind. The event-side index is simple, computable from public event feeds, and released with code; the paper is explicit about its limitations and reports honest nulls (e.g., the redeemed-junk refinement). The falsifiable outcome correlations and the transparent disclosure of corpus-wide fitted constants are strengths. However, the central quantitative claim of non-reducibility to on-ball value rests on an unvalidated VAEP implementation, and the spatial-layer results depend on an unevaluated off-screen imputation method; both are load-bearing for the paper's stated contributions.

major comments (4)
  1. [§4.4, Table 3] The central claim that the junk flag 'is not reducible to on-ball possession value' is supported only by a same-match OLS regression using a self-implemented VAEP whose own predictive validity is never reported. The paper does not report VAEP's univariate association with points or xG, nor any cross-fitted or out-of-sample VAEP comparison. If the implemented VAEP is too noisy, the conditional non-significance of VAEP (p=0.34) is expected, and the conclusion that the junk index 'carries outcome-relevant variance that this on-ball action-value model does not capture' is not established. The manuscript should validate the VAEP implementation (e.g., report its standalone correlation with outcomes and compare it against a published VAEP baseline) and, ideally, repeat the VAEP-controlled analysis in an out-of-sample or cross-fitted setting. As written, Section 7's own acknowledgment that junk-open is constructed from live score state makes the same-match regression descriptive, so the predictive weight falls on the splits in §4.5, which do not include VAEP—leaving the non-reducibility claim without a predictive test.
  2. [§4.4 and §7] The paper reports conditional significance of junk-open while VAEP is not significant, but it does not report a formal incremental-validity test, such as a nested-model comparison or a test that the junk-open coefficient differs from the VAEP coefficient. The conclusion that the index 'adds information beyond' VAEP requires demonstrating that adding junk-open to a model that already contains VAEP and tilt improves fit or that the coefficient difference is statistically significant; a non-significant VAEP coefficient in the joint model is not equivalent to such evidence. The manuscript should provide this test, preferably with match-clustered errors, and interpret it explicitly.
  3. [§5 and §7] The spatial-layer results—most importantly the 74% non-space-creating verdict across 31 flagged windows—depend on pitch-control estimates computed over off-screen player positions imputed by the 'ghosting layer' from companion work. The paper states that it did not filter windows on imputation sensitivity and does not evaluate the imputation's accuracy in this manuscript. If the ghosted positions are biased in even a subset of windows, the SCI classifications and the central off-ball message are unreliable. Since this is structurally different from the event-side validation, the manuscript should provide a sensitivity analysis (e.g., perturbing ghosted positions, comparing against a tracking-data subset, or reporting confidence intervals for SCI under imputation uncertainty) or at least quantify the potential impact on the 31-window classifications.
  4. [§4.5 and §7] The leave-one-match-out and half-split analyses are presented as the 'genuinely predictive evidence,' but the paper acknowledges that the q-normalizing 90th percentile, the q<0.15 threshold, and the box weight are corpus-wide constants, and the xG model is fit in-sample on the same corpus's shot–goal labels. This means the out-of-sample splits are cross-fitted only up to these global constants and that goal information enters the index weakly through the xG fit. The paper should state more explicitly how much of the reported out-of-sample correlation could be driven by these corpus-wide fitted components, and ideally report a version with the xG model trained out-of-fold for the split analyses.
minor comments (5)
  1. [Eq. 1] Equation (1) contains a corrupted placeholder ('bracehtipupleft/bracehtipdownright/...') that obscures the formula; the weights 0.7 and 0.10 should be presented in a clean, rendered form.
  2. [§4.1] The 'junk dominators' versus 'efficient dominators' comparison (n=8 vs n=63) is correctly labeled descriptive, but the text should avoid phrasing like 'the eight averaged 28% of the efficient group's goals' without emphasizing the wide uncertainty; a confidence interval or a nonparametric test would help contextualize the tiny sample.
  3. [Table 2] The xG and xG-difference columns are acknowledged as partly mechanically coupled, but a table footnote should explicitly repeat this caveat so readers do not misread the r=-0.51 coefficient as independent evidence.
  4. [§5] The SCI bins (SCI>=+12, +4<=SCI<+12, SCI<+4) are described as 'prespecified operational bins from earlier internal work,' but no reference or rationale is provided; because these bins determine the 74%/19%/6% percentages, the paper should either cite the internal work or provide a justification for the cutoffs.
  5. [§6] Table 5 would benefit from listing the reasons for the four excluded windows (three kit contrast, one clip quality) directly in the table caption or a footnote, rather than only in the text.

Circularity Check

2 steps flagged · score 4.0 of 10

Partial circularity: the 'out-of-sample' splits are cross-fitted up to corpus-wide constants and an in-sample xG model, the VAEP non-reducibility test is a same-match descriptive regression with an unvalidated self-implemented comparator, and the spatial layer leans on an unvalidated companion imputation pipeline; independent empirical content remains.

  1. fitted input called prediction [Section 7 (Limitations), applied to Section 4.5 out-of-sample claims]
    "the q-normalising 90th percentile, the q <0.15 threshold, and the learned box weight are all corpus-wide quantities, so the out-of-sample splits are cross-fitted only up to these global constants."

    Section 4.5 presents first-to-second-half and leave-one-match-out splits as the genuinely predictive evidence, but the index's normalisation constant, q threshold, and box weight are estimated from the entire corpus, and the embedded xG model is fit in-sample on the same corpus's shot-goal labels without an out-of-fold protocol. The 'held-out' matches and halves therefore enter the index construction through these fitted quantities, so the reported predictions are not strictly out-of-sample; the associations are partly re-descriptions of the fitting sample rather than independent forecasts.

  2. self citation load bearing [Section 5 (GSR pipeline), with limitation in Section 7]
    "Off-screen players are imputed with a training-free ghosting layer (companion work)."

    The spatial layer's central verdicts (74% non-creating, 19% weak progression, 6% space-creating) are computed from pitch-control values over imputed off-screen positions supplied by an unvalidated 'companion work' by the same author. Section 7 concedes 'we did not filter windows on imputation sensitivity.' The off-ball 'space-creating versus dead' distinction therefore rests on a self-referential, unevaluated imputation pipeline, so the spatial classifications are not independent of the author's own prior assumptions; the paper itself flags this only as a limitation, not as a validation gap in the load-bearing chain.

full rationale

The paper is unusually candid: Section 7 openly concedes the in-sample xG fit, corpus-wide normalisation constants, score-state feedback, and the descriptive nature of the same-match VAEP regression. I find no step in which the junk index is defined as the outcome it predicts, and the VAEP comparison is not mathematically forced by construction; it is a conditional association with an external, though self-implemented, benchmark. The main circularity-adjacent issue is statistical leakage: the index's global constants and xG model are fit on the same corpus used for the 'held-out' validation, so the predictive evidence is cross-fitted rather than cleanly out-of-sample. In addition, the spatial layer's load-bearing imputation is referenced only as 'companion work' and explicitly not sensitivity-filtered, so the off-ball verdicts rest on an unvalidated self-referential pipeline. These are partial circularity and leakage problems rather than full equivalence-by-construction; the central points correlation is not built into the index's formula, and the paper's explicit caveats prevent the derivation from collapsing entirely into its inputs. Score 4 reflects 'some self-citation and partial leakage; central claim still has independent content,' not the 6-8 range reserved for cases where the central prediction reduces by construction.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The central claims rest on several fitted constants and domain assumptions: the sequence value weights and thresholds are operational values partly determined by the same corpus, the xG model is fit in-sample, and the spatial layer depends on an unvalidated off-screen imputation method. No new physical entities are postulated.

free parameters (6)
  • Sequence value weight for xG term = 0.7
    Weight on the realised-shot xG term in Eq. 1; chosen as an operational value, not formally calibrated.
  • Sequence value weight for box touch = 0.10 (initially 0.03)
    Reward for any touch in the 18-yard box; raised from 0.03 to 0.10 after a fitted logistic model showed a large box-touch coefficient (Section 4.2).
  • Junk sequence threshold q = 0.15
    Sequences with q < 0.15 are flagged junk; an operational value.
  • SCI classification thresholds = +12 and +4
    Bins for space-creating (>=+12), weak progression (+4 to +12), and non-space-creating (<+4).
  • q normalization percentile = 90th percentile
    Sequence values are divided by the corpus-wide 90th percentile before clipping.
  • xG model hyperparameters = not specified
    Gradient-boosted xG model trained on the event corpus; specification released in the repository but not detailed in the paper.
assumptions (5)
  • domain assumption Expected threat grid (Karun-Singh-style) is a valid measure of pitch location scoring potential
    Used in Eq. 1 to price sequence threat gain; treated as given from prior work.
  • domain assumption Pitch-control model of Spearman et al. provides valid control shares
    SCI is computed from pitch-control shares; the paper consumes this model as a metric substrate.
  • ad hoc to paper Off-screen player imputation (ghosting layer) yields accurate enough positions
    The spatial layer relies on imputed positions for off-screen players; the method is from companion work and not validated here.
  • ad hoc to paper The in-sample gradient-boosted xG model is a valid shot-quality estimator
    Substituted for the constant placeholder shot-xG field; trained on the same corpus without out-of-fold protocol.
  • domain assumption Live scoreline reconstruction from goal events is accurate
    Gates junk-open by game state; relies on event feed goal data.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Space-Creating versus Dead Possession: An Off-Ball Possession-Quality Index for Broadcast Football." pith.science (2026). https://pith.science/paper/7RHTY5V7

@misc{pith2026260809887,
  author       = {Pith},
  title        = {Pith review of: Space-Creating versus Dead Possession: An Off-Ball Possession-Quality Index for Broadcast Football},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7RHTY5V7}},
  note         = {Machine review of arXiv:2608.09887}
}
read the original abstract

Ball possession is the most-cited and most-misleading number in football: 60% recycled in one's own half is not 60% spent pinning the opponent back. Existing event-based possession-value frameworks (expected threat, VAEP, on-ball value) price on-ball actions but ignore the off-ball question a sterile possession poses: did holding the ball create space, or was the circulation dead? We answer this in two layers. First, an event-side junk-possession index prices each possession sequence by its peak threat gain under an expected-threat grid and -- after reconstructing the live scoreline to exclude lead-protecting circulation -- flags low-threat sequences in tied-or-losing states. On the 2026 FIFA World Cup (103 matches, 206 team-matches) the flag correlates negatively with points (r=-0.37) and xG difference (r=-0.51, partly index-coupled). It is not a repackaging of on-ball value: with team offensive VAEP and field tilt held fixed, the junk flag stays strongly negatively associated with points (p<0.0001, also match-clustered) while VAEP is not significant -- in this same-match (descriptive) regression it adds information beyond this on-ball action-value model. Second, for a flagged window we resolve whether it was spatially dead or space-creating by projecting broadcast video to pitch coordinates and measuring a Space-Creation Index (SCI): a net pitch-control change capturing whether the possession seized space or pushed the opponent's block back. Across 31 of 35 flagged windows from nine World Cup matches (a purposive sample), 74% are spatially non-space-creating, 19% weak progression, and 6% space-creating windows the event flag alone would score as failure -- including a side with 73% of the ball that exited on penalties (two non-creating windows). The two layers separate space-creating-but-unconverted from sterile possession, a distinction event-only on-ball value cannot make.

Figures

Figures reproduced from arXiv: 2608.09887 by the authors.

Figure 1
Figure 1. Among the 71 team-match observations over 55% possession, splitting by corpus-wide [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Two flagged junk windows from the same match (Germany–Paraguay), opposite verdicts. [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

15 extracted references · 12 canonical work pages

  1. [1]

    BoT-SORT: Robust associations multi- pedestrian tracking.arXiv preprint arXiv:2206.14651, 2022

    Nir Aharon, Roy Orfaig, and Ben-Zion Bobrovsky. BoT-SORT: Robust associations multi- pedestrian tracking.arXiv preprint arXiv:2206.14651, 2022

  2. [2]

    TranSPORTmer: A Holistic Approach to Trajectory Understanding in Multi-Agent Sports

    Guillem Capellera, Luis Ferraz, Antonio Rubio, Antonio Agudo, and Francesc Moreno-Noguer. TranSPORTmer: A holistic approach to trajectory understanding in multi-agent sports. In Asian Conference on Computer Vision (ACCV), 2024. arXiv:2410.17785

  3. [3]

    Actions speak louder than goals: Valuing player actions in soccer

    Tom Decroos, Lotte Bransen, Jan Van Haaren, and Jesse Davis. Actions speak louder than goals: Valuing player actions in soccer. InProceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 1851–1861, 2019

  4. [4]

    Inferring Player Location in Sports Matches: Multi-Agent Spatial Imputation from Limited Observations

    Gregory Everett, Ryan J. Beal, Tim Matthews, Joseph Early, Timothy J. Norman, and Sarva- pali D. Ramchurn. Inferring player location in sports matches: Multi-agent spatial imputation from limited observations. InProceedings of the 22nd International Conference on Autonomous Agents and Multiagent Systems (AAMAS), pages 1643–1651, 2023. arXiv:2302.06569

  5. [5]

    Falaleev and Ruilong Chen

    Nikolay S. Falaleev and Ruilong Chen. Enhancing soccer camera calibration through keypoint exploitation. In7th ACM International Workshop on Multimedia Content Analysis in Sports (MMSports), 2024

  6. [6]

    A framework for the fine-grained evaluation of the instantaneous expected value of soccer possessions.Machine Learning, 110(6):1389–1427, 2021

    Javier Fernández, Luke Bornn, and Daniel Cervone. A framework for the fine-grained evaluation of the instantaneous expected value of soccer possessions.Machine Learning, 110(6):1389–1427, 2021. 10

  7. [7]

    PnLCalib: Sports field registration via points and lines optimization.arXiv preprint arXiv:2404.08401, 2024

    Marc Gutiérrez-Pérez and Antonio Agudo. PnLCalib: Sports field registration via points and lines optimization.arXiv preprint arXiv:2404.08401, 2024

  8. [8]

    Real time quantification of dangerousity in football using spatiotemporal tracking data.PloS one, 11(12):e0168768, 2016

    Daniel Link, Steffen Lang, and Philipp Seidenschwarz. Real time quantification of dangerousity in football using spatiotemporal tracking data.PloS one, 11(12):e0168768, 2016

Show all 15 references
  1. [9]

    Connor, Paul Muller, Natalie Mackraz, Kris Cao, Pol Moreno, Pablo Sprechmann, Demis Hassabis, Ian Graham, William Spearman, Nicolas Heess, and Karl Tuyls

    Shayegan Omidshafiei, Daniel Hennes, Marta Garnelo, Zhe Wang, Adria Recasens, Eugene Tarassov, Yi Yang, Romuald Elie, Jerome T. Connor, Paul Muller, Natalie Mackraz, Kris Cao, Pol Moreno, Pablo Sprechmann, Demis Hassabis, Ian Graham, William Spearman, Nicolas Heess, and Karl T...

  2. [10]

    Not all passes are created equal: Objectively measuring the risk and reward of passes in soccer from tracking data

    Paul Power, Hector Ruiz, Xinyu Wei, and Patrick Lucey. Not all passes are created equal: Objectively measuring the risk and reward of passes in soccer from tracking data. InProceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, page...

  3. [11]

    Introducing expected threat (xt)

    Karun Singh. Introducing expected threat (xt). 2019. https://karun.in/blog/ expected-threat.html

  4. [12]

    SoccerNet game state reconstruction: End-to-end athlete tracking and identification on a minimap

    Vladimir Somers, Victor Joos, Anthony Cioppa, Silvio Giancola, Seyed Abolfazl Ghasemzadeh, Floriane Magera, Baptiste Standaert, Amir Mohammad Mansourian, Xin Zhou, Shohreh Kasaei, Bernard Ghanem, Alexandre Alahi, Marc Van Droogenbroeck, and Christophe De Vleeschouwer. SoccerNe...

  5. [13]

    Beyond expected goals

    William Spearman. Beyond expected goals. InMIT Sloan Sports Analytics Conference, 2018

  6. [14]

    Physics-based modeling of pass probabilities in soccer

    William Spearman, Austin Basye, Greg Dick, Ryan Hotovy, and Paul Pop. Physics-based modeling of pass probabilities in soccer. InMIT Sloan Sports Analytics Conference, 2017

  7. [15]

    On-ball value (obv), 2021

    StatsBomb. On-ball value (obv), 2021. https://statsbomb.com/articles/soccer/ introducing-on-ball-value-obv/. 11

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.