Pith. sign in

REVIEW 2 major objections 5 minor 1 cited by

Dyadic data with ordered outcome variables

T0 review · 2 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper proves that a pooled tetrad logit estimator consistently estimates $\beta_0$ in ordered dyadic models with category-specific sender and receiver fixed effects, under only pooled identification across outcome categories.

desk verdict Solid, useful extension of tetrad-differencing CML to ordered dyadic outcomes; PTLE is a real contribution, but the inference section is conjectural and the abstract overstates the empirical contrast. read the letter →

arxiv 2507.16689 v1 pith:62STQ46Q submitted 2025-07-22 econ.EM

classification econ.EM MSC 62P2062F12
keywords orderedlogitdyadicdatanetworkformationfixedeffectsincidentalparameterstetraddifferencingconditionalmaximumlikelihoodhomophily
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper shows that ordered relationship data between pairs of agents can be estimated without estimating the many sender and receiver effects as parameters. The device is to binarize every outcome threshold and compare two senders against two receivers in a four-node tetrad; for a given threshold, the conditional probability of a contrasting tetrad pattern is logistic and free of all fixed effects. The paper proposes two estimators that aggregate these tetrad contributions across thresholds: an equally weighted version (ETLE) and a pooled version (PTLE). The central result is that PTLE is consistent under the weak condition that the thresholds together supply enough identifying information, even when no single outcome category does. This matters because networks are often sparse in high categories such as "best friendship," where equal weighting and single-cutoff binary methods become unstable.

What carries the argument

The machinery is tetrad-differencing conditional maximum likelihood. For each threshold $m$, define $D_{ij}(m)=1\{Y_{ij}\ge m\}$ and form the tetrad statistic $Z_\sigma(m)=\tfrac12\big((D_{i_1j_1}(m)-D_{i_1j_2}(m))-(D_{i_2j_1}(m)-D_{i_2j_2}(m))\big)$. Tetrads with $Z_\sigma(m)=\pm 1$ are informative, and their conditional probability is $\Lambda(r_\sigma'\beta_0)$, where $r_\sigma=(X_{i_1j_1}-X_{i_1j_2})-(X_{i_2j_1}-X_{i_2j_2})$. The additive decomposition $\lambda^*_{ijm}=\lambda_{im}+\delta_{jm}$ makes the fixed effects cancel in the tetrad odds ratio, and cancellation requires the same cutoff $m$ on all four dyads. PTLE pools every informative tetrad-threshold pair into one binary logit objective; ETLE instead normalizes each threshold's contribution by its own number of informative tetrads, which is what makes ETLE vulnerable to category-specific sparsity.

What would settle it

Simulate the ordered dyadic model with $\lambda^*_{ijm}=\lambda_{im}+\delta_{jm}+\rho_{ijm}$ for nonzero dyad-specific $\rho_{ijm}$; if the PTLE estimates do not track $\beta_0$ as $N$ grows, the sufficiency theorem is specific to the additive class. In the additive class, a design with $Np_N\to\infty$ but $Np_{mN}\not\to\infty$ for every threshold $m$ should show PTLE centering on $\beta_0$ while ETLE drifts.

Watch

Extended reading notes

Core claim

The core claim is that the incidental-parameter problem in ordered network logit models can be solved by extending tetrad-differencing conditional maximum likelihood from binary to ordered outcomes. After writing $D_{ij}(m)=1\{Y_{ij}\ge m\}$ and imposing the additive threshold structure $\lambda^*_{ijm}=\lambda_{im}+\delta_{jm}$, conditioning on informative tetrads gives $P(Z_\sigma(m)=1\mid Z_\sigma(m)\in\{-1,1\}, X_\sigma)=\Lambda(r_\sigma'\beta_0)$, with no fixed effects appearing. The pooled tetrad logit estimator, which maximizes the sum of these conditional log-likelihood contributions over all informative tetrad-threshold pairs, is proved to converge in probability to $\beta_0$ under Assumptions 1 through 4 and Assumption 6, requiring only that the pooled information $Np_N$ diverge and that the pooled Hessian have full rank. By contrast, the equally weighted estimator requires Assumption 5, namely that each cutoff individually provide enough information, which fails when some outcome categories are rare.

Load-bearing premise

The load-bearing premise is that every dyad's threshold is additively separable, $\lambda^*_{ijm}=\lambda_{im}+\delta_{jm}$, with no dyad-specific interaction: if real thresholds contain terms such as $\rho_{ijm}$ that depend on the specific pair, the fixed effects no longer cancel in the tetrad odds ratio and the estimator can be inconsistent.

Editorial extensions

If this is right

  • Researchers can estimate ordered network formation models without estimating the $2NM$ incidental sender and receiver effects, so sparse networks and rare outcome categories no longer force a binary cutoff choice.
  • The pooled estimator PTLE should be preferred over the equally weighted ETLE in applications, since PTLE remains consistent even when identification comes only from pooling information across thresholds.
  • Single-cutoff binary tetrad estimators are inefficient and highly sensitive to which cutoff is selected; pooling all informative tetrad-cutoff pairs avoids discarding information from the rest of the outcome distribution.
  • The dyad-clustered sandwich variance estimator corrects the severe under-coverage that naive standard errors produce in both dense and sparse networks, making the method usable for inference.
  • In the Dutch friendship application, the fixed-effects-adjusted PTLE finds significant positive homophily in gender, smoking, and academic program, whereas methods without fixed effects can produce counterintuitive negative signs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An extension the paper leaves implicit is that the same tetrad-pooling logic should transfer to other ordered link-strength scales, such as alliance levels or rating data, whenever the additive threshold decomposition holds.
  • A natural refinement the paper does not derive is an information-weighted pooling scheme that reweights tetrad-cutoff pairs by their precision, which could improve finite-sample efficiency beyond PTLE's implicit weighting by number of informative tetrads.
  • If empirically relevant thresholds contain dyad-specific interactions, the sufficiency result fails, so applied users should treat the additive decomposition as a substantive assumption rather than a normalization.
  • Because PTLE only requires pooled information across categories, it opens the door to distribution-regression or semi-parametric versions in which coefficients vary by threshold, a direction the paper mentions only through its discussion of concurrent work.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This paper develops estimation methods for ordered logit models of directed dyadic/network data with category-specific sender and receiver fixed effects. The authors extend tetrad-differencing conditional maximum likelihood from binary network models to ordered outcomes by binarizing at each threshold. They propose two estimators: ETLE, which weights each threshold equally, and PTLE, which pools informative tetrad-threshold pairs. Theorem 1 derives a conditional probability that eliminates the fixed effects under an additive threshold structure. Theorem 2 claims consistency of ETLE under Assumption 5 (sufficient information at each threshold) and of PTLE under the weaker Assumption 6 (sufficient pooled information). The paper includes Monte Carlo simulations, an empirical application to friendship networks, and extensions to more restrictive fixed-effects structures.

Significance. If the results hold, the paper offers a practical and novel solution to a relevant problem: estimating ordered network formation models with flexible fixed effects under sparsity and rare outcome categories. The PTLE's consistency under weaker pooled-information conditions is a useful theoretical contribution, and the simulation evidence supports the practical preference for PTLE. The empirical application is illustrative, though based on a small network. However, the main consistency proof has a technical gap, and the inference results are explicitly conjectural; these issues currently limit the strength of the paper's claims.

major comments (2)
  1. [Appendix A.2, proof of Theorem 2] The proof asserts E[(Σ_{m,σ} ℓ̃_{mσ})²] = O(N³ q_N p_N) after only deriving the per-tetrad bound E[ℓ̃²_{mσ}] = O(1). From the displayed inequalities and this per-tetrad bound, the best that follows is O(N³ M q_N) = O(N³ q_N), not O(N³ q_N p_N). The sharper order is essential: with only O(N³ q_N), the variance of the scaled PTLE objective is O(1/(N p_N²)), which need not vanish under Assumption 6 when p_N → 0 (e.g., p_N ~ 1/√N). The same issue affects the ETLE proof in part (a), where the analogous assertion for each cutoff m would require E[ℓ̃²_{mσ}] = O(p_{mN}). The authors should either prove a sharper per-tetrad second-moment bound under the stated assumptions (for instance, using bounded covariates or a fourth-moment condition with an additional argument) or revise the assumptions and proof so that the claimed variance rate follows.
  2. [Section 4.2] The asymptotic normality of PTLE and the consistency of the sandwich variance estimator (Eq. 18) are presented only as a conjecture, with 'suitable regularity conditions' left unspecified. Despite this, the empirical application in Section 6 reports standard errors and significance levels based on Eq. (18), and Section 5.3 evaluates its coverage properties. Without a formal theorem establishing the asymptotic distribution, the inferential claims in the empirical section are not rigorously justified. The authors should either provide a complete asymptotic normality result with explicit conditions (e.g., the sixth-moment condition mentioned) or clearly label the standard errors as heuristic and temper the conclusions drawn from them.
minor comments (5)
  1. [Section 4.2 and Section 4.3] The sandwich estimator Υ_N sums over all ordered dyads (i,j) of v_{ij} v_{ij}', where each tetrad contributes to four different dyad-level sums. The paper does not explain whether or how this double-counting is accounted for in the sandwich formula, which would be helpful for readers implementing the method.
  2. [Section 4.3] The computation requires enumerating all q_N = O(N^4) tetrads, which becomes prohibitive for networks much larger than the N=32 example. The authors should discuss computational feasibility or possible subsampling schemes for larger networks.
  3. [Section 5.2] In the heterogeneous-threshold simulation, the thresholds are defined using node-specific components λ_{im} = δ_{im} (imposing symmetry between sender and receiver effects), but the main model allows λ_{im} and δ_{jm} to differ. The text should clarify that the simulation only covers a symmetric special case and does not explore asymmetric sender/receiver effects.
  4. [Table 14] The standard errors for the Binary (3) column appear misaligned; for instance, the Common gender row shows five parenthetical values for six coefficient columns. The table should be reformatted for clarity.
  5. [Abstract and Section 6.2] The abstract states that 'standard methods without fixed effects produce counterintuitive results,' but the ordered logit without fixed effects in Table 14 yields positive coefficients for all three homophily variables, which are not counterintuitive. The claim should be reconciled with the reported results or clarified.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the conditional-logit derivation is algebraic and the consistency proof builds on an external benchmark.

full rationale

Theorem 1's conditional probability Λ(r′_σ β_0) is obtained by direct algebra from Assumption 1: the proof computes the odds ratio P(Z_σ(m)=+1)/P(Z_σ(m)=−1) and shows that the fixed-effects term Δλ(m) vanishes exactly when m_{11}=m_{12}=m_{21}=m_{22}=m. This is a derivation, not an assumption of the result, and it does not import any conclusion from the authors' prior work. Theorem 2's consistency argument is modeled on Jochmans (2018), an external benchmark, and Assumptions 5 and 6 are explicit growth and rank conditions rather than restatements of the consistency claim. The self-citations to Muris (2017), Abrevaya and Muris (2020), and Botosaru et al. (2023) appear in Section 7 as contextual references to related fixed-effects ordered-choice strategies; Theorems 3 and 4 are proved in the appendix by the same direct odds-ratio algebra, so the citations are not load-bearing. The paper openly conjectures the asymptotic normality result and does not provide a formal proof of the sandwich variance estimator, but this is an acknowledged incompleteness, not a circular step. No fitted parameter is renamed as a prediction, and no uniqueness theorem is imported from the authors' own work. The possible moment-bound issue in the proof of Theorem 2 (the asserted O(N^3 q_N p_N) variance) is a question of proof rigor, not circularity: the bound follows from E[S_{σm}]=p_{σm} and the boundedness of the logit contribution, and in any event the claim does not reduce to its inputs by construction. The paper is therefore self-contained with respect to its central identification and consistency claims.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No free parameters are fitted; the incidental fixed effects are eliminated by construction. The proofs use standard logistic algebra, i.i.d. node sampling, compactness, and sparsity-growth conditions, none of which are ad hoc to this paper.

assumptions (5)
  • domain assumption Conditional independence and logistic errors (Assumption 1)
    The likelihood is a product over dyads of logistic ordered probabilities; this is the primitive for Theorem 1.
  • domain assumption Independent and identically distributed node sampling (Assumption 2)
    Controls the dependence structure of overlapping tetrads in the consistency proof.
  • domain assumption Additive threshold decomposition lambda*_ijm = lambda_im + delta_jm (Section 2)
    Makes the fixed effects cancel in tetrad differencing; if false, the central identification result fails.
  • standard math Compact parameter space and bounded moments (Assumptions 3 and 4)
    Regularity conditions for M-estimator consistency.
  • domain assumption Pooled informative tetrad growth N p_N to infinity and full-rank pooled Hessian (Assumption 6)
    Needed for PTLE consistency; fails in extremely sparse networks.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Dyadic data with ordered outcome variables." pith.science (2026). https://pith.science/paper/62STQ46Q

@misc{pith2026250716689,
  author       = {Pith},
  title        = {Pith review of: Dyadic data with ordered outcome variables},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/62STQ46Q}},
  note         = {Machine review of arXiv:2507.16689}
}
read the original abstract

We consider ordered logit models for directed network data that allow for flexible sender and receiver fixed effects that can vary arbitrarily across outcome categories. This structure poses a significant incidental parameter problem, particularly challenging under network sparsity or when some outcome categories are rare. We develop the first estimation method for this setting by extending tetrad-differencing conditional maximum likelihood (CML) techniques from binary choice network models. This approach yields conditional probabilities free of the fixed effects, enabling consistent estimation even under sparsity. Applying the CML principle to ordered data yields multiple likelihood contributions corresponding to different outcome thresholds. We propose and analyze two distinct estimators based on aggregating these contributions: an Equally-Weighted Tetrad Logit Estimator (ETLE) and a Pooled Tetrad Logit Estimator (PTLE). We prove PTLE is consistent under weaker identification conditions, requiring only sufficient information when pooling across categories, rather than sufficient information in each category. Monte Carlo simulations confirm the theoretical preference for PTLE, and an empirical application to friendship networks among Dutch university students demonstrates the method's value. Our approach reveals significant positive homophily effects for gender, smoking behavior, and academic program similarities, while standard methods without fixed effects produce counterintuitive results.

Figures

Figures reproduced from arXiv: 2507.16689 by the authors.

Figure 1
Figure 1. Mean degrees across λ2 by cutoff category for CN = 0 and CN = log(N1/2 ). 23 [PITH_FULL_IMAGE:figures/full_fig_p023_1.png] view at source ↗
Figure 2
Figure 2. Comparison of mean estimates across values of the maximum threshold, [PITH_FULL_IMAGE:figures/full_fig_p024_2.png] view at source ↗
Figure 3
Figure 3. Mean PTLE and ETLE estimates by maximum threshold [PITH_FULL_IMAGE:figures/full_fig_p030_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Estimated coefficients across time points for (a) [PITH_FULL_IMAGE:figures/full_fig_p038_4.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Pairwise Differencing Distribution Regression Approach for Network Models

    econ.EM 2026-08 conditional novelty 6.0 of 10

    A conditional maximum likelihood estimator for distribution regression in dyadic networks with two-way fixed effects is developed, with joint inference across thresholds.

Reference graph

Works this paper leans on

14 extracted references · 14 canonical work pages · cited by 1 Pith paper

  1. [1]

    Das, M., & van Soest, A. (1999). A panel data model for subjective information on household income growth. Journal of Economic Behavior & Organization, 40 (4), 409–426

  2. [2]

    Lai, B., & Reiter, D. (2000). Democracy, Political Similarity, and International Al- liances, 1816-1992. The Journal of Conflict Resolution , 44 (2), 203–227

  3. [3]

    Johnson, E. (2004). Panel Data Models with Discrete Dependent Variables [Doctoral dissertation, Stanford University]. Van Duijn, M. A., Gile, K. J., & Handcock, M. S. (2009). A framework for the com- parison of maximum pseudo-likelihood and maximum likelihood estimation of exponential family random graph models. Social Networks, 31 (1), 52–62

  4. [4]

    Baetschmann, G. (2012). Identification and estimation of thresholds in the fixed effects ordered logit model. Economics Letters, 115 (3), 416–418

  5. [5]

    E., & Winkelmann, R

    Baetschmann, G., Staub, K. E., & Winkelmann, R. (2015). Consistent estimation of the fixed effects ordered logit model. Journal of the Royal Statistical Society A , 178 (3), 685–703

  6. [6]

    Charbonneau, K. B. (2017). Multiple fixed effects in binary response panel data models. The Econometrics Journal , 20 (3), S1–S13

  7. [7]

    Graham, B. S. (2017). An econometric model of network formation with degree hetero- geneity. Econometrica, 85 (4), 1033–1063

  8. [8]

    Muris, C. (2017). Estimation in the Fixed-Effects Ordered Logit Model. Review of Economics & Statistics , 99 (3), 465–477

Show all 14 references
  1. [9]

    Jochmans, K. (2018). Semiparametric Analysis of Network Formation. Journal of Busi- ness & Economic Statistics , 36 (4), 705–713. 52

  2. [10]

    Dzemski, A. (2019). An Empirical Model of Dyadic Link Formation in a Network with Unobserved Heterogeneity. The Review of Economics and Statistics , 101 (5), 763–776

  3. [11]

    Abrevaya, J., & Muris, C. (2020). Interval censored regression with fixed effects.Journal of Applied Econometrics , 35 (2), 198–216

  4. [12]

    Botosaru, I., Muris, C., & Pendakur, K. (2023). Identification of time-varying transfor- mation models with fixed effects, with an application to unobserved heterogene- ity in resource shares. Journal of Econometrics , 232 (2), 576–597

  5. [13]

    Hughes, D. W. (2023, March). Estimating Nonlinear Network Data Models with Fixed Effects

  6. [14]

    Szini, G. M. M. (2025). A Pairwise Differencing Distribution Regression Approach for Network Models. 53

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.