{"id":"cb588b89-aa0a-4f1e-bde5-da878fda0469","arxiv_id":"2602.18201","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"High-capacity self-organizing maps recover age and income orderings from unsupervised tabular data with Spearman correlations up to 0.85, but the comparison is weakened by feature selection that uses the withheld attributes.","lead":"A new audit method built on self-organizing maps claims that purely unsupervised embeddings can strongly encode withheld details like age and income, with correlation up to 0.85, far above PCA, UMAP, t-SNE, and autoencoders. But the paper uses the withheld attributes to pick which features to keep, so the 'withheld' claim doesn't hold as stated.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Sensitive attribute is used in preprocessing, contradicting the 'withheld' premise; reported leakage may not be emergent.","rationale":"The paper's central claim is that withheld sensitive attributes emerge in purely unsupervised SOM embeddings, with correlations far exceeding baselines. For this to be a demonstration of emergent leakage, the sensitive attribute must not influence any step of the representation pipeline. The paper's own §4.2 explicitly violates this: Census features were removed based on their correlation with age/income, and WVS question selection appears to be guided by Table 1 correlations ('justifying our choice for their inclusion'). This directly contradicts §3.1's claim that sensitive attributes are used only for post-hoc auditing. The reader's weakest_assumption identifies exactly this concern, and I agree. The concern is load-bearing because if feature selection uses s, the high correlations could be an artifact of choosing features that are individually predictive of s; the claim that the SOM topology itself amplifies sensitive signal is not isolated. Even though removing correlated columns (Census) makes recovery harder, the protocol is not a clean test of 'withheld'; the investigator's knowledge of s is baked into the input. A concrete re-run with s-agnostic preprocessing would settle whether the effect is real. Credit: the paper provides code and averages over runs, but that does not address the preprocessing contamination. Thus the current REJECT verdict stands; no change needed.","tokens_in":15014,"tokens_out":6055,"duration_ms":52003,"concrete_test":"Re-run the full SOMtime and baseline pipelines on (a) the original 40-column Census-Income dataset without any column removal based on age/income correlations, and (b) a pre-registered WVS question set chosen without reference to age (e.g., all ethics/morals questions, or a random subset of them). If SOMtime's Spearman correlations with age/income remain above ~0.5 under both conditions, the preprocessing is not the driver; if they fall to baseline levels, the reported leakage is an artifact of sensitive-attribute-informed feature selection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that sensitive attributes be withheld from every step that shapes the representation. §3.1 states 'Sensitive attributes are used only for post-hoc auditing,' but §4.2 directly contradicts this: for Census, columns with strong correlations to age/income were removed because they would make recovery 'trivial,' and for WVS, Table 1 reports correlations of selected questions with age as 'justifying our choice for their inclusion.' This means the input feature set is chosen using knowledge of the sensitive attribute. Even if removing correlated columns makes the task harder, the representation is not purely unsupervised with respect to s—the investigator's knowledge of s has shaped the input. The paper's own admission 'We only ever use knowledge of the sensitive feature in this context' (preprocessing) violates the stated 'withheld' condition. Without a pre-registered, s-agnostic feature selection, the headline comparison (SOMtime 0.85 vs baselines <0.23) may reflect feature engineering rather than emergent leakage, and the conclusion that fairness-through-unawareness fails at the representation level is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SOMtime, an auditing method based on high-capacity Self-Organizing Maps, and claims that purely unsupervised topology-preserving embeddings can recover withheld ordinal sensitive attributes (age, income, capital gains) as monotonic latent axes. On the WVS and Census-Income datasets, the authors report Spearman correlations up to 0.85 for SOMtime versus at most 0.34 for PCA, UMAP, t-SNE, and autoencoders. They conclude that fairness-through-unawareness fails at the representation level and that unsupervised representations should be routinely audited. The paper includes a trajectory-extraction algorithm (Algorithm 1) and reports recovery accuracy of age-group orderings.","tokens_in":15285,"tokens_out":2526,"duration_ms":24807,"significance":"If the central claim were established, the paper would make a useful contribution to fairness auditing by showing that a classical topology-preserving method can expose global monotonic structure aligned with sensitive attributes, a form of leakage that correlation-based probing may miss. The paper also provides code and reports stability across five runs with low standard deviations, which are strengths. However, the validity of the headline comparison rests on two conditions: that the sensitive attribute is truly withheld from every step shaping the representation, and that the evaluation is symmetric across methods. Both conditions are violated in the current manuscript, so the significance of the empirical findings is not yet demonstrated.","major_comments":[{"comment":"The central premise that sensitive attributes are 'withheld from all representation learning' is contradicted by the dataset preprocessing. §3.1 states that sensitive attributes are used only for post-hoc auditing, but §4.2 describes removing Census columns with strong correlations to age/income to avoid 'trivial' recovery, and Table 1 justifies WVS question selection by correlation with age ('justifying our choice for their inclusion'). This means the input feature set was chosen using knowledge of the sensitive attribute. Consequently, the reported SOMtime correlations may reflect feature engineering rather than emergent leakage from purely unsupervised learning, and the abstract's claim that sensitive attributes 'emerge ... even when explicitly excluded from the input' is not supported.","section":"§4.2, §3.1"},{"comment":"The evaluation is asymmetric. The Table 2 caption states 'Best values from ablations used,' meaning baseline hyperparameters were selected to maximize their reported correlations, while SOMtime uses fixed hyperparameters. Additionally, §4.3 states that for SOMtime only the Z-axis (activation) is examined, whereas for the other methods the maximum over every embedding dimension is reported. Since Table 2 reports 'maximum correlation between any single embedding axis,' the comparison is not apples-to-apples: baselines are given the advantage of 50 axes plus ablated hyperparameters, while SOMtime is restricted to one axis. This undermines the quantitative claim of a 3–8× improvement.","section":"§4.3, Table 2 caption"},{"comment":"The method description is internally inconsistent about which axis is used. §3.5 says the maximum absolute correlation across the three embedding dimensions is reported, but Experiment 2 (§5.2) says for SOMtime 'only looking at the Z-axis,' while baselines use the 'strongest principal component.' The reader cannot determine whether Table 2 reports the max over (x,y,z) for SOMtime or only z, nor how the 'dominant 1D axis' was chosen for baselines. This needs clarification and a consistent protocol.","section":"§3.5, §5.2"},{"comment":"The abstract and discussion claim that 'unsupervised segmentation of SOMtime embeddings produces demographically skewed clusters,' but no clustering experiment or quantitative cluster-skew result appears in the paper. Table 3 reports recovery of age-group ordering edges, not cluster demographic skew. This claim is unsupported and should either be removed or backed by an actual clustering experiment with a measurable imbalance metric.","section":"Abstract, §5.3, §6"}],"minor_comments":[{"comment":"The description of the adjacency-graph construction is hard to follow: 'values of the graph are boolean identifiers of edge existence from row to column vertices' is vague, and the scaling operation in lines 14–15 uses sums over centroids without explaining why this normalization is appropriate. Pseudocode would benefit from concrete notation and a worked example.","section":"§3.4, Algorithm 1"},{"comment":"Reference [31] has a malformed author list ('Yutaka Matsuo Yusuke Iwasawa, Kotaro Nakayama' appears to be three authors concatenated). Please correct.","section":"References"},{"comment":"The figure shows a qualitative visualization of age ordering but does not include a colorbar legend or axis labels. Since the paper emphasizes quantitative correlation, a version with the SOM grid coordinates and the recovered trajectory overlaid would be more informative.","section":"Figure 2"},{"comment":"The sentence 'A selection criteria implemented here for determining questions to be used was that the data does not contain missing values' has subject-verb agreement errors and is unclear; presumably it means 'questions were selected for having no missing values.'","section":"§4.1"},{"comment":"The limitation paragraph correctly notes that only ordinal attributes are studied, but it does not mention the more serious limitation that feature selection used the sensitive attribute; this should be acknowledged.","section":"Section 7.1"}],"recommendation":"reject","confidential_remarks":"The paper has a plausible core idea—using SOMs as an auditing tool for global monotonic leakage—but the current execution invalidates the headline result. The §4.2 admission that feature selection used the sensitive attribute, combined with the asymmetric evaluation in Table 2, means the central claim is not established. A full re-run with a pre-registered, sensitive-attribute-agnostic feature-selection protocol and a symmetric evaluation would be needed to make the result credible. I see no evidence of deliberate misconduct, but the current manuscript's framing overstates what the experiments can support. The KDD 2025 submission date with an arXiv 2026 listing is also a metadata inconsistency worth checking."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is about a genuinely interesting tool: SOMtime, a high-capacity SOM plus a hand-specified trajectory-adjacency algorithm, which recovers global monotonic orderings of withheld ordinal attributes (age, income) better than PCA, UMAP, t-SNE, or autoencoders on WVS and Census data. That specific method and the empirical setup are new, and the autoencoder capacity ablation is a nice control: more capacity does not by itself explain the gap. The paper is also honest in places, admitting it doesn't know whether SOMs preferentially amplify sensitive structure or just preserve more structure overall, and limiting its claims to ordinal attributes and tabular data.\n\nBut the central claim is not established as stated. The paper says in §3.1 that sensitive attributes are used only for post-hoc auditing, but §4.2 openly says the opposite: correlated Census columns were removed because they would make recovery \"trivial,\" and WVS questions were selected based on their correlation with age. That means the input feature set itself was shaped by knowledge of the sensitive attribute. The representation is not purely unsupervised with respect to s, so the \"withheld\" premise fails. This is not a minor quibble—the 0.85 vs 0.23 comparison is the whole paper. The stress-test note is right.\n\nThere are other soft spots. Table 2's caption \"best values from ablations used\" means baselines were tuned, but SOMtime's hyperparameters are fixed; that tilts the comparison. For baselines the reported value is the max over all axes, while SOMtime is reported only on the Z-axis, which was chosen after seeing the data. The abstract claims \"demographically skewed clusters,\" but I can't find any experiment in the text that measures cluster demographic skew. That claim is simply unsupported. Table 3's \"recovery accuracy\" is an unusual edge-recovery metric, though the Spearman correlations carry the argument anyway.\n\nThis is not a fatal idea. The phenomenon is plausible, and the method is concrete. But the current protocol cannot support the conclusion that fairness through unawareness fails at the representation level. A revised version with a genuinely s-agnostic feature selection, a pre-registered axis, and a real cluster-skew analysis could change the verdict.\n\nWho is this for? Fairness auditors and anyone building on unsupervised embeddings. It deserves a serious referee—not a desk reject—but the referee should treat the current version as a preliminary report, not an established result. My private verdict: reject in current form, but encourage a major revision along the lines above.","headline":"The SOMtime idea is worth a referee, but the headline claim is undermined by the paper's own use of the sensitive attribute in preprocessing, and the abstract promises a cluster-skew result the paper never reports.","tokens_in":15761,"tokens_out":2699,"would_cite":false,"duration_ms":73590,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Purely unsupervised self-organizing maps recover withheld age and income as dominant embedding axes, with Spearman correlations up to 0.85.","keywords":["fairness through unawareness","self-organizing maps","sensitive attribute leakage","unsupervised representation learning","global geometric ordering","dimensionality reduction","representation auditing","ordinal sensitive attributes"],"falsifier":"Train SOMtime on data where the sensitive attribute is genuinely independent of all input features, or on synthetic data with known generative structure, with no feature selection or preprocessing informed by the attribute. If the Spearman correlation between the embedding axes and the withheld attribute stays near zero, the claim that SOMs systematically expose such attributes would be refuted. Alternatively, rerun the Census experiment on the full 40 columns without removing age- and income-correlated features; if correlations do not rise above baseline, the reported gap depends on preproces","tokens_in":14868,"feed_emoji":"🗺️","tokens_out":3920,"duration_ms":37022,"temperature":0.7,"pith_summary":"This paper claims that purely unsupervised, topology-preserving representations—specifically high-capacity self-organizing maps—arrange data along monotonic axes aligned with sensitive attributes such as age and income, even when those attributes are withheld from training. On the World Values Survey and Census-Income datasets, the method achieves Spearman correlations up to 0.85, while PCA, UMAP, t-SNE, and autoencoders stay below 0.34. This would mean the common assumption of fairness through unawareness fails at the representation level: any downstream use of such an embedding, whether clustering, visualization, or feature extraction, inherits demographic skew before a supervised task exists. The paper positions SOMtime as an audit tool that reveals this leakage, not as a fairness mitigation.","feed_headline":"SOMs leak withheld age and income at 0.85 correlation","feed_subtitle":"Withheld age and income emerge as dominant axes in self-organizing maps, so fairness audits must include embeddings.","key_machinery":"The central object is a self-organizing map (SOM): a square lattice of prototype vectors trained by competitive learning with a Gaussian neighborhood function, scaled to high capacity with lattice size K = 5·N^0.54. Each observation is embedded in three dimensions using its best-matching unit's (x, y) lattice coordinates plus its quantization error z = ||x - w_bmu||, the distance to its prototype. A trajectory-adjacency algorithm links cluster centroids in this 3D space to recover the path along which a sensitive attribute is monotonically ordered, and leakage is quantified as the maximum absolute Spearman correlation between any single embedding axis and the withheld attribute.","core_discovery":"On two real-world tabular datasets, with age, income, and capital gains withheld from all training, a high-capacity self-organizing map arranges observations so that the withheld attributes vary monotonically along the learned lattice, reaching Spearman correlations of 0.85 for age on the WVS Canada subset and 0.83 for age on Census-Income. PCA, UMAP, t-SNE, and autoencoders, including a capacity-matched 1.4M-parameter autoencoder with near-perfect reconstruction, stay below 0.34. The authors interpret this as evidence that sensitive attributes are not merely extractable by a probing classifier but are dominant organizing axes of the representation, a form of ambient leakage that affects all","pith_inferences":["The paper's own preprocessing (Section 4.2) removes Census columns that correlate strongly with age and income and selects WVS questions based on their correlation with age, meaning the reported leakage is partly guided by the sensitive attribute. A cleaner test that withholds the attribute before any feature selection would make the emergence claim much stronger.","The monotonic-ordering definition targets ordinal attributes like age and income; the same mechanism may produce categorical separation (e.g., race or gender) as discontinuous gradients, and a categorical auditing metric would test whether the effect extends beyond ordinals.","SOMs have a fixed lattice resolution, so the magnitude of leakage likely depends on sample size and lattice size; a systematic sensitivity analysis across map dimensions would show how robust the effect is.","If representation-level audits become standard, the paper's suggested mitigations—such as demographic entropy balancing per neuron or penalizing sensitive-attribute gradients across the lattice—could be tested within the same SOM framework, turning the audit tool into a fairness-aware representation learner."],"forward_implications":["Fairness auditing must extend beyond supervised predictors to unsupervised embeddings, because structured leakage exists before any task is defined.","Probing classifiers and SOMtime are complementary: probes measure worst-case extractability, while SOMtime measures whether a sensitive attribute is a dominant organizing principle of the representation.","Unsupervised segmentation of such embeddings produces demographically skewed clusters, so downstream clustering and recommendation systems inherit fairness risks without any supervised objective.","A capacity-matched autoencoder with near-perfect reconstruction does not reproduce the leakage, suggesting the SOM's topological inductive bias, not parameter count, drives the effect.","Practitioners should run representation-level leakage audits, including correlation checks and global ordering analysis, before deploying any unsupervised embedding in decision-making pipelines."],"fun_headline_variants":["SOMs expose withheld age and income at 0.85 correlation","Unsupervised SOMs leak sensitive data: fairness fails","Self-organizing maps reveal hidden age, income axes","Fairness through unawareness fails in SOM embeddings","SOMs beat PCA, UMAP, t-SNE at leaking hidden attributes"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The sensitive attribute is truly withheld from every step that shapes the representation; however, Section 4.2 admits that Census columns strongly correlated with age and income were removed and WVS questions were selected based on their age correlation, so the reported leakage is not purely emergent from unsupervised learning.","fun_headline_variants_meta":{"raw":{"variants":["SOMs expose withheld age and income at 0.85 correlation","Unsupervised SOMs leak sensitive data: fairness fails","Self-organizing maps reveal hidden age, income axes","Fairness through unawareness fails in SOM embeddings","SOMs beat PCA, UMAP, t-SNE at leaking hidden attributes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000356,"raw_usage":{"total_tokens":1782,"prompt_tokens":767,"completion_tokens":1015,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":511,"completion_tokens_details":{"reasoning_tokens":929}},"tokens_in":511,"tokens_out":1015,"duration_ms":7267,"temperature":1.0,"reasoning_tokens":929,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T21:58:50.155070+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train SOMtime on data where the sensitive attribute is genuinely independent of all input features, or on synthetic data with known generative structure, with no feature selection or preprocessing informed by the attribute. If the Spearman correlation between the embedding axes and the withheld attribute stays near zero, the claim that SOMs systematically expose such attributes would be refuted. Alternatively, rerun the Census experiment on the full 40 columns without removing age- and income-correlated features; if correlations do not rise above baseline, the reported gap depends on preproces","supporting_citations":[],"review_version":1}