{"id":"f5cff59b-1ac8-4d8a-8e53-2e2435052ca7","arxiv_id":"2504.18491","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"An autoencoder and t-SNE pipeline on S-PLUS DR4 12-band photometry selects about 19,000 SED outliers and identifies candidate carbon stars, white dwarf subtypes, and active low-mass stars for spectroscopic follow-up.","lead":"This paper uses an autoencoder and t-SNE to find stars with unusual colors in the S-PLUS southern sky survey, yielding about 19,000 candidates. It proposes 69 likely carbon-rich stars, groups white dwarfs by atmosphere type, and flags active low-mass stars for follow-up.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"t-SNE perplexity was tuned to maximize the carbon-star candidate count, so the reported 69/96 candidates may be a selected maximum rather than a stable population; cluster-to-class mapping needs quantitative validation.","rationale":"The paper's central claim is that an autoencoder plus t-SNE can select SED outliers and that the resulting t-SNE clusters correspond to real stellar populations. The most load-bearing component is the link between t-SNE cluster structure and astrophysical class membership, because all candidate counts and population labels are read off from specific regions of the t-SNE map. The reported choice of perplexity, tuned to maximize the number of carbon-star candidates, is a concrete post-hoc selection that directly affects the headline number of 69 carbon-rich candidates. This is not a fatal flaw: candidate identification for spectroscopic follow-up remains useful, and the paper is transparent about the hyperparameter dependence and about the fact that population sizes can vary. However, the absence of a quantitative stability analysis means the central number is not yet established as robust. The reader's verdict was CONDITIONAL, which already captures the need to address this circularity, so my analysis does not move the verdict. I only partially agree with the reader's weakest_assumption: the quality-cut contamination is a related but distinct concern, and my chosen attack is specifically the t-SNE perplexity optimization rather than the training-set purity, although both are worth checking.","tokens_in":19755,"tokens_out":6774,"duration_ms":74344,"concrete_test":"Rerun the full t-SNE pipeline at perplexities 25, 30, 35, 40, and 45 with at least 5 random seeds each. For every run, recompute the carbon-star candidate set using a fixed, reproducible criterion (e.g., a DBSCAN/HDBSCAN cluster enclosing the SIMBAD C*/ChemPec* reference objects, or the same manual region if its boundary is specified). Report the candidate counts and the pairwise Jaccard index between candidate lists across runs. If the candidate list has Jaccard index below 0.5 or the count varies by more than ~20% across perplexity/seed choices, the reported 69/96 candidates are a hyperparameter artifact and should be reframed as a range rather than a point estimate. If the overlap is high, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.1 states that perplexity was set to 40 largely because it yielded the highest number of carbon-star candidates, and Section 5.1 then reports 96 C-star candidates (69 newly identified). Because t-SNE is stochastic and its local/global geometry depends on perplexity, optimizing this hyperparameter against the target population's count means the reported carbon-star abundance is not an unbiased measurement but a maximum over the tested configurations. The same embedding is used to define the other populations, so the claimed WD and active-M-dwarf clusters inherit this post-hoc selection. The paper's statement that variations are 'subtle' is qualitative and not supported by a quantitative measure of candidate-list stability; the authors do not report the candidate-count variation across perplexity values or random seeds. If the candidate list changes substantially under modest hyperparameter changes, the central claim that t-SNE clusters correspond to genuine astrophysical classes is weakened. Independent support from the Lucey et al. (2023) cross-match strengthens the CEMP association for known objects, but it does not validate the 69 newly counted candidates, since the counting procedure itself was tuned to maximize that number. The WD separation is also partly aided by adding MG as an input feature, which should be disclosed as a luminosity-dependent rather than purely color-based separation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents an unsupervised machine-learning pipeline for identifying spectral energy distribution (SED) outliers in S-PLUS DR4 photometry. An autoencoder is trained on roughly 1.5 million stars selected by a series of quality cuts, and about 19,000 objects whose SEDs are poorly reconstructed are flagged as anomalous. t-SNE is then applied to 66 S-PLUS colors for the anomalous sample, and clusters are interpreted using cross-matches with Gaia, SIMBAD, eROSITA, APOGEE, and external carbon-star catalogs. The authors report 96 carbon-star candidates (69 newly identified, 4 likely dwarf carbon stars), three white-dwarf subgroups (DA, DB, WD+MS binaries), active low-mass stars detected in X-rays, and a large number of binary systems. The results are framed as candidate lists for spectroscopic follow-up, and a machine-readable candidate table is provided.","tokens_in":19954,"tokens_out":4522,"duration_ms":47425,"significance":"If the method is robust, the paper offers a practical demonstration that unsupervised anomaly detection plus nonlinear dimensionality reduction on narrow-band S-PLUS photometry can produce useful spectroscopic follow-up targets. The manuscript is grounded in real survey data and makes appropriate use of external cross-matches: the Lucey et al. (2023) match strengthens the CEMP association for known objects, the eROSITA data support the active-star interpretation, and the comparison with Li et al. (2024) helps place the carbon-star candidates in context. The code and exact hyperparameters are described in enough detail to reproduce the pipeline. The principal weakness is that the headline candidate counts are partly determined by a t-SNE hyperparameter chosen to maximize the carbon-star count, so the quantitative claims need a robustness analysis before they can be taken at face value.","major_comments":[{"comment":"The choice of t-SNE perplexity is made after inspecting the outcome: the text states that perplexity 40 'was mostly motivated because it yielded the highest number of carbon stars candidates.' The subsequently reported number of 96 C-star candidates (69 newly identified) is therefore a selected maximum over the tested configurations, not an unbiased estimate. Because t-SNE is stochastic and its geometry depends on perplexity, this post-hoc selection also affects the white-dwarf and M-star clusters defined on the same map. Please provide a quantitative stability analysis: for each tested perplexity value (25, 30, 35, 40, 45) and, ideally, for several random seeds, report the total number of carbon-star candidates, the overlap or Jaccard index between the perplexity-40 list and the other lists, and the membership stability of the carbon-star and white-dwarf regions. The statement that variations are 'subtle' should be supported by these numbers, not only by visual inspection.","section":"§4.1 and §5.1"},{"comment":"The separation of white dwarfs from hot subdwarfs is achieved only after adding the absolute magnitude MG as an additional input feature to t-SNE. Since MG is a luminosity-dependent quantity, the resulting separation is not purely based on SED colors, and the manuscript should explicitly state that the WD sub-population segregation is partly driven by absolute magnitude. In addition, the assignment of overdensities 1, 2, and 3 to WD+MS binaries, DB/DC stars, and DA stars rests on visual inspection of SIMBAD spectral types. Please add a quantitative validation, for example a contingency table, purity/completeness of each region with respect to SIMBAD classes, or a significance test of the spatial segregation. Without such validation, the claim that t-SNE 'reliably' segregates WD sub-populations is not fully supported.","section":"§5.2"},{"comment":"The training set is designed to contain 'well-behaved' stars, but no diagnostic is shown that the autoencoder reconstruction errors on the training set are approximately Gaussian, which is the basis for the 3σ anomaly threshold. Since all subsequent populations are derived from this threshold, even a small fraction of contaminants in the training set could propagate into the outlier sample and into every population label. Please show the distribution of reconstruction errors and report the sensitivity of the anomalous sample size, and of the carbon-star and white-dwarf candidate lists, to the threshold (for example 2.5σ and 3.5σ). The discussion already acknowledges threshold dependence qualitatively, but the paper would be much stronger with a quantitative test.","section":"§3.2 and §6"}],"minor_comments":[{"comment":"The abstract reports 69 carbon-rich star candidates, while §5.1 says that 96 C-star candidates were identified, 27 of them previously reported. Please clarify explicitly that 69 is the number of newly identified candidates, and state whether the 73-object cross-match with Lucey et al. (2023) includes previously known objects.","section":"Abstract and §5.1"},{"comment":"The phrase 'the number of iteration hyperparameters' should read 'the number of iterations hyperparameter'.","section":"§4.1"},{"comment":"'Counter parts' should be 'counterparts' in the last paragraph of Section 5.3.","section":"§5.3"},{"comment":"The sentence 'The ChemPec* label encompass a variety...' should be 'The ChemPec* label encompasses a variety...'.","section":"Figure 5 caption"},{"comment":"The claim that photometric errors are 'much smaller' than the reconstruction-error threshold is asserted without numbers. Please quote typical reconstruction errors and typical photometric errors to support this statement.","section":"§3.2"}],"recommendation":"major_revision","confidential_remarks":"This is a useful applications paper that fits the journal's scope. The main issue is the post-hoc selection of t-SNE perplexity based on the carbon-star count; this is transparently disclosed but still undermines the headline number. I recommend major revision rather than rejection because the issue is fixable with a robustness analysis, and the paper otherwise provides a well-described pipeline and valuable candidate lists. I would also ask the authors to strengthen the WD sub-population validation, as the current support is largely visual."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read: this is a solid, honest application of a standard anomaly-detection pipeline (vanilla autoencoder plus t-SNE) to S-PLUS DR4, and it produces genuinely useful candidate lists. The method isn't new — the paper says so — but the S-PLUS outlier catalog, the 69 new carbon-star candidates, and the demonstration that WD subtypes separate in an MG-augmented t-SNE map are new data products. The carbon-star work is the strongest part: 73 of the candidates cross-match Lucey et al. (2023) CEMP candidates, which is independent external support, and the kinematic analysis (Toomre diagram, Z versus R) is consistent with a thick-disk/halo CH/CEMP population. The WD and active-M-dwarf sections are more suggestive than conclusive, but they are properly framed as candidate selection.\n\nThe main soft spot is the one the stress test flagged: t-SNE perplexity was chosen partly because it maximized the carbon-star candidate count (Section 4.1). That makes the reported 69/96 numbers a selected maximum over the tested hyperparameter grid, not an unbiased estimate. The authors say the map variations are 'subtle' but don't quantify candidate-list stability across perplexity values or random seeds. That should be fixable: report the count variance, or at least show the carbon-star region is stable under different perplexities. The 69 versus 96 number isn't actually an inconsistency — 96 total, 27 previously known, 69 new — so I'd let that go.\n\nSecond issue: the WD subtype separation uses MG (absolute magnitude) as an extra t-SNE input, not just the 66 S-PLUS colors. The paper discloses this, but the claim that t-SNE on S-PLUS colors alone reliably segregates WD subpopulations is overstated. The luminosity information is doing real work. That should be reworded and, ideally, tested with colors only.\n\nMinor: the autoencoder training cut at mag < 19 means fainter sources are biased toward being flagged anomalous; the authors acknowledge this, but it would be good to quantify how many of the 19,000 outliers are simply faint. Also, the binary section is a null result, reported honestly.\n\nBottom line: this deserves a serious referee. The pipeline is reproducible, the caveats are mostly acknowledged, and the candidate lists are useful for follow-up. I'd ask for a stability analysis of the t-SNE clusters and a clearer separation of color-driven versus MG-driven WD claims before accepting.","headline":"A useful, transparent candidate-selection pipeline for S-PLUS; the carbon-star candidates have external support, but the t-SNE perplexity tuning and the MG-augmented WD claim need harder stability numbers.","tokens_in":20642,"tokens_out":2232,"would_cite":true,"duration_ms":21641,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An unsupervised autoencoder plus t-SNE pipeline on S-PLUS photometry flags roughly 19,000 anomalous SEDs and isolates 69 carbon-rich candidates, along with distinct white-dwarf and active low-mass star populations.","keywords":["autoencoder","t-SNE","S-PLUS","spectral energy distribution","anomaly detection","carbon stars","white dwarfs","low-mass stars"],"falsifier":"Spectroscopically observe the 69 carbon-rich candidates: if most turn out to lack enhanced carbon features such as CH or C2 bands, and if their space velocities place them in the thin disk rather than the halo or thick disk, then the claimed CH/CEMP identification collapses.","tokens_in":19528,"feed_emoji":"🔭","tokens_out":7583,"duration_ms":73722,"temperature":0.7,"pith_summary":"This paper tries to establish that a purely unsupervised pipeline, an autoencoder followed by t-SNE, can use 12-band photometry from the S-PLUS survey to produce a short list of genuinely peculiar stars worth spectroscopic follow-up. The authors train an autoencoder on about 1.5 million stars they believe are normal, then flag the roughly 19,000 sources whose spectral energy distributions the network cannot reconstruct beyond a $3\\sigma$ error threshold. Running t-SNE on the 66 colors of the anomalous sample, they identify 69 carbon-rich candidates (likely CH or carbon-enhanced metal-poor stars, four possibly dwarf carbon stars), separate DA, DB, and WD-plus-main-sequence white dwarf groups, and isolate very active low-mass stars using X-ray data. If correct, the method turns a broad-band plus narrow-band photometric survey into a discovery engine for chemically peculiar stars, compact remnants, and active stars, independent of spectroscopy.","feed_headline":"Unsupervised pipeline finds 69 carbon-star candidates in S-PLUS","feed_subtitle":"Autoencoder outliers plus t-SNE clustering separate carbon stars, white dwarf types, and active low-mass stars from photometry alone.","key_machinery":"The machinery is a two-stage unsupervised pipeline. First, a vanilla autoencoder with a seven-dimensional latent code is trained on about 1.5 million stars that pass strict astrometric and photometric quality cuts; its mean squared reconstruction error on the 12 scaled magnitudes defines normality, and a $3\\sigma$ threshold selects roughly 19,000 anomalies. Second, t-SNE is run on the 66 color differences among the 12 S-PLUS filters, with perplexity set to 40, to lay the anomalies out on a 2D map where known populations form dense islands; for white dwarfs, an absolute magnitude from an external astrometric catalog is added as an extra feature to separate them from hot subdwarfs.","core_discovery":"The central discovery is that SED outliers selected by reconstruction error are not a random collection: in the t-SNE projection they form coherent islands that match established stellar classes. The autoencoder's $3\\sigma$ anomalies, when projected through t-SNE on 66 color indices, cluster at positions occupied by known carbon stars, DA and DB white dwarfs, WD-plus-main-sequence binaries, and X-ray-active low-mass stars. The paper's key population claim is that 69 of the carbon candidates belong, by kinematics and space position, to the CH/CEMP family, with four being dwarf carbon stars. It also reports that binary systems are abundant among anomalies but show no clean relation between t-SNE overdensities and orbital parameters.","pith_inferences":["Because the t-SNE perplexity was chosen to maximize the carbon-star candidate count, other cluster boundaries on the same map may be partly optimized for that population; re-running t-SNE across several perplexities and asking which clusters persist would test whether the white-dwarf and active-star groupings are stable.","The null result on binary orbital properties suggests SED shape is dominated by photospheric parameters rather than orbital architecture; a forward model that generates synthetic SEDs from binary parameters could quantify how many SB1 systems this pipeline is intrinsically blind to.","The same autoencoder-plus-t-SNE recipe should transfer directly to surveys sharing the S-PLUS filter set, so a cross-survey run would reveal whether the identified populations are survey-independent or artifacts of S-PLUS depth and footprint."],"forward_implications":["The 69 carbon-rich candidates form a concrete target list for medium- and high-resolution spectroscopy; confirmed CH or CEMP stars would expand the known chemically peculiar population in the southern sky.","The clean separation of DA, DB, and WD+MS groups implies that S-PLUS colors, plus a distance, can pre-select white-dwarf subtypes for follow-up without needing spectra.","The X-ray-active low-mass stars occupy a distinct t-SNE region, so the pipeline can flag active stars from photometry alone, even before epoch photometry or activity-line indices are examined.","Population sizes and the anomaly sample itself depend on the $3\\sigma$ threshold and the t-SNE perplexity, so the outputs are candidate lists rather than complete censuses."],"supporting_citations":[{"why":"Supplies the t-SNE algorithm, the probability formulation, and the perplexity guidance used to build the 2D map.","marker":"van der Maaten & Hinton 2008"},{"why":"Establishes autoencoders for dimensionality reduction and underpins the anomaly-detection approach.","marker":"Hinton & Salakhutdinov 2006"},{"why":"Defines the S-PLUS filter system and survey design that provide all input photometry.","marker":"Mendes de Oliveira et al. 2019"},{"why":"Provides the star/galaxy/quasar classification used to build the clean training sample.","marker":"Nakazono et al. 2021"},{"why":"Supplies the astrometry, photometry, ruwe, variability flags, and non-single-star flags used in the quality cuts.","marker":"Gaia Collaboration et al. 2023a"},{"why":"Provides photogeometric distances used to compute absolute magnitudes and color-magnitude diagrams.","marker":"Bailer-Jones et al. 2021"},{"why":"Gives an independent catalog of CEMP candidates used to cross-match and reinforce the carbon-star classification.","marker":"Lucey et al. 2023"},{"why":"Supplies a spectroscopically classified carbon-star sample used for the C-N, C-H, and C-R comparison diagrams.","marker":"Li et al. 2024"},{"why":"Provides the eROSITA X-ray catalog used to identify active low-mass stars.","marker":"Merloni et al. 2024"},{"why":"Defines the X-ray main-sequence diagram that separates active stars from accreting compact objects.","marker":"Rodriguez 2024"}],"fun_headline_variants":["Autoencoder + t-SNE spot 69 carbon-star candidates in S-PLUS","69 carbon-star candidates found by S-PLUS autoencoder and t-SNE","S-PLUS outliers cluster into 69 carbon-star candidates via ML"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result depends on the quality cuts leaving a training set of about 1.5 million genuinely normal, single, non-variable stars; if unresolved binaries, variables, or misclassified extragalactic sources survive, the autoencoder learns a biased idea of normal and every anomaly label inherits that bias.","fun_headline_variants_meta":{"raw":{"variants":["Autoencoder + t-SNE spot 69 carbon-star candidates in S-PLUS","69 carbon-star candidates found by S-PLUS autoencoder and t-SNE","S-PLUS outliers cluster into 69 carbon-star candidates via ML"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000641,"raw_usage":{"total_tokens":2980,"prompt_tokens":1002,"completion_tokens":1978,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":618,"completion_tokens_details":{"reasoning_tokens":1911}},"tokens_in":618,"tokens_out":1978,"duration_ms":15433,"temperature":1.0,"reasoning_tokens":1911,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:15:08.607544+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Spectroscopically observe the 69 carbon-rich candidates: if most turn out to lack enhanced carbon features such as CH or C2 bands, and if their space velocities place them in the thin disk rather than the halo or thick disk, then the claimed CH/CEMP identification collapses.","supporting_citations":[{"cited_title":"2008, Journal of Machine Learning Research, 9, 2579","cited_arxiv_id":null,"evidence_quote":"Supplies the t-SNE algorithm, the probability formulation, and the perplexity guidance used to build the 2D map."}],"review_version":1}