{"id":"9da37c40-0c49-4baa-99f9-dcac7831fa08","arxiv_id":"2607.07708","paper_version":1,"verdict":"CONDITIONAL","confidence":"UNKNOWN","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":6,"one_line_summary":"A single autoregressive foundation model unifies protein, molecule, and crystal structures into a shared token vocabulary and generates inspectable reasoning traces, achieving SOTA on 67 of 86 scientific tasks.","lead":"SciReasoner is a multimodal AI model that reasons about protein, molecule, and crystal structures by converting them into a unified token vocabulary, achieving state-of-the-art on 67 of 86 scientific benchmarks. A generalist might read it to see whether structure-aware tokens improve AI scientific reasoning beyond text-only LLMs.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Materials benchmarks use random splits without structural-similarity filtering, and several results show extraordinary magnitudes (JARVIS-QETB MAD/MAE=108.98 vs 0.84 for next-best) that may reflect leakage through isostructural compounds in train and test.","rationale":"The reader correctly identified the absence of released code/weights and lack of error bars as serious gaps, and the cross-domain consolidation concern is reasonable. However, the reader did not identify the more specific and testable issue: the materials evaluation uses random splits without structural-similarity filtering, despite the same databases (JARVIS-QETB, GNoME, etc.) being used for both pretraining and benchmarking. This is the weakest link in the evidence chain because (1) it directly affects the quantitative SOTA claim, (2) the magnitude of several materials improvements is extraordinary enough to warrant specific scrutiny, and (3) it is a concrete, fixable methodological gap — proteins get identity filtering, molecules get canonicalized exclusion, but materials get only exact-sample removal. The concern does not move the verdict below CONDITIONAL because the non-materials results (proteins, chemistry) may still hold, the architecture and training methodology appear sound, and the concern is testable rather than fatal. But it should be explicitly addressed before any upgrade to ACCEPT. The reader's consolidation concern is secondary: even if consolidation is imperfect, the benchmark results would still need to hold, and if they do, minor consolidation degradation doesn't invalidate the central claim.","tokens_in":44101,"tokens_out":7028,"duration_ms":338682,"concrete_test":"Re-run all 10 materials sub-tasks using a composition- and structure-aware split: group materials by (chemical formula, space group) and ensure no group spans train and test. Additionally, report MAE in absolute units alongside MAD/MAE so the metric is not inflated by target dispersion. If JARVIS-QETB MAD/MAE drops below ~10 and GNoME below ~5 under the structure-aware split, the original random-split results were inflated by structural similarity leakage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of SOTA on 67/86 benchmarks relies substantially on materials tasks (14 tasks in Table A3). For materials, §4.1.3 describes an '80/10/10 random split at the material-sample level' with only exact-sample removal from pretraining. Critically, JARVIS-QETB is listed as both a pretraining data source (§4.1.3: 'The primary sources include Materials Project, JARVIS-DFT, SNUMAT, hMOF, QMOF, OQMD, OMDB, JARVIS-QETB, GNoME, and Cantor HEA') and as a benchmark target (Table A3). Unlike proteins, where >30% sequence identity filtering is applied (§4.1.1), or molecules, where canonicalized identifiers are excluded (§4.1.2), the materials split does not control for structural or compositional similarity. Polymorphs, isostructural compounds, or materials sharing the same space group and similar composition can appear in both train and test. The JARVIS-QETB result (MAD/MAE = 108.98 vs 0.73 for Opus-4.7, a ~149x improvement) and GNoME (21.91 vs 1.94, ~11x) are far beyond typical ML improvements and could reflect the model exploiting structural similarity rather than learning genuine structure-property relationships. The MAD/MAE metric itself can amplify this: if the target distribution has large dispersion, even moderate MAE yields a high ratio. Without error bars on any benchmark and with no released code, data, or weights, these results cannot be independently verified. If the materials results are inflated, the '67 of 86' headline weakens by up to 14 tasks.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"The manuscript introduces SciReasoner, a multimodal foundation model for scientific reasoning across proteins, small molecules, and inorganic crystals. The core architectural idea is a structure-aware vocabulary that discretizes 3D coordinates, molecular topologies, and crystallographic lattices into domain-native tokens (via Foldseek, ConfSeq, and SLICES encoders), integrated into a Qwen3-14B backbone through a dedicated embedding layer. A multi-stage pretraining curriculum (warm-up, full-parameter, annealing) is followed by a two-stage post-training framework (intra-domain structural evidence grounding via RL, then cross-domain reasoning consolidation). The paper evaluates on 86 benchmarks spanning biology, chemistry, and materials science, reporting SOTA on 67 tasks. Key qualitative claims include: (1) GO prediction gains are largest in low-homology regimes, (2) retrosynthesis traces are auditable and chemically grounded, (3) materials representations separate polymorphs and band-gap regimes, and (4) double-blind expert evaluation (N=1776 judgments) prefers SciReasoner over DeepSeek-V4-Pro in 98% of comparisons.","tokens_in":44487,"tokens_out":1891,"duration_ms":410172,"significance":"The paper addresses a genuine gap: unifying structure-native representation with inspectable reasoning traces across multiple scientific domains in a single autoregressive model. The structure-aware vocabulary design—treating structural tokens as addressable evidence units rather than auxiliary descriptors—is a clean and well-motivated architectural contribution. The homology-stratified GO analysis (Fig. 2A) and the structure-ablation experiments (Fig. 5) provide falsifiable evidence that performance gains are not driven by sequence-similarity shortcuts. The retrosynthesis case studies (Fig. 3B-C) with atom-by-auditable SMILES fragments in the reasoning trace are a concrete strength. The double-blind human evaluation protocol (Appendix B) with a structured rubric and ground-truth fact sheets is more rigorous than typical LLM-judge-only evaluations. The self-bootstrapped post-training framework that avoids pooling heterogeneous CoT supervision in a single pass is a reasonable methodological contribution to the RL-for-reasoning literature.","major_comments":[{"comment":"§4.1.3 and Table A3: The materials data split uses '80/10/10 random split at the material-sample level' with only exact-sample removal from pretraining. Unlike proteins (>30% sequence identity filtering, §4.1.1) or molecules (canonicalized identifier exclusion, §4.1.2), the materials split does not control for structural or compositional similarity. This is a load-bearing concern because JARVIS-QETB is listed as both a pretraining source (§4.1.3) and a benchmark target (Table A3), and the reported MAD/MAE = 108.98 vs 0.73 for the next-best model is a ~149x improvement—far beyond typical ML gains. Similarly, GNoME shows 21.91 vs 1.94 (~11x). These magnitudes could reflect the model exploiting isostructural or compositionally similar compounds in train and test rather than learning genuine structure-property relationships. The MAD/MAE metric itself can amplify this: if the target variance,","section":null},{"comment":"§2.5, Fig. 6E: The double-blind human evaluation compares SciReasoner only against DeepSeek-V4-Pro, a general-purpose LLM, rather than against domain specialists or structure-aware baselines (e.g., SaProt for proteins, RSGPT for chemistry). The 98% preference rate (8.7/10 vs 4.3/10) is so one-sided that it raises the question of whether the comparison is informative about SciReasoner's scientific reasoning quality or primarily about the weakness of general-purpose LLMs on structure-intensive tasks. The per-axis scores for DeepSeek-V4-Pro (e.g., 3.9/10 on reasoning coherence) suggest a baseline that may not be competitive enough to establish that SciReasoner's traces meet a high scientific standard. Including at least one domain-specialist baseline in the human evaluation would substantially strengthen the claim that the reasoning traces are genuinely useful to experts, not merely better-","section":null},{"comment":"§4.4, Fig. 6B-C: The cross-domain reasoning consolidation step pools expert-generated traces and performs a single RL pass, but the paper does not isolate whether consolidation preserves or degrades individual expert capabilities relative to the experts themselves. Fig. 6B shows reward dynamics and Fig. 6C shows pass@1/pass@10, but neither directly compares the unified model against the domain experts on their respective domains. If the unified model underperforms the experts on some domains, this would indicate destructive interference during consolidation—a risk the paper acknowledges (§4.4: 'joint training induces destructive interference') but does not empirically rule out. A table comparing per-domain performance of the unified model vs. each domain expert would address this.","section":null},{"comment":"Tables A2-A4: No error bars, confidence intervals, or significance tests are reported for any of the 86 benchmark results. Given that several margins are narrow (e.g., GO-MF 0.66 vs 0.67 for SaProt in Table A1, Non-coding RNA family 0.90 vs 0.89 for RNA-MSM), it is unclear whether these differences are statistically meaningful. The absence of variance estimates is particularly consequential for the headline claim of 'SOTA on 67/86 tasks,' as some of these 67 may not be statistically distinguishable from the second-best model.","section":null}],"minor_comments":[{"comment":"§2.2.1, Fig. 2D-E: The reasoning-quality evaluation relies on GPT-5.5 as an LLM judge. The paper does not report inter-rater reliability between GPT-5.5 and the human evaluators, making it difficult to assess whether the LLM judge scores generalize to expert assessments.","section":null},{"comment":"Table A3: The Cantor HEA result (MAD/MAE = 7.79 for SciReasoner vs 8.40 for LLM-Prop) is the only materials task where SciReasoner underperforms the specialist baseline. This is not discussed in the text.","section":null},{"comment":"Fig. 1C: The tokenizer comparison shows compression ratio but does not report vocabulary size for the structure-aware vocabulary, which is relevant for assessing the embedding parameter count (W_v in Eq. 1).","section":null},{"comment":"§4.1.3: The footnote numbering for materials databases (footnotes 7-11) appears to use a different numbering scheme than the main text footnotes.","section":null},{"comment":"Fig. 4A: The y-axis label and metric for the 10 materials sub-tasks are not clearly specified in the figure caption; the reader must cross-reference Table A3 to determine whether each bar is MAE or AUC.","section":null},{"comment":"Table A1: Several entries show SciReasoner underperforming the specialist (LIPO RMSE 0.80 vs 0.65, Human PPI ACC 0.73 vs 0.77, Solubility ACC 0.72 vs 0.77, GO-MF 0.66 vs 0.67). These are not discussed in the main text, which states SciReasoner 'matches or exceeds' specialists.","section":null},{"comment":"§2.2.3: The DUD-E evaluation extracts embeddings from 'the 10 tokens generated immediately after the prompt.' The sensitivity of the AUC/EF results to this arbitrary choice of 10 tokens is not discussed.","section":null},{"comment":"The paper does not mention whether code, data, or model weights will be released, which limits independent reproducibility of the 86 benchmark results.","section":null}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about materials data leakage is well-founded and is the most serious issue. The JARVIS-QETB result (108.98 vs 0.73) is not plausibly a genuine ~149x improvement; either the MAD/MAE metric is behaving pathologically (e.g., near-zero MAE due to test set memorization, or extreme target variance inflating the ratio), or there is structural leakage through the random split. The paper's own protein and molecule sections implement similarity-based leakage controls, making the absence of analogous controls for materials conspicuous. I recommend the authors either (a) implement structural-similarity filtering for materials (e.g., removing compounds with the same space group and similar composition from training), or (b) provide a convincing explanation for the extraordinary magnitudes, or (c) re-run materials benchmarks with a scaffold-split-style protocol. If the materials results are inflated, the '67/86' headline weakens by up to 14 tasks, which would change the paper's contribution profile substantially. The human evaluation concern is also worth flagging to the authors: comparing only against DeepSeek-V4-Pro (a generalist LLM scoring 4.3/10) does not establish that SciReasoner's reasoning is expert-quality, only that it is better than a non-specialist baseline."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive report. The referee correctly identifies the core contributions of SciReasoner and raises four substantive concerns: (1) the materials data split lacks structural/compositional similarity filtering, and the MAD/MAE gains on JARVIS-QETB and GNoME are suspiciously large; (2) the human evaluation compares only against a general-purpose LLM rather than domain specialists; (3) the cross-domain consolidation step does not empirically rule out destructive interference relative to per-domain experts; and (4) no error bars or significance tests are reported across the 86 benchmarks. We agree that all four points identify genuine gaps in the current manuscript and will address each in revision. Below we respond point by point.","responses":[{"response":"The referee is correct on all counts. The materials data split in the current manuscript uses only exact-sample removal, which is inconsistent with the stricter leakage controls applied for proteins (>30% sequence identity filtering) and molecules (canonicalized identifier exclusion). This is a genuine weakness in our experimental design. Furthermore, the MAD/MAE magnitudes on JARVIS-QETB (108.98 vs 0.73) and GNoME (21.91 vs 1.94) are far beyond what is typically observed in materials ML and could plausibly be inflated by structural or compositional similarity between training and test compounds. We acknowledge that without similarity-based filtering, we cannot rule out this explanation. In the revision, we will: (1) implement structural similarity filtering for the materials split, using composition-based (e.g., SMACT-style elemental overlap) and structure-based (e.g., pymatgen StructureMatcher with default distance/angle tolerances) deduplication across train/test boundaries, applied consistently to all materials benchmarks including JARVIS-QETB and GNoME; (2) re-run all materials benchmarks with the corrected split and report updated MAD/MAE values; (3) add an explicit discussion of why the original magnitudes were suspicious and how the corrected numbers compare; and (4) clarify in §4.1.3 that JARVIS-QETB and GNoME entries in the pretraining corpus had their test-split material samples removed but that this was insufficient without similarity filtering. If the corrected numbers remain substantially above baselines, we will provide additional analysis (e.g., stratification by compositional novelty) to support the claim. If they do not, we will report the corrected numbers honestly and adjust our claims accordingly.","revision_made":"yes","referee_comment":"§4.1.3 and Table A3: The materials data split uses 80/10/10 random split at the material-sample level with only exact-sample removal from pretraining, unlike proteins (>30% sequence identity filtering) or molecules (canonicalized identifier exclusion). JARVIS-QETB is listed as both a pretraining source and benchmark target, and the reported MAD/MAE = 108.98 vs 0.73 (~149x) and GNoME 21.91 vs 1.94 (~11x) could reflect isostructural or compositionally similar compounds leaking between train and test."},{"response":"The referee raises a valid concern. The 98% preference rate against a single general-purpose LLM baseline does not by itself establish that SciReasoner's reasoning traces meet a high scientific standard—it could partly reflect the difficulty general-purpose LLMs face on structure-intensive tasks. We agree that including at least one domain-specialist baseline would substantially strengthen the evaluation. In the revision, we will expand the human evaluation to include a structure-aware domain specialist (e.g., SaProt for protein GO prediction, and a chemistry-specialist model for retrosynthesis) as an additional comparison arm. We will retain the DeepSeek-V4-Pro comparison as a general-purpose reference but will frame the results as 'SciReasoner vs. general-purpose LLM' rather than as an absolute quality claim. We will also add discussion acknowledging that the one-sided margin against DeepSeek-V4-Pro is partly attributable to the absence of structural tokenization in general-purpose LLMs, and that the specialist comparison is the more informative test of whether the reasoning traces are genuinely useful to experts. We note that recruiting qualified domain experts for double-blind evaluation is logistically constrained, and we may not be able to add all three domain specialists in time for the next revision, but we commit to including at least one.","revision_made":"partial","referee_comment":"§2.5, Fig. 6E: The double-blind human evaluation compares SciReasoner only against DeepSeek-V4-Pro, a general-purpose LLM, rather than against domain specialists or structure-aware baselines. The 98% preference rate is so one-sided that it raises the question of whether the comparison is informative about SciReasoner's scientific reasoning quality or primarily about the weakness of general-purpose LLMs on structure-intensive tasks."},{"response":"This is a fair and important point. The current manuscript shows reward dynamics (Fig. 6B) and pass@1/pass@10 (Fig. 6C) but does not directly compare the unified model against the per-domain experts on their respective domains. Without this comparison, we cannot empirically demonstrate that consolidation preserves expert capabilities rather than degrading them. We acknowledge this gap. In the revision, we will add a table comparing the unified SciReasoner against each domain-structure expert (protein, molecule, material) on the representative tasks used in Fig. 6C (GO prediction, DUD-E, QMOF, Retrosynthesis USPTO-50K). If the unified model underperforms any expert on its domain, we will report this transparently and discuss it as evidence of partial destructive interference, consistent with the caveat we already note in §4.4. If the unified model matches or exceeds all experts, this will strengthen the consolidation claim. Either way, the comparison will be included.","revision_made":"yes","referee_comment":"§4.4, Fig. 6B-C: The cross-domain reasoning consolidation step pools expert-generated traces and performs a single RL pass, but the paper does not isolate whether consolidation preserves or degrades individual expert capabilities relative to the experts themselves. A table comparing per-domain performance of the unified model vs. each domain expert would address this."},{"response":"The referee is correct. The absence of variance estimates is a significant omission, especially given that several of the 67 claimed SOTA results have narrow margins where statistical distinguishability is uncertain. We will address this in the revision by: (1) reporting confidence intervals or standard errors for all benchmark results where multiple evaluation runs or bootstrap resampling is feasible; (2) applying paired significance tests (e.g., bootstrap or McNemar's test for classification, bootstrap CI for regression metrics) for the narrow-margin comparisons the referee identifies, including GO-MF (0.66 vs 0.67) and Non-coding RNA family (0.90 vs 0.89); and (3) revising the headline claim to distinguish between tasks where SciReasoner is statistically significantly better than the second-best model and tasks where the difference is within noise. We will report the revised count as 'SOTA on N of 86 tasks with statistical significance (at p<0.05)' alongside the original count, and will list the borderline cases explicitly. We acknowledge that for some benchmarks, only a single evaluation run was performed due to computational cost, and in those cases we will use bootstrap resampling over the test set to estimate confidence intervals rather than requiring full re-runs.","revision_made":"yes","referee_comment":"Tables A2-A4: No error bars, confidence intervals, or significance tests are reported for any of the 86 benchmark results. Several margins are narrow (e.g., GO-MF 0.66 vs 0.67 for SaProt, Non-coding RNA family 0.90 vs 0.89 for RNA-MSM), and it is unclear whether these differences are statistically meaningful, particularly for the headline claim of 'SOTA on 67/86 tasks.'"}],"tokens_in":44433,"tokens_out":1671,"duration_ms":276209,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"The headline: SciReasoner is a genuine integration effort — unifying Foldseek 3Di, ConfSeq, and SLICES tokenizers into a single Qwen-based LLM with a two-stage RL post-training pipeline — and the protein and chemistry results are solid. The materials results are not yet trustworthy enough to support the headline claim of SOTA on 67/86 benchmarks without further scrutiny. No code, data, or weights are released, which compounds the problem. It deserves a serious referee because the core idea and the non-materials results are real, but the materials evaluation has a load-bearing flaw that must be addressed before publication. The stress-test concern about materials leakage is the single most important issue. The paper applies >30% sequence-identity filtering for proteins (§4.1.1) and canonicalized identifier exclusion for molecules (§4.1.2), but materials use only an 80/10/10 random split with exact-sample removal (§4.1.3). No structural or compositional similarity filtering is applied. JARVIS-QETB appears in both the pretraining corpus and the benchmark targets, and the reported MAD/MAE of 108.98 (vs. 0.73 for the next-best model) is roughly 149× — a magnitude that is not plausible for genuine generalization. GNoME at 21.91 vs. 1.94 is similarly suspicious. These are not incremental improvements; they are the kind of numbers you see when isostructural compounds or polymorphs leak across train and test. If 14 materials tasks are inflated, the 67/86 headline weakens substantially. The protein work is the strongest part. The low-homology GO prediction result — F_max from 0.42 to 0.55 in the ≤30% identity bin — is meaningful and the homology stratification is the right experimental design. The attention analysis on DNA-binding residues (AUROC 0.78–0.91) is a good check that the model is attending to the right structural regions. The retrosynthesis result (0.72 vs. 0.63 for RSGPT) is credible and the five-case qualitative comparison in Fig. 3C is well-chosen. The self-bootstrapped RL framework is reasonable in principle, but the reader's concern about destructive interference during cross-domain consolidation is valid: the paper shows reward curves but does not isolate whether the unified model preserves individual expert capabilities. This is a moderate concern, not a fatal one. The GPT-5.5-as-judge evaluation for reasoning quality is a weak spot. Using one frontier LLM to grade another is inherently limited, and the human evaluation (N=1776, still described as a pilot) only compares against DeepSeek-V4-Pro, not against domain specialists. No error bars on any benchmark is a presentation gap that a referee should flag. The ablation experiments (Fig. 5) are the right idea and show that structural tokens matter, but without error bars the effect sizes are hard to calibrate. This paper is for researchers working on scientific foundation models, structure-aware representation learning, and interdisciplinary AI for science. The integration idea is worth engaging with even if the materials results need reworking. Recommend serious peer review with a requirement that the authors (1) apply structural-similarity filtering to materials splits and re-report, (2) release code and at least a data split protocol, and (3) add error bars. If the materials results survive proper leakage controls, this is a strong paper. If they do not, the protein and chemistry contributions may still stand on their own.","headline":"Ambitious integration of structure-aware tokens across three scientific domains with strong protein/chemistry results, but materials benchmarks have a plausible leakage problem and reproducibility is currently zero.","tokens_in":45350,"tokens_out":825,"would_cite":false,"duration_ms":203212,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"One model reasons across proteins, molecules, and crystals by citing structural evidence","keywords":["structural reasoning","scientific foundation model","structure-aware vocabulary","autoregressive reasoning","multimodal","protein function prediction","retrosynthesis","crystal properties"],"falsifier":"If structural tokens are ablated and performance does not drop, the tokens are not load-bearing for reasoning.","tokens_in":44182,"feed_emoji":"🧬","tokens_out":625,"duration_ms":208999,"temperature":0.7,"pith_summary":"SciReasoner is a foundation model that converts 3D coordinates, molecular topologies, and crystal lattices into a unified token vocabulary, then treats those tokens as addressable evidence within a single autoregressive reasoning trajectory. The central claim is that when structural information is represented as discrete tokens that preserve domain-native semantics—rather than as text strings or black-box inputs—a single model can reason across proteins, small molecules, and inorganic crystals while producing inspectable chains of thought that cite specific residues, molecular fragments, or crystallographic features. The paper argues this is the first model to unify sequence, structure, and language reasoning across all four scientific modalities (proteins, DNA/RNA, small molecules, crystals) within one autoregressive pass, achieving state-of-the-art on 67 of 86 benchmarks. The key mechanism is a two-stage post-training procedure: first, domain-specific experts learn to use structural tokens as reasoning evidence via reinforcement learning; then, expert-generated reasoning traces are pooled and consolidated into a single unified model, avoiding the destructive interference that would arise from jointly training all domains at once.","feed_headline":"Structural tokens let one model reason across proteins, molecules, and crystals","feed_subtitle":"SciReasoner converts 3D structures into citable evidence tokens, achieving SOTA on 67 of 86 benchmarks while producing inspectable reasoning","key_machinery":"Structure-aware tokens + two-stage post-training (intra-domain expert grounding via RL, then cross-domain consolidation)","core_discovery":"The paper's central object is the structure-aware token—a discrete representation of local geometry, bonding, symmetry, or conformation that preserves scientific semantics and can be cited as evidence within a reasoning chain. The core discovery is that by discretizing structural information into such tokens (using Foldseek for proteins, ConfSeq for molecules, SLICES for crystals) and integrating them into a language model's vocabulary, structural features become inspectable reasoning substrates rather than opaque input descriptors. The authors demonstrate this through three probe results: (1) protein function prediction improves most in low-homology regimes where sequence similarity is un帮助","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Structure-aware tokens turn 3D coordinates into inspectable reasoning evidence","One model reasons across proteins, molecules, and crystals via structural tokens","Discretized structural units let SciReasoner cite geometry as reasoning evidence","Structure-property reasoning improves when models cite discrete structural tokens","Proteins, molecules, and crystals unified by addressable structure-aware tokens"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The cross-domain consolidation step assumes that reasoning patterns learned independently by domain experts (proteins, molecules, materials) are compatible when merged into a single model, but the paper provides no formal guarantee against destructive interference during this merge.","fun_headline_variants_meta":{"raw":{"variants":["Structure-aware tokens turn 3D coordinates into inspectable reasoning evidence","One model reasons across proteins, molecules, and crystals via structural tokens","Discretized structural units let SciReasoner cite geometry as reasoning evidence","Structure-property reasoning improves when models cite discrete structural tokens","Proteins, molecules, and crystals unified by addressable structure-aware tokens"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":743,"prompt_tokens":653,"completion_tokens":90,"prompt_tokens_details":null},"tokens_in":653,"tokens_out":90,"duration_ms":76994,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T01:29:28.288367+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If structural tokens are ablated and performance does not drop, the tokens are not load-bearing for reasoning.","supporting_citations":[],"review_version":1}