REVIEW 4 major objections 6 minor 34 references
Nanbeige4.2-3B claims a 3B-parameter model can outdo 9B and 12B models on agentic benchmarks while staying competitive on reasoning.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 05:49 UTC pith:2WLTCE2R
load-bearing objection A competent system report for a 3B agentic model whose headline benchmark claims rest too heavily on unchecked, partially self-referential evaluation. the 4 major comments →
Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery, stated on the paper's own terms, is that a 3B non-embedding model can be a general agent that outperforms substantially larger models (9B and 12B) across code-agent, office-agent, and complex tool-use benchmarks while remaining competitive on mathematical, coding, and scientific reasoning. The authors attribute this to three interacting ingredients: a Looped Transformer that increases effective depth without adding parameters, pre-training from scratch on a 28T-token corpus that includes a small share of agentic trajectories, and a post-training pipeline whose SFT data is produced by closed-loop environment synthesis and whose RL stages progressively stabilize generati
What carries the argument
The Looped Transformer — a transformer that passes hidden states through the same layer stack for a second pass, increasing effective depth at no parameter cost — is the architectural carrier of the argument. Around it, the post-training recipe does the heavy lifting: a three-stage SFT curriculum shifting from reasoning to agentic tokens, turn-level loss masking that keeps recovery context without training on bad turns, two-stage RLHF for think and non-think responses, length-controlled reasoning RL with a difficulty-aware penalty, and agentic RL with action-centric process rewards.
Load-bearing premise
The reported benchmark scores faithfully compare the models — that they are unbiased, not inflated by training-data overlap with the test tasks, and obtained with equivalent scaffolds and protocols.
What would settle it
An independent replication that re-runs the seven general-agent and three code-agent benchmarks under the stated settings and finds the 3B model no longer ranks first, or a contamination audit showing that synthesized training trajectories overlap with held-out test tasks.
If this is right
- A locally deployable 3B model can carry daily-assistant, office-automation, and research workflows that today are assumed to need much larger models.
- Agentic capability in small models may come primarily from execution-grounded data and reward design rather than parameter count.
- The reported cross-task and cross-mode generalization of RLHF implies that behavior-level regularization can improve reasoning and agentic performance simultaneously.
- The closed-loop data-synthesis pipelines suggest a path to keep raising task difficulty as a model improves, potentially extending to other domains.
- If the results reproduce, the gap between small and large models in agentic settings is smaller than commonly assumed.
Where Pith is reading between the lines
- The paper's benchmarking rests partly on in-house harnesses and an in-house benchmark; a neutral third-party replication would be the quickest test of the headline claim.
- A natural extension would be ablating the loop depth to isolate how much of the gain comes from the architecture versus the data and RL recipe.
- The difficulty-aware length penalty suggests a generic method for balancing reasoning effort and cost that could transfer to other small models without modification.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents Nanbeige4.2-3B, a 3B non-embedding parameter model pretrained from scratch on 28T tokens with a Looped Transformer architecture that reuses a shared layer stack. Post-training combines a three-stage SFT curriculum over reasoning, general, and agentic data with a multi-stage RL pipeline: two-stage RLHF for Think/Non-Think response quality, length-controlled reasoning RL, and agentic RL with action-centric process rewards. The central empirical claim is that the model outperforms substantially larger open models, including Qwen3.5-9B and Gemma4-12B, across general-agent, code-agent, reasoning, and alignment benchmarks (Table 3), and that it transfers to a local personal-assistant setting under the OpenClaw framework (Table 4). The paper also reports base-model improvements from the looped architecture and updated data mix (Table 1).
Significance. If the reported results hold up, this would be a notable demonstration that a compact model can serve as a general-purpose agent across code, office, and tool-use settings while retaining competitive reasoning, with practical implications for local deployment. The paper's strengths include a detailed account of the training recipe, a transparently described data-synthesis pipeline, release of the checkpoint, and same-scaffold comparisons in the OpenClaw evaluation (Table 4). The claimed gains at 3B scale are significant for the agentic-small-model line of work. However, the central evidence is entirely benchmark-based, and the evaluation infrastructure is partly author-controlled: several benchmarks or harnesses are in-house or co-authored by team members, no contamination audit is provided, and no variance information is reported despite 8-run averaging. These issues place the headline outperformance claim on an insecure footing that needs to be addressed before the result can be credited.
major comments (4)
- [§3.1.1, Fig. 2, and Table 3] The repository-to-task SFT pipeline mines real GitHub repositories, selects reference patches, and reconstructs executable tasks from historical development activities. SWE-bench Verified, SWE-bench Pro, and Terminal-Bench 2.0 are likewise derived from real code repositories/issues. The paper never states that the training pipeline excluded the benchmark repositories, patches, or near-duplicates, nor does it provide any overlap statistics or decontamination analysis. Since the headline outperformance on code-agent benchmarks rests on these numbers, the absence of a contamination audit is load-bearing. A concrete overlap analysis, a release of task provenance, or an evaluation on a held-out, non-GitHub-derived benchmark is required.
- [§3.3.1, §3.3.2, and Appendix B.3] Several evaluation components are author-affiliated or in-house: Recruit-Bench is described as 'our in-house benchmark'; GDPval and AgentIF-Oneday are run with 'our in-house harness' and an agent judge; ClawGym (ref. [3]) has current Nanbeige team members among its authors. Combined with LLM judges used for scoring, this creates a risk that favorable scoring conventions, prompt formats, or judge behavior advantage the proposed model. The manuscript should provide exact evaluation prompts, judge outputs or human-audit subsets, and ideally third-party or publicly hosted harness runs for at least the non-author-affiliated benchmarks. Without this, the cross-model comparisons in Table 3 and Table 4 are hard to verify independently.
- [Appendix B.2 and Table 3] The paper reports that SWE-bench Verified, SWE-bench Pro, and Terminal-Bench 2.0 results are averaged over 8 independent runs, but Table 3 reports only single point values with no standard deviation, confidence interval, or significance test. On several rows the gaps to the next-best baseline are modest (e.g., OfficeQA-Pro 21.1 vs 15.8, HLE 17.8 vs 14.8, IF-Bench 54.6 vs 55.3). Without variance information it is impossible to know whether the claimed superiority over Qwen3.5-9B and Gemma4-12B is statistically meaningful on those metrics. Per-run scores and significance testing should be reported, especially for the headline agentic benchmarks.
- [§2.1 and Table 1] The architecture section claims that training the looped transformer from scratch performs 'significantly better' than upcycling, that a two-pass loop retains approximately 75% of token efficiency, and that deeper loops give only marginal gains. No experimental table, learning curves, or ablations are shown to support these claims. Since the Looped Transformer is a key component of the paper's contribution, these claims should be documented with concrete numbers and controlled comparisons; otherwise the base-model improvement in Table 1 cannot be attributed to the architecture as opposed to the data mix.
minor comments (6)
- [§2.1] Please specify the exact number of layers, hidden width, loop pass count, and inference-time FLOPs/ latency figures; the current prose leaves the configuration underspecified.
- [Table 3] Use consistent notation for pass rates: 'Claw-Eval Pass^3' appears with a superscript but is not defined in the table caption or Appendix B; state whether this is pass@3 or another aggregation.
- [Table 3] Missing entries for AgentIF-Oneday under Gemma4-E4B and Gemma4-12B are shown as dashes; add a footnote explaining why these baselines were not evaluated.
- [Appendix B.1 and B.2] General inference settings list temperature 0.6, but code-agent evaluations use temperature 1.0. Clarify why the temperature differs and whether the general agent evaluations ever use temperature 1.0.
- [§3.2.3, Eq. (1)] The penalty term uses p_q, the fraction of fully correct responses in the 'current rollout group.' The group size and how p_q is computed are not defined; also, the claim that α<1 guarantees a correct response is always preferred assumes binary base rewards. State this assumption explicitly.
- [§3.3.1] Recruit-Bench is described as in-house, but no description of its size, task distribution, or construction protocol is provided. At minimum, include a benchmark card or public release to allow scrutiny.
Circularity Check
No construction-level circularity; one minor self-referential benchmark does not drive the headline.
specific steps
-
other
[§3.3.1 (General Agents), Table 3, ref [3]]
"[3] Fei Bai, Huatong Song, Shuang Sun, Daixuan Cheng, Yike Yang, Chuan Hao, Renyuan Li, Feng Chang, Yuan Wei, Ran Tao, Bryan Dai, Jian Yang, Wayne Xin Zhao, et al. Clawgym: A scalable framework for building effective claw agents. arXiv preprint arXiv:2604.26904, 2026."
ClawGym is used as one of the evidence columns for the headline agentic-capability claim, and its author list overlaps with Nanbeige4.2-3B's author list (Huatong Song and Shuang Sun). Scores on a benchmark built by the team are not fully independent of the team's choices. However, ClawGym is one of several benchmarks; SWE-bench Verified/Pro, Terminal-Bench 2.0, HLE, GPQA, SciCode, and LiveCodeBench are external and would remain in Table 3 even if ClawGym were removed. Thus the self-citation is minor and not load-bearing.
full rationale
The paper has no derivation chain in which an output is equal by construction to an input. The Looped Transformer, SFT pipeline, RLHF stages, and RL reward in Eq. 1 are design/optimization procedures; the reported length reductions are consequences of the length-control objective, not predictions. The main comparative evaluation includes multiple independent external benchmarks, so the central claim does not reduce to in-house artifacts. The in-house Recruit-Bench and the in-house GDPval/AgentIF harness/judge are benchmark-integrity caveats (lack of contamination audit, no variance), and ClawGym has author overlap; these justify a low non-zero circularity score but not a 6+ finding. Score 2.
Axiom & Free-Parameter Ledger
free parameters (5)
- Loop depth (number of passes through shared layer stack) =
2
- Length budget b_q per problem =
median length of successful historical rollouts
- Penalty magnitude α and constrained/free phase schedules =
not specified numerically
- SFT stage mixture percentages =
82.7/11.6/5.7, 47.8/22.7/29.5, 22.4/68.9/8.7
- Agentic RL difficulty filter =
not specified
axioms (4)
- domain assumption Benchmarks measure the capabilities they claim and are free of contamination.
- ad hoc to paper General RLHF behavioral regularization transfers across tasks and response modes.
- domain assumption Synthesized trajectories and simulated tools transfer to real-world deployment.
- domain assumption Evaluation judge models and same-environment scoring are neutral across compared models.
read the original abstract
We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-agent, and complex tool-use tasks while maintaining highly competitive reasoning capabilities in mathematics, coding, and science. Nanbeige4.2-3B is pretrained from scratch on 28T tokens with a Looped Transformer that reuses the layer stack to increase capacity without adding parameters. For SFT data and trajectory construction, we expand the diversity of executable environments, task assets, and agentic scaffolds through real-world deployment and large-scale synthesis. Our RL pipeline applies mixed-mode RLHF over Think and Non-Think responses to improve overall model quality and reduce failure cases, length-controlled reasoning RL to balance accuracy and reasoning efficiency, and agentic RL with outcome and process rewards to stabilize long-horizon training. Extensive evaluations show that Nanbeige4.2-3B outperforms larger models, including Qwen3.5-9B and Gemma4-12B, across diverse agentic benchmarks while remaining competitive on reasoning and alignment tasks. Performance with OpenClaw further supports its use as a compact local personal assistant.
Figures
Reference graph
Works this paper leans on
-
[1]
Program synthesis with large language models.arXiv preprint arXiv:2108.07732, 2021
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. Program synthesis with large language models.arXiv preprint arXiv:2108.07732, 2021
Pith/arXiv arXiv 2021
-
[2]
Mixture-of-recursions: Learning dy- namic recursive depths for adaptive token-level computation.Advances in Neural Information Processing Systems, 38:96572–96617, 2026
Sangmin Bae, Yujin Kim, Reza Bayat, Sungnyun Kim, Jiyoun Ha, Tal Schuster, Adam Fisch, Hrayr Harutyunyan, Ziwei Ji, Aaron Courville, et al. Mixture-of-recursions: Learning dy- namic recursive depths for adaptive token-level computation.Advances in Neural Information Processing Systems, 38:96572–96617, 2026
2026
-
[3]
Fei Bai, Huatong Song, Shuang Sun, Daixuan Cheng, Yike Yang, Chuan Hao, Renyuan Li, Feng Chang, Yuan Wei, Ran Tao, Bryan Dai, Jian Yang, Wayne Xin Zhao, et al. Clawgym: A scalable framework for building effective claw agents.arXiv preprint arXiv:2604.26904, 2026
Pith/arXiv arXiv 2026
-
[4]
Chaithanya Bandi, Ben Hertzberg, Geobio Boo, Tejas Polakam, Jeff Da, Sami Hassaan, Manasi Sharma, Andrew Park, Ernesto Hernandez, Dan Rambado, Ivan Salazar, Rafael M. O. Cruz, Chetan Rane, Ben Levin, Brad Kenstler, and Bing Liu. Mcp-atlas: A large-scale benchmark for tool-use competency with real mcp servers.arXiv preprint arXiv:2602.00933, 2026
Pith/arXiv arXiv 2026
-
[5]
Agentif-oneday: A task-level instruction-following benchmark for general ai agents in daily scenarios, 2026
Kaiyuan Chen, Qimin Wu, Taiyu Hou, Tianhao Tang, Xueyu Hu, et al. Agentif-oneday: A task-level instruction-following benchmark for general ai agents in daily scenarios, 2026
2026
-
[6]
Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021
Pith/arXiv arXiv 2021
-
[7]
Jasper Dekoninck, Nikola Jovanovi´c, Tim Gehrunger, Kári Rögnvaldsson, Ivo Petrov, Chenhao Sun, and Martin Vechev. Beyond benchmarks: Matharena as an evaluation platform for mathematics with llms.arXiv preprint arXiv:2605.00674, 2026
Pith/arXiv arXiv 2026
-
[8]
Xiang Deng et al. Swe-bench pro: Can ai agents solve long-horizon software engineering tasks? arXiv preprint arXiv:2509.16941, 2025
Pith/arXiv arXiv 2025
-
[9]
Supergpqa: Scaling llm evaluation across 285 graduate disciplines.Advances in Neural Information Processing Systems, 38, 2026
Xeron Du, Yifan Yao, Kaijing Ma, Bingli Wang, Tianyu Zheng, Minghao Liu, Yiming Liang, Xiaolong Jin, Zhenlin Wei, Chujie Zheng, et al. Supergpqa: Scaling llm evaluation across 285 graduate disciplines.Advances in Neural Information Processing Systems, 38, 2026
2026
-
[10]
Adam Fry, Boaz Barak, Chris Painter, Elizabeth Proehl, Grace Kim, Jason Kwon, Michele Wang, Olivia Watkins, Rachel Dias, Ronnie Chatterji, Samuel Miserendino, Tejal Patwardhan, and Tracy Yang. Gdpval: Evaluating ai model performance on real-world economically valuable tasks.arXiv preprint arXiv:2510.04374, 2025
Pith/arXiv arXiv 2025
-
[11]
Tmax: A simple recipe for terminal agents.arXiv preprint arXiv:2606.23321, 2026
Hamish Ivison, Junjie Oscar Yin, Rulin Shao, Teng Xiao, Nathan Lambert, and Hannaneh Hajishirzi. Tmax: A simple recipe for terminal agents.arXiv preprint arXiv:2606.23321, 2026
Pith/arXiv arXiv 2026
-
[12]
Livecodebench: Holistic and contamination free evaluation of large language models for code, 2024
Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Ar- mando Solar-Lezama, Koushik Sen, and Ion Stoica. Livecodebench: Holistic and contamination free evaluation of large language models for code, 2024
2024
-
[13]
Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R. Narasimhan. Swe-bench: Can language models resolve real-world github issues? In International Conference on Learning Representations, 2024
2024
-
[14]
Ruizhe Li, Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, and Zhendong Mao. Deepresearch bench ii: Diagnosing deep research agents via rubrics from expert reports.arXiv preprint arXiv:2601.08536, 2026
arXiv 2026
-
[15]
Towards robust mathematical reasoning
Minh-Thang Luong, Dawsen Hwang, Hoang H Nguyen, Golnaz Ghiasi, Yuri Chervonyi, Insuk Seo, Junsu Kim, Garrett Bingham, Jonathan Lee, Swaroop Mishra, et al. Towards robust mathematical reasoning. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 35406–35430, 2025. 13
2025
-
[16]
Mike A. Merrill et al. Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces.arXiv preprint arXiv:2601.11868, 2026
Pith/arXiv arXiv 2026
-
[17]
Krista Opsahl-Ong, Arnav Singhvi, Jasmine Collins, Ivan Zhou, Cindy Wang, Ashutosh Baheti, Owen Oertell, Jacob Portes, Sam Havens, Erich Elsen, Michael Bendersky, Matei Zaharia, and Xing Chen. Officeqa pro: An enterprise benchmark for end-to-end grounded reasoning.arXiv preprint arXiv:2603.08655, 2026
arXiv 2026
-
[18]
Humanity’s last exam, 2025
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, and et al. Humanity’s last exam, 2025
2025
-
[19]
Generalizing verifiable instruction following
Valentina Pyatkin, Saumya Malik, Victoria Graf, Hamish Ivison, Shengyi Huang, Pradeep Dasigi, Nathan Lambert, and Hannaneh Hajishirzi. Generalizing verifiable instruction following. InAdvances in Neural Information Processing Systems, volume 38, 2025
2025
-
[20]
Qwen3.5: Towards native multimodal agents, February 2026
Qwen Team. Qwen3.5: Towards native multimodal agents, February 2026
2026
-
[21]
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. Gpqa: A graduate-level google-proof q&a benchmark.arXiv preprint, 2023
2023
-
[22]
Hendryx, Brad Kenstler, and Bing Liu
Manasi Sharma, Chen Bo Calvin Zhang, Chaithanya Bandi, Clinton Wang, Ankit Aich, Huy Nghiem, Tahseen Rabbani, Ye Htet, Brian Jang, Sumana Basu, Aishwarya Balwani, Denis Peskoff, Marcos Ayestaran, Sean M. Hendryx, Brad Kenstler, and Bing Liu. Researchrubrics: A benchmark of prompts and rubrics for evaluating deep research agents.arXiv preprint arXiv:2511.0...
arXiv 2025
-
[23]
Openclaw: Your own personal ai assistant, 2026
Peter Steinberger and contributors. Openclaw: Your own personal ai assistant, 2026
2026
-
[24]
Challenging big- bench tasks and whether chain-of-thought can solve them
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc Le, Ed H Chi, Denny Zhou, et al. Challenging big- bench tasks and whether chain-of-thought can solve them. InFindings of the Association for Computational Linguistics: ACL 2023, pages 13003–13051, 2023
2023
-
[25]
Gemma 4 technical report.arXiv preprint arXiv:2607.02770, 2026
Gemma Team, Sherif El Abd, Vaibhav Aggarwal, Robin Algayres, Alek Andreev, Olivier Bachem, Ian Ballantyne, Cormac Brick, Victor C ˘arbune, Michelle Casbon, et al. Gemma 4 technical report.arXiv preprint arXiv:2607.02770, 2026
Pith/arXiv arXiv 2026
-
[26]
Scicode: A research coding benchmark curated by scientists
Minyang Tian, Luyu Gao, Shizhuo Dylan Zhang, Xinan Chen, Cunwei Fan, Xuefei Guo, Roland Haas, Pan Ji, Kittithat Krongchon, Yao Li, et al. Scicode: A research coding benchmark curated by scientists. InAdvances in Neural Information Processing Systems, volume 37, 2024
2024
-
[27]
Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290, 2024
Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, et al. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290, 2024
2024
-
[28]
Sen Xu, Shixi Liu, Wei Wang, Jixin Min, Yingwei Dai, Zhibin Yin, Yirong Chen, Xin Zhou, and Junlin Zhang. Vibethinker-3b: Exploring the frontier of verifiable reasoning in small language models.arXiv preprint arXiv:2606.16140, 2026
arXiv 2026
-
[29]
Toolmind technical report: A large-scale, reasoning-enhanced tool-use dataset
Chen Yang, Ran Le, Yun Xing, Zhenwei An, Zongchao Chen, Wayne Xin Zhao, Yang Song, and Tao Zhang. Toolmind technical report: A large-scale, reasoning-enhanced tool-use dataset. arXiv preprint arXiv:2511.15718, 2025
arXiv 2025
-
[30]
Chen Yang, Guangyue Peng, Jiaying Zhu, Ran Le, Ruixiang Feng, Tao Zhang, Wei Ruan, Xiaoqi Liu, Xiaoxue Cheng, Xiyun Xu, et al. Nanbeige4-3b technical report: Exploring the frontier of small language models.arXiv preprint arXiv:2512.06266, 2025
arXiv 2025
- [31]
-
[32]
Iquest-coder-v1 technical report.arXiv preprint arXiv:2603.16733, 2026
Jian Yang, Wei Zhang, Shawn Guo, Zhengmao Ye, Lin Jing, Shark Liu, Yizhi Li, Jiajun Wu, Cening Liu, X Ma, et al. Iquest-coder-v1 technical report.arXiv preprint arXiv:2603.16733, 2026
arXiv 2026
-
[33]
Claw-eval: Toward trustworthy evaluation of autonomous agents.arXiv preprint arXiv:2604.06132, 2026
Bowen Ye, Rang Li, Qibin Yang, Yuanxin Liu, Linli Yao, Hanglong Lv, Zhihui Xie, Chenxin An, Lei Li, Lingpeng Kong, et al. Claw-eval: Toward trustworthy evaluation of autonomous agents.arXiv preprint arXiv:2604.06132, 2026
Pith/arXiv arXiv 2026
-
[34]
Scaling latent reasoning via looped language models.arXiv preprint arXiv:2510.25741, 2025
Rui-Jie Zhu, Zixuan Wang, Kai Hua, Tianyu Zhang, Ziniu Li, Haoran Que, Boyi Wei, Zixin Wen, Fan Yin, He Xing, et al. Scaling latent reasoning via looped language models.arXiv preprint arXiv:2510.25741, 2025. 15 A Author List Authors are listed inalphabetical order by first name. Names marked with an asterisk (*) denote individuals who were previously affi...
Pith/arXiv arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.