{"id":"1053cb20-fd28-447e-8432-4493315fef01","arxiv_id":"2608.09819","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"The released Macaron-V1-Venti model uses a frozen 744B base plus four per-turn-routed LoRA specialists and reports high internal benchmark scores, but it does not demonstrate cross-generation continual-learning gains.","lead":"Macaron-V1 is an open family of agent models that keep a large frozen base and use several small add-on modules, selecting one per conversation turn through a routing layer. The report describes the architecture, the training and improvement loop, and benchmark results, while explicitly leaving continual-learning and collective-intelligence gains as open questions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Internal benchmark validity is the load-bearing risk: no item-level overlap audit and a same-family judge leave the headline UI4A and Personal Intelligence scores unestablished.","rationale":"The reader's weakest assumption identifies the same load-bearing risk that I would flag: the internal evaluation instruments are the primary evidence for the strongest claims, and their coupling to the training process is admitted but not quantified. I agree with the conditional verdict because the paper is unusually transparent, explicitly scopes what it does and does not claim, and releases weights and harness code. The concern is not that the authors are hiding a known flaw; it is that the validation section does not include the overlap audits, judge-calibration studies, or intervals needed to convert the admitted limitation into a bounded error bar. The UI4A-Bench result deserves particular attention because it is the largest reported separation and the GenUI claim is presented as a direct system-level conclusion. If the L3 specialist was trained on benchmark-derived cases, the 87.8 score is expected and says little about generalization; if the ChatBench judge favors GLM-family responses, the Personal Intelligence lead is similarly uninterpretable. These are concrete, testable risks rather than vague concerns about benchmark quality. A conditional accept remains the right posture: the architecture and infrastructure contributions stand on their own, but the performance claims should be re-tested on frozen external benchmarks and with independent judges before being credited.","tokens_in":44520,"tokens_out":6934,"duration_ms":59866,"concrete_test":"Run a held-out validation of the two headline benchmarks: (1) compute exact, near-duplicate, and semantic overlap between the 161 UI4A-Bench prompts and the L3 training corpus, and likewise between the 46 ChatBench transcripts and the L0/L1 training corpora; (2) re-score a random 30-case UI4A-Bench subset and a 20-case ChatBench subset using a cross-family judge (e.g., Claude Opus 4.6 or GPT-5.5) under the same runtime and scoring policy, with human rating on a further subset if feasible. If the 87.8 versus 67.1 UI4A gap or the 58.3 ChatBench lead narrows by more than about five points or reverses sign, the internal-benchmark validation is inflated and the working-implementation claim needs support from frozen external benchmarks.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that Macaron-V1-Venti is a working MoL system is supported by the three internal evaluation suites, which Section 5.1 itself describes as 'less independent of the training process than a frozen external test set.' The most consequential result is the 87.8 UI4A-Bench Final Score against 67.1 for GLM-5.2, a 20.7-point gap. The paper does not report an item-level overlap audit between the 161 UI4A-Bench cases and the L3 training data, and Section 5.1 states that evaluation artifacts can supply candidates for later training. If L3 was trained on cases drawn from the same product distribution as UI4A-Bench, the gap measures in-distribution fit rather than a general capability. The same structural issue applies to ChatBench and LivingBench, whose seeds come from product failure taxonomies targeted by the RSI loop, as Appendix B.1 concedes. ChatBench additionally uses a private GLM-5.2 judge, the same base family as Venti, so the 58.3 ChatBench score may reward GLM-family stylistic mimicry rather than the behaviors the axioms are meant to capture. The paper is transparent about these limits, but the strongest claim inherits them: absent an overlap audit and cross-family or human judge calibration, the headline numbers are not independent evidence for the architecture. This is a correctness risk from missing evidence, not a disagreement with external consensus.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Macaron-V1, an open agent-model family for \"experiential intelligence,\" organized around two goals: adaptation via recursive improvement of versioned model-harness pairs, and collaboration via Mixture-of-LoRA (MoL) composition on a frozen base. The flagship Macaron-V1-Venti pairs a frozen 744B GLM-5.2 base with four LoRA specialists (chat, agent, coding, GenUI) selected per user turn by an L0-emitted route through a Proxy; a 50B Qwen3.6-based variant (Tall) uses the same design. The algorithmic half describes Model-Harness Co-design (UI4A component-native GenUI harness, REPL agent harness, Harness Context Protocol) and a three-stage MindForge recursive self-improvement loop (Discovery, Expansion, Update), with an isolated Expansion experiment covering 122/122 selected base-failure tasks through HCP-carried configuration changes alone. Infrastructure (MinT, LongStraw, sparse-base rollout-mismatch controls) is reported largely from companion studies. Evaluation covers three internal suites (ChatBench, LivingBench, UI4A-Bench) plus twelve general-capability rows against six frontier baselines; headline numbers are 58.3/64.0 on Personal Intelligence and 87.8 on UI4A-Bench. The paper explicitly disclaims cross-generation continual-learning gains and collective-intelligence emergence, and marks imported versus reproduced values throughout.","tokens_in":44836,"tokens_out":21019,"duration_ms":122405,"significance":"If the evidence were fully established, the contribution would be a credible, released instantiation of a modular continual-learning architecture: the MoL separation (frozen base, registered adapters, per-turn routing with per-adapter conversation views and emergent KV reuse) is a clean engineering substrate, and the RSI Expansion study is a falsifiable, well-scoped measurement that many scored base failures can be recovered purely by harness configuration. The paper is exemplary in its scoping: it marks imported values, admits the internal benchmarks are in-distribution (Section 5.1), states the missing overlap audit and judge calibration (Appendix B.1), disclaims routing-quality equivalence (Section 2.3), and explicitly does not claim cross-generation gains (Section 7.2). It ships open weights and a harness, retains versioned UI4A run manifests, and verifies trace identity by sample ID and input hash in the routing study.","major_comments":[{"comment":"The headline support for the abstract's \"results validate the current system\" is the three internal suites (ChatBench 58.3, LivingBench 64.0, UI4A-Bench 87.8 in Table 8). The manuscript itself states in Section 5.1 that these benchmarks are \"less independent of the training process than a frozen external test set,\" that \"evaluation artifacts can in turn supply candidates for later training,\" and in Appendix B.1 that \"This release does not report a frozen data cutoff, an item-level overlap audit against post-training data, human-judge agreement, or cross-judge sensitivity for either benchmark.\" Because Section 5.5.1 sources the 161 UI4A-Bench cases from \"curated examples, de-identified production traffic, and coverage-gap sampling,\" the 20.7-point UI4A gap against the GLM-5.2 base (87.8 vs 67.1) may measure in-distribution fit of the L3 specialist rather than general UI-generation ability, and the Section 6.4 phrasing that this \"support[s] a direct conclusion about clear, accurate, interactive UI generation\" is accordingly stronger than the current evidence. The paper should add an item-level overlap audit between UI4A-Bench, ChatBench, and LivingBench items and the training corpora, support the conclusions with a held-out external benchmark and a human-scored sample, or downgrade the claims to in-distribution characterization.","section":"§5.1, §5.5.1, §B.1, Table 8"},{"comment":"The ChatBench row is scored by \"a privately deployed GLM-5.2 judge\" (Appendix B.1), the same model family as the Macaron-V1-Venti base. Section 6.2 acknowledges that \"sharing that model family may favor GLM-derived responses\" and that no human or cross-family judge calibration is available, yet the 58.3 score is still presented as a 2.8-point lead over GPT-5.5 (55.5) and a 3.8-point lead over the GLM-5.2 base (54.5). With an unquantified same-family judge, these differences cannot be attributed to conversational quality rather than stylistic mimicry. A scored subset with a second, non-GLM judge or with human raters is needed to bound the effect; otherwise the ChatBench row should be presented as an internal diagnostic rather than a comparative result.","section":"§B.1, §6.2"},{"comment":"The evidence base for the headline comparisons is entirely point estimates. ChatBench and LivingBench average three runs per case, but no intervals, bootstrap, or significance tests are reported for any of the twelve rows in Table 8, and Section 6.4 states that scores lack \"interval or judge-sensitivity analysis.\" The paper applies the right caution to the 0.2-point LivingBench gap (\"should not be interpreted as established superiority\") but not to the 2.8-point ChatBench lead or the UI4A Layer-Score leads. The same statistical weakness affects the routing-quality claim: Table 3 compares five seed-level aggregates per arm with seed identifiers not retained (hence unpaired), and Section 2.3 concedes the comparison \"does not establish equivalence,\" yet Section 2.8 nonetheless lists \"no detected reuse-related degradation\" among the practical consequences of the design. The central MoL claim that routing does not harm task quality needs a paired, adequately powered comparison with retained seed IDs, and the benchmark leads need interval estimates or a uniformly hedged framing.","section":"§2.3 (Table 3), §6.2, §6.4"},{"comment":"The Expansion experiment is cleanly designed in one respect: the model is frozen, no optimizer step is taken, and every change is an HCP-carried edit, so the 122/122 coverage is a genuine demonstration of configuration-search reach on the selected slice. The manuscript also states the key limitation that this is \"a coverage ceiling under adaptive configuration selection, not a held-out estimate of how any single configuration generalizes.\" What the manuscript does not provide is any way to assess the structure of the set: it states outright that \"the retained artifact does not include the per-family task counts,\" so the reader cannot tell whether 122 tasks from 29 TerminalBench 2.1 families are spread evenly or dominated by a few families, which materially changes the interpretation of full coverage. I ask for the per-family counts (and per-family coverage trajectories if available) to be reported, and for a statement of how many of the 122 tasks are passed by configurations that also pass previously covered tasks, so the reach result can be separated from per-task overfitting of the search.","section":"§3.2.5, Table 4"},{"comment":"The abstract's \"Our results validate the current system\" and Section 6.4's \"direct conclusion about clear, accurate, interactive UI generation\" are stronger than the evidence the paper itself describes. Section 7.4 concludes that \"the current results document execution checks for parts of this stack and one model snapshot,\" and Section 7.2 states that the internal suites \"target the same Personal Intelligence distribution that informs the RSI loop\" and \"remain focused in size and scope.\" Given the admissions in Sections 5.1 and B.1 (in-distribution benchmarks, missing overlap audit, unquantified same-family judge), the scoped phrasing of Section 7.4 is the one the abstract and Section 6.4 should use; otherwise the headline claim exceeds what the reported experiments establish.","section":"Abstract; §6.4; §7.4"}],"minor_comments":[{"comment":"The confusion-matrix entries run together (e.g., \"97719 4 0\" should read \"977, 19, 4, 0\"); add column separators so the per-class counts are legible.","section":"Table 2"},{"comment":"The percentage labels on the upper panel of Figure 7 (12.0%, 6.1%, 64.5%, 81.2%) appear to be the pooled pass rates of Table 4, but the caption describes that panel as cumulative coverage; clarify what the labels annotate.","section":"Fig. 7 / Table 4"},{"comment":"Section 1 states that \"MoL demonstrates modular collaboration,\" which is stronger than Section 7.1's \"Neither bet is settled by Macaron-V1\"; this release tests only the four shipped specialists, so the Section 1 sentence should be aligned with the scoped statement.","section":"§1, §7.1"},{"comment":"The -52.99-point MME perception drop for Macaron-V1-Tall is reported without discussion; given the text-only training of the adapters, a sentence interpreting this large negative delta (e.g., routing behavior or adapter interference on vision-language inputs) would help the reader.","section":"Table 10"},{"comment":"The data-governance gap is acknowledged but remains material: the paper evaluates on de-identified product conversations and traffic without documenting the de-identification procedure, residual re-identification audit, retention controls, or consent basis. This should either be documented or the applicability of the results to non-product settings should be further qualified.","section":"§7.2"},{"comment":"The final term of Equation (3) is rendered ambiguously; clarify that the score-memory term scales with the summed response length.","section":"Eq. (3)"}],"recommendation":"major_revision","confidential_remarks":"This is an industrial systems report in which a large share of the load-bearing measurements (MinT catalog scale, LongStraw execution receipts, R3/DSA mismatch controls, REPL-harness substrate comparisons) are executed in companion reports and only summarized here; that is acceptable if the journal treats the paper as a system overview, but the core evaluation instruments are internal to the same lab and the same product distribution the training loop targets, and the benchmark-overlap question will require an audit the authors have, by their own admission, not yet performed. The citation pattern is also heavily self-referential (Mind Lab 2026a-g plus companion reports), which is normal for a system paper but makes external verification harder. I do not see an internal-consistency error that would warrant rejection; the gap is missing evidence, and the authors' own limitation statements are unusually candid."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should read this for the MoL architecture, not for the benchmark scores. The per-turn L0-routed loop with own-view summaries is a clean idea, and they have actually shipped it: weights on HF, harness on GitHub, honest latency numbers, and an expansion study that covers 122/122 base-failure tasks through harness changes alone. That last result is the most interesting thing here—it shows configuration search can elicit behaviors the frozen base already has, without any weight update. What the paper does well is unusual honesty about scope. It explicitly says the internal ChatBench, LivingBench, and UI4A-Bench suites are \"less independent of the training process than a frozen external test set.\" It defers continual learning and collective intelligence claims to future work. It distinguishes imported public numbers from its own measurements, and it presents routing accuracy and KV-reuse data as implementation diagnostics, not generalization claims. The soft spot is exactly where the stress-test points: the headline UI4A and Personal Intelligence scores are load-bearing and not independently established. There is no item-level overlap audit between UI4A-Bench and L3 training data; ChatBench uses a private GLM-5.2 judge from the same family as Venti; there are no confidence intervals or human-judge calibration. The 20.7-point UI4A gap could be in-distribution fit. That said, the authors flag all of this themselves in Section 5.1 and Appendix B.1, so it is a correctness risk from missing evidence rather than a hidden flaw. Two smaller caveats: the paper leans heavily on companion reports from the same lab (MinT, LongStraw, UI4A, R3), and the abstract's \"current system validation\" is more upbeat than the body's careful hedging. Neither is disqualifying, but the reader should check whether those companion reports are separately verified. I would send this to peer review. The architecture and released infrastructure deserve referee time, but the evaluation sections need major work: external benchmarks, released eval artifacts, human or cross-family judge calibration, and overlap audits. A serious editor should not desk-reject, but should expect heavy revision on the evidence side.","headline":"A transparent systems report with a genuinely novel MoL serving design, but the headline numbers rest on lab-built benchmarks that the paper itself admits are in-distribution; worth engaging, not worth taking as independent validation.","tokens_in":45645,"tokens_out":2094,"would_cite":false,"duration_ms":19078,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Macaron-V1 claims a frozen 744B base with four per-turn-routed LoRA specialists is a working Mixture-of-LoRA system: 87.8 on UI4A-Bench and 122/122 base-failure tasks covered by harness search alone.","keywords":["Mixture-of-LoRA","continual learning","recursive self-improvement","model-harness co-design","generative UI","agentic reinforcement learning","long-context RL","experiential intelligence"],"falsifier":"Re-run the 46 ChatBench cases and 40 LivingBench scenarios with human judges, or with a judge model from a different family than GLM-5.2, and check whether Macaron-V1-Venti's leads of 58.3 over 55.5 and 64.0 over 63.8 survive; if they collapse, the Personal Intelligence claims are judge-family artifacts. Separately, run the HCP configurations discovered in the 122-task Expansion study on a fresh, unselected slice of TerminalBench-family tasks; if coverage drops toward the 11/122 single-configuration baseline, the 122/122 result reflects adaptive search over a curated failure set rather than a general property of harness search.","tokens_in":44287,"feed_emoji":"🧩","tokens_out":15390,"duration_ms":103808,"temperature":0.7,"pith_summary":"Macaron-V1 is a bid to make continual learning an architectural property rather than a training add-on: a frozen base model carries general ability, specialised LoRA adapters (small plug-in training matrices) carry differentiated behaviour, and a per-turn routing loop chooses which adapter answers. The paper presents Macaron-V1-Venti, a 744B GLM-5.2 base with four such adapters for chat, agent, coding, and GenUI, as a working instance of this Mixture-of-LoRA design, reporting 87.8 on UI4A-Bench, 58.3 on ChatBench, and 64.0 on LivingBench. It also reports that harness-side configuration search alone, with the model frozen, covers all 122 tasks the base fails outright, where the best single configuration covered 11, so many apparent model failures are unelicited behaviours rather than missing skills. The deeper claim is that the harness is a first-class optimisation target: tools, prompts, skills, and UI substrates can be revised, evaluated, and shipped without touching weights, which is what makes the continual-learning story concrete. The paper is explicit that cross-generation compounding and collective-intelligence gains are open questions, so a sympathetic reader should read it as a systems characterisation with two open bets.","feed_headline":"Frozen 744B base plus four LoRA specialists posts frontier UI scores","feed_subtitle":"Harness-only search covers all 122 tasks the frozen base fails, and cross-generation learning stays an open bet.","key_machinery":"The load-bearing object is the Mixture-of-LoRA (MoL) serving layer: a frozen base, a small registry of specialist LoRA adapters, and a Proxy that treats adapter selection as a first-class per-turn action. The route label is emitted by the chat adapter L0 under a constrained-decoding grammar, so the router is not a separate model but a property of the chat specialist's understanding of the request. Two mechanisms make the loop cheap: the own-view, which rebuilds each specialist's conversation deterministically from an append-only timeline (own turns verbatim, other specialists collapsed to 192-token summaries) and thereby gives emergent per-adapter KV-prefix reuse; and the summary hop, which caps cross-adapter state at 192 tokens. On the learning side, the carried object is the model-harness pair, written $\\pi_\\varphi(a_t \\mid o_{\\le t}; \\theta, c)$ with $\\theta$ the frozen base, $\\varphi$ the trainable LoRA parameters, and $c$ a versioned harness configuration: the recursive self-improvement cycle (Discovery, Expansion, Update) alternates configuration search over $c$ with GRPO updates to $\\varphi$, and the reported Expansion experiment isolates the $c$-search half. The harness side is carried by three named substrates: UI4A, a component-native generative-UI harness where the model writes ordinary frontend code under runtime-enforced boundaries; the REPL agent harness with executable composition and validated helper reuse; and the Harness Context Protocol, a versioned TOML contract that makes a run reconstructable at the configuration boundary.","core_discovery":"On the paper's own terms, the central discovery is that four specialist LoRA adapters on a frozen 744B base, selected per user turn by the chat adapter's own constrained-decoded label, form a workable substitute for a single monolithic post-trained model. The MoL Proxy runs a three-hop loop: L0 routes the request in 24 tokens, the chosen specialist answers from its own conversation view, and a 192-token summary preserves continuity across specialists, with routing accuracy of 99.12% on a 6,448-sample trace and a measured route-plus-summary overhead of about 32% of per-turn latency. The same system reports 87.8 on UI4A-Bench against 75.9 for the strongest published baseline, alongside leading scores on its own Personal Intelligence benchmarks. The second claimed result is that the harness is improvable as well as the model: across 69 jobs and 450 attempts on 122 TerminalBench-family tasks that the frozen GLM-5.2 base fails under the official reward, adaptive configuration search reaches 122/122 cumulative coverage without a single optimizer step. The paper deliberately stops short of claiming that this demonstrates continual learning, defined as compounded gains across model generations, or collective intelligence, defined as complementary gains from independently trained specialists.","pith_inferences":["If the 122/122 coverage generalises beyond the curated failure slice, configuration search becomes the cheap first line of continual improvement, with LoRA updates reserved for behaviours that configuration cannot elicit; the paper's own 13x per-attempt yield ratio between targeted search and full-set sweeps hints at this but is not a controlled comparison.","The 192-token summary is an information bottleneck between specialists; a natural test is varying summary length, or replacing summaries with the shared-L0-KV substrate the paper sketches, and measuring cross-specialist task quality.","The routing accuracy figure of 99.12% comes from LoRA training data, so a held-out routing audit is the cleanest next check of whether L0's routing generalises.","Because the internal benchmarks are judged by LLMs from the same model families as the systems under test, an external human-judge or cross-family-judge calibration would settle whether the Personal Intelligence leads reflect assistance quality or in-distribution mimicry."],"forward_implications":["New capabilities can ship as adapter registrations on a frozen base, so the base, the specialists, and the harness each move on their own release clock.","Configuration search is a first-pass improvement path: 122/122 coverage versus 11/122 for the best single full-set configuration implies that many apparent model failures are unelicited behaviours rather than missing skills.","The MoL resident layout stores about 26% of the replicated-base parameter count, a 74% reduction, which is what makes multi-specialist long-context serving feasible on fixed hardware.","Routing by the chat adapter itself means routing quality improves for free as the base or the chat specialist improves, at a measured cost of roughly one third of per-turn latency.","Because only adapters receive gradients, the base cannot drift as a side effect of specialisation, although routing and harness changes can still alter end-to-end behaviour."],"supporting_citations":[{"why":"Supplies LoRA, the low-rank adapter mechanism that MoL freezes the base around and layers specialists on top.","marker":"Hu et al., 2022"},{"why":"The open-sourced MoL serving harness: the reference implementation of the routing loop, own-view construction, and summary hop whose costs and accuracy the paper measures.","marker":"Mind Lab, 2026d"},{"why":"The REPL agent harness study: establishes that the action substrate changes success rate and token cost, and supplies the executable-composition and validated-reuse mechanisms behind the L1 specialist.","marker":"Wu et al., 2026"},{"why":"Introduces UI4A, the component-native generative-UI harness, and the token-length comparison (672 versus 1,224 output tokens) that underpins the L3 specialist and UI4A-Bench.","marker":"Zhuang et al., 2026"},{"why":"MinT, the post-training platform that gives adapter revisions their lifecycle, lineage, and the measured million-entry catalog assumed by the MoL registry.","marker":"Lu et al., 2026"},{"why":"GRPO, the reinforcement-learning update used to train the LoRA specialists from selected trajectories in the RSI Update stage.","marker":"Shao et al., 2024"},{"why":"Source of the 29 families from which the 122-task base-failure set is drawn for the Expansion coverage experiment.","marker":"TerminalBench Authors, 2026"},{"why":"R3 router replay, one of the sparse-base rollout-training mismatch controls that keep the agentic RL on-policy.","marker":"Ma et al., 2025"},{"why":"Documents GLM-5.2, the frozen 744B base that MoL's design depends on and against which the judge-family overlap risk is assessed.","marker":"GLM-5 Team, 2026"}],"fun_headline_variants":["Frozen 744B base plus four LoRA specialists beats monolithic","LoRA routing hits 99.12% on trace, tops UI4A-Bench","Harness-only search covers all 122 tasks base fails","Four LoRA specialists on frozen base: top UI scores, no optimizer","Macaron-V1: self-improvement without model retraining, but not continual"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the internal evaluation instruments — ChatBench's judge drawn from the same GLM-5.2 family as the Venti base, LivingBench's simulated sandbox with LLM judges, and UI4A-Bench's scoring policy — measure genuine assistance quality rather than rewarding outputs that resemble the training distribution; the paper itself concedes that these benchmarks are 'less independent of the training process than a frozen external test set'.","fun_headline_variants_meta":{"raw":{"variants":["Frozen 744B base plus four LoRA specialists beats monolithic","LoRA routing hits 99.12% on trace, tops UI4A-Bench","Harness-only search covers all 122 tasks base fails","Four LoRA specialists on frozen base: top UI scores, no optimizer","Macaron-V1: self-improvement without model retraining, but not continual"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000238,"raw_usage":{"total_tokens":1596,"prompt_tokens":1116,"completion_tokens":480,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":732,"completion_tokens_details":{"reasoning_tokens":382}},"tokens_in":732,"tokens_out":480,"duration_ms":5154,"temperature":1.0,"reasoning_tokens":382,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:14:48.792250+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the 46 ChatBench cases and 40 LivingBench scenarios with human judges, or with a judge model from a different family than GLM-5.2, and check whether Macaron-V1-Venti's leads of 58.3 over 55.5 and 64.0 over 63.8 survive; if they collapse, the Personal Intelligence claims are judge-family artifacts. Separately, run the HCP configurations discovered in the 122-task Expansion study on a fresh, unselected slice of TerminalBench-family tasks; if coverage drops toward the 11/122 single-configuration baseline, the 122/122 result reflects adaptive search over a curated failure set rather than a general property of harness search.","supporting_citations":[],"review_version":1}