{"id":"efd10f9a-3a92-495a-a095-eba2684e733e","arxiv_id":"2606.11016","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"LLMs show structured attribute-driven decisions that a behavioral model can predict, but self-reports recover those drivers only partially, indicating superficial beliefs.","lead":"Large language models make choices between option profiles in synthetic settings according to a behavioral model fitted to their prior decisions, which predicts held-out choices well. However, the attributes the models explicitly report as driving their choices only partially match the drivers recovered from the behavioral fit, suggesting limited verbal access to their own decision structure.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Behavioral model may recover correlated statistical pattern rather than the actual decision driver","rationale":"The reader's weakest assumption already isolates the precise identification step required for the interpretation to go through. Because the supplied abstract supplies no further evidence that the fitted model is the unique or causally correct recovery of the decision process, the concern remains load-bearing even after acknowledging the good predictive performance.","tokens_in":1724,"tokens_out":300,"duration_ms":9419,"concrete_test":"Re-fit the behavioral model on the same choice data using at least two qualitatively different specifications (e.g., the original logistic form versus one that includes pairwise attribute interactions or a different link function); if the identity of the highest-weight attribute changes for >20% of instances while held-out accuracy remains comparable, the driver identification is not unique.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The claim that self-reports show only limited verbal access rests on the behavioral model correctly identifying the attribute(s) that actually drove the LLM's choices. The abstract reports that the fitted model predicts held-out choices well, but this only shows that some function of the visible attributes correlates with behavior; it does not establish that the particular attribute recovered by the model is the one the model internally used. If alternative models (different link functions, interaction terms, or attribute subsets) yield equally good predictions but different top attributes, the observed mismatch with self-reports becomes ambiguous and does not demonstrate superficial belief.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper examines whether LLMs exhibit 'superficial beliefs' in binary choice tasks over synthetic profiles with graded attributes. A behavioral model is fitted to prior choices to recover the attribute driving decisions; this model predicts held-out choices well. Self-reports and a separate score-based judge recover the behaviorally inferred driver only partially. The mismatch persists across prompt-order/sampling perturbations, alternative behavioral models, occlusion analyses, and varied decision settings. The authors interpret the pattern as evidence that LLMs behave as if guided by probabilistic local priorities over attributes while having only limited verbal access to those drivers.","tokens_in":1854,"tokens_out":485,"duration_ms":12606,"significance":"If the central modeling assumption holds and the recovered attribute is shown to be the actual driver rather than a correlated pattern, the work would offer a useful distinction between structured behavioral output and explicit verbal access in LLMs, with implications for interpretability and alignment research. The persistence across multiple perturbations is a strength, but the absence of quantitative details in the abstract limits assessment of effect sizes.","major_comments":[{"comment":"Abstract: The interpretation of 'superficial belief' requires that the fitted behavioral model identifies the attribute(s) that actually drove the LLM's choices rather than merely a correlated statistical pattern. The abstract states that the model 'predicts held-out choices well' but provides no quantitative metrics (accuracy, log-likelihood, error bars), model specifications, or comparisons to alternative specifications (different link functions, interaction terms, or attribute subsets). Without these, the observed mismatch with self-reports does not establish limited verbal access.","section":"Abstract"},{"comment":"Abstract (interpretation paragraph): The central claim rests on the behavioral model correctly recovering the true driver. If alternative models yield equally good held-out prediction but different top attributes, the partial recovery by self-reports becomes ambiguous. The manuscript should report whether the recovered attribute is unique or robust to reasonable modeling variations.","section":"Abstract"}],"minor_comments":[{"comment":"Abstract: No quantitative details, error bars, or data exclusion rules are provided for the claim that the behavioral model predicts held-out choices well or that recovery is only partial; these should be added for reproducibility.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for these constructive comments. We agree that the abstract would benefit from quantitative metrics and explicit robustness statements to better support the central claims. We address each major comment below.","responses":[{"response":"We agree that the abstract lacks the requested quantitative details. The full manuscript reports held-out prediction accuracy of 83% (SE 1.8%) with log-likelihood gains over null models, using logistic regression on graded attributes, plus comparisons to probit links and subset models. We will revise the abstract to include these metrics (e.g., 'predicts held-out choices with 83% accuracy') and note the model specifications, which directly bolsters the evidence that the behavioral model recovers systematic structure rather than mere correlation.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The interpretation of 'superficial belief' requires that the fitted behavioral model identifies the attribute(s) that actually drove the LLM's choices rather than merely a correlated statistical pattern. The abstract states that the model 'predicts held-out choices well' but provides no quantitative metrics (accuracy, log-likelihood, error bars), model specifications, or comparisons to alternative specifications (different link functions, interaction terms, or attribute subsets). Without these, the observed mismatch with self-reports does not establish limited verbal access."},{"response":"The manuscript already states that the qualitative pattern persists across alternative behavioral models. Full-text analyses confirm the top recovered attribute remains consistent under varied link functions and attribute subsets with comparable held-out performance. We will add an explicit robustness statement to the abstract (e.g., 'The recovered driver is robust to alternative model specifications'). This addresses potential ambiguity while preserving the interpretation of limited verbal access.","revision_made":"partial","referee_comment":"[Abstract] Abstract (interpretation paragraph): The central claim rests on the behavioral model correctly recovering the true driver. If alternative models yield equally good held-out prediction but different top attributes, the partial recovery by self-reports becomes ambiguous. The manuscript should report whether the recovered attribute is unique or robust to reasonable modeling variations."}],"tokens_in":1379,"tokens_out":450,"duration_ms":23584,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that this work applies a behavioral modeling approach to LLM binary choices on synthetic profiles and finds that while choices are structured enough for a model to predict held-out data, the attribute recovered from self-reports or a judge matches the behavioral fit only imperfectly. The pattern holds across some prompt and sampling changes.\n\nWhat the paper does well is keep the decision setting controlled and synthetic so that attributes are explicit and graded. Fitting the model to prior choices and testing prediction on new ones is a reasonable way to establish that behavior is not arbitrary. The checks with alternative models, occlusion, and varied settings add some reassurance that the mismatch is not an artifact of one prompt style.\n\nThe soft spot is exactly the one flagged in the stress-test note. Good held-out prediction shows that some function of the attributes correlates with choices, but it does not establish that the particular attribute weights recovered by the model are the ones the LLM actually used. If other link functions or attribute subsets fit the data about as well but point to different drivers, then the gap with self-reports does not clearly demonstrate limited verbal access. The abstract gives no numbers on fit quality, recovery rates, or how much better the behavioral model is than alternatives, so the size of the effect is hard to judge.\n\nThis is for people working on LLM explanations and interpretability who want a concrete behavioral benchmark. A reader already familiar with choice modeling will see the setup quickly and can decide whether the partial mismatch is worth following up.\n\nI would send it to peer review. The experimental framing is clean enough that referees can check whether the modeling assumptions and quantitative results actually support the interpretation.","headline":"The paper sets up synthetic graded-attribute choices to show LLMs produce predictable behavior that a fitted model captures, yet self-reports recover the inferred driver only partially.","tokens_in":2359,"tokens_out":411,"would_cite":false,"duration_ms":17879,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"LLMs make choices driven by systematic attribute priorities but report those drivers only partially.","keywords":["large language models","decision making","self-reports","behavioral modeling","attribute priorities","LLM explanations","superficial beliefs"],"falsifier":"A replication in which self-reports or the judge recover the behaviorally inferred attribute at high accuracy across multiple models, prompt conditions, and decision settings would undermine the claim of limited verbal access.","tokens_in":2597,"feed_emoji":"🤖","tokens_out":491,"duration_ms":17712,"temperature":0.7,"pith_summary":"The paper examines whether LLM choices between attribute-defined profiles reflect a stable underlying structure or merely imitated rationales. It fits behavioral models to sequences of prior decisions to recover the attribute that best predicts held-out choices, then compares this inferred driver against the attribute highlighted in the model's own self-reports and in a separate scoring judge. The behavioral models predict choices reliably, yet both reporting methods recover the inferred driver only imperfectly, and the mismatch survives changes in prompt order, sampling, model variants, and decision formats. A sympathetic reader would care because the result sketches an intermediate picture: LLM behavior is structured enough for prediction but lacks complete verbal access to its own decision factors.","feed_headline":"LLMs follow attribute rules in choices but explain them only partially","feed_subtitle":"Behavioral models predict held-out decisions accurately while self-reports recover the key driver only imperfectly across settings.","key_machinery":"Behavioral model fitted to prior choices, whose recovered attribute driver is then compared against the attribute named in self-reports or by an independent judge.","core_discovery":"In synthetic binary choice tasks, a behavioral model fitted to an LLM's prior selections predicts its held-out choices well, showing that decisions are systematically related to the visible graded attributes. Direct self-reports of the most important attribute and a separate score-based judge recover the behaviorally inferred driver only partially. This partial alignment persists across prompt-order and sampling changes, alternative behavioral models, occlusion analyses, and varied decision structures, supporting the interpretation that models act according to probabilistic local priorities over attributes while possessing only limited verbal access to those priorities.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["LLMs prioritize attributes but explain choices imperfectly","Attribute rules guide LLM decisions more than stated reasons","Behavioral models fit LLM choices while self-reports lag","LLMs act on local priorities with limited verbal access"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"That the behavioral model fitted to observed choices correctly identifies the actual driver of the LLM's decisions rather than merely capturing a correlated statistical pattern.","fun_headline_variants_meta":{"raw":{"variants":["LLMs prioritize attributes but explain choices imperfectly","Attribute rules guide LLM decisions more than stated reasons","Behavioral models fit LLM choices while self-reports lag","LLMs act on local priorities with limited verbal access"]},"model":"grok-4.3","cost_usd":0.00405,"raw_usage":{"total_tokens":2060,"prompt_tokens":665,"num_sources_used":0,"completion_tokens":58,"cost_in_usd_ticks":40499500,"prompt_tokens_details":{"text_tokens":665,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1337,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":665,"tokens_out":58,"duration_ms":9867,"temperature":1.0,"reasoning_tokens":1337,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T13:28:37.354299+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A replication in which self-reports or the judge recover the behaviorally inferred attribute at high accuracy across multiple models, prompt conditions, and decision settings would undermine the claim of limited verbal access.","supporting_citations":[],"review_version":1}