{"id":"39a14c7b-9640-4769-813a-42addff1da58","arxiv_id":"2606.06388","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"ALMANAC is a new dataset of 2,987 annotated dyadic collaboration actions from the Map Task, each with theory-informed mental model annotations for self-reasoning, partner intent, and team goal, used to benchmark six LLMs on predicting next-turn behavior and mental models.","lead":"The paper creates ALMANAC, a dataset of 2,987 human collaboration actions from a map routing task, each annotated with details on self-reasoning, perceived partner intent, and shared team goals. Smart generalists might read it to understand how such data could improve AI agents that work alongside humans rather than just completing tasks alone.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Annotation fidelity to actual participant mental models remains unvalidated, undermining claims of utility for mental-model inference benchmarks.","rationale":"The reader's weakest_assumption directly identifies the same annotation-validity gap that blocks interpretation of the benchmarking results; full-text details on collection protocol would be needed to move beyond this.","tokens_in":1669,"tokens_out":295,"duration_ms":12782,"concrete_test":"Select 200 random action-annotation triples; have the original participant (or a matched independent rater given the same audio/video) re-annotate the three mental-model fields; compute Cohen's kappa per dimension. If mean kappa < 0.6 the central utility claim is weakened.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The headline result ('ALMANAC's utility in evaluating models' ability to ... infer their underlying mental models') requires that the 2,987 theory-informed annotations are faithful proxies for participants' self-reasoning, perceived partner intent, and team goals. The construction uses a classic Map Task but supplies no reported checks (inter-annotator agreement on the three mental-model dimensions, post-hoc participant confirmation of their own annotations, or correlation with observable behavior such as route deviations) that would establish this fidelity. Without such evidence the benchmark scores on next-turn prediction and mental-model inference cannot be interpreted as measuring the intended capability rather than surface-level pattern matching to the annotation scheme itself.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces ALMANAC, a dataset of 2,987 action-level annotations collected from dyadic Map Task collaborations. Each action is paired with theory-informed labels for self-reasoning, perceived partner intent, and perceived team goals. The authors benchmark six LLMs on next-turn behavior prediction and mental-model inference, claiming the results demonstrate the dataset's utility for evaluating models' ability to simulate human collaborative behaviors and infer underlying mental models.","tokens_in":1805,"tokens_out":375,"duration_ms":23693,"significance":"If the annotations are shown to be faithful to participants' actual mental models, ALMANAC would provide a rare resource of process-level human collaboration data that could support development of agents capable of maintaining aligned mental models rather than optimizing solely for task completion. The benchmarking protocol offers a concrete evaluation framework that the community could extend.","major_comments":[{"comment":"§3 (ALMANAC Dataset Construction) and §5 (Benchmarking Experiments): the manuscript reports the annotation scheme and the 2,987 instances but supplies no inter-rater agreement statistics, participant self-validation, or correlation with observable behavior (e.g., route deviations). Because the headline utility claim rests on these annotations serving as accurate proxies for mental models, the absence of such checks is load-bearing; benchmark scores on mental-model inference cannot be interpreted as measuring the intended capability without them.","section":"§3 and §5"}],"minor_comments":[{"comment":"The abstract states the dataset size and benchmarking setup but does not preview any quantitative results (e.g., accuracy numbers or statistical comparisons), making it difficult for readers to gauge the strength of the utility demonstration at first reading.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. We address the major comment below.","responses":[{"response":"The annotations in ALMANAC are collected directly from participants as self-reports of their self-reasoning, perceived partner intent, and perceived team goals immediately following each action. As participant-provided data rather than external labels, inter-rater agreement statistics do not apply. Self-validation is inherent to the collection method. We agree that correlation with observable behaviors (e.g., route deviations) would strengthen claims about the annotations as faithful proxies and will add such an analysis in the revision.","revision_made":"partial","referee_comment":"[§3 and §5] §3 (ALMANAC Dataset Construction) and §5 (Benchmarking Experiments): the manuscript reports the annotation scheme and the 2,987 instances but supplies no inter-rater agreement statistics, participant self-validation, or correlation with observable behavior (e.g., route deviations). Because the headline utility claim rests on these annotations serving as accurate proxies for mental models, the absence of such checks is load-bearing; benchmark scores on mental-model inference cannot be interpreted as measuring the intended capability without them."}],"tokens_in":1322,"tokens_out":260,"duration_ms":30597,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core takeaway is that this paper releases ALMANAC, a dataset of 2,987 action-level annotations drawn from the classic Map Task, each tagging self-reasoning, perceived partner intent, and team goals. That specific framing for LLM agent evaluation is new relative to the task-completion focus in most prior agent work.\n\nThe paper does a clean job identifying the gap: current LLMs are optimized for outcomes rather than maintaining aligned mental models during joint activity, and it supplies a concrete resource built on established social-science task data. Benchmarking six models on next-turn behavior and mental-model inference is a reasonable first use of the data.\n\nThe soft spot is exactly the one flagged in the stress-test note. The abstract states the annotation count and the benchmarking setup but gives zero information on how the annotations were produced, whether multiple raters agreed, whether participants confirmed the labels matched their actual thinking, or any correlation with observable behavior like route changes. Without those checks the benchmark numbers cannot be read as measuring mental-model inference rather than pattern-matching to the annotation scheme. That is a load-bearing issue for the utility claim.\n\nThis paper is aimed at researchers building or testing collaborative LLM agents who need process-level data. A reader already working on mental models or human-AI teaming would find the resource direction useful once the validation details are filled in.\n\nIt deserves a serious referee if the full manuscript contains the missing annotation protocol and reliability numbers; otherwise the central claim stays unsupported.","headline":"ALMANAC adds action-level mental model annotations to the Map Task for LLM collaboration eval, but the abstract supplies no validation evidence for those annotations.","tokens_in":2304,"tokens_out":375,"would_cite":false,"duration_ms":19217,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"ALMANAC supplies 2,987 human collaboration actions each labeled with self-reasoning, partner intent, and team goals.","keywords":["ALMANAC dataset","mental model annotations","agent collaboration","Map Task","LLM evaluation","human collaboration data","theory-informed annotations"],"falsifier":"A follow-up study in which original participants review the annotations for their own actions and report systematic mismatches with what they actually thought at the time, or an experiment showing that LLMs fine-tuned on ALMANAC produce no measurable gain in next-turn prediction accuracy over untuned baselines.","tokens_in":2588,"feed_emoji":"","tokens_out":703,"duration_ms":32953,"temperature":0.7,"pith_summary":"The paper creates ALMANAC, a dataset drawn from the classic Map Task, to fill the absence of action-level mental model data for training collaborative agents. Every action receives theory-informed annotations that record what each participant reasoned about themselves, inferred about their partner, and understood as the shared team goal. Six LLMs are then tested on whether they can predict the next human action and recover those same mental-model labels from the data. A sympathetic reader would care because agents optimized only for task success rarely maintain aligned models of reasoning and intent, and the dataset offers a concrete way to measure and improve that process-level competence.","feed_headline":"Dataset labels 2,987 actions with humans' mental models","feed_subtitle":"ALMANAC records self-reasoning, partner intent and team goals so models can be tested on simulating collaborative turns.","key_machinery":"The ALMANAC dataset of theory-informed mental model annotations (self-reasoning, perceived partner intent, perceived team goal) paired with each of the 2,987 Map Task collaboration actions.","core_discovery":"ALMANAC is a dataset of Action-Level Mental model ANnotations for Agent Collaboration built from the Map Task that contains 2,987 collaboration actions, each paired with theory-informed mental model annotations recording the participants' self-reasoning, perceived partner intent, and perceived team goal. Benchmarking six LLMs on the data demonstrates its utility for evaluating models' ability to simulate human collaborative behaviors and infer their underlying mental models.","pith_inferences":["The same annotation protocol could be applied to other dyadic tasks to test whether mental-model patterns generalize beyond route-finding.","If models improve on ALMANAC, they might sustain longer multi-turn collaborations without explicit goal reminders.","The dataset could support training objectives that penalize divergence between a model's inferred partner model and the human annotations.","Patterns in how humans update their annotations across turns might reveal timing regularities that current agents ignore."],"forward_implications":["LLMs can be benchmarked on predicting humans' next-turn behavior directly from the annotated actions.","Models can be tested for their ability to infer the three mental-model components from observed collaboration turns.","The dataset supplies a concrete signal for moving agents from task-completion optimization toward process-level collaborative competence.","Researchers gain an authentic human baseline against which to measure whether agents align on shared goals and partner intent."],"fun_headline_variants":["Action-level mental models annotated in 2,987 human collaboration steps","ALMANAC links 2,987 actions to self-reasoning, partner intent and goals","Dataset of 2,987 annotated actions tests LLMs on collaboration simulation","Human mental models captured at each step in ALMANAC collaboration data"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The theory-informed annotations accurately capture participants' actual mental models of self-reasoning, partner intent, and team goals during the Map Task collaboration.","fun_headline_variants_meta":{"raw":{"variants":["Action-level mental models annotated in 2,987 human collaboration steps","ALMANAC links 2,987 actions to self-reasoning, partner intent and goals","Dataset of 2,987 annotated actions tests LLMs on collaboration simulation","Human mental models captured at each step in ALMANAC collaboration data"]},"model":"grok-4.3","cost_usd":0.00732,"raw_usage":{"total_tokens":3367,"prompt_tokens":663,"num_sources_used":0,"completion_tokens":80,"cost_in_usd_ticks":73199500,"prompt_tokens_details":{"text_tokens":663,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2624,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":663,"tokens_out":80,"duration_ms":27826,"temperature":1.0,"reasoning_tokens":2624,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T01:16:44.694885+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A follow-up study in which original participants review the annotations for their own actions and report systematic mismatches with what they actually thought at the time, or an experiment showing that LLMs fine-tuned on ALMANAC produce no measurable gain in next-turn prediction accuracy over untuned baselines.","supporting_citations":[],"review_version":1}