{"id":"5406b9a2-5988-4f4c-a3af-9dcb09a5995d","arxiv_id":"2608.05729","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A compact carried state of engagement evidence, stated facts, and the standing request lets one agent answer device-unspecified requests later, outperforming full-context and memory/multi-agent baselines on the authors' new UA-BENCH.","lead":"The paper proposes Unified Agent, a system that keeps a compact memory of engagement evidence, stated facts, and the user's standing request so one AI agent can act across a user's devices and over time. The authors build a synthetic benchmark of 200 cross-device episodes and report that this state design beats adapted published baselines and a full-context control across several multimodal LLMs.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The large gains over published baselines may be an input-feature artifact: Appendix D.2 restricts baselines to a subset of the shared perception record, so near-floor Eng/Inf could reflect missing visual cues rather than state design; the only full-record control (Full context) shows a 0.055 gap.","rationale":"The reader's weakest assumption is about benchmark circularity: UA-BENCH's ground truths are generated from the same categories (engagement facts, stated facts, standing request) that Unified Agent stores, so the benchmark may reward the method by construction. That is a serious external-validity concern, and the single real-photo pair is thin. However, the more immediate internal-validity risk is the baseline protocol. If the four published baselines are evaluated on a restricted subset of the shared perception record, the large Table 1 gaps do not compare state designs; they compare which input fields are available. The main text's 'shared per-frame perception output' language gives the impression that all MLLM methods receive identical inputs, and Appendix D.2's 'restricted to the input fields their designs specify' is easy to miss. The paper does not list which fields each baseline received. Mem0's Eng=0.082 and MM-DST's Inf=0.025 are consistent with missing visual engagement and fact-binding information. Full context, which receives everything, beats or nearly matches Unified Agent on Eng, Int, and Rsp, and is only 0.055 behind overall, so the state-design advantage over raw history is real but small. The central claim that state representation, not raw history or model strength, is decisive therefore rests mainly on a 0.055 controlled gap, while the large baseline gaps may be artifacts. My proposed check, rerunning baselines with the full record, settles this directly. If the gaps persist, the baseline comparison is sound; if they collapse, the conclusion should be restricted to 'a compact explicit aggregate helps over raw history,' not that Unified Agent's specific categories outperform other state designs. This is consistent with the reader's CONDITIONAL verdict, but for a different, more specific reason.","tokens_in":18056,"tokens_out":12478,"duration_ms":109113,"concrete_test":"Obtain the code and data from the authors and re-run Mem0, MM-DST, Mixture-of-Agents, and Debate-or-vote on the default GPT-5.6-Luna setting while giving each baseline the same full per-frame shared perception record that Full context receives (all fields, including attended-device, activity counts, pointing cues, stated values), keeping the baseline's own state mechanism, decoder, and prompt template otherwise unchanged. Compare Overall scores and paired-gap confidence intervals to Table 1. If gaps remain at least 0.10 with Holm-adjusted p < 0.05, the advantage is not an input artifact; if Eng/Inf/Nxt scores rise and gaps approach the Full-context gap of about 0.055, the reported large gains are caused by withheld input fields.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Appendix D.2 states that compared systems 'begin from the shared perception record, restricted to the input fields their designs specify.' In the main text (Section 5) this is described as a 'shared per-frame perception output,' which implies all MLLM-based methods receive the same record. If Mem0 and MM-DST are restricted to text/fact fields and are not given the attended-device, activity-count, or pointing entries, their Eng scores (0.082 and 0.195) and Inf scores (0.395 and 0.025) are near floor for a reason unrelated to state design: the information the ground truth is computed from (Appendix B.3) is withheld. The controlled comparison against Full context, which does receive the complete record, yields a much smaller gap (0.055, CI +0.035 to +0.075). Thus the headline gaps of 0.380-0.408 do not isolate 'carried state representation'; they may only show that a system without visual engagement input cannot identify the engaged device. The paper does not report what fields each baseline received, so the reader cannot verify that the comparison is fair. If the restriction is confirmed, the central claim is supported only by the 0.055 full-context gap, which is below the paper's own 0.10 effect-size criterion for the published-baseline family.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Unified Agent, a stateful agent for cross-device, cross-time interaction that carries a compact state composed of three streams: per-device engagement evidence, stated facts, and the standing request. The authors introduce UA-BENCH, a rendered 3D benchmark with 100 matched pairs (200 episodes) in which later device-unspecified requests depend on earlier interaction evidence, plus a small real-photo matched-pair case study. The agent is evaluated on five downstream decision tasks (Eng, Int, Inf, Rsp, Nxt) under four MLLM settings. The authors report that Unified Agent significantly outperforms adapted published baselines (gaps 0.194-0.408) and the full-context control (gap 0.055), with statistical significance assessed by clustered bootstrap and Holm correction, and that the advantage persists across MLLM families, capabilities, and reasoning efforts.","tokens_in":18277,"tokens_out":5499,"duration_ms":48948,"significance":"If the central claim holds, the paper makes a useful and simple contribution: a compact, inspectable carried-state design that is action-ready and bounded, in contrast to full-history retention. The strengths are the explicit state specification, the carefully controlled matched-pair evaluation, clustered bootstrap uncertainty quantification, Holm-adjusted multiple comparisons, the inclusion of strong state controls such as Full context, and the promise of public code and data. The paper also presents ablations separating the roles of engagement counts, pointing, prior state, and device abilities, which are informative. However, the significance is conditional on resolving two issues: the claimed large advantages over published baselines may be driven by asymmetric input features rather than state design, and the benchmark's ground truth is generated from the same categories that Unified Agent's state explicitly tracks, which may make the measured advantage partly construction-internal. The external evidence is limited to one author-prepared real-photo matched pair.","major_comments":[{"comment":"The manuscript states in Section 5 that all MLLM-based methods 'are evaluated using a shared per-frame perception output,' but Appendix D.2 clarifies that compared systems 'begin from the shared perception record, restricted to the input fields their designs specify.' The paper never reports the specific input-field subset given to each baseline. If Mem0 and MM-DST are restricted to text/fact fields and are not given the attended-device, activity-count, or pointing entries, their near-floor Eng scores (0.082 and 0.195) and Inf scores (0.395 and 0.025) are explained by missing inputs rather than by a failure of state design. The only comparison that truly holds inputs fixed is Unified Agent versus Full context, and that gap is 0.055 (CI +0.035 to +0.075), which is below the 0.10 effect-size threshold the paper itself uses for deciding that a published-baseline gap supports a conclusion. Please provide the exact fields each baseline received, justify each restriction as a faithful adaptation of the published design, and ideally re-run the baselines on the complete shared record. Without this, the headline gaps of 0.380-0.408 do not isolate the contribution of carried-state representation.","section":"Section 5 / Appendix D.2 / Table 1"},{"comment":"The benchmark's ground truths for Eng, Inf, Rsp, and Nxt are, by the authors' own description, deterministic functions of the construction script, device metadata, and controlled role assignment, which specify engagement schedules, stated facts, and standing requests. Unified Agent's state is defined to store exactly these categories: engagement counts, pointing cues, first stated facts per device-topic, and the standing request. Consequently, the evaluation targets are generated from the same conceptual categories that Unified Agent tracks, and its strong performance may reflect agreement with the construction protocol rather than a general property of stateful agents. The real-photo matched pair (Appendix F) is a single author-prepared example and does not establish external validity. I request either an evaluation on episodes whose ground truth is not derived from the method's state categories (e.g., human-annotated or independently generated targets) or a clear and explicit scoping of the claims as benchmark-internal rather than general.","section":"Appendix B.3 / Appendix A.1"},{"comment":"The advantage over the full-record control is small and concentrated in one decision. Comparing Unified Agent with Full context: Eng +0.012, Int -0.031, Inf +0.010, Rsp -0.023, and Nxt +0.305. Thus the overall gap of 0.055 is almost entirely driven by Nxt, while Full context is actually slightly better on Int and Rsp. Because Full context receives all the same information, the paper should analyze whether the Nxt difference is a robust state-design effect or an artifact of the specific Nxt output space and scoring rule (exact action-device match on designated frames). Otherwise the broader statement that a well-organized carried state provides a performance advantage is stronger than the evidence supports.","section":"Table 1 / Section 6.2"}],"minor_comments":[{"comment":"The caption and axis labels are difficult to parse ('131 1,442'); please clarify whether the numbers are characters at specific frames or a range, and ensure the figure is readable in grayscale.","section":"Figure 5"},{"comment":"The Limitations section reads partly as a defense of the design rather than a limitation statement; please acknowledge more directly that an explicit carried state is itself a privacy-sensitive artifact and that the real-photo evidence is a single matched pair with no uncertainty quantification.","section":"Section 8 (Limitations)"},{"comment":"The phrase 'significant outperforms' is used in the abstract and contributions; in the main text significance is defined only in terms of bootstrap p-values, not an effect-size criterion for the state controls. Please align the wording with the statistical definitions used in Section 6 and Appendix E.","section":"Abstract / Section 1"},{"comment":"The baseline adaptations are described only briefly; please add a sentence for each baseline about which input fields it receives and how the adaptation maps the original method's interface to the shared perception record, so that the fairness of the comparison is verifiable.","section":"Appendix D.2"},{"comment":"The paper states that code and data will be publicly released but provides no repository link or availability artifact; include a URL or an availability statement at the time of publication.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper's core idea is interesting and the experimental framework is unusually careful in its use of matched pairs, bootstrap, and multiple-testing correction. However, the input-feature asymmetry in the published-baseline comparisons is potentially decisive, and the benchmark's ground truth is closely aligned with the method's state categories. If the authors can supply the per-system input field specifications and re-run the comparisons with matched inputs, or provide an independent evaluation that breaks the construction-method alignment, the paper could become a solid contribution. At present, the evidence for the headline claims is not fully convincing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this paper has a real new problem formulation and a carefully built benchmark, but the strongest numbers do not isolate what the authors claim. The 0.38–0.41 gaps over Mem0 and MM-DST are probably input-feature artifacts. Appendix D.2 says compared systems receive the shared perception record \"restricted to the input fields their designs specify,\" and the paper never reports which fields each baseline actually got. If Mem0 and MM-DST never see the attended-device, activity-count, or pointing entries, their near-floor Eng and Inf scores tell us nothing about state representation. The only control that receives the complete record is Full context, and the gap there is 0.055—below the paper's own 0.10 effect-size criterion for the published-baseline family. So the robust state-design advantage is, on current evidence, a modest advantage over full raw history, not the large advantage the abstract implies.\n\nWhat is genuinely good: the observe–think–act task decomposition for cross-device, cross-time interaction is useful and well-motivated. UA-BENCH's matched-pair construction—shared request/recall frames with role-swapped engagement—is a clever way to isolate carried evidence, and the statistical protocol (clustered bootstrap, Holm correction, aligned endpoints) is more careful than most of this literature. The ablations coherently show that engagement counts, pointing, prior state, and device abilities each contribute to their targeted decisions. The real-photo case study is thin—one prepared pair, no human agreement on the intent judge—but at least it is an attempt at external validation, and the authors admit the limitation.\n\nThe circularity concern is real but not fatal. The benchmark's ground truths are deterministic functions of the same three streams—engagement evidence, stated facts, standing request—that Unified Agent's state stores. That tests whether an agent explicitly built to track those streams can exploit them, which is a legitimate design question, but it inflates the apparent advantage over methods not built that way. The authors need to report exactly what each baseline received, add a full-record baseline that is not just raw history (or justify why raw history is the right control), and either pre-register or justify the 0.10 criterion. The intent judge being a single LLM without human agreement is a minor issue for relative comparison but should be disclosed.\n\nThe citation pattern looks fine; the related work is current and relevant. Code and data are promised but not yet available, so independent verification is impossible right now.\n\nBottom line: a serious paper worth engaging, and the benchmark could become a useful resource once code and data are out. I would send it to peer review, expecting major revision: fix the baseline-input reporting, temper the central claim, and address the benchmark's built-in advantage for the proposed state design. It is not a paradigm shift, but it is a solid step for agent-state research.","headline":"A genuinely new problem and a careful benchmark, but the headline gains over published baselines are likely input-feature artifacts; the honest effect size vs the full-record control is small, and the benchmark's targets encode the method's own state categories.","tokens_in":18838,"tokens_out":2466,"would_cite":false,"duration_ms":22960,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A compact carried state, organized as engagement evidence, stated facts, and the standing request, is the decisive factor for cross-device, cross-time agent requests.","keywords":["cross-device agents","carried state","engagement evidence","standing request","agent benchmark","multimodal large language models","device-unspecified requests"],"falsifier":"Build a version of UA-BENCH where the five downstream targets come from an independent protocol (e.g., human annotators watching the full device-camera history) rather than from the construction script, and check whether Unified Agent's margin over the Full-context control persists; if it shrinks or vanishes, the reported advantage reflects the benchmark's construction rather than a general property of state design.","tokens_in":17830,"feed_emoji":"🤖","tokens_out":15389,"duration_ms":104746,"temperature":0.7,"pith_summary":"The paper argues that an agent serving one user across devices and moments fails not because the model is weak or the history is missing, but because the state it carries forward is badly organized. It proposes Unified Agent, which compresses each observation into three streams — engagement evidence, stated facts, and the standing request — and reads that compact state together with the current view to act. On the new UA-BENCH benchmark, this design beats adapted versions of four published agent designs by 0.194 to 0.408 points and the full-history control by 0.055, with the advantage holding across four multimodal large-language-model settings. The point of the paper is that the representation of carried state, not raw history or model strength, is the decisive factor for cross-device, cross-time requests.","feed_headline":"Carried state, not full history, wins cross-device agent tests","feed_subtitle":"An agent that tracks engagement, facts, and the pending request answers device-unspecified asks across time.","key_machinery":"The carried state $S_t$ is the load-bearing mechanism: a compact tuple $(C_t, P_t, K_t, r_t)$ updated by folding each observation ($S_t = U(S_{t-1}, O_t)$) and read together with the current observation to act ($a_t = D(S_t, O_t)$). Engagement evidence accumulates per-device activity and attended-device cues as tallies; stated facts keep the first fact per device–topic pair; the standing request stores the latest actionable request. The state is what lets a later decision recover evidence after the original cue is no longer visible; the paper's claim is that this specific organization — not raw history, not decoded answers, not per-agent memory — is what carries the cross-device advantage.","core_discovery":"Unified Agent is a stateful agent design for the setting where one agent serves one user across multiple devices over time, with observations arriving one device-camera view at a time and no replay of earlier moments. At each step it folds the current observation into a carried state $S_t = (C_t, P_t, K_t, r_t)$, where $C_t$ and $P_t$ store engagement evidence (activity and attended-device cues, with per-device tallies), $K_t$ stores the first stated fact for each device–topic (or topic-only) pair, and $r_t$ stores the latest standing request. The agent then acts from the updated state together with the current observation, producing the five downstream decisions: identify the engaged device, infer intent, recall relevant information, pick the responding device, and decide the next action. On UA-BENCH, 100 matched pairs (200 episodes) of rendered 3D interactions in which the later device-unspecified request and recall questions depend on swapped earlier engagement roles, Unified Agent achieves the highest overall score in the default GPT-5.6-Luna low-effort setting (0.668), leading the full-context control (0.613) and four adapted published designs (0.474, 0.474, 0.288, 0.260), with paired-bootstrap gaps whose 95% intervals are all above zero; it remains ahead in all four MLLM settings tested. The paper presents this as evidence that the way carried state is represented and used, not added system complexity or model strength, is what makes cross-device, cross-time requests answerable.","pith_inferences":["Beyond the paper: the same state-design principle may apply to single-device long-horizon tasks (e.g., a browser agent that must remember a fact from an earlier tab), where the deciding cue also passes before the later request; the compact three-stream state is a candidate generalization.","Beyond the paper: a testable extension is to replace UA-BENCH's scripted engagement cues with real behavioral signals (gaze, touch, app usage) and check whether the engagement-tally mechanism still resolves device-unspecified references; the current real-photo case uses hand-on-device contact as a stand-in.","Beyond the paper: the paper's benchmark construction defines ground truth directly from the same categories the method stores; an independent, human-annotated variant of UA-BENCH would clarify whether the measured advantage is a property of the state design or of the benchmark's construction rules.","Beyond the paper: if the state-design advantage holds in production, one implication is that agent memory systems should expose a small, revision-friendly state to the user (what the agent believes about engagement, facts, and pending requests) rather than a raw transcript, which would also make privacy review and data minimization more tractable."],"forward_implications":["Designers of cross-device agents should prioritize a compact, action-ready carried state over retaining the full interaction transcript; the paper shows a bounded state (11 times smaller than full context after 12 frames) yields higher overall accuracy.","The three-stream organization (engagement evidence, stated facts, standing request) sets the agenda for what an agent should persist: evidence for device attribution, topic-keyed facts, and the pending request, rather than raw frames or cached answers.","The advantage generalizes across MLLM families, capabilities, and reasoning efforts, so the state-design lesson transfers to different foundation models without fine-tuning or extra machinery.","Ablations predict which state element matters for which decision: engagement counts and pointing for identifying the engaged device, prior state for recall, and per-device ability clauses for routing.","Because the state is explicit, it can be inspected and selectively revised, which is a privacy-relevant property: the agent keeps a minimal record rather than the full interaction history."],"supporting_citations":[{"why":"Supplies the Mem0 baseline, which maintains retrieved fact memory as an alternative to compact carried state.","marker":"Chhikara et al., 2025"},{"why":"Supplies the MM-DST baseline, a slot–value multimodal dialogue state tracking approach the method is compared against.","marker":"Le et al., 2022"},{"why":"Extends the multimodal dialogue-state baseline with attention-based video-grounded embeddings.","marker":"Abdessaied et al., 2024"},{"why":"Supplies the Mixture-of-Agents baseline, which synthesizes per-device proposals from running summaries.","marker":"Wang et al., 2025"},{"why":"Supplies the Debate-or-vote baseline, which coordinates per-device agents through peer revision and voting.","marker":"Choi et al., 2025"},{"why":"Provides the precedent for exact semantic ground truth and controlled visual variation that UA-BENCH follows.","marker":"Johnson et al., 2017"},{"why":"Supplies the 3D co-habitat simulator used to render UA-BENCH scenes.","marker":"Puig et al., 2024"},{"why":"Supplies the HSSD-200 dataset of realistic 3D scenes used for the benchmark's home environments.","marker":"Khanna et al., 2024"},{"why":"Supplies the MuJoCo Menagerie humanoid robot used in episodes that field a fetch robot.","marker":"Zakka et al., 2022"}],"fun_headline_variants":["Compact state, not full history, wins agent tests","Stateful agent bests replays in cross-device tasks","Unified Agent: carried state outperforms full logs","Cross-device agents thrive on action-ready state","State design, not model might, lifts agent scores"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"If the benchmark's ground-truth rules already encode the very categories the method stores, the reported advantage may come from the benchmark's design rather than from the state design itself.","fun_headline_variants_meta":{"raw":{"variants":["Compact state, not full history, wins agent tests","Stateful agent bests replays in cross-device tasks","Unified Agent: carried state outperforms full logs","Cross-device agents thrive on action-ready state","State design, not model might, lifts agent scores"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00034,"raw_usage":{"total_tokens":1952,"prompt_tokens":1102,"completion_tokens":850,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":718,"completion_tokens_details":{"reasoning_tokens":773}},"tokens_in":718,"tokens_out":850,"duration_ms":7389,"temperature":1.0,"reasoning_tokens":773,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:34:57.288970+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build a version of UA-BENCH where the five downstream targets come from an independent protocol (e.g., human annotators watching the full device-camera history) rather than from the construction script, and check whether Unified Agent's margin over the Full-context control persists; if it shrinks or vanishes, the reported advantage reflects the benchmark's construction rather than a general property of state design.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Extends the multimodal dialogue-state baseline with attention-based video-grounded embeddings."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Debate-or-vote baseline, which coordinates per-device agents through peer revision and voting."},{"cited_title":"Lawrence Zitnick, and Ross B","cited_arxiv_id":null,"evidence_quote":"Provides the precedent for exact semantic ground truth and controlled visual variation that UA-BENCH follows."},{"cited_title":"Chang, and Manolis Savva","cited_arxiv_id":null,"evidence_quote":"Supplies the HSSD-200 dataset of realistic 3D scenes used for the benchmark's home environments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the MuJoCo Menagerie humanoid robot used in episodes that field a fetch robot."}],"review_version":1}