{"id":"14b33131-c102-4838-a502-980404222e9f","arxiv_id":"2411.19921","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"SIMS couples retrieval-augmented LLM scripts with a text-conditioned, physics-based control policy to generate stylized human-scene interactions.","lead":"SIMS is a system that plans long, story-like action sequences with a large language model and then drives a physically simulated character through those actions with expressive styles such as sad, tired, or drunk. It combines retrieval-augmented script generation with a text-conditioned physics controller, and reports improved diversity and physical plausibility over prior work.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No direct evidence that text conditions actually produce the named styles; APD and user-study ratings show only diversity/emotional resonance, leaving 'stylized control' potentially indistinguishable from unlabeled diversity.","rationale":"The reader's weakest_assumption precisely identifies the gap I consider load-bearing: text-to-style control is asserted but never directly measured. I agree with the CONDITIONAL verdict because the system is well-engineered and the ViconStyle dataset is a useful contribution, but the central stylized-control claim needs a direct fidelity check. My concrete test would settle whether the conditional discriminator actually maps language to style. If it fails, the paper's headline contribution is unsupported; if it passes, the current CONDITIONAL verdict can be upgraded. Therefore no verdict change is needed; the reader's conditionality already captures this.","tokens_in":17716,"tokens_out":6221,"duration_ms":50587,"concrete_test":"Run a per-style fidelity evaluation for Walk or Sit. For each style in {happy, sad, drunk, tired, neutral}, generate N motions from fixed initial states with text 'walk/sit <style>'. (a) Human raters classify style without seeing labels; report accuracy/confusion matrix. (b) Compute FID between generated motions of each style and held-out real motions of that style. (c) Report per-style APD for the w/o-text ablation. If classification is near chance or per-style FID does not separate styles, the text-conditioned discriminator is not implementing style control.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central novelty is stylized text-conditioned control via a conditional discriminator (Sec. 3.3, 'we employ a conditional discriminator to inject text-based style control'). The style reward r_t^style uses CLIP text embeddings z of captions. However, no experiment checks whether a specific text condition (e.g., 'walk sadly') produces the corresponding style in the generated motion. Table 4's APD is averaged over randomly sampled text conditions and hence only measures diversity, not style-text alignment. Table 11's 'w/o text' ablation shows APD changes that are small for Sit and Lie (16.52±0.47 vs 16.29±0.22; 16.99±1.28 vs 16.59±0.28) with overlapping error bars, so text has at best a weak effect on output diversity for these skills. Table 5's user study rates 'emotional resonance' of whole videos, which is confounded by script/story and does not attribute style to motion control. There is no per-style FID, style classification accuracy, or forced-choice human evaluation of whether each text label yields its named style. If the discriminator uses z mainly to increase stochasticity, the high APD in Table 4 would be achieved without semantic style control, and the central claim of 'stylized' interactions collapses into generic diversity. This gap is the most load-bearing because the title and abstract make text-conditioned style the core contribution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents SIMS, a hierarchical framework for generating long-term stylized human-scene interactions in a physics simulator. A high-level LLM-based planner with retrieval-augmented generation (RASG) creates coherent scripts with keyframes; a low-level multi-condition control policy executes these keyframes, using CLIP text embeddings as style conditions, egocentric heightmaps for scene awareness, and task-specific goal conditions. The authors contribute a short-script database, a new ViconStyle motion capture dataset, and quantitative evaluations reporting improvements in FID, APD, success rate, and contact error over several baselines.","tokens_in":18034,"tokens_out":6301,"duration_ms":51419,"significance":"If the claims hold, SIMS advances physics-based character animation by integrating language-driven planning with controllable low-level policies and by releasing a reusable stylized motion dataset. The system is end-to-end automatic, and the reported FID/APD improvements over UniHSI and InterPhys are substantial for several skills. The paper also demonstrates scalability to new skills and generalization to unseen objects. However, the central claim of text-conditioned style control is not directly verified: the current evidence is partially consistent with the alternative that language conditioning mostly increases stochastic diversity rather than aligning with the specified style. The dataset release is a valuable community contribution.","major_comments":[{"comment":"The paper's core claim is that the language embedding z conveys stylistic cues to the low-level policy, but no experiment directly verifies that a given text condition produces the corresponding motion style. Table 4's APD is averaged over randomly sampled text conditions, so it measures diversity, not text-style alignment. The ablation in Table 11 shows that removing text changes APD for Sit and Lie by amounts within the reported error bars (16.52±0.47 vs. 16.29±0.22 for Sit; 16.99±1.28 vs. 16.59±0.28 for Lie), while only Carry shows a clear drop (14.92±0.23 vs. 12.41±0.19). The qualitative examples in Fig. 4 are suggestive but not quantified. Please add a direct per-style evaluation: for instance, train a style classifier on held-out stylized motion, report per-style FID against style-specific references, or run a forced-choice user study in which raters select which text label matches a generated motion clip with the script/story context removed. Without such evidence, the 'stylized' component of the central contribution is not established.","section":"Sec. 3.3, Sec. 4.3.2, Tables 4 and 11"},{"comment":"The abstract states that SIMS 'significantly outperforms previous methods,' but Table 3 shows that SIMS's Reach success rate (95.2) is below UniHSI's (97.5) and its Carry contact error (0.099) is worse than InterPhys's (0.08). Although the text in Sec. 4.3.1 acknowledges that these two results are slightly lower, the abstract and the contribution list present an unqualified dominance claim. Please qualify the claim to reflect the subset of metrics and skills where the advantage holds, or extend the experiments/training data to close these gaps.","section":"Abstract and Table 3"},{"comment":"The ablation study that supports the text-conditioning claim is under-specified. The row labeled 'SIMS(ours)' reports success rates (96.9 for Sit, 89.7 for Lie) that match the dataset-augmented results in Table 10, while the 'w/o text' row may have been trained on a different data mixture; the paper does not state the training data for each row. Since additional data changes both success rate and APD (Tables 8-10), the difference between 'w/o text' and 'SIMS(ours)' may conflate data volume with text conditioning. Please report the exact training set for each ablation row and, ideally, add error bars over multiple seeds for success rates as well as APD.","section":"Sec. 4.4.4, Table 11"},{"comment":"The FID evaluation in Table 4 uses SAMP as the reference distribution for Sit and Lie, but the generated motions for UniHSI are conditioned on the ScenePlan chain of contacts, not on SAMP-style motions; the comparison may be biased because SIMS is trained on SAMP while UniHSI is not. Please specify the reference motion set used for each method and skill, and consider reporting FID against a common held-out motion set.","section":"Sec. 4.2, Table 4"}],"minor_comments":[{"comment":"The word 'Smultating' appears to be a typo for 'Simulating' in the introduction.","section":"Sec. 1, Introduction"},{"comment":"The heading 'Furture Work' should be 'Future Work'.","section":"Sec. 6, heading"},{"comment":"The caption contains 'Comparision' and 'Physical-Plausibe' typos; the table layout appears misaligned, with the skill columns (Walk, Sit, Lie, GetUp, Reach, Idle, Carry) running together with the property columns, making the checkmarks difficult to interpret.","section":"Table 1 caption and layout"},{"comment":"The dataset abbreviations in the captions are inconsistent: for example, '100S' is defined as '100Style' but the table only reports two rows, and the text does not state whether 'A+100S' includes all of AMASS or a subset. Please clarify the exact data composition in each ablation row.","section":"Sec. 4.4.3, Tables 8-10"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of the journal and the proposed system is interesting, but the verification of the key 'stylized control' claim is currently insufficient. The authors should be asked to add a direct per-style evaluation, and to qualify the dominance claims in the abstract. The dataset release is a useful contribution. I do not see evidence of circularity or fabrication; the main issue is experimental scope rather than soundness of the machinery."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the useful parts. The paper builds a hierarchical system that combines RAG-based script generation with a multi-condition physics policy, and it works well enough to produce long, varied, physically plausible interactions in indoor scenes. The RASG ablation is credible: compared to direct LLM generation it gives lower SBERT similarity and faster generation. The ViconStyle capture is a real contribution, especially for stylized carrying and lying, which existing datasets lack. The generalization test on PartNet is a nice touch.\n\nThe soft spot is the one the stress-test note identifies. The paper claims text-conditioned style control, but the evaluation never shows that a specific text condition produces the intended style. APD is a diversity metric, not an alignment metric; the user study rates emotional resonance of full videos, so it is confounded by script and story. The w/o text ablation is revealing: for Sit and Lie the APD differences are small and within error bars (16.29±0.22 vs 16.52±0.47 and 16.59±0.28 vs 16.99±1.28). Only Carry shows a clear increase. That means the conditional discriminator may be adding stochasticity rather than semantic style. If so, the central claim of 'stylized' control collapses into unlabeled diversity. This is not a fatal flaw, but it needs to be addressed with a per-style evaluation, such as a forced-choice user study or style classification accuracy.\n\nTwo smaller complaints. The abstract says 'significantly outperforming previous methods,' but Table 3 shows Reach success rate below UniHSI and Carry contact error above InterPhys. The authors explain this by limited training data, but the abstract should not overstate. Also, no error bars are reported for the main interaction metrics, and no code or data release is mentioned. These are addressable.\n\nOverall, the paper is coherent and the thinking is honest. It is not circular and the citations look appropriate. The contribution is incremental but real: a scalable integration with a new dataset. A serious referee should engage with it; the right outcome is likely major revision with the style evaluation added. I would bring this to a reading group if the goal is to discuss how to evaluate text-conditioned control in physics-based animation.","headline":"A solid system paper with a useful new dataset, but the text-to-style control claim lacks direct evidence and the abstract overreaches.","tokens_in":18573,"tokens_out":3087,"would_cite":false,"duration_ms":28841,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Physics-based characters can act out long, stylized stories from a text theme","keywords":["human-scene interaction","physics-based character control","retrieval-augmented generation","text-conditioned control policy","style diversity","long-term motion planning","finite state machine","motion capture dataset"],"falsifier":"Run the trained Sit or Carry policy many times under a fixed scene and fixed task goal with two text conditions that are semantically opposite (for example, 'sad' versus 'excited') and measure the distribution of generated joint angles. If the two distributions overlap almost completely, the claim that text embeddings provide style control collapses; diversity alone would not separate conditions.","tokens_in":17543,"feed_emoji":"🎬","tokens_out":4494,"duration_ms":39469,"temperature":0.7,"pith_summary":"This paper tries to show that a simulated human character can follow a long, everyday narrative—walking to a sofa, sitting sadly, carrying a box, getting up—while moving in a believable style and staying physically plausible in a cluttered 3D room. It proposes a two-level system: a language model that plans the story as a sequence of short, retrievable scripts, and a physics-based controller that executes each scripted action with a style drawn from text descriptions. The claim is that combining retrieval-augmented script planning with a multi-condition control policy solves a problem earlier long-horizon interaction methods left open, namely producing both diverse style and stable physical contact in one pipeline. If the claim holds, character animation and embodied simulation can move from short, neutral motion clips toward long, emotionally expressive performances without hand-authored choreography.","feed_headline":"Simulated humans act out long, styled stories in 3D scenes","feed_subtitle":"A retrieval-augmented script planner plus a text-conditioned physics controller keeps long interactions diverse and physically plausible.","key_machinery":"The load-bearing mechanism is the pairing of Retrieval-Augmented Script Generation (RASG) with a multi-condition physics control policy. The short-script database converts open-ended narratives into executable keyframe tuples, each specifying skill, target object, caption, and style label; the retrieval step keeps the story grounded in actions the controller can actually perform. At the low level, the policy is conditioned on proprioception, an egocentric heightmap, a task goal, and a style embedding, and is trained with a text-conditioned adversarial reward, so a desired style is injected through language rather than through dense reference motion.","core_discovery":"SIMS is a hierarchical framework that separates what to do from how to do it. High-level planning is handled by retrieval-augmented script generation: a library of short scripts, each a few keyframes with skill, object, caption, and emotion label, is built once; given a user theme, the planner retrieves the most semantically similar short scripts and asks a language model to concatenate them into a coherent long script. Low-level execution is handled by a set of physics-based policies that observe an egocentric heightmap of the scene, a task goal, and a language embedding of the desired style, and are trained with a text-conditioned adversarial discriminator plus task rewards. A finite state machine switches policies at keyframe boundaries. The central discovery is that this division makes stylized, long-term human-scene interaction trainable and controllable: style comes through text, scene geometry comes through the heightmap, and physical plausibility comes through reinforcement learning in simulation.","pith_inferences":["A sharper test of the style mechanism than the paper reports would compare repeated rollouts under the same text embedding: if within-condition diversity is nearly as high as across-condition diversity, then much of the reported diversity could come from task randomness rather than from style control.","The same retrieval-augmented planning structure could be reused for other embodied agents, such as robot arms or virtual humans with finger articulation, by swapping the skill set and the low-level policy, since the planner only reasons about keyframe tuples.","The short-script database could become a lightweight planning benchmark independent of any physics simulator, allowing language planners to be evaluated on narrative coherence and executability before motion control is trained."],"forward_implications":["A user can type a theme such as \"a person gets fired and drinks at home\" and receive a physically simulated animation that chains walking, carrying, sitting, and lying with matching emotional styles.","Because styles enter through language embeddings, new styles can be added by collecting or captioning more motion data and retraining the relevant policy, without rewriting the planner.","The egocentric heightmap lets policies trained on one furniture set generalize to unseen objects, so the same controllers can be dropped into new room layouts.","The ViconStyle dataset plus the short-script database gives later work a common benchmark for stylized long-horizon interaction, not just single-action generation."],"supporting_citations":[{"why":"supplies the retrieval-augmented generation technique the paper adapts for long-term script planning","marker":"[18]"},{"why":"provides the adversarial motion-prior training framework used to train the physics-based policies","marker":"[29]"},{"why":"supplies the pretrained text encoder used for short-script retrieval and for aligning motion embeddings with language","marker":"[31]"},{"why":"is the main physics-based long-term HSI baseline and the source of the goal-conditioning and contact-planning design the paper extends with style control","marker":"[47]"},{"why":"gives the finite state machine and scene-conditioned controller structure that SIMS inherits and augments with heightmaps and text","marker":"[26]"},{"why":"supplies the SAMP scene-interaction motion data used as a fair-comparison training set for sit, lie, and reach policies","marker":"[12]"},{"why":"is the InterPhys baseline for physics-based character-scene interaction and is re-implemented for comparison","marker":"[13]"},{"why":"provides the 3D-FRONT furniture and room layouts used for training and generalization tests","marker":"[7]"}],"fun_headline_variants":["SIMS: AI scripts give 3D humans style and staying power","Scripted AI humans move with style and physics in 3D","Retrieval scripts guide physics-based human motion","SIMS: Long stylized interactions from text and geometry","Style-aware scripts drive physically plausible human motion"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that a text-conditioned adversarial style reward trained with aligned text and motion embeddings reliably turns a language description into the intended motion style; if that mapping is weak, the reported diversity is just unlabeled variety rather than controllable style.","fun_headline_variants_meta":{"raw":{"variants":["SIMS: AI scripts give 3D humans style and staying power","Scripted AI humans move with style and physics in 3D","Retrieval scripts guide physics-based human motion","SIMS: Long stylized interactions from text and geometry","Style-aware scripts drive physically plausible human motion"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000469,"raw_usage":{"total_tokens":2334,"prompt_tokens":941,"completion_tokens":1393,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":557,"completion_tokens_details":{"reasoning_tokens":1312}},"tokens_in":557,"tokens_out":1393,"duration_ms":10099,"temperature":1.0,"reasoning_tokens":1312,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:40:36.218019+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained Sit or Carry policy many times under a fixed scene and fixed task goal with two text conditions that are semantically opposite (for example, 'sad' versus 'excited') and measure the distribution of generated joint angles. If the two distributions overlap almost completely, the claim that text embeddings provide style control collapses; diversity alone would not separate conditions.","supporting_citations":[{"cited_title":"Retrieval-augmented generation for knowledge-intensive nlp tasks","cited_arxiv_id":null,"evidence_quote":"supplies the retrieval-augmented generation technique the paper adapts for long-term script planning"},{"cited_title":"Amp: Adversarial motion priors for styl- ized physics-based character control","cited_arxiv_id":null,"evidence_quote":"provides the adversarial motion-prior training framework used to train the physics-based policies"},{"cited_title":"Learn- ing transferable visual models from natural language super- vision","cited_arxiv_id":null,"evidence_quote":"supplies the pretrained text encoder used for short-script retrieval and for aligning motion embeddings with language"},{"cited_title":"Unified human-scene interaction via prompted chain-of-contacts","cited_arxiv_id":null,"evidence_quote":"is the main physics-based long-term HSI baseline and the source of the goal-conditioning and contact-planning design the paper extends with style control"},{"cited_title":"Synthesizing phys- ically plausible human motions in 3d scenes","cited_arxiv_id":null,"evidence_quote":"gives the finite state machine and scene-conditioned controller structure that SIMS inherits and augments with heightmaps and text"},{"cited_title":"Stochastic scene-aware motion prediction","cited_arxiv_id":null,"evidence_quote":"supplies the SAMP scene-interaction motion data used as a fair-comparison training set for sit, lie, and reach policies"},{"cited_title":"Synthesizing physi- cal character-scene interactions","cited_arxiv_id":null,"evidence_quote":"is the InterPhys baseline for physics-based character-scene interaction and is re-implemented for comparison"},{"cited_title":"3d-front: 3d furnished rooms with layouts and semantics","cited_arxiv_id":null,"evidence_quote":"provides the 3D-FRONT furniture and room layouts used for training and generalization tests"}],"review_version":1}