{"id":"a0686bae-c030-467f-9b9b-d548f36b49ee","arxiv_id":"2509.09804","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A FrameNet Brasil team annotated 48 turn-organizing gestures in a Brazilian travel series and proposed pragmatic frames, with mental spaces and blending, to explain them.","lead":"This paper adds a new annotation layer to a Portuguese multimodal dataset, tagging hand and head gestures that pass, take, or confirm conversational turns. It frames these gestures as evoking 'pragmatic frames' built from mental spaces and conceptual blending.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim of previously undocumented gesture variants rests on treating an edited TV series as naturalistic dialogue; without unedited footage, the variants may be camera or editing artifacts.","rationale":"The paper's contribution is the first annotation of turn-organizing gestures with pragmatic frames and the observation of variants beyond lab descriptions. For that contribution to hold, the gestures must be genuine interactive gestures in ordinary conversation and the variants must be real form differences, not artifacts of the recording. The manuscript itself flags the edited nature of the corpus, but only as a limitation that leaves some gestures unrecorded; it does not consider that editing, camera angle, and staging could actively distort or fabricate the variants. Because the raw footage is unavailable, every observed form is filtered through the production. This is load-bearing because the strongest claim in the abstract is about what communicators \"make use of\" in face-to-face conversation and about \"variations... not documented before.\" If the variants disappear in unedited footage, the paper's main novelty evaporates. The missing inter-annotator agreement is a real methodological gap, but it would still leave the possibility that the documented video examples are genuine; the corpus artifact undercuts the entire naturalistic inference. The abstract's \"keeping\" claim is a clear overstatement, but correcting it does not save the variant claim. Therefore, the safest concrete check is to inspect unedited or multi-angle footage. Since the authors' permission for the series presumably allows access, this check is feasible. We agree with the reader's identification of the weakest assumption; the verdict should remain conditional pending such evidence.","tokens_in":14402,"tokens_out":4908,"duration_ms":45486,"concrete_test":"Obtain the unedited rushes or full-frame takes for the scenes underlying Figs. 5-7 and the 48 annotations; have independent coders, blind to the paper's labels, identify turn-relevant gestures in the continuous unstaged stretches using Bavelas et al.'s interactive-gesture criteria. If the \"less prototypical\" variants are still produced and recognized when no cuts or camera redirection occur, the artifact interpretation fails; if they appear only in edited or b-roll shots, the naturalistic-variation claim must be retracted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2 concedes that \"the camera was not always focused on the communicators\" and \"sometimes the scene was cut before possible gestures were performed,\" yet the abstract's central claim is that the results confirm naturalistic face-to-face use and reveal undocumented variants. These concessions reveal a deeper problem: the corpus is a professionally edited travel series, not a record of spontaneous interaction. Editing does not merely occlude gestures; it can create apparent variants. A turn-passing gesture filmed from the side while walking (Figs. 5-6) may look less prototypical simply because of camera angle or because a cut removes the gesture's onset or apex; the \"new\" turn-taking form in Fig. 7 could be an artifact of framing rather than a novel interactive gesture. The authors treat the format as an advantage (\"outside a laboratory\"), but a TV production is still a staged, camera-aware setting. Moreover, the dataset and tools are not released, so readers cannot check whether the variants appear in unedited stretches. The abstract's \"keeping\" claim is contradicted by Section 3's finding of zero Turn_keeping gestures, but that overstatement is secondary; the artifact threat undermines the main novel result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a framework for modeling multimodal conversational turn organization by annotating co-speech interactive gestures with pragmatic frames in the FrameNet Brasil paradigm. The authors enrich the existing Frame2 multimodal dataset (ten episodes of the Brazilian TV series Pedro Pelo Mundo) with new pragmatic frames: Organization_of_conversation, Turn_passing, Turn_taking, Turn_keeping, and Turn_confirmation. They report identifying 48 interactive gestures: 30 Turn_passing, 16 Turn_confirmation, 2 Turn_taking, and 0 Turn_keeping. They further claim that the corpus reveals previously undocumented variations of known interactive gestures, and they offer a theoretical account based on mental spaces theory, conceptual blending, the Basic Communicative Spaces Network, and the metaphor SPEECH TURN IS AN OBJECT. The paper concludes that pragmatic frames can be evoked by gestures as well as language, and that the annotation enriches the multimodal dataset for future machine-learning use.","tokens_in":14678,"tokens_out":3443,"duration_ms":29314,"significance":"If the claims were fully supported, the paper would make a useful contribution by extending FrameNet-style annotation to pragmatic, gesture-evoked frames and by pointing to variants of interactive gestures outside laboratory settings. The frame definitions and the explicit counts (48 gestures with a breakdown) are falsifiable and a useful starting point. The prototype-based account of gesture variation is also a testable idea. However, the current evidence base is preliminary: the corpus is an edited TV series, the annotation was performed without reported reliability measures, and the frames were defined in advance by the same team that annotated. These factors materially limit the strength of the central claims, though they do not invalidate the value of the annotation schema as a proposal.","major_comments":[{"comment":"The abstract states that the results 'confirmed that communicators involved in face-to-face conversation make use of gestures as a tool for passing, taking and keeping conversational turns,' but Section 3 reports that no gestures evoking the Turn_keeping frame were identified (48 gestures total: 30 Turn_passing, 16 Turn_confirmation, 2 Turn_taking, 0 Turn_keeping). The claim about 'keeping' is therefore unsupported by the reported data. The abstract overstates the findings and should be corrected, and the discussion in Section 3 should explicitly address the zero-count result rather than merely describing the corpus as not conducive to Turn_keeping gestures.","section":"Abstract; Section 3"},{"comment":"The main novel result—previously undocumented variants of turn-passing and turn-taking gestures—is derived from a professionally edited television series. Section 2 acknowledges that 'the camera was not always focused on the communicators' and 'sometimes the scene was cut before possible gestures were performed.' These conditions mean that the 'less prototypical' variants shown in Figs. 5–7 could be artifacts of camera angle, framing, or editorial cuts rather than naturalistic behavior of the communicators. The paper treats the TV format as an advantage without addressing the alternative explanation that the variants are products of the production process. Because the novelty claim rests on these variants, the manuscript must either provide evidence that the gestures occur in unedited or minimally edited stretches, or substantially qualify the claim that the observations reflect naturalistic face-to-face dialogue.","section":"Section 2"},{"comment":"The pragmatic frames were defined before annotation based on 'the turn organization gestures we expected to find in the corpus,' and the same team that defined the frames performed the annotation. No inter-annotator agreement or other reliability measure is reported. As a result, the 'confirmation' that gestures evoke these frames is partly self-fulfilling: the annotators applied their own pre-existing categories to the data. For an annotation-methodology paper, a reliability measure (for example, Cohen's kappa or agreement on a held-out subset annotated by independent coders) is necessary to support the claim that the annotation is reproducible. This issue affects the evidentiary value of the reported counts and should be addressed before the results are presented as confirmed.","section":"Section 2; Section 3"}],"minor_comments":[{"comment":"The in-text cross-reference 'as seen in section 1.6.1' does not correspond to any subsection in the manuscript; the reference should be corrected to the actual section number where the Bavelas et al. gesture descriptions are presented.","section":"Section 3"},{"comment":"The citation 'Subirats-Rüggeberg; You; Liu, 2005' is inconsistent with the reference list, which lists 'Subirats-Rüggeberg, C., & Petruck, M. R. (2003)' and 'You, L., & Liu, K. (2005)' separately; the in-text citation should identify the correct sources.","section":"Section 1.2"},{"comment":"The paper notes that FrameNet Brasil obtained rights to use the episodes, but it does not state whether the annotated dataset or the annotation tool configuration will be publicly released; stating the availability of the dataset would help reproducibility and future work.","section":"Section 2"},{"comment":"The discussion of Fig. 8 links the greeting example to the BCSN blending account, but the connection to the gesture-evoked pragmatic frames is not explicitly spelled out; adding a concrete step-by-step mapping for one of the observed gestures (e.g., the turn-passing gesture in Fig. 4) would strengthen the argument.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a real gap and the proposed frame set is a reasonable extension of FrameNet practice, but the current evidence does not support the abstract's strong claim of 'keeping' turns, and the edited-TV corpus raises a fundamental question about the naturalistic status of the observed variants. I would advise the editor that the paper may be publishable after substantial revision that qualifies the claims, adds reliability evidence, and releases or otherwise documents the data."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a modest, honest annotation study. It does something genuinely new — encoding turn-organizing gestures as pragmatic frames in FrameNet Brasil's WebTool, extending Czulo et al.'s pragmatic frames to gesture. For that reason it's a useful data point for conversational AI and multimodal FrameNet work. The counts (48 gestures: 30 turn-passing, 16 turn-confirmation, 2 turn-taking) are internally consistent and the identification criteria are made explicit.\n\nThe soft spots are real but not fatal to the core claim. The abstract says the data confirm gestures for 'passing, taking and keeping' turns, but Section 3 reports zero Turn_keeping gestures. That's an overstatement; easy fix. More important, the corpus is an edited TV travel series. The authors acknowledge the camera doesn't always capture communicators and scenes get cut, but they don't engage with the possibility that cuts or camera angles produce the 'new' variants. A turn-passing gesture filmed from the side while walking may look less prototypical because of framing; a 'turn-taking' gesture may be an artifact. This is the main threat to the novelty claim, and it's not addressed beyond an appeal to 'outside a laboratory.' The dataset and tools are not released, so readers can't check. There's also no inter-annotator agreement reported, which matters for a new annotation scheme, and the frame definitions were built before annotation by the same team that annotated, which introduces a moderate circularity concern.\n\nThe theoretical section (mental spaces, blending, SPEECH TURN IS AN OBJECT) is well-grounded in the cognitive linguistics literature but is speculative post-hoc explanation, not tested prediction. That's fine for a linguistics paper if presented as interpretation, and it mostly is.\n\nWho is this for? Researchers working on multimodal FrameNet, gesture annotation, and turn organization. It's a serious contribution to a niche subfield. I'd send it to peer review — it deserves referee time — but I'd want the abstract fixed, a discussion of corpus artifacts, and at least a sample of the annotated data or a plan to release it.","headline":"A modest, honest annotation study that adds pragmatic frames for turn-organizing gestures to FrameNet Brasil, but the headline claim of previously undocumented variants rests on an edited TV corpus and the abstract overstates the data.","tokens_in":15152,"tokens_out":1887,"would_cite":false,"duration_ms":16893,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The authors show that gestures for passing, taking, confirming, and keeping conversational turns can be systematically encoded as pragmatic frames, and that naturalistic settings reveal undocumented variants of known gestures.","keywords":["pragmatic frames","interactive gestures","turn organization","multimodality","mental spaces","conceptual blending","conversation analysis"],"falsifier":"Compare the broadcast episodes against the raw, unedited interview footage: if the newly described gesture variants appear only after editing or disappear when the camera is on the communicators throughout, the claim that these are natural interactive gestures would collapse. A complementary test would ask naive viewers whether the side-by-side turn-passing gesture is recognized as passing the turn without verbal context.","tokens_in":14142,"feed_emoji":"👋","tokens_out":9099,"duration_ms":449567,"temperature":0.7,"pith_summary":"This paper claims that gestures used to organize conversational turns—passing, taking, confirming, or keeping the floor—can be systematically encoded as pragmatic frames in a multimodal, frame-semantic annotation scheme. The authors annotate 48 interactive gestures in ten episodes of a Brazilian travel-interview series, adding a pragmatic layer to an existing multimodal dataset. They report that speakers do use gestures for turn management in non-laboratory conversation, and that some documented gestures take new forms when interlocutors walk side by side rather than face each other. The paper argues these recognitions are not arbitrary: they arise from blending processes in a mental-spaces network and from a conceptual metaphor that treats a speech turn as an object to be handed over.","feed_headline":"48 gestures show how speakers pass, take, and confirm turns","feed_subtitle":"Frame-based annotation of a Brazilian travel series reveals side-by-side turn hand-offs that lab studies missed.","key_machinery":"The machinery is the pragmatic frame: a frame-semantic unit defined for interactive conversational functions, with frame elements such as Utterer, Comprehender, and Communicators. New frames—Organization_of_conversation, Turn_passing, Turn_taking, Turn_keeping, Turn_confirmation—are implemented in the authors' annotation tool and used to tag gesture instances with bounding boxes. The interpretive mechanism that carries the argument is conceptual blending: an integration network whose base space and speech-act space are blended into the communicators' social roles, and whose SPEECH TURN IS AN OBJECT metaphor licenses the passing, taking, and holding of a conversational floor.","core_discovery":"The central discovery is that turn-organizational gestures evoke a small set of pragmatic frames—Turn_passing, Turn_taking, Turn_keeping, and Turn_confirmation—that can be added to a frame-semantic annotation of video. In the authors' corpus, 30 of the 48 gestures realize Turn_passing, 16 realize Turn_confirmation, 2 realize Turn_taking, and none realize Turn_keeping. The naturalistic setting also yields previously undocumented variants: a turn-passing gesture performed while walking side by side, for example, points forward rather than directly at the addressee, yet remains recognizable as the same category. The authors explain this stability through prototype categorization and through a blend of the Basic Communicative Spaces Network with the metaphor of a speech turn as a manipulable object.","pith_inferences":["If the same gesture-to-frame mapping holds across speakers and languages, automatic dialogue systems could use hand-motion features as early cues for transition-relevance places, complementing prosody and syntax.","The side-by-side turn-passing variant suggests that prototype-based gesture categories may be even wider in mobile and multi-party settings than a face-to-face laboratory setup would indicate.","Because the corpus is an edited travel series, the zero Turn_keeping count could reflect either the interview genre or editorial cuts; annotating unedited dyadic interaction would separate those explanations."],"forward_implications":["A multimodal dataset annotated this way becomes usable for machine learning: gesture-to-frame mappings could train models to predict turn-management acts from video.","The gesture categories known from laboratory studies are broader than their original descriptions; prototype-based recognition accounts for the observed variants.","Pragmatic frames can be evoked by non-verbal behavior alone, so frame-based multimodal annotation should include a pragmatic layer rather than only semantic frames.","The absence of Turn_keeping gestures in this corpus is consistent with an interview genre in which the host encourages the guest to keep talking; larger datasets may reveal when turn-keeping gestures appear."],"supporting_citations":[{"why":"Defines interactive gestures, the face-to-face criteria used to recognize them, and the three turn-management gestures the paper extends.","marker":"Bavelas et al. (1992)"},{"why":"Provides the model of turn-taking and transition-relevance places used to locate gestural turn organization.","marker":"Sacks, Schegloff, and Jefferson (1974)"},{"why":"Establishes frame semantics and the notion that frames are activated as whole structures, grounding the extension to pragmatic frames.","marker":"Fillmore (1982)"},{"why":"Argues for a pragmatic layer in FrameNet and supplies the model for pragmatic frames such as Communicative_context.","marker":"Czulo, Ziem and Torrent (2020)"},{"why":"Provides the Frame2 multimodal dataset and the ten-episode corpus that was extended with gesture annotations.","marker":"Belcavello et al. (2024)"},{"why":"Supplies mental-spaces and conceptual-blending theory used to explain how pragmatic frames are conceptualized.","marker":"Fauconnier and Turner (2002)"},{"why":"Supplies prototype theory used to explain how non-prototypical gesture variants remain recognizable as the same category.","marker":"Rosch (1973)"},{"why":"First FrameNet to define pragmatic frames, whose model the authors followed for the Communicative_context frame.","marker":"Ziem; Willich; Triesch, 2023"},{"why":"Defines face-to-face dialogue and co-speech gesture as the multimodal setting the corpus is claimed to exemplify.","marker":"Bavelas (2022)"},{"why":"Introduces conduit metaphors that motivate gestures in which a speech turn is handed over as an object.","marker":"McNeill & Levy (1982)"}],"fun_headline_variants":["48 gestures reveal turn hand-offs missed in labs","Gestures pass turns: 30 pass, 16 confirm, 2 take","Side-by-side turn pass points forward, still readable","Pragmatic frames from gestures: turn passing, taking, confirming"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument treats the gestures filmed in an edited television travel series as spontaneous face-to-face interaction, even though the camera sometimes missed the speakers and edits could have cut or staged the relevant moments.","fun_headline_variants_meta":{"raw":{"variants":["48 gestures reveal turn hand-offs missed in labs","Gestures pass turns: 30 pass, 16 confirm, 2 take","Side-by-side turn pass points forward, still readable","Pragmatic frames from gestures: turn passing, taking, confirming"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000114,"raw_usage":{"total_tokens":1078,"prompt_tokens":968,"completion_tokens":110,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":584,"completion_tokens_details":{"reasoning_tokens":39}},"tokens_in":584,"tokens_out":110,"duration_ms":2059,"temperature":1.0,"reasoning_tokens":39,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:57:39.884361+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the broadcast episodes against the raw, unedited interview footage: if the newly described gesture variants appear only after editing or disappear when the camera is on the communicators throughout, the claim that these are natural interactive gestures would collapse. A complementary test would ask naive viewers whether the side-by-side turn-passing gesture is recognized as passing the turn without verbal context.","supporting_citations":[],"review_version":2}