Pith. sign in

REVIEW 3 major objections 5 minor 130 references

Social Processes: Probabilistic Meta-learning for Adaptive Multiparty Interaction Forecasting

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper tries to show that treating every conversation group as a meta-learning task, and conditioning forecasts on a short context of the same group's behavior, lets a model adapt to groups it never saw during training.

desk verdict Honest synthetic analysis showing SP models interpolate but don't extrapolate; the 'unseen groups' claim overreaches, but the paper deserves review. read the letter →

arxiv 2501.01915 v1 pith:RHJ7XHC4 submitted 2025-01-03 cs.LG

classification cs.LG
keywords meta-learningsocialcueforecastingneuralprocessesmultipartyinteractionconversationdynamicslatentvariablemodelhumanbehaviorgeneralization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is trying to establish that group-level social forecasting can adapt to unseen conversation groups if each group is treated as its own meta-learning task. The proposed Social Process models learn to condition a probabilistic forecast of all participants' future low-level cues on a short context set of that same group's observed-future behavior pairs, so that at evaluation a new group can be handled by supplying a context, with no retraining. The authors argue this matters because even the same person behaves differently across different groups, and standard supervised social forecasting models struggle to generalize to new groups. Their synthetic experiments support the mechanism but also bound it: the model can interpolate between known social behavior types and learns meaningful latent representations when contexts are informative, yet it extrapolates poorly to out-of-distribution dynamics, so the claimed adaptation to unseen groups works only when the unseen group's dynamics resemble a subset or combination of training dynamics.

What carries the argument

The central mechanism is the context-conditioned predictive distribution p(Y|X,C), built on the Neural Process latent-variable setup: a group-level latent z sampled from q(z|C) injects the group's identity into decoding, and an optional deterministic path r_C, with cross-attention in the Attentive Social Process variant, carries context directly. Around this core, sequence encoders produce per-participant embeddings by combining a self-encoder with a partner encoder that pools, from the target participant's frame of reference, relative quaternion orientation, relative location, and relative speaking status; offset encodings based on sinusoidal positional encodings inject the time gap between observed and future windows, and decoding is autoregressive with a Gaussian observation model plus geometric auxiliary losses for pose and speaking status.

What would settle it

Within the paper's own synthetic speaking-turn setup, create a group whose conversation rule changes halfway through an interaction, then measure whether the Social Process model's forecast log-likelihood on the post-switch target window stays high when the context set is drawn only from the pre-switch period; if it collapses, the stationarity assumption that underlies the group-level latent variable is violated.

Watch

Extended reading notes

Core claim

The paper claims to introduce and formalize Social Cue Forecasting (SCF), jointly predicting a distribution over future multimodal cues (pose, head orientation, speaking status) for all members of a conversation group from their preceding cues, and to solve it with Social Process models, a meta-learning family that predicts p(Y|X,C) by conditioning on a context set C of the same group's past observed-future sequence pairs. The experiments on synthetic glancing and speaking-turn data show that the proposed models, particularly the GRU-based Social Process, give better log-likelihood and uncertainty estimates than non-meta-learning and Neural Process baselines. The paper also demonstrates that when contexts carry information about the behavior type, the model maps different group types to separated latent distributions, learns a semantic latent axis, and can interpolate between known behaviors; however, the same experiments show that generalization to unseen groups is limited to interpolation, since models trained on narrower behavior sets fail to produce sensible forecasts for a new, more complex set of group dynamics.

Load-bearing premise

The load-bearing premise, stated in Section V, is that the stochastic process generating a group's social behavior does not evolve over time, so a single context set and a single latent variable can keep representing the group's future.

Editorial extensions

If this is right

  • A social robot or agent could adapt to a new group's interaction style after observing a short context of that group, without training a separate model for the group.
  • Forecasts come with calibrated uncertainty estimates, which the paper argues is necessary because one observed sequence can lead to multiple socially valid futures.
  • When the context is informative, the model's latent space organizes groups by behavior type and supports interpolation between known behaviors, as shown in the separated-context glancing experiment.
  • Generalization to unseen groups is bounded by training diversity: models trained on a wider variety of group dynamics adapt better to a new dynamic than models trained on narrow dynamics.
  • The Social Cue Forecasting formulation, with non-contiguous observed and future windows and explicit offset encodings, is designed to support social-science tasks such as forecasting lagged synchrony, mimicry, and disengagement.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the interpolation result generalizes, the context set is best understood not as a source of new dynamics but as a selector over a learned library of dynamics; a direct test would be to measure whether the latent axis learned on synthetic glancing behavior appears on real conversations with graded head-turn amplitudes.
  • The stationarity assumption is the most exposed point: real group interactions drift, so a natural extension is to make the latent variable time-dependent and re-estimate it from a rolling context window, which the paper's fixed-behavior synthetic setting cannot distinguish.
  • The partner-encoding design, which transforms partners' cues into the target participant's frame before pooling, is a transferable building block that could be tested in other multi-agent forecasting tasks such as traffic or team sports, where each agent's future depends on how it perceives others' positions and headings.
  • The paper's interpolation conclusion suggests a practical deployment rule: before trusting forecasts for a new group, verify that the group's observed dynamics fall within the convex hull of training dynamics, for example by checking the posterior distance of its context encoding to training encodings.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper formalizes Social Cue Forecasting (SCF), a meta-learning task in which each conversation group is treated as a separate task, and proposes Social Process (SP) models, Neural-Process-style sequence models that condition forecasts on a context set of observed-future pairs from the same group. An Attentive Social Process variant adds cross-attention over context sequences, and the method includes an offset encoding for non-contiguous observed/future windows and auxiliary geometric losses. The authors evaluate on two families of synthetic datasets (glancing behavior and speaking turns), reporting log-likelihood and mean errors and analyzing latent-space structure. The central claim is that conditioning on a group's context set yields generalization to unseen groups at evaluation.

Significance. If the central claim were established, the formulation would be a useful bridge between meta-learning and multiparty interaction forecasting, and the paper contains several genuine strengths: the synthetic tasks are carefully constructed so that the latent type is identifiable, the ELBO objective is evaluated on held-out meta-samples, the authors are explicit that extrapolation fails, and code, data, and trained models are released. However, the validation does not currently support the advertised real-world generalization. The only cross-distribution experiment shows interpolation, and the introduction promises real-world experiments that do not appear in the manuscript. The contribution is therefore best assessed as a well-executed synthetic analysis with an overstated generalization claim.

major comments (3)
  1. [Section I and Section VI] The Introduction states that this paper extends previous work with 'considerably stronger real-world experiments with larger and more expressive datasets,' but Section VI contains only synthetic experiments; no real-world results appear anywhere in the manuscript. This is a direct mismatch between the stated contribution and the reported content. The authors should either supply the promised real-world experiments or revise the contribution statement to describe the synthetic generalization analysis accurately.
  2. [Section VI.B, Table II, Figure 11] The abstract and Section V claim that conditioning on a context set leads to 'generalization to unseen groups,' but the only cross-distribution experiment, Table II, shows that models trained on Dual and Dual-random fail on the Dominating dataset and that the Full-random model succeeds only because Dominating dynamics are a subset of Full-random dynamics. The paper's own conclusion in Section VI.B is that the model can interpolate between known social behaviors but has difficulty extrapolating to out-of-distribution data. The unqualified claim of generalization to unseen groups is therefore not supported for dynamics outside the training distribution; either the claims should be qualified or the evaluation should include held-out groups generated by genuinely novel dynamics.
  3. [Section V] The load-bearing assumption stated in Section V, that 'the underlying stochastic process generating social behaviors does not evolve over time,' justifies the use of a single group-level latent z and a static context set C. This assumption is not tested anywhere in the manuscript, and the synthetic datasets satisfy it by construction because each group's generative rule is fixed throughout. The paper should either provide an experiment with time-varying group dynamics or explicitly scope the method's validity to stationary interactions, since the meta-learning conditioning mechanism would otherwise fail under drift.
minor comments (5)
  1. [Section I] There is a typo in 'simultaneusly'; it should be 'simultaneously'.
  2. [Section II] There is a typo in the heading 'monanidc'; it should be 'monadic'.
  3. [Section VI.B, Table I] The table headers 'Mixed context', 'Type I context', and 'Type III Context' are used without an explicit explanation of what each column represents; adding a sentence in the text or a footnote would improve readability.
  4. [Appendix A] The phrase 'under the random context regime and no-pool configuration' is used without defining these configurations in the main text or appendix; please define them or remove the reference.
  5. [Section V, Equation 10] The notation in Equation 10 uses \hat{s}_l and \hat{s}_q before explaining that \hat{s} := \log \hat{\sigma}^2; the explanation should be moved before or integrated into the equation.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the SP forecast is the standard Neural Process ELBO applied to sequence windows, evaluated on target windows excluded from the context; the paper's own Section VI.B explicitly bounds the claimed generalization to interpolation.

full rationale

The paper's derivation chain is self-contained. Equations 2-5 define SP as a Neural Process with a Seq2Seq encoder/decoder: the predictive distribution is p(Y|X,C) over target pairs D := (X,Y), with C a disjoint set of the same group's observed-future pairs, and training maximizes the ELBO (Eq. 3). Evaluation (Tables I-II, Figures 4-11) scores log p(Y|X,C) on target windows that are not in C. There is no fitted parameter that is later renamed a prediction; the latent variable z is trained by variational inference, and the 'predicted' futures are never used to define the training objective. The synthetic datasets are authored for the paper, but the forecasting task is a genuine out-of-sample draw: the target sequence type is not copied from the context, and the paper's own failure analysis (Dual and Dual-random models failing on Dominating; Section VI.B, Figure 11) shows the evaluation can reject the method, so it is not forced by construction. The self-citation to [35] is contextual, since the method is re-derived in full here, and is not load-bearing. The only flagged gap is a support/overclaim issue, not circularity: the Introduction promises 'considerably stronger real-world experiments with larger and more expressive datasets,' while Section VI reports only synthetic experiments, so the broad 'generalization to unseen groups' claim is under-supported; but under-support is a correctness risk, not equivalence of output to input.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The central claim that conditioning on a group's past context enables adaptation to unseen groups rests on four premises: (1) the social stochastic process is stationary within an interaction, so a single group-level latent is sufficient; (2) relative partner features (quaternion and position differences, speaking-status differences) capture the relevant social influence; (3) each conversation group is a well-defined meta-learning task; (4) the synthetic datasets approximate real social dynamics closely enough to support the generalization claim. The experiments also depend on hand-chosen hyperparameters such as latent dimensionality and context set size.

free parameters (5)
  • latent_dim_glancing = 1
    Section VI.B: 'we use a 1-dimensional latent variable, as the necessary latent information is inherently binary.' This dimensionality is chosen by hand for the synthetic glancing experiment.
  • latent_dim_speaking = 64
    Section VI.B: 'We use RNN encoders and decoders and have a 64-dimensional latent space.' Chosen by hand for the speaking-turn experiments.
  • context_size_glancing = 25% of a 100-sequence batch; 785 sequences at evaluation
    Appendix B: 'We train with batches of 100 sequences, using a randomly sampled 25% of the batch as context. For evaluation, we fix 785 randomly sampled phase values as context.' These context set sizes are chosen by the authors and affect the meta-learning conditioning.
  • context_size_speaking = 8 context sequences per meta-sample
    Section VI.B: 'Each meta-sample contains interactions from a single group from which we sample 8 sequences for the context and 11 sequences for the target.' The context size is chosen by hand; results may depend on it.
  • speaking_turn_duration = 2 timesteps
    Section VI.B: 'In all datasets, each speaking turn lasts exactly two timesteps.' This synthetic data design choice determines the difficulty of the forecasting task.
assumptions (5)
  • domain assumption The stochastic process generating social behaviors does not evolve over time within an interaction.
    Section V, paragraph on modeling assumptions: 'our modeling assumption is that the underlying stochastic process generating social behaviors does not evolve over time.' This justifies a single group-level latent z and a static context set; if false, the meta-learning conditioning would be invalid.
  • domain assumption Relative partner features transformed to an individual's frame (quaternion differences, position differences, speaking-status differences) capture the influence of partners on that individual's future behavior.
    Section V, Equation 6a-c defines qrel, lrel, srel. The partner encoding assumes these transformed features contain the information needed to predict the individual's future; this is a modeling assumption not independently verified.
  • domain assumption Each conversation group is a well-defined meta-learning task whose identity is captured by a latent variable.
    Section IV and V: treating every group as a separate meta-learning task. This assumes group identity is a sufficient delimiter for task structure, which is plausible for synthetic data but unverified for real conversations.
  • standard math Neural Process generative model and evidence lower bound (ELBO).
    Section IV, Equations 2-3. The method inherits the NP latent variable formulation and optimizes the ELBO; this is standard machine learning background.
  • domain assumption The synthetic datasets are representative of real social interaction dynamics.
    Section VI.B presents synthetic glancing and speaking-turn datasets as proxies for real social behaviors. The central claim of generalization to unseen groups is tested only on these synthetic simulations, so the claim's real-world relevance depends on this representativeness assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Social Processes: Probabilistic Meta-learning for Adaptive Multiparty Interaction Forecasting." pith.science (2026). https://pith.science/paper/RHJ7XHC4

@misc{pith2026250101915,
  author       = {Pith},
  title        = {Pith review of: Social Processes: Probabilistic Meta-learning for Adaptive Multiparty Interaction Forecasting},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RHJ7XHC4}},
  note         = {Machine review of arXiv:2501.01915}
}
read the original abstract

Adaptively forecasting human behavior in social settings is an important step toward achieving Artificial General Intelligence. Most existing research in social forecasting has focused either on unfocused interactions, such as pedestrian trajectory prediction, or on monadic and dyadic behavior forecasting. In contrast, social psychology emphasizes the importance of group interactions for understanding complex social dynamics. This creates a gap that we address in this paper: forecasting social interactions at the group (conversation) level. Additionally, it is important for a forecasting model to be able to adapt to groups unseen at train time, as even the same individual behaves differently across different groups. This highlights the need for a forecasting model to explicitly account for each group's unique dynamics. To achieve this, we adopt a meta-learning approach to human behavior forecasting, treating every group as a separate meta-learning task. As a result, our method conditions its predictions on the specific behaviors within the group, leading to generalization to unseen groups. Specifically, we introduce Social Process (SP) models, which predict a distribution over future multimodal cues jointly for all group members based on their preceding low-level multimodal cues, while incorporating other past sequences of the same group's interactions. In this work we also analyze the generalization capabilities of SP models in both their outputs and latent spaces through the use of realistic synthetic datasets.

Figures

Figures reproduced from arXiv: 2501.01915 by the authors.

Figure 1
Figure 1. Illustration of the two forecasting approaches on a real-world situation from the MatchNMingle dataset [36]. The top part of the figure illustrates a high-order group leaving event [37], where the individual leaves from one group (in tobs) to another (tfut). The bottom part depicts the low-level social cues b i t : head pose (solid normal), body pose (hollow normal), and speaking status (speaker in orange), which ar… view at source ↗
Figure 2
Figure 2. Architecture of the SP and ASP family [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 4
Figure 4. Ground truths and predictions for the mixed context glancing behavior task. All models learn to average over the possible futures. Our SP models learn a better fit than the NP model, SP-GRU being the best (see zoomed insets). 10 11 12 13 14 15 16 17 18 19 timestep 1.1 0.6 0.1 0.4 0.9 1.4 LL NP-latent (Baseline) SP-latent, MLP (Ours) SP-latent, GRU (Ours) Ground-truth LL [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figures from the paper (7 more)
Figure 5
Figure 5. Figure 5: Mean per timestep LL over the sequences in the synthetic glancing mixed context dataset. Higher is better. TABLE I: Mean (Std.) Metrics on the Synthetic Glancing Behavior Dataset.. All models are latent variants. The metrics are averaged over timesteps; mean and std. a…
Figure 7
Figure 7. Figure 7: Two different meta samples, representing synthetic Type I and Type III glances. Both subfigures contain two columns: one for the context and one for the target. The context column depicts 2 sequences (out of 4) and the context representation q(z|C). We observe that the…
Figure 8
Figure 8. Figure 8: A visual representation of latent space in the synthetic glancing experiment. Instead of encoding some context to z, we sample z uniformly from the interval [0.25; 1.75] (the ends of the interval represent Type I and Type III contexts respectively). We inject these val…
Figure 9
Figure 9. Figure 9: Encodings of different contexts with 3 different models. Each plot shows 20 different curves – each of them represent an encoding of some meta sample’s context. The meta samples are taken from the Dominating dataset. This allows looking into the latent space in order t…
Figure 11
Figure 11. Figure 11: Comparison of three models trained on the Dual, Dual-random and Full-random datasets. Dark blue cells denote the observed sequences provided to the models, while the brighter blue cells denote the model’s predicted sequences and the number inside of each cell represen…
Figure 12
Figure 12. Figure 12: Mean Per Timestep Metrics over the Sequences in the Synthetic Glancing Dataset. [PITH_FULL_IMAGE:figures/full_fig_p015_12.png]
Figure 13
Figure 13. Figure 13 [PITH_FULL_IMAGE:figures/full_fig_p015_13.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

130 extracted references · 53 canonical work pages

  1. [35]

    Social processes: Self-supervised meta-learning over conversational groups for 12 forecasting nonverbal social cues

    Chirag Raman, Hayley Hung, and Marco Loog. Social processes: Self-supervised meta-learning over conversational groups for 12 forecasting nonverbal social cues. In Lecture Notes in Computer Science, Lecture notes in computer science, pages 639–659. Springer Nature Switzerland, Cham, 2023. 2

  2. [1]

    Conducting Interaction: Patterns of Behavior in F ocused Encounters

    Adam Kendon. Conducting Interaction: Patterns of Behavior in F ocused Encounters. Number 7 in Studies in Interactional Sociolinguistics. Cambridge University Press, Cambridge ; New York, 1990. ISBN 978-0-521-38036-2 978-0-521-38938-9. 1, 2, 3, 4

  3. [2]

    Managing Human-Robot Engage- ment with Forecasts and

    Dan Bohus and Eric Horvitz. Managing Human-Robot Engage- ment with Forecasts and. . . um. . . Hesitations. Proceedings of the 16th International Conference on Multimodal Interaction , page 8, 2014. 2, 3

  4. [3]

    Models for multiparty engagement in open-world dialog

    Dan Bohus and Eric Horvitz. Models for multiparty engagement in open-world dialog. In Proceedings of the SIGDIAL 2009 10 p1p2p3p4p5 Person p1p2p3p4p5 Person 1 2 3 4 5 6 Time p1p2p3p4p5 Person 1 2 3 4 5 6 Time (a) Context. An example of context that was provided to the models (we show 6 out of 8 sequences here). This same context was provided to all 3 mode...

  5. [4]

    Prediction of Next-Utterance Timing using Head Movement in Multi-Party Meetings

    Ryo Ishii, Shiro Kumano, and Kazuhiro Otsuka. Prediction of Next-Utterance Timing using Head Movement in Multi-Party Meetings. In Proceedings of the 5th International Conference on Human Agent Interaction , HAI ’17, pages 181–187, New York, NY , USA, October 2017. Association for Computing Machinery. ISBN 978-1-4503-5113-3. doi: 10.1145/3125739.3125765. 1

  6. [5]

    The use of intonation for turn anticipation in observed conversations without visual signals as source of information

    Anne Keitel and Moritz M Daum. The use of intonation for turn anticipation in observed conversations without visual signals as source of information. Frontiers in psychology , 6:108, 2015. 1, 2, 3

  7. [6]

    The use of content and timing to predict turn transitions

    Simon Garrod and Martin J Pickering. The use of content and timing to predict turn transitions. Frontiers in psychology , 6: 751, 2015. 1

  8. [7]

    Take a breath and take the turn: how breathing meets turns in spontaneous dialogue

    Amélie Rochet-Capellan and Susanne Fuchs. Take a breath and take the turn: how breathing meets turns in spontaneous dialogue. Philosophical Transactions of the Royal Society B: Biological Sciences , 369(1658):20130399, 2014

Show all 130 references
  1. [8]

    Wlodarczak and M

    M. Wlodarczak and M. Heldner. Respiratory turn-taking cues. In INTERSPEECH, 2016. 1, 2

  2. [9]

    Predictive language processing: integrating comprehension and production, and what atypical populations can tell us

    Simone Gastaldon, Noemi Bonfiglio, Francesco Vespignani, and Francesca Peressotti. Predictive language processing: integrating comprehension and production, and what atypical populations can tell us. Front. Psychol., 15:1369177, May 2024. 1

  3. [10]

    Audiovisual Detection of Behavioural Mimicry

    Sanjay Bilakhia, Stavros Petridis, and Maja Pantic. Audiovisual Detection of Behavioural Mimicry. In 2013 Humaine Association Conference on Affective Computing and Intelligent Interaction , pages 123–128, Geneva, Switzerland, September 2013. IEEE. ISBN 978-0-7695-5048-0. doi: ...

  4. [11]

    team player

    G Klein, D D Woods, J M Bradshaw, R R Hoffman, and P J Feltovich. Ten challenges for making automation a “team player” in joint human-agent activity. IEEE Intell. Syst. , 19(06):91–95, November 2004. 1

  5. [12]

    Artificial cognition for social human–robot interaction: An implementation

    Séverin Lemaignan, Mathieu Warnier, E Akin Sisbot, Aurélie Clodic, and Rachid Alami. Artificial cognition for social human–robot interaction: An implementation. Artif. Intell. , 247: 45–69, June 2017

  6. [13]

    Towards human-aware cognitive robots

    R Alami, R Chatila, A Clodic, S Fleury, M Herrb, V Montreuil, and E A Sisbot. Towards human-aware cognitive robots. 2006. 1

  7. [14]

    Toward a theory of situation awareness in dynamic systems

    Mica R Endsley. Toward a theory of situation awareness in dynamic systems. Hum. Factors, 37(1):32–64, March 1995. 1

  8. [15]

    A situation awareness perspective on human-AI interaction: Tensions and opportunities

    Jinglu Jiang, Alexander J Karran, Constantinos K Coursaris, Pierre-Majorique Léger, and Joerg Beringer. A situation awareness perspective on human-AI interaction: Tensions and opportunities. Int. J. Hum. Comput. Interact. , 39(9):1789–1806, May 2023

  9. [16]

    Special issue on situation awareness in intelligent human-computer interaction for time critical decision making

    Wei Wei, Jinsong Wu, and Chunsheng Zhu. Special issue on situation awareness in intelligent human-computer interaction for time critical decision making. IEEE Intell. Syst. , 35:3–5, January 2020. 1

  10. [17]

    Human motion trajectory prediction: A survey

    Andrey Rudenko, Luigi Palmieri, Michael Herman, Kris M Kitani, Dariu M Gavrila, and Kai O Arras. Human motion trajectory prediction: A survey. The International Journal of Robotics Research, 39(8):895–935, 2020. 1, 6, 9

  11. [18]

    DIALOGPT : Large-scale generative pre-training for conversational response generation

    Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. DIALOGPT : Large-scale generative pre-training for conversational response generation. In Proceedings of the 58th Annual Meeting of the Association for C...

  12. [19]

    GPT-4 technical report

    OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, 11 Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Je...

  13. [20]

    Early prediction for physical human robot collaboration in the operating room

    Tian Zhou and Juan Pablo Wachs. Early prediction for physical human robot collaboration in the operating room. Auton. Robots, 42(5):977–995, June 2018. 1

  14. [21]

    NRDF: Neural riemannian distance fields for learning articulated pose priors

    Yannan He, Garvita Tiwari, Tolga Birdal, Jan Eric Lenssen, and Gerard Pons-Moll. NRDF: Neural riemannian distance fields for learning articulated pose priors. arXiv [cs.CV] , March 2024. 1

  15. [22]

    NeMF: Neural motion fields for kinematic animation

    Chengan He, Jun Saito, James Zachary, Holly Rushmeier, and Yi Zhou. NeMF: Neural motion fields for kinematic animation. arXiv [cs.CV] , June 2022

  16. [23]

    Human motion diffusion as a generative prior

    Yonatan Shafir, Guy Tevet, Roy Kapon, and Amit H Bermano. Human motion diffusion as a generative prior. arXiv [cs.CV] , March 2023

  17. [24]

    InterControl: Zero-shot human interaction generation by controlling every joint

    Zhenzhi Wang, Jingbo Wang, Yixuan Li, Dahua Lin, and Bo Dai. InterControl: Zero-shot human interaction generation by controlling every joint. arXiv [cs.CV] , November 2023. 1

  18. [25]

    Can language models learn to listen? arXiv [cs.CV] , August 2023

    Evonne Ng, Sanjay Subramanian, Dan Klein, Angjoo Kanazawa, Trevor Darrell, and Shiry Ginosar. Can language models learn to listen? arXiv [cs.CV] , August 2023. 1

  19. [26]

    The GENEA challenge 2023: A large scale evaluation of gesture generation models in monadic and dyadic settings

    Taras Kucherenko, Rajmund Nagy, Youngwoo Yoon, Jieyeon Woo, Teodor Nikolov, Mihail Tsakov, and Gustav Eje Henter. The GENEA challenge 2023: A large scale evaluation of gesture generation models in monadic and dyadic settings. arXiv [cs.HC], August 2023

  20. [27]

    Learning to listen: Mod- eling non-deterministic dyadic facial motion

    Evonne Ng, Hanbyul Joo, Liwen Hu, Hao Li, Trevor Darrell, Angjoo Kanazawa, and Shiry Ginosar. Learning to listen: Mod- eling non-deterministic dyadic facial motion. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, June 2022

  21. [28]

    ChaLearn LAP challenges on self-reported personality recognition and non-verbal behavior forecasting during social dyadic interactions: Dataset, design, and results

    Cristina Palmero, Germán Barquero, Julio C Jacques Junior, Albert Clapés, Johnny Núñez, David Curto, Sorina Smeureanu, Javier Selva, Zejian Zhang, David Saeteros, D Gallardo-Pujol, G Guilera, D Leiva, Feng Han, Xiaoxue Feng, Jennifer He, Wei-Wei Tu, T Moeslund, Isabelle M Guyo...

  22. [29]

    To react or not to react: End-to-end visual pose forecasting for personalized avatar during dyadic conversations

    Chaitanya Ahuja, Shugao Ma, Louis-Philippe Morency, and Yaser Sheikh. To react or not to react: End-to-end visual pose forecasting for personalized avatar during dyadic conversations. In 2019 International Conference on Multimodal Interaction , New York, NY , USA, October 2019. ACM. 1

  23. [30]

    Nina-Jo Moore, Hickson Mark III, and W Don. Stacks. Nonverbal communication: Studies and applications. 2013. 1, 4

  24. [31]

    Socially and contextually aware human motion and pose forecasting

    Vida Adeli, Ehsan Adeli, Ian Reid, Juan Carlos Niebles, and Hamid Rezatofighi. Socially and contextually aware human motion and pose forecasting. IEEE Robotics and Automation Letters, 5(4):6033–6040, 2020. 1, 3, 9

  25. [32]

    Social diffusion: Long-term multiple human motion anticipation

    Julian Tanke, Linguang Zhang, Amy Zhao, Chen Tang, Yujun Cai, Lezi Wang, Po-Chen, Wu, Juergen Gall, and Cem Keskin. Social diffusion: Long-term multiple human motion anticipation. ICCV, pages 9567–9577, October 2023. 3

  26. [33]

    Multi-person 3d motion prediction with multi-range transformers

    Jiashun Wang, Huazhe Xu, Medhini Narasimhan, and Xi- aolong Wang. Multi-person 3d motion prediction with multi-range transformers. In M. Ranzato, A. Beygelzimer, Y . Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems , vol- ...

  27. [34]

    Social Signal Processing: Understanding social interactions through nonverbal behavior analysis (PDF)

    Alessandro Vinciarelli, H Salamin, and M Pantic. Social Signal Processing: Understanding social interactions through nonverbal behavior analysis (PDF). 2009 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2009 , June 2009. doi: 10.1109/CVPRW.2009.5204290. 1, 3

  28. [36]

    The matchnmingle dataset: a novel multi-sensor resource for the analysis of social inter- actions and group dynamics in-the-wild during free-standing conversations and speed dates

    Laura Cabrera-Quiros, Andrew Demetriou, Ekin Gedik, Leander van der Meij, and Hayley Hung. The matchnmingle dataset: a novel multi-sensor resource for the analysis of social inter- actions and group dynamics in-the-wild during free-standing conversations and speed dates. IEEE ...

  29. [37]

    Rituals of Leaving: Predictive Modelling of Leaving Behaviour in Conversation

    Felix van Doorn. Rituals of Leaving: Predictive Modelling of Leaving Behaviour in Conversation. Master of Science Thesis, Delft University of Technology , 2018. 2, 3

  30. [38]

    DeepPose: Human pose estimation via deep neural networks

    Alexander Toshev and Christian Szegedy. DeepPose: Human pose estimation via deep neural networks. In 2014 IEEE Conference on Computer Vision and Pattern Recognition . IEEE, June 2014. 2

  31. [39]

    OpenPose: Realtime multi-person 2D pose estimation using part affinity fields

    Zhe Cao, Gines Hidalgo, Tomas Simon, Shih-En Wei, and Yaser Sheikh. OpenPose: Realtime multi-person 2D pose estimation using part affinity fields. IEEE Trans. Pattern Anal. Mach. Intell. , 43(1):172–186, January 2021. 2

  32. [40]

    SMPL: A skinned multi- person linear model

    Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. SMPL: A skinned multi- person linear model. In Seminal Graphics Papers: Pushing the Boundaries, V olume 2, pages 851–866. ACM, New York, NY , USA, August 2023. 2

  33. [41]

    Expressive body capture: 3D hands, face, and body from a single image

    Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A Osman, Dimitrios Tzionas, and Michael J Black. Expressive body capture: 3D hands, face, and body from a single image. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE...

  34. [42]

    Embod- ied hands: Modeling and capturing hands and bodies together

    Javier Romero, Dimitrios Tzionas, and Michael J Black. Embod- ied hands: Modeling and capturing hands and bodies together. arXiv [cs.GR] , January 2022. 2

  35. [43]

    Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image

    Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Peter Gehler, Javier Romero, and Michael J Black. Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image. In Computer Vision – ECCV 2016 , Lecture notes in computer science, pages 561–578. Springer I...

  36. [44]

    Learning to reconstruct 3D human pose and shape via model-fitting in the loop

    Nikos Kolotouros, Georgios Pavlakos, Michael Black, and Kostas Daniilidis. Learning to reconstruct 3D human pose and shape via model-fitting in the loop. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) . IEEE, October 2019. 2

  37. [45]

    RoHM: Robust human motion reconstruction via diffusion

    Siwei Zhang, Bharat Lal Bhatnagar, Yuanlu Xu, Alexander Winkler, Petr Kadlecek, Siyu Tang, and Federica Bogo. RoHM: Robust human motion reconstruction via diffusion. arXiv [cs.CV], January 2024. 2

  38. [46]

    Github, 2021

    Easymocap - make human motion capture easier. Github, 2021. URL https://github.com/zju3dv/EasyMocap. 2

  39. [47]

    DeepMoCap: Deep optical motion capture using multiple depth sensors and retro-reflectors

    Anargyros Chatzitofis, Dimitrios Zarpalas, Stefanos Kollias, and Petros Daras. DeepMoCap: Deep optical motion capture using multiple depth sensors and retro-reflectors. Sensors (Basel) , 19 (2):282, January 2019. 2

  40. [48]

    Panoptic studio: A massively multiview system for social interaction capture

    Hanbyul Joo, Tomas Simon, Xulong Li, Hao Liu, Lei Tan, Lin Gui, Sean Banerjee, Timothy Scott Godisart, Bart Nabbe, Iain Matthews, Takeo Kanade, Shohei Nobuhara, and Yaser Sheikh. Panoptic studio: A massively multiview system for social interaction capture. IEEE Transactions on...

  41. [49]

    NRDF - neural region descriptor fields as implicit ROI representation for robotic 3D surface processing

    Anish Pratheepkumar, Markus Ikeda, Michael Hofmann, Fabian Widmoser, Andreas Pichler, and Markus Vincze. NRDF - neural region descriptor fields as implicit ROI representation for robotic 3D surface processing. In 2024 IEEE/RSJ International Conference on Intelligent Robots and...

  42. [50]

    HuMoR: 3D human motion model for robust pose estimation

    Davis Rempe, Tolga Birdal, Aaron Hertzmann, Jimei Yang, Srinath Sridhar, and Leonidas J Guibas. HuMoR: 3D human motion model for robust pose estimation. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) . IEEE, October 2021

  43. [51]

    GFPose: Learning 3D human pose prior with gradient fields

    Hai Ci, Mingdong Wu, Wentao Zhu, Xiaoxuan Ma, Hao Dong, Fangwei Zhong, and Yizhou Wang. GFPose: Learning 3D human pose prior with gradient fields. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, June 2023

  44. [52]

    Adversarial parametric pose prior

    Andrey Davydov, Anastasia Remizova, Victor Constantin, Sina Honari, Mathieu Salzmann, and Pascal Fua. Adversarial parametric pose prior. arXiv [cs.CV] , December 2021. 2

  45. [53]

    Pose transformers (POTR): Human motion prediction with non-autoregressive transformers

    Angel Martinez-Gonzalez, Michael Villamizar, and Jean-Marc Odobez. Pose transformers (POTR): Human motion prediction with non-autoregressive transformers. In 2021 IEEE/CVF Inter- national Conference on Computer Vision Workshops (ICCVW) . IEEE, October 2021. 2

  46. [54]

    Nakano, and Louis-Philippe Morency

    Chaitanya Ahuja, Dong Won Lee, Yukiko I. Nakano, and Louis-Philippe Morency. Style Transfer for Co-Speech Gesture Animation: A Multi-Speaker Conditional-Mixture Approach. arXiv:2007.12553 [cs] , July 2020. 2

  47. [55]

    AI choreographer: Music conditioned 3D dance generation with AIST++

    Ruilong Li, Shan Yang, David A Ross, and Angjoo Kanazawa. AI choreographer: Music conditioned 3D dance generation with AIST++. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) . IEEE, October 2021. 2

  48. [56]

    Multi-person extreme motion prediction

    Wen Guo, Xiaoyu Bie, Xavier Alameda-Pineda, and Francesc Moreno-Noguer. Multi-person extreme motion prediction. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, June 2022. 2

  49. [57]

    Prediction of turn-taking by combining prosodic and eye-gaze information in poster conversations

    Tatsuya Kawahara, Takuma Iwatate, and Katsuya Takanashi. Prediction of turn-taking by combining prosodic and eye-gaze information in poster conversations. In Interspeech, pages 727– 730, 2012. 2

  50. [58]

    Response- conditioned turn-taking prediction

    Bing’er Jiang, Erik Ekstedt, and Gabriel Skantze. Response- conditioned turn-taking prediction. In Findings of the Association for Computational Linguistics: ACL 2023 , Stroudsburg, PA, USA, 2023. Association for Computational Linguistics. 2

  51. [59]

    To React or not to React: End-to-End Visual Pose Forecasting for Personalized Avatar during Dyadic Conversations

    Chaitanya Ahuja, Shugao Ma, Louis-Philippe Morency, and Yaser Sheikh. To React or not to React: End-to-End Visual Pose Forecasting for Personalized Avatar during Dyadic Conversations. arXiv:1910.02181 [cs] , October 2019. 2, 3

  52. [60]

    Chalearn lap challenges on self-reported personality recognition and non- verbal behavior forecasting during social dyadic interactions: Dataset, design, and results

    Cristina Palmero, German Barquero, Julio CS Jacques Junior, Albert Clapés, Johnny Núnez, David Curto, Sorina Smeureanu, Javier Selva, Zejian Zhang, David Saeteros, et al. Chalearn lap challenges on self-reported personality recognition and non- verbal behavior forecasting duri...

  53. [61]

    Context-aware human behaviour forecasting in dyadic interactions

    Nguyen Tan Viet Tuyen and Oya Celiktutan. Context-aware human behaviour forecasting in dyadic interactions. In Un- derstanding Social Behavior in Dyadic and Small Group Interactions, pages 88–106. PMLR, 2022. 2

  54. [62]

    Anticipating averted gaze in dyadic interactions

    Philipp Müller, Ekta Sood, and Andreas Bulling. Anticipating averted gaze in dyadic interactions. In ACM Symposium on Eye Tracking Research and Applications , New York, NY , USA, June

  55. [63]

    Real-time eye-gaze based interaction for human intention prediction and emotion analysis

    Hao He, Yingying She, Jianbing Xiahou, Junfeng Yao, Jun Li, Qingqi Hong, and Yingxuan Ji. Real-time eye-gaze based interaction for human intention prediction and emotion analysis. In Proceedings of Computer Graphics International 2018 , New York, NY , USA, June 2018. ACM. 2

  56. [64]

    Emotion-aware human attention prediction

    Macario O Cordel, Shaojing Fan, Zhiqi Shen, and Mohan S Kankanhalli. Emotion-aware human attention prediction. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, June 2019. 2

  57. [65]

    Behavior in Public Places: Notes on the Social Organization of Gatherings

    Erving Goffman. Behavior in Public Places: Notes on the Social Organization of Gatherings . The Free Press, 1. paperback ed.,

  58. [66]

    Social Force Model for Pedes- trian Dynamics

    Dirk Helbing and Peter Molnar. Social Force Model for Pedes- trian Dynamics. Physical Review E, 51(5):4282–4286, May 1995. ISSN 1063-651X, 1095-3787. doi: 10.1103/PhysRevE.51.4282. 13 2

  59. [67]

    ISBN 978-0-02-911940-2

    printing edition, 1966. ISBN 978-0-02-911940-2. 2

  60. [68]

    Discrete Choice Models for Pedestrian Walking Behavior

    Gianluca Antonini, Michel Bierlaire, and Mats Weber. Discrete Choice Models for Pedestrian Walking Behavior. Transportation Research Part B: Methodological , 40:667–687, September 2006. doi: 10.1016/j.trb.2005.09.006

  61. [69]

    Matuszyk

    Jarosław W ˛ as, Bartłomiej Gudowski, and Paweł J. Matuszyk. Social Distances Model of Pedestrian Dynamics. In Cellular Automata, volume 4173, pages 492–501. Springer Berlin Hei- delberg, Berlin, Heidelberg, 2006. ISBN 978-3-540-40929-8 978-3-540-40932-8. doi: 10.1007/11861201\_57

  62. [70]

    Learning Social Etiquette: Human Trajectory Understanding In Crowded Scenes

    Alexandre Robicquet, Amir Sadeghian, Alexandre Alahi, and Silvio Savarese. Learning Social Etiquette: Human Trajectory Understanding In Crowded Scenes. In Computer Vision – ECCV 2016, volume 9912, pages 549–565. Springer International Publishing, Cham, 2016. ISBN 978-3-319-464...

  63. [71]

    Continuum crowds

    Adrien Treuille, Seth Cooper, and Zoran Popovi ´c. Continuum crowds. ACM Transactions on Graphics / SIGGRAPH 2006 , 25 (3):1160–1168, July 2006

  64. [72]

    Modelling Smooth Paths Using Gaussian Processes

    Christopher Tay and Christian Laugier. Modelling Smooth Paths Using Gaussian Processes. In Proc. of the Int. Conf. on Field and Service Robotics , 2007

  65. [73]

    J. M. Wang, D. J. Fleet, and A. Hertzmann. Gaussian Process Dynamical Models for Human Motion. IEEE Transactions on Pattern Analysis and Machine Intelligence , 30(2):283–298, February 2008. ISSN 1939-3539. doi: 10.1109/TPAMI.2007. 1167

  66. [74]

    Social LSTM: Human Trajectory Prediction in Crowded Spaces

    Alexandre Alahi, Kratarth Goel, Vignesh Ramanathan, Alexan- dre Robicquet, Li Fei-Fei, and Silvio Savarese. Social LSTM: Human Trajectory Prediction in Crowded Spaces. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 961–971, Las Vegas, NV , USA...

  67. [75]

    Intent-Aware Probabilistic Trajectory Estimation for Collision Prediction with Uncertainty Quantification

    Andrew Patterson, Arun Lakshmanan, and Naira Hovakimyan. Intent-Aware Probabilistic Trajectory Estimation for Collision Prediction with Uncertainty Quantification. arXiv:1904.02765 [cs, math] , April 2019

  68. [76]

    Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks

    Agrim Gupta, Justin Johnson, Li Fei-Fei, Silvio Savarese, and Alexandre Alahi. Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks. arXiv:1803.10892 [cs] , March 2018. 4

  69. [77]

    SR-LSTM: State Refinement for LSTM towards Pedestrian Trajectory Prediction

    Pu Zhang, Wanli Ouyang, Pengfei Zhang, Jianru Xue, and Nanning Zheng. SR-LSTM: State Refinement for LSTM towards Pedestrian Trajectory Prediction. arXiv:1903.02793 [cs] , March 2019

  70. [78]

    STGAT: Modeling Spatial-Temporal Interactions for Human Trajectory Prediction

    Yingfan Huang, Huikun Bi, Zhaoxin Li, Tianlu Mao, and Zhaoqi Wang. STGAT: Modeling Spatial-Temporal Interactions for Human Trajectory Prediction. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) , pages 6271–6280, Seoul, Korea (South), October 2019. IEEE. IS...

  71. [79]

    Forecasting People Trajectories and Head Poses by Jointly Reasoning on Tracklets and Vislets

    Irtiza Hasan, Francesco Setti, Theodore Tsesmelis, Vasileios Be- lagiannis, Sikandar Amin, Alessio Del Bue, Marco Cristani, and Fabio Galasso. Forecasting People Trajectories and Head Poses by Jointly Reasoning on Tracklets and Vislets. arXiv:1901.02000 [cs], January 2019

  72. [80]

    TNT: Target-driveN Trajectory Prediction

    Hang Zhao, Jiyang Gao, Tian Lan, Chen Sun, Benjamin Sapp, Balakrishnan Varadarajan, Yue Shen, Yi Shen, Yuning Chai, Cordelia Schmid, Congcong Li, and Dragomir Anguelov. TNT: Target-driveN Trajectory Prediction. arXiv:2008.08294 [cs] , August 2020

  73. [81]

    Social-STGCNN: A Social Spatio-Temporal Graph Convolutional Neural Network for Human Trajectory Prediction

    Abduallah Mohamed, Kun Qian, Mohamed Elhoseiny, and Christian Claudel. Social-STGCNN: A Social Spatio-Temporal Graph Convolutional Neural Network for Human Trajectory Prediction. arXiv:2002.11927 [cs] , February 2020

  74. [82]

    Group Split and Merge Prediction With 3D Convolutional Networks

    Allan Wang and Aaron Steinfeld. Group Split and Merge Prediction With 3D Convolutional Networks. IEEE Robotics and Automation Letters , 5(2):1923–1930, April 2020. ISSN 2377-3766. doi: 10.1109/LRA.2020.2969947. 2

  75. [83]

    THOMAS: Tra- jectory Heatmap Output with learned Multi-Agent Sampling

    Thomas Gilles, Stefano Sabatini, Dzmitry Tsishkou, Bog- dan Stanciulescu, and Fabien Moutarde. THOMAS: Tra- jectory Heatmap Output with learned Multi-Agent Sampling. arXiv:2110.06607 [cs] , January 2022. 2

  76. [84]

    SocialInteractionGAN: Multi-person Interaction Se- quence Generation

    Louis Airale, Dominique Vaufreydaz, and Xavier Alameda- Pineda. SocialInteractionGAN: Multi-person Interaction Se- quence Generation. arXiv:2103.05916 [cs, stat] , March 2021. 2, 3

  77. [85]

    The roundtable: An abstract model of conversation dynamics

    Massimo Mastrangeli, Martin Schmidt, and Lucas Lacasa. The roundtable: An abstract model of conversation dynamics. arXiv:1010.2943 [physics] , October 2010. 2

  78. [86]

    A survey on contrastive self-supervised learning

    Ashish Jaiswal, Ashwin Ramesh Babu, Mohammad Zaki Zadeh, Debapriya Banerjee, and Fillia Makedon. A survey on contrastive self-supervised learning. Technologies (Basel), 9(1):2, December

  79. [87]

    Mgpi: A computational model of multiagent group perception and interaction

    Navyata Sanghvi, Ryo Yonetani, and Kris Kitani. Mgpi: A computational model of multiagent group perception and interaction. arXiv preprint arXiv:1903.01537 , 2019. 2, 3

  80. [88]

    Toward a histology of social behavior: Judgmental accuracy from thin slices of the behavioral stream

    Nalini Ambady, Frank J Bernieri, and Jennifer A Richeson. Toward a histology of social behavior: Judgmental accuracy from thin slices of the behavioral stream. In Advances in experimental social psychology, volume 32, pages 201–271. Elsevier, 2000. 3

  81. [89]

    Self-supervised learning for videos: A survey

    Madeline C Schiappa, Yogesh S Rawat, and Mubarak Shah. Self-supervised learning for videos: A survey. ACM Comput. Surv., 55(13s):1–37, December 2023. 3

  82. [90]

    Some signals and rules for taking speaking turns in conversations

    Starkey Duncan. Some signals and rules for taking speaking turns in conversations. Journal of Personality and Social Psychol- ogy, 23(2):283–292, 1972. ISSN 1939-1315(Electronic),0022- 3514(Print). doi: 10.1037/h0033031. 3, 6

  83. [91]

    Pauses, gaps and overlaps in conversations

    Mattias Heldner and Jens Edlund. Pauses, gaps and overlaps in conversations. Journal of Phonetics , 38(4):555–568, October

  84. [92]

    Gazing in triads: A powerful signal in floor apportionment

    Akko Kalma. Gazing in triads: A powerful signal in floor apportionment. British Journal of Social Psychology , 31(1): 21–39, March 1992. 3

  85. [93]

    Levinson and Francisco Torreira

    Stephen C. Levinson and Francisco Torreira. Timing in turn- taking and its implications for processing models of language. Frontiers in Psychology , 6, June 2015. ISSN 1664-1078. doi: 10.3389/fpsyg.2015.00731. 3

  86. [94]

    Monica M. Moore. Nonverbal courtship patterns in women: Context and consequences. Ethology and Sociobiology , 6(4): 237–247, January 1985. ISSN 0162-3095. doi: 10.1016/ 0162-3095(85)90016-0. 3, 6

  87. [95]

    Multiple Granularity Group Interaction Prediction

    Taiping Yao, Minsi Wang, Bingbing Ni, Huawei Wei, and Xiaokang Yang. Multiple Granularity Group Interaction Prediction. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 2246–2254, Salt Lake City, UT, June 2018. IEEE. ISBN 978-1-5386-6420-9. doi: 1...

  88. [96]

    Forecasting Human Dynamics from Static Images

    Yu-Wei Chao, Jimei Yang, Brian Price, Scott Cohen, and Jia Deng. Forecasting Human Dynamics from Static Images. arXiv:1704.03432 [cs] , April 2017. 3

  89. [97]

    Towards Social Artificial Intelligence: Nonverbal Social Signal Prediction in a Triadic Interaction

    Hanbyul Joo, Tomas Simon, Mina Cikara, and Yaser Sheikh. Towards Social Artificial Intelligence: Nonverbal Social Signal Prediction in a Triadic Interaction. In 2019 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR) , pages 10865–10875, Long Beach, CA, US...

  90. [98]

    The Pose Knows: Video Forecasting by Generating Pose Futures

    Jacob Walker, Kenneth Marino, Abhinav Gupta, and Martial Hebert. The Pose Knows: Video Forecasting by Generating Pose Futures. arXiv:1705.00053 [cs] , April 2017

  91. [99]

    A Recurrent Variational Autoencoder for Human Motion Synthesis

    Ikhsanul Habibie, Daniel Holden, Jonathan Schwarz, Joe Years- ley, and Taku Komura. A Recurrent Variational Autoencoder for Human Motion Synthesis. In Procedings of the British Machine Vision Conference 2017 , page 119, London, UK, 2017. British Machine Vision Association. ISB...

  92. [100]

    Recurrent Network Models for Human Dynamics

    Katerina Fragkiadaki, Sergey Levine, Panna Felsen, and Jitendra Malik. Recurrent Network Models for Human Dynamics. arXiv:1508.00271 [cs] , September 2015

  93. [101]

    Inter- personal Synchrony: A Survey of Evaluation Methods across Dis- ciplines

    Emilie Delaherche, Mohamed Chetouani, Ammar Mahdhaoui, Catherine Saint-Georges, Sylvie Viaux, and David Cohen. Inter- personal Synchrony: A Survey of Evaluation Methods across Dis- ciplines. IEEE Transactions on Affective Computing , 3(3):349– 365, July 2012. ISSN 1949-3045. d...

  94. [102]

    Meta-Learning in Neural Networks: A Survey

    Timothy Hospedales, Antreas Antoniou, Paul Micaelli, and Amos Storkey. Meta-Learning in Neural Networks: A Survey. arXiv:2004.05439 [cs, stat] , November 2020. 3

  95. [103]

    QuaterNet: A Quaternion-based Recurrent Model for Human Motion

    Dario Pavllo, David Grangier, and Michael Auli. QuaterNet: A Quaternion-based Recurrent Model for Human Motion. arXiv:1805.06485 [cs] , July 2018. 3

  96. [104]

    Sequence to Sequence Learning with Neural Networks

    Ilya Sutskever, Oriol Vinyals, and Quoc V Le. Sequence to Sequence Learning with Neural Networks. In Z. Ghahramani, M. Welling, C. Cortes, N. D. Lawrence, and K. Q. Weinberger, editors, Advances in Neural Information Processing Systems 27 , pages 3104–3112. Curran Associates, ...

  97. [105]

    Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation

    Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation. arXiv:1406.1078 [cs, stat] , September 2014. 4

  98. [106]

    Rezende, S

    Marta Garnelo, Jonathan Schwarz, Dan Rosenbaum, Fabio Viola, Danilo J. Rezende, S. M. Ali Eslami, and Yee Whye Teh. Neural Processes. arXiv:1807.01622 [cs, stat] , 2018. 3, 6

  99. [107]

    Sequential Neural Processes

    Gautam Singh, Jaesik Yoon, Youngsung Son, and Sungjin Ahn. Sequential Neural Processes. Advances in Neural Information Processing Systems , 32, 2019. URL http://arxiv.org/abs/1906. 10264. 4, 6

  100. [108]

    Robustifying Sequential Neural Processes

    Jaesik Yoon, Gautam Singh, and Sungjin Ahn. Robustifying Sequential Neural Processes. In International Conference on Machine Learning , pages 10861–10870. PMLR, November 2020

  101. [109]

    Attentive Neural Processes

    Hyunjik Kim, Andriy Mnih, Jonathan Schwarz, Marta Garnelo, Ali Eslami, Dan Rosenbaum, Oriol Vinyals, and Yee Whye Teh. Attentive Neural Processes. arXiv:1901.05761 [cs, stat] , July

  102. [110]

    Spatiotemporal Modeling using Recurrent Neural Processes

    Sumit Kumar. Spatiotemporal Modeling using Recurrent Neural Processes. Master of Science Thesis, Carnegie Mellon University , page 43, 2019. 4

  103. [111]

    Stephanie Tan, David M. J. Tax, and Hayley Hung. Multimodal Joint Head Orientation Estimation in Interacting Groups via Proxemics and Interaction Dynamics. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies , 5(1):1–22, March 2021. ISSN 2474-95...

  104. [112]

    Recurrent Neural Processes

    Timon Willi, Jonathan Masci, Jürgen Schmidhuber, and Christian Osendorfer. Recurrent Neural Processes. arXiv:1906.05915 [cs, stat], November 2019

  105. [113]

    On Social Involvement in Mingling Scenarios: Detecting Associates of F-formations in Still Images

    Lu Zhang and Hayley Hung. On Social Involvement in Mingling Scenarios: Detecting Associates of F-formations in Still Images. IEEE Transactions on Affective Computing , 2018. 5

  106. [114]

    Geometric Loss Functions for Camera Pose Regression with Deep Learning

    Alex Kendall and Roberto Cipolla. Geometric Loss Functions for Camera Pose Regression with Deep Learning. arXiv:1704.00390 [cs], May 2017. 5

  107. [115]

    Analyzing Free-standing Conversational Groups: A Multimodal Approach

    Xavier Alameda-Pineda, Yan Yan, Elisa Ricci, Oswald Lanz, and Nicu Sebe. Analyzing Free-standing Conversational Groups: A Multimodal Approach. In Proceedings of the 23rd ACM interna- tional conference on Multimedia , pages 5–14. ACM Press, 2015. ISBN 978-1-4503-3459-4. doi: 10...

  108. [116]

    A Neural Representation of Sketch Drawings

    David Ha and Douglas Eck. A Neural Representation of Sketch Drawings. arXiv:1704.03477 [cs, stat] , May 2017. 6

  109. [117]

    Bowman, Luke Vilnis, Oriol Vinyals, Andrew M

    Samuel R. Bowman, Luke Vilnis, Oriol Vinyals, Andrew M. Dai, Rafal Jozefowicz, and Samy Bengio. Generating Sentences from a Continuous Space. arXiv:1511.06349 [cs] , May 2016. 6

  110. [118]

    Gomez, Lukasz Kaiser, and Illia Polo- sukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polo- sukhin. Attention Is All You Need. arXiv:1706.03762 [cs], June

  111. [119]

    Schegloff, and Gail Jefferson

    Harvey Sacks, Emanuel A. Schegloff, and Gail Jefferson. A Simplest Systematics for the Organization of Turn-Taking for Conversation. 50(4):40, 1974. 6

  112. [120]

    Evelyn Z. McClave. Linguistic functions of head movements in the context of speech. Journal of Pragmatics , 32(7):855– 878, 2000. ISSN 0378-2166. doi: https://doi.org/10.1016/ S0378-2166(99)00079-X. URL https://www.sciencedirect.com/ science/article/pii/S037821669900079X. 6

  113. [121]

    Gaze and mutual gaze

    Michael Argyle, Mark Cook, and Duncan Cramer. Gaze and mutual gaze. The British Journal of Psychiatry , 165(6):848–850,

  114. [122]

    Turn-taking in conversational systems and human-robot interaction: A review

    Gabriel Skantze. Turn-taking in conversational systems and human-robot interaction: A review. Comput. Speech Lang. , 67 (101178):101178, May 2021. 6

  115. [123]

    Gaze and turn-taking behavior in casual conversational interactions

    Kristiina Jokinen, Hirohisa Furukawa, Masafumi Nishida, and Seiichi Yamamoto. Gaze and turn-taking behavior in casual conversational interactions. ACM Trans. Interact. Intell. Syst. , 3 (2):1–30, July 2013

  116. [124]

    Some relationships between body motion and speech

    Adam Kendon. Some relationships between body motion and speech. Studies in dyadic communication , 7(177):90, 1972. 6

  117. [125]

    Understanding posterior collapse in generative latent variable models, 2019

    James Lucas, George Tucker, Roger Grosse, and Mohammad Norouzi. Understanding posterior collapse in generative latent variable models, 2019. URL https://openreview.net/forum?id= r1xaVLUYuE. 7

  118. [126]

    Auto-encoding variational bayes

    Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv [stat.ML] , December 2013. 7

  119. [127]

    Optimizing the turn-taking behavior of task-oriented spoken dialog systems

    Antoine Raux and Maxine Eskenazi. Optimizing the turn-taking behavior of task-oriented spoken dialog systems. ACM Trans. Speech Lang. Process. , 9(1):1–23, May 2012. 6

  120. [130]

    A modular approach for synchronized wireless multimodal multisensor data acquisition in highly dynamic social settings

    Chirag Raman, Stephanie Tan, and Hayley Hung. A modular approach for synchronized wireless multimodal multisensor data acquisition in highly dynamic social settings. arXiv preprint arXiv:2008.03715, 2020. 9 1 Social Processes: Probabilistic Meta-learning for Adaptive Multipart...

  121. [2010]

    doi: 10.1016/j.wocn.2010.08.002

    ISSN 0095-4470. doi: 10.1016/j.wocn.2010.08.002. 3

  122. [2020]

    Association for Computational Linguistics. 1

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.