REVIEW 3 major objections 5 minor 46 references
Mind the Context: Continual Learning of Socially Appropriate Robot Actions via Environmental-Social Disentanglement
T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Explicitly separating environmental and social cues improves continual learning of socially appropriate robot actions.
desk verdict Solid engineering and a genuinely new task formulation, but the paper's central claim about disentanglement rests on differences smaller than the noise; needs significance testing before that conclusion should be believed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the semantic context decomposition, a pair of deterministic preprocessing transforms $x \mapsto (x_E, x_S)$ that turn one scene into an environmental view and a social view using masks from a pretrained panoptic segmentation model. The environmental view blanks out detected agents and samples extra bounding-box occlusions from the current domain's mask pool; the social view keeps only agent silhouettes against a uniform background. Two separate encoders produce feature vectors $h_E$ and $h_S$, which are concatenated and passed to a regression head; a replay buffer mixes current-domain and past-domain samples during each training step. The structural assumption this machinery is built to satisfy is $p(y|x)\approx p(y|x_E, x_S)$: the approximation that appropriateness depends on environment and social configuration separately.
What would settle it
Take two scenes from different domains in which the panoptic segmentation model fails to detect one person, so that person remains visible in the environmental view. If EDD's predictions on those failure cases degrade no more than on correctly detected scenes, the social branch alone is carrying the signal; if they degrade sharply, the decomposition's benefit depends on near-perfect segmentation. A cleaner test: find or construct pairs of scenes with identical environmental masks and identical social masks but different human appropriateness ratings; any such pair would show that $p(y|x)\approx p(y|x_E, x_S)$ is incomplete.
Extended reading notes
Core claim
Domain-incremental continual learning of socially appropriate robot actions is improved by making the two sources of appropriateness information explicit at the input level. Instead of feeding full scene images to the network, EDD masks detected agents out of the environmental view, adding decoy bounding boxes to hide their locations, and masks the background out of the social view, then trains a dual-branch network whose fused representation predicts nine appropriateness scores. On the test domains, this decomposition reduces RMSE to 0.78 and raises Pearson and concordance correlation to 0.55 and 0.38 relative to single-branch and no-decomposition controls, while backward transfer stays near 0.02, meaning forgetting is small. The paper interprets these results as supporting the assumption that action appropriateness can be approximated as a function of environmental context and social configuration separately, with their combination recovered by concatenation.
Load-bearing premise
The argument stands on the claim that hiding people to build the environmental view, and hiding the room to build the social view, throws away none of the information needed to judge whether an action is appropriate; if appropriateness depends on the joint interaction between who is present and where they are, this decomposition could destroy the very signal the model needs.
Editorial extensions
If this is right
- A robot can keep adapting to new rooms with a replay buffer of about five percent of the training data, staying within a small margin of the joint-training upper bound in both error and correlation.
- Using bounding-box masks for the environmental view outperforms silhouette masks, robot close-ups, and no decomposition, so how the separation is done matters more than the fact of separation.
- Domain ordering has only modest effects on overall performance; the largest visible effect is that presenting the home domain last increases home-domain predictions, suggesting a mild curriculum effect.
- Zero-shot vision-language models can reach competitive error rates but correlate weakly with human appropriateness ratings, indicating that foundation-model priors alone are not enough for this task.
Reading between the lines
- Inference: A natural extension the paper does not run is a stress test with deliberately corrupted or missing segmentation masks; if EDD degrades gracefully as segmentation quality drops, the decomposition principle transfers to real robots with imperfect perception.
- Inference: The equal per-domain quota in the replay buffer implicitly controls class balance; an order-aware replay policy that samples more from domains most similar to the current one could amplify the small curriculum effects reported.
- Inference: The latency breakdown suggests that for real-time deployment, the bottleneck is the panoptic segmentation step at about 7.2 seconds per image, not the dual-branch network at about 20 milliseconds, so a faster segmentation backbone would make the approach practical.
- Inference: Because the two branches are trained on visually distinct inputs, the fused representation offers a natural route to explainability: one could measure how much each branch changes the action scores, which the paper lists as future work but does not test.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses domain-incremental continual learning (CL) for socially appropriate robot actions. It proposes the Explicit Disentanglement Dual-Branch (EDD) framework, which uses panoptic segmentation to separate a scene into environmental and social-agent views, processes them through two branches of a shared network, and trains with experience replay. The authors evaluate EDD on six indoor domains from the merged OfficeDB/MannersDB+ datasets, compare it against robotics CL baselines, domain-incremental methods, and vision-language models, and additionally study the effect of different disentanglement strategies and domain orderings. The central claim is that explicitly separating environmental and social cues improves domain-incremental CL of action appropriateness.
Significance. If the central claim is supported, the paper addresses a relevant and underexplored gap: domain-incremental continual learning for socially appropriate robot actions. The framework is clearly described, the code is publicly available, and the evaluation includes standard 5-fold cross-validation, multiple regression metrics, a forgetting metric, ablations, and a latency analysis. The comparison against several baselines, including zero-shot VLMs, is a useful contribution. The main weakness is that the claimed benefit of explicit disentanglement over a no-decomposition control is quantitatively very small and not supported by statistical tests, which undermines the paper's primary research conclusion. The manuscript's own discussion acknowledges that all decomposition strategies perform relatively similarly, creating an internal inconsistency with the RQ1 conclusion.
major comments (3)
- [Section V-B / Table VII] The evidence does not support the RQ1 conclusion that explicit environmental-social separation supports CL outcomes. The load-bearing comparison is EDD (Bbox) versus the NoDec control: overall RMSE 0.78 vs 0.79, PCC 0.55 vs 0.53, CCC 0.38 vs 0.36, with standard deviations of 0.03/0.02, 0.05/0.05, and 0.04/0.05 respectively. Every advantage is under one standard deviation of the 5-fold cross-validation noise, and no significance test, confidence interval, or per-fold breakdown is provided. Table VI further shows that the single-branch CL control achieves 0.79 RMSE, 0.53 PCC, and 0.36 CCC, nearly indistinguishable from EDD CL, while single-branch joint training achieves 0.75 RMSE and 0.58 PCC, which is better than EDD CL. The paper's own RQ2 discussion states that 'all methods perform relatively similarly,' yet the RQ1 conclusion asserts that explicit separation supports CL outcomes. Please add paired statistical tests across the five folds, report effect sizes or confidence intervals, and temper the central claim accordingly.
- [Section III-C / Eq. (2)] The decomposition premise p(y|x) ≈ p(y|x_E, x_S) is load-bearing for the entire method, but it is not validated under imperfect panoptic segmentation. The paper's future-work section acknowledges that robustness to imperfect panoptic segmentation is untested. If the segmentation model misses agents or produces false detections, the environmental and social views constructed by the transformations T_E and T_S discard information the network needs. The manuscript provides no evaluation of segmentation accuracy on the six domains and no sensitivity analysis. Please report segmentation quality on the evaluation data or introduce controlled noise experiments (e.g., dropping or adding masks) to demonstrate that the decomposition is sufficiently lossless for the proposed benefit to be plausible.
- [Section IV-B / Table V] The baseline comparison is weakened by the extremely poor performance of FedLGR (RMSE 2.091, near-zero PCC and CCC), which suggests that the adaptation of this federated continual-learning method to the proposed domain-incremental setup may be suboptimal. The paper states that all baselines shared identical training procedures, but it does not describe in sufficient detail how FedLGR's client-based formulation was extended to this setting or whether its hyperparameters were tuned comparably. Please provide the adaptation details and, if possible, a stronger or more fairly tuned robotics CL baseline; otherwise, the claim that EDD outperforms state-of-the-art baselines rests partly on an unfavorable baseline configuration.
minor comments (5)
- [Section VI / Discussion] The discussion sentence 'all methods perform relatively similarly' (Section V-B) conflicts with the later conclusion that EDD supports CL outcomes; please revise the wording to make the distinction between descriptive results and conclusive claims clear.
- [Section IV-D / Table IV] The Default order is described as the 'OfficeDB default order,' but no citation is provided for this ordering; please cite the source or clarify how it was chosen.
- [References] Several references contain formatting errors, such as 'Y .-C. Hsu' and 'P. ¨Ogren'; please fix the spacing and accent glyphs throughout the bibliography.
- [Appendix / Table IX] The text in Section V-F states that semantic decomposition has a per-image latency of '~255 ms,' while Table IX reports a mean of 255.02 ms with SD 79.26; consider reporting the SD in the text or stating that it is the mean.
- [Algorithm 1 / Step 8] In Algorithm 1, the assignments h_E = phi_E(x_E; theta_E) and h_S = phi_S(x_S; theta_S) are written on a single line without an explicit separator; please format them as separate lines or add a semicolon for readability.
Circularity Check
No circularity: EDD's claims rest on held-out empirical comparisons, not on fitted inputs or load-bearing self-citations.
full rationale
The paper's central contribution is an empirical machine-learning pipeline: it proposes an input-level decomposition (Eq. 2), a dual-branch architecture, and replay-based rehearsal, then evaluates on held-out test folds with human-annotated appropriateness scores. The decomposition in Eq. (2) is explicitly introduced as an assumption ('We formalise this as follows'), not as a consequence of the method, and the model is trained and tested on disjoint data. No fitted parameter is renamed as a prediction; no uniqueness theorem is imported from prior work; no ansatz is smuggled in via self-citation. The ablations in Tables VI and VII are genuine experiments against single-branch and alternative decomposition controls, even though the reported differences are often within one standard deviation and the NoDec control is single-branch, which conflates architecture with decomposition; those are statistical-validity concerns, not circularity. Self-citations to [10], [11], and [15] supply the dataset and baselines, but the annotations are external human judgments and the baselines are executed rather than assumed, so the citations are not load-bearing. The paper's own limitations (RQ2: 'all methods perform relatively similarly'; RQ3: 'indicative pattern rather than strong evidence'; future-work admission that robustness to imperfect panoptic segmentation is untested) weaken the strength of the claims but do not amount to circular reasoning. No step in the derivation chain reduces to its own input by construction.
Assumptions & free parameters
free parameters (2)
- m (number of decoy bbox masks) =
10
- Replay buffer size B =
60 (5% of training set)
assumptions (3)
- domain assumption p(y|x) ≈ p(y|xE, xS): action appropriateness is determined by separable environmental and social views.
- domain assumption Grounded-SAM panoptic segmentation reliably detects all social agents in every domain.
- domain assumption Human rater annotations averaged per image are an adequate ground truth for social appropriateness.
Cite this review
Pith. "Pith review of Mind the Context: Continual Learning of Socially Appropriate Robot Actions via Environmental-Social Disentanglement." pith.science (2026). https://pith.science/paper/EBDYWCTC
@misc{pith2026260813448,
author = {Pith},
title = {Pith review of: Mind the Context: Continual Learning of Socially Appropriate Robot Actions via Environmental-Social Disentanglement},
year = {2026},
howpublished = {\url{https://pith.science/paper/EBDYWCTC}},
note = {Machine review of arXiv:2608.13448}
}
read the original abstract
Social robots are expected to operate across diverse environments, where similar arrangements can imply different socially appropriate actions, e.g., starting a conversation may be acceptable in a crowded home but disruptive in an office meeting. Because such norms and environments cannot all be anticipated in advance, robots require continual learning (CL) to adapt from sequential experience while retaining previously acquired knowledge. Prior work has studied CL for generating socially appropriate robot actions, but it has not addressed domain-incremental settings in which the robot incrementally encounters diverse contexts (e.g., living room, meeting room, office, hallway), where both environmental (e.g., whether the space is open or cluttered with furniture) and social cues (e.g., how people or other agents are positioned around the robot) jointly shape the appropriateness of robot actions. We address this gap with the Explicit Disentanglement Dual-Branch (EDD) framework. EDD explicitly separates environmental and social-agent related knowledge and uses replay-based rehearsal to mitigate forgetting while learning the appropriateness of robot actions (e.g., cleaning, serving, starting a conversation) across several indoor domains. Experiments show that EDD outperforms several state-of-the-art baselines, and ablation studies further evaluate different disentanglement strategies and the sensitivity to domain ordering. Our code is publicly available at https://github.com/Cambridge-AFAR/Mind-the-Context.git.
Figures
Reference graph
Works this paper leans on
-
[1]
M. Colledanchise and P. Ogren,Behavior Trees in Robotics and Al: An Introduction, 1st ed. USA: CRC Press, Inc., June 2018
work page 2018
-
[2]
A survey of Behavior Trees in robotics and AI,
M. Iovino, E. Scukins, J. Styrud, P. ¨Ogren, and C. Smith, “A survey of Behavior Trees in robotics and AI,”Robotics and Autonomous Systems, vol. 154, p. 104096, August 2022
work page 2022
-
[3]
T. Lesort, V . Lomonaco, A. Stoian, D. Maltoni, D. Filliat, and N. D´ıaz- Rodr´ıguez, “Continual learning for robotics: Definition, framework, learning strategies, opportunities and challenges,”Information Fusion, vol. 58, pp. 52–68, June 2020
work page 2020
-
[4]
Continual learning and catastrophic forgetting,
G. M. Van de Ven, N. Soures, and D. Kudithipudi, “Continual learning and catastrophic forgetting,”arXiv preprint arXiv:2403.05175, 2024
arXiv 2024
-
[5]
Continual learning for real-world autonomous systems: Algorithms, challenges and frameworks,
K. Shaheen, M. A. Hanif, O. Hasan, and M. Shafique, “Continual learning for real-world autonomous systems: Algorithms, challenges and frameworks,”Journal of Intelligent & Robotic Systems, vol. 105, no. 1, p. 9, 2022
work page 2022
-
[6]
Advancements and challenges in continual reinforcement learning: A comprehensive review,
A. Zuffer, M. Burke, and M. Harandi, “Advancements and challenges in continual reinforcement learning: A comprehensive review,”arXiv preprint arXiv:2506.21899, 2025
arXiv 2025
-
[7]
E. Bartoli, F. I. Do ˘gan, and I. Leite, “Streak: Streaming network for continual learning of object relocations under household context drifts,” in2025 34th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), 2025, pp. 1550–1557
work page 2025
-
[8]
A Deep Incremental Boltzmann Machine for Modeling Context in Robots,
F. I. Do ˘gan, H. C ¸ elikkanat, and S. Kalkan, “A Deep Incremental Boltzmann Machine for Modeling Context in Robots,” in2018 IEEE International Conference on Robotics and Automation (ICRA), May 2018, pp. 2411–2416
work page 2018
Show all 46 references
-
[9]
Cinet: A learning based approach to incremental context modeling in robots,
F. Irmak Do ˘gan, I. Bozcan, M. Celik, and S. Kalkan, “Cinet: A learning based approach to incremental context modeling in robots,” in2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2018, pp. 4641–4646
2018
-
[10]
Mind Your Manners! A Dataset and a Continual Learning Approach for Assessing Social Appropriateness of Robot Actions,
J. Tjomsland, S. Kalkan, and H. Gunes, “Mind Your Manners! A Dataset and a Continual Learning Approach for Assessing Social Appropriateness of Robot Actions,”Frontiers in Robotics and AI, vol. 9, March 2022
2022
-
[11]
Feature aggregation with latent generative replay for federated contin- ual learning of socially appropriate robot behaviours,
N. Churamani, S. Checker, F. I. Dogan, H.-T. L. Chiang, and H. Gunes, “Feature aggregation with latent generative replay for federated contin- ual learning of socially appropriate robot behaviours,”arXiv preprint arXiv:2405.15773, 2024
2024 arXiv
-
[12]
The social context of human–robot interactions,
S. Thompson, K. Candon, and M. V ´azquez, “The social context of human–robot interactions,”Annual Review of Control, Robotics, and Autonomous Systems, vol. 9, 2025
2025
-
[13]
The effect of task ordering in continual learning,
S. J. Bell and N. D. Lawrence, “The effect of task ordering in continual learning,”arXiv preprint arXiv:2205.13323, 2022
2022 arXiv
-
[14]
CLIFER: Continual Learning with Imagination for Facial Expression Recognition,
N. Churamani and H. Gunes, “CLIFER: Continual Learning with Imagination for Facial Expression Recognition,” in2020 15th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2020), November 2020, pp. 322–328
2020
-
[15]
Grace: Generating socially appropriate robot actions leveraging llms and human expla- nations,
F. I. Do ˘gan, U. Ozyurt, G. Cinar, and H. Gunes, “Grace: Generating socially appropriate robot actions leveraging llms and human expla- nations,” in2025 IEEE International Conference on Robotics and Automation (ICRA), 2025, pp. 4330–4336
2025
-
[16]
Large language models as zero-shot human models for human-robot interaction,
B. Zhang and H. Soh, “Large language models as zero-shot human models for human-robot interaction,” in2023 IEEE/RSJ international conference on intelligent robots and systems (IROS). IEEE, 2023, pp. 7961–7968
2023
-
[17]
Norm learning with reward models from instructive and evaluative feedback,
E. Rosen, E. Hsiung, V . B. Chi, and B. F. Malle, “Norm learning with reward models from instructive and evaluative feedback,” inIEEE RO- MAN, 2022, pp. 1634–1640
2022
-
[18]
Are large lan- guage models aligned with people’s social intuitions for human–robot interactions?
L. Wachowiak, A. Coles, O. Celiktutan, and G. Canal, “Are large lan- guage models aligned with people’s social intuitions for human–robot interactions?” inIEEE/RSJ IROS, 2024, pp. 2520–2527
2024
-
[19]
Federated Learning of Socially Appropriate Agent Behaviours in Simulated Home Environ- ments,
S. Checker, N. Churamani, and H. Gunes, “Federated Learning of Socially Appropriate Agent Behaviours in Simulated Home Environ- ments,”arXiv preprint arXiv.2403.07586, 2024
2024 arXiv
-
[20]
Replay-based domain incremen- tal learning for cross-user gesture recognition in robot task allocation,
K. K. Podder, P. Dutta, and J. Zhang, “Replay-based domain incremen- tal learning for cross-user gesture recognition in robot task allocation,” Electronics, vol. 14, no. 19, p. 3946, 2025
2025
-
[21]
Uncertainty-aware domain incremental learning for gesture recognition in robotic iot systems,
K. K. Podder, J. Zhang, S. Mao, and Z. Yu, “Uncertainty-aware domain incremental learning for gesture recognition in robotic iot systems,” in 2025 IEEE 11th World Forum on Internet of Things (WF-IoT). IEEE, 2025, pp. 1–7
2025
-
[22]
Domain-adversarial training of neural networks,
Y . Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Lavi- olette, M. March, and V . Lempitsky, “Domain-adversarial training of neural networks,”Journal of machine learning research, vol. 17, no. 59, pp. 1–35, 2016
2016
-
[23]
Incremental Adversarial Domain Adaptation for Continually Changing Environments,
M. Wulfmeier, A. Bewley, and I. Posner, “Incremental Adversarial Domain Adaptation for Continually Changing Environments,” in2018 IEEE International Conference on Robotics and Automation (ICRA), May 2018, pp. 4489–4495
2018
-
[24]
Adversarial Continuous Learning in Unsupervised Domain Adaptation,
Y . Zhang and B. D. Davison, “Adversarial Continuous Learning in Unsupervised Domain Adaptation,” inProceedings of the Interna- tional Conference on Pattern Recognition. Workshops and Challenges. Springer, January 2021, pp. 672–687
2021
-
[25]
Experience Replay for Continual Learning,
D. Rolnick, A. Ahuja, J. Schwarz, T. Lillicrap, and G. Wayne, “Experience Replay for Continual Learning,”Advances in Neural Information Processing Systems, vol. 32, 2019
2019
-
[26]
Challenges in task incremental learning for assistive robotics,
F. Feng, R. H. M. Chan, X. Shi, Y . Zhang, and Q. She, “Challenges in task incremental learning for assistive robotics,”IEEE Access, vol. 8, pp. 3434–3441, 2020
2020
-
[27]
The role of social norms in human–robot interaction: A systematic review,
S. Lawrence, M. Jouaiti, J. Hoey, C. L. Nehaniv, and K. Dautenhahn, “The role of social norms in human–robot interaction: A systematic review,”ACM Transactions on Human-Robot Interaction, vol. 14, no. 3, pp. 1–44, 2025
2025
-
[28]
Socially aware motion planning with deep reinforcement learning,
Y . F. Chen, M. Everett, M. Liu, and J. P. How, “Socially aware motion planning with deep reinforcement learning,” in2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), September 2017, pp. 1343–1350
2017
-
[29]
SEAN: Social Environment for Autonomous Navigation,
N. Tsoi, M. Hussein, J. Espinoza, X. Ruiz, and M. V ´azquez, “SEAN: Social Environment for Autonomous Navigation,” inProceedings of the 8th International Conference on Human-Agent Interaction, November 2020, pp. 281–283
2020
-
[30]
SocNav1: A Dataset to Benchmark and Learn Social Navigation Conventions,
L. J. Manso, P. Nu ˜nez, L. V . Calderita, D. R. Faria, and P. Bachiller, “SocNav1: A Dataset to Benchmark and Learn Social Navigation Conventions,”Data, vol. 5, no. 1, p. 7, March 2020
2020
-
[31]
Socially CompliAnt Navigation Dataset (SCAND): A Large-Scale Dataset of Demonstrations for Social Nav- igation,
H. Karnan, A. Nair, X. Xiao, G. Warnell, S. Pirk, A. Toshev, J. Hart, J. Biswas, and P. Stone, “Socially CompliAnt Navigation Dataset (SCAND): A Large-Scale Dataset of Demonstrations for Social Nav- igation,”IEEE Robotics and Automation Letters, vol. 7, no. 4, pp. 11 807–11 81...
2022
-
[32]
Learning Socially Appropriate Robot Approaching Behavior Toward Groups using Deep Reinforcement Learning,
Y . Gao, F. Yang, M. Frisk, D. Hemandez, C. Peters, and G. Castellano, “Learning Socially Appropriate Robot Approaching Behavior Toward Groups using Deep Reinforcement Learning,” in2019 28th IEEE International Conference on Robot and Human Interactive Commu- nication (RO-MAN),...
2019
-
[33]
Robotic etiquette: Results from user studies involving a fetch and carry task,
M. L. Walters, K. Dautenhahn, S. N. Woods, and K. L. Koay, “Robotic etiquette: Results from user studies involving a fetch and carry task,” inProceedings of the ACM/IEEE International Conference on Human- robot Interaction, March 2007, pp. 317–324
2007
-
[34]
A Review of Intent Detection, Arbitration, and Communication Aspects of Shared Control for Physical Human–Robot Interaction,
D. P. Losey, C. G. McDonald, E. Battaglia, and M. K. O’Malley, “A Review of Intent Detection, Arbitration, and Communication Aspects of Shared Control for Physical Human–Robot Interaction,”Applied Mechanics Reviews, vol. 70, no. 010804, February 2018
2018
-
[35]
Fully Automatic Analysis of Engagement and Its Relationship to Personality in Human-Robot Interactions,
H. Salam, O. C ¸ eliktutan, I. Hupont, H. Gunes, and M. Chetouani, “Fully Automatic Analysis of Engagement and Its Relationship to Personality in Human-Robot Interactions,”IEEE Access, vol. 5, pp. 705–721, 2017
2017
-
[36]
A spatial model of engagement for a social robot,
M. Michalowski, S. Sabanovic, and R. Simmons, “A spatial model of engagement for a social robot,” in9th IEEE International Workshop on Advanced Motion Control, 2006., March 2006, pp. 762–767
2006
-
[37]
Toward understanding social cues and signals in human–robot interaction: effects of robot gaze and proxemic behavior,
S. M. Fiore, T. J. Wiltshire, E. J. Lobato, F. G. Jentsch, W. H. Huang, and B. Axelrod, “Toward understanding social cues and signals in human–robot interaction: effects of robot gaze and proxemic behavior,” Frontiers in psychology, vol. 4, p. 859, 2013
2013
-
[38]
Re-evaluating continual learning scenarios: A categorization and case for strong baselines,
Y .-C. Hsu, Y .-C. Liu, A. Ramasamy, and Z. Kira, “Re-evaluating continual learning scenarios: A categorization and case for strong baselines,”arXiv preprint arXiv:1810.12488, 2018
2018 arXiv
-
[39]
Dual cognitive architecture: Incorporating biases and multi-memory systems for lifelong learning,
S. Gowda, B. Zonooz, and E. Arani, “Dual cognitive architecture: Incorporating biases and multi-memory systems for lifelong learning,” Transactions on Machine Learning Research, 2023
2023
-
[40]
Gradual divergence for seamless adaptation: a novel domain incremental learning method,
K. Jeeveswaran, E. Arani, and B. Zonooz, “Gradual divergence for seamless adaptation: a novel domain incremental learning method,” inProceedings of the 41st International Conference on Machine Learning, ser. ICML’24. JMLR.org, 2024
2024
-
[41]
Qwen2.5-vl technical report,
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y . Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y . Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin, “Qwen2.5-vl technical report,”arXiv prepri...
2025 arXiv
-
[42]
Llava-onevision: Easy visual task transfer,
B. Li, Y . Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, Y . Li, Z. Liu, and C. Li, “Llava-onevision: Easy visual task transfer,” 2024
2024
-
[43]
Deepseek-vl: Towards real-world vision-language understanding,
H. Lu, W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, Y . Sun, C. Deng, H. Xu, Z. Xie, and C. Ruan, “Deepseek-vl: Towards real-world vision-language understanding,” 2024
2024
-
[44]
A Concordance Correlation Coefficient to Evaluate Reproducibility,
L. I.-K. Lin, “A Concordance Correlation Coefficient to Evaluate Reproducibility,”Biometrics, vol. 45, no. 1, pp. 255–268, 1989
1989
-
[45]
Mobilenetv2: Inverted residuals and linear bottlenecks,
M. Sandler, A. G. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,”2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4510–4520, 2018
2018
-
[46]
Curriculum learning,
Y . Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum learning,” inProceedings of the 26th Annual International Conference on Machine Learning, 2009, p. 41–48. APPENDIX LATENCYANALYSIS To assess the real-world deployability of our design, we measure the end-to-end...
2009
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.