Pith. sign in

REVIEW 3 cited by

Robot-Led Vision Language Model Wellbeing Assessment of Children

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.02765 v1 pith:SZCMUL6N submitted 2025-04-03 cs.RO

classification cs.RO
keywords childrenwellbeingassessmentsmodelrobot-ledassessmentgenderlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study presents a novel robot-led approach to assessing children's mental wellbeing using a Vision Language Model (VLM). Inspired by the Child Apperception Test (CAT), the social robot NAO presented children with pictorial stimuli to elicit their verbal narratives of the images, which were then evaluated by a VLM in accordance with CAT assessment guidelines. The VLM's assessments were systematically compared to those provided by a trained psychologist. The results reveal that while the VLM demonstrates moderate reliability in identifying cases with no wellbeing concerns, its ability to accurately classify assessments with clinical concern remains limited. Moreover, although the model's performance was generally consistent when prompted with varying demographic factors such as age and gender, a significantly higher false positive rate was observed for girls, indicating potential sensitivity to gender attribute. These findings highlight both the promise and the challenges of integrating VLMs into robot-led assessments of children's wellbeing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Relational Positioning as a Measurable Risk Object: History-Carried Lock-in and Self-Confabulation in Multi-Turn Human-AI Dialogue

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Relational positioning in multi-turn LLM dialogue forms a history-carried lock-in (~60-point separation under identical neutral continuations) and models self-confabulate autobiographies on ~40% of reciprocity-eliciti...

  2. FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment

    cs.AI 2026-04 unverdicted novelty 5.0 of 10

    Zero-shot vision-language models are unreliable and vary widely for depression screening, and explainability-based fairness interventions often trade away accuracy without reliable fairness gains.

  3. Critical Insights about Robots for Mental Wellbeing

    cs.RO 2025-06 conditional novelty 3.0 of 10

    Social robots for mental wellbeing can work as coaches or via video, and should be designed with clinicians and studied long term.

Pith tools