Pith. sign in

REVIEW 1 cited by

LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.09390 v3 pith:LMU4XMHB submitted 2024-06-13 cs.CV cs.LG

classification cs.CVcs.LG
keywords llavidaladl-xvideoactivitiesbenchmarkscomplexdailydatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current Large Language Vision Models (LLVMs) trained on web videos perform well in general video understanding but struggle with fine-grained details, complex human-object interactions (HOI), and view-invariant representation learning essential for Activities of Daily Living (ADL). This limitation stems from a lack of specialized ADL video instruction-tuning datasets and insufficient modality integration to capture discriminative action representations. To address this, we propose a semi-automated framework for curating ADL datasets, creating ADL-X, a multiview, multimodal RGBS instruction-tuning dataset. Additionally, we introduce LLAVIDAL, an LLVM integrating videos, 3D skeletons, and HOIs to model ADL's complex spatiotemporal relationships. For training LLAVIDAL a simple joint alignment of all modalities yields suboptimal results; thus, we propose a Multimodal Progressive (MMPro) training strategy, incorporating modalities in stages following a curriculum. We also establish ADL MCQ and video description benchmarks to assess LLVM performance in ADL tasks. Trained on ADL-X, LLAVIDAL achieves state-of-the-art performance across ADL benchmarks. Code and data will be made publicly available at: https://adl-x.github.io/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living

    cs.CV 2025-02 conditional novelty 6.0 of 10

    SKI models distill skeleton-language knowledge into video-language encoders, boosting zero-shot ADL action recognition accuracy by up to 7.8 percentage points while discarding skeletons at inference.

Pith tools