Pith. sign in

REVIEW 3 cited by

Masked Diffusion with Task-awareness for Procedure Planning in Instructional Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.07409 v1 pith:GEVXEOHH submitted 2023-09-14 cs.CV

classification cs.CV
keywords actiondiffusiontypestaskvisualchallengedecisionembedding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A key challenge with procedure planning in instructional videos lies in how to handle a large decision space consisting of a multitude of action types that belong to various tasks. To understand real-world video content, an AI agent must proficiently discern these action types (e.g., pour milk, pour water, open lid, close lid, etc.) based on brief visual observation. Moreover, it must adeptly capture the intricate semantic relation of the action types and task goals, along with the variable action sequences. Recently, notable progress has been made via the integration of diffusion models and visual representation learning to address the challenge. However, existing models employ rudimentary mechanisms to utilize task information to manage the decision space. To overcome this limitation, we introduce a simple yet effective enhancement - a masked diffusion model. The introduced mask acts akin to a task-oriented attention filter, enabling the diffusion/denoising process to concentrate on a subset of action types. Furthermore, to bolster the accuracy of task classification, we harness more potent visual representation learning techniques. In particular, we learn a joint visual-text embedding, where a text embedding is generated by prompting a pre-trained vision-language model to focus on human actions. We evaluate the method on three public datasets and achieve state-of-the-art performance on multiple metrics. Code is available at https://github.com/ffzzy840304/Masked-PDPP.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    InsTALL feeds a procedural graph mined from training videos into a multimodal LLM, improving task recognition, action recognition, next-action and plan prediction, and error detection on instructional videos.

  2. VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting

    cs.CV 2024-12 conditional novelty 5.0 of 10

    VG-TVP enriches LLM-generated text plans with captions from instructional videos and produces a short video per step; human raters prefer it over text-only baselines.

  3. A Survey on Diffusion Models for Anomaly Detection

    cs.LG 2025-01 conditional novelty 4.0 of 10

    A taxonomy and literature review of diffusion-model-based anomaly detection, covering methods, tasks, benchmarks, and open challenges.

Pith tools