Pith. sign in

REVIEW 1 cited by

Stacked Dense U-Nets with Dual Transformers for Robust Face Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1812.01936 v1 pith:5IFSAQ32 submitted 2018-12-05 cs.CV

classification cs.CV
keywords facestackedu-netsdensemodelaggregationalignmentimages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Facial landmark localisation in images captured in-the-wild is an important and challenging problem. The current state-of-the-art revolves around certain kinds of Deep Convolutional Neural Networks (DCNNs) such as stacked U-Nets and Hourglass networks. In this work, we innovatively propose stacked dense U-Nets for this task. We design a novel scale aggregation network topology structure and a channel aggregation building block to improve the model's capacity without sacrificing the computational complexity and model size. With the assistance of deformable convolutions inside the stacked dense U-Nets and coherent loss for outside data transformation, our model obtains the ability to be spatially invariant to arbitrary input face images. Extensive experiments on many in-the-wild datasets, validate the robustness of the proposed method under extreme poses, exaggerated expressions and heavy occlusions. Finally, we show that accurate 3D face alignment can assist pose-invariant face recognition where we achieve a new state-of-the-art accuracy on CFP-FP.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SVFR: A Unified Framework for Generalized Video Face Restoration

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A unified diffusion model jointly trained on video face restoration, colorization, and inpainting outperforms single-task baselines on VFHQ-test.

Pith tools