Pith. sign in

REVIEW 17 cited by

Resolution-robust Large Mask Inpainting with Fourier Convolutions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.07161 v2 pith:KQAM3OCX submitted 2021-09-15 cs.CV eess.IV

classification cs.CVeess.IV
keywords inpaintinglargefieldlamanetworkreceptiveachievesconvolutions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern image inpainting systems, despite the significant progress, often struggle with large missing areas, complex geometric structures, and high-resolution images. We find that one of the main reasons for that is the lack of an effective receptive field in both the inpainting network and the loss function. To alleviate this issue, we propose a new method called large mask inpainting (LaMa). LaMa is based on i) a new inpainting network architecture that uses fast Fourier convolutions (FFCs), which have the image-wide receptive field; ii) a high receptive field perceptual loss; iii) large training masks, which unlocks the potential of the first two components. Our inpainting network improves the state-of-the-art across a range of datasets and achieves excellent performance even in challenging scenarios, e.g. completion of periodic structures. Our model generalizes surprisingly well to resolutions that are higher than those seen at train time, and achieves this at lower parameter&time costs than the competitive baselines. The code is available at \url{https://github.com/saic-mdal/lama}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FantasyID: A dataset for detecting digital manipulations of ID-documents

    cs.CV 2025-07 conditional novelty 7.0 of 10

    The new FantasyID benchmark shows that current forgery detectors miss roughly half of face-swapped ID cards at a 10% false positive rate, while performing better on text edits.

  2. FocusView: Understanding and Customizing Informational Video Watching Experiences for Viewers with ADHD

    cs.HC 2025-07 conditional novelty 7.0 of 10

    FocusView, a customizable video interface, improved self-reported viewability of informational videos for 12 participants with ADHD.

  3. Fast Fourier Convolutional GAN for 30 m Clear-Sky Land Surface Temperature Gap-Free Reconstruction

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An FFC-based GAN guided by SAR and topographic data reconstructs 30 m clear-sky land-surface temperature in cloud-covered Landsat pixels with typical RMSE 0.8–1.8 K.

  4. LENS: LLM-guided Environment Simplification for Planning and Control in Clutter

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A vision-language-model-based prune-and-merge abstraction improves success and runtime for TAMP, contact-implicit MPC, and a VLA policy in cluttered tabletop manipulation.

  5. What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Erasing objects from front-camera images shows Alpamayo 1's trajectories depend most on large vehicles, pedestrians, and traffic lights, but attributions are seed-unstable and some effects reach the output without tou...

  6. ROSE: Remove Objects with Side Effects in Videos

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A video inpainting model trained on 3D-rendered pairs removes objects together with their shadows, reflections, and other side effects, plus a new benchmark.

  7. DreamPainter: Image Background Inpainting for E-commerce Scenarios

    cs.CV 2025-08 conditional novelty 6.0 of 10

    DreamPainter introduces a two-stage diffusion framework trained on a new synthetic e-commerce dataset, DreamEcom-400K, that outperforms open-source inpainting baselines on background generation with text and reference...

  8. Scalable and Realistic Virtual Try-on Application for Foundation Makeup with Kubelka-Munk Theory

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A virtual try-on framework approximates Kubelka-Munk optical blending with a Taylor expansion, enabling realistic foundation previews using only e-commerce product data.

  9. InstaInpaint: Instant 3D-Scene Inpainting with Masked Large Reconstruction Model

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A masked fine-tuning strategy adapts large reconstruction models to perform feed-forward 3D scene inpainting in 0.4 seconds with competitive state-of-the-art quality.

  10. Restoration of contaminated data in an Intensity Mapping survey using deep neural networks

    eess.SP 2025-06 conditional novelty 6.0 of 10

    Restoring RFI-contaminated pixels with the LaMa inpainting network reduces post-foreground-removal RMS and improves recovery of the large-scale 21-cm power spectrum in simulations.

  11. 3D-GIMP: When 3D Gaussian Inpainting Meets PatchMatch

    cs.CV 2026-07 conditional novelty 5.0 of 10

    3D-GIMP removes objects from 3D Gaussian Splatting scenes by inpainting one reference view and propagating it to all views via a 3D-aware PatchMatch field, cutting optimization time from ~1 hour to ~6 minutes.

  12. PhysOmni: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction

    cs.GR 2026-05 unverdicted novelty 5.0 of 10

    TelePhysics is a training-free pipeline that builds a unified 3D scene model from one photo and then runs decoupled physics simulation to produce controllable, penetration-free multi-object videos.

  13. STEMTOX: From Collaborative Tags to Fine-Grained Toxic Meme Detection via Entropy-Guided Multi-Task Learning

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    A new dataset and an entropy-guided multi-task model show that collaborative tags improve fine-grained toxic meme detection.

  14. CA-World: Multi-Object Counterfactual Alignment for Efficient Interactive-Ready Reconstruction

    cs.CV 2026-05 unverdicted novelty 4.0 of 10

    The paper's stated CA-World counterfactual claim is absent from the body, which instead describes the SAM3D-Phys pipeline for multi-object interactive reconstruction and simulation.

  15. Neural Field Representations of Mobile Computational Photography

    cs.CV 2025-08 conditional novelty 4.0 of 10

    Fitting neural fields directly to raw phone bursts reconstructs depth, separates reflections and occluders, and stitches panoramas, outperforming the compared baselines on the thesis's benchmarks.

  16. WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration

    cs.SD 2025-08 conditional novelty 3.0 of 10

    WaveLLDM, a lightweight latent diffusion model with a neural codec, achieves low spectral distortion (LSD 0.48-0.60) on speech restoration but scores far below SOTA on PESQ and STOI.

  17. Enhancing non-Rigid 3D Model Deformations Using Mesh-based Gaussian Splatting

    cs.GR 2025-07 reject novelty 2.0 of 10

    A proposal to combine 3D Gaussian splatting, SAM segmentation, GS2Mesh conversion, LLM-based material assignment, and XPBD physics into a mesh-based editing pipeline, with no experimental validation.

Pith tools