REVIEW 17 cited by
Resolution-robust Large Mask Inpainting with Fourier Convolutions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Modern image inpainting systems, despite the significant progress, often struggle with large missing areas, complex geometric structures, and high-resolution images. We find that one of the main reasons for that is the lack of an effective receptive field in both the inpainting network and the loss function. To alleviate this issue, we propose a new method called large mask inpainting (LaMa). LaMa is based on i) a new inpainting network architecture that uses fast Fourier convolutions (FFCs), which have the image-wide receptive field; ii) a high receptive field perceptual loss; iii) large training masks, which unlocks the potential of the first two components. Our inpainting network improves the state-of-the-art across a range of datasets and achieves excellent performance even in challenging scenarios, e.g. completion of periodic structures. Our model generalizes surprisingly well to resolutions that are higher than those seen at train time, and achieves this at lower parameter&time costs than the competitive baselines. The code is available at \url{https://github.com/saic-mdal/lama}.
Forward citations
Cited by 17 Pith papers
-
FantasyID: A dataset for detecting digital manipulations of ID-documents
The new FantasyID benchmark shows that current forgery detectors miss roughly half of face-swapped ID cards at a 10% false positive rate, while performing better on text edits.
-
FocusView: Understanding and Customizing Informational Video Watching Experiences for Viewers with ADHD
FocusView, a customizable video interface, improved self-reported viewability of informational videos for 12 participants with ADHD.
-
Fast Fourier Convolutional GAN for 30 m Clear-Sky Land Surface Temperature Gap-Free Reconstruction
An FFC-based GAN guided by SAR and topographic data reconstructs 30 m clear-sky land-surface temperature in cloud-covered Landsat pixels with typical RMSE 0.8–1.8 K.
-
LENS: LLM-guided Environment Simplification for Planning and Control in Clutter
A vision-language-model-based prune-and-merge abstraction improves success and runtime for TAMP, contact-implicit MPC, and a VLA policy in cluttered tabletop manipulation.
-
What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning
Erasing objects from front-camera images shows Alpamayo 1's trajectories depend most on large vehicles, pedestrians, and traffic lights, but attributions are seed-unstable and some effects reach the output without tou...
-
ROSE: Remove Objects with Side Effects in Videos
A video inpainting model trained on 3D-rendered pairs removes objects together with their shadows, reflections, and other side effects, plus a new benchmark.
-
DreamPainter: Image Background Inpainting for E-commerce Scenarios
DreamPainter introduces a two-stage diffusion framework trained on a new synthetic e-commerce dataset, DreamEcom-400K, that outperforms open-source inpainting baselines on background generation with text and reference...
-
Scalable and Realistic Virtual Try-on Application for Foundation Makeup with Kubelka-Munk Theory
A virtual try-on framework approximates Kubelka-Munk optical blending with a Taylor expansion, enabling realistic foundation previews using only e-commerce product data.
-
InstaInpaint: Instant 3D-Scene Inpainting with Masked Large Reconstruction Model
A masked fine-tuning strategy adapts large reconstruction models to perform feed-forward 3D scene inpainting in 0.4 seconds with competitive state-of-the-art quality.
-
Restoration of contaminated data in an Intensity Mapping survey using deep neural networks
Restoring RFI-contaminated pixels with the LaMa inpainting network reduces post-foreground-removal RMS and improves recovery of the large-scale 21-cm power spectrum in simulations.
-
3D-GIMP: When 3D Gaussian Inpainting Meets PatchMatch
3D-GIMP removes objects from 3D Gaussian Splatting scenes by inpainting one reference view and propagating it to all views via a 3D-aware PatchMatch field, cutting optimization time from ~1 hour to ~6 minutes.
-
PhysOmni: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction
TelePhysics is a training-free pipeline that builds a unified 3D scene model from one photo and then runs decoupled physics simulation to produce controllable, penetration-free multi-object videos.
-
STEMTOX: From Collaborative Tags to Fine-Grained Toxic Meme Detection via Entropy-Guided Multi-Task Learning
A new dataset and an entropy-guided multi-task model show that collaborative tags improve fine-grained toxic meme detection.
-
CA-World: Multi-Object Counterfactual Alignment for Efficient Interactive-Ready Reconstruction
The paper's stated CA-World counterfactual claim is absent from the body, which instead describes the SAM3D-Phys pipeline for multi-object interactive reconstruction and simulation.
-
Neural Field Representations of Mobile Computational Photography
Fitting neural fields directly to raw phone bursts reconstructs depth, separates reflections and occluders, and stitches panoramas, outperforming the compared baselines on the thesis's benchmarks.
-
WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration
WaveLLDM, a lightweight latent diffusion model with a neural codec, achieves low spectral distortion (LSD 0.48-0.60) on speech restoration but scores far below SOTA on PESQ and STOI.
-
Enhancing non-Rigid 3D Model Deformations Using Mesh-based Gaussian Splatting
A proposal to combine 3D Gaussian splatting, SAM segmentation, GS2Mesh conversion, LLM-based material assignment, and XPBD physics into a mesh-based editing pipeline, with no experimental validation.
Discussion (0). Continue with ORCID to comment.