{"id":"f3b9328b-785e-407a-b845-88d6dc1eae41","arxiv_id":"2412.00955","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"WAFFLE is a diverse, automatically curated dataset of nearly 20,000 in-the-wild floorplans with LLM-extracted pseudo-labels, enabling new building understanding and generation tasks.","lead":"This paper introduces WAFFLE, a dataset of nearly 20,000 floorplan images from Wikimedia Commons paired with text metadata such as building type, country, and labeled architectural regions. Fine-tuning standard vision and generation models on WAFFLE improves building type retrieval, semantic segmentation, and text-to-floorplan generation compared with the same models without this data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Building-type retrieval may be inflated by text leakage: floorplan images contain OCR-readable text, and the retrieval experiment does not mask it, so CLIP may be reading words rather than understanding floorplan layouts.","rationale":"The reader's conditional verdict is appropriate, and this stress-test pass does not move it. The reader identified pseudo-label accuracy as the weakest assumption; I agree that this matters, but I find a more specific and arguably more load-bearing gap: the retrieval experiment does not control for text visible in the input image. This can be settled with a concrete masking experiment. If the concern lands, the building-type understanding result weakens, but the dataset contribution and the generative results are not necessarily invalidated; the paper would need an additional condition, not a rejection. If the concern does not land, the retrieval evidence stands. Since the current verdict is already CONDITIONAL, the appropriate target verdict remains unchanged. I mark agreement as 'partial' because the reader's pseudo-label concern and my text-leakage concern are related but not identical: my proposed test does not depend on whether the pseudo-labels are accurate, only on whether the model can exploit text in the image. The paper should either add OCR-masked retrieval evaluation or explicitly reframe the task as floorplan-with-text understanding rather than floorplan-layout understanding.","tokens_in":21501,"tokens_out":6855,"duration_ms":71170,"concrete_test":"Re-evaluate the fine-tuned CLIP model from Section 4.1 on the test set after masking all OCR-text regions (using the existing Google Vision OCR boxes and the same Stable Diffusion inpainting procedure introduced in Section 4.2). If R@1, R@5, or MRR drop by more than half relative to Table 2, text leakage is the dominant source of the reported retrieval improvement; if retrieval performance is unchanged, the visual-semantics interpretation survives this objection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline discriminative evidence is the building-type retrieval experiment in Section 4.1, but it never controls for text inside the input images. The pipeline already collects OCR detections (Section 3.1), and the segmentation experiment explicitly inpaints OCR regions to prevent text leakage (Section 4.2), showing the authors are aware that raw floorplan images carry readable text. For retrieval, however, the images are fed to CLIP with no such masking. Many in-the-wild floorplans contain titles, labels, and legends; CLIP's vision-language pretraining gives it some ability to read text, and fine-tuning can amplify this shortcut. If the model raises R@1 from 1.5% to 11.8% (Table 2) by recognizing words like 'cathedral' or 'castle' in the image, the result does not demonstrate floorplan-layout understanding. This is distinct from, but adjacent to, the reader's concern about pseudo-label accuracy: even if the Llama-2 labels are correct, the retrieval evaluation could be measuring text recognition rather than visual semantics. The paper's own Figure 1 highlights images 'lacking textual descriptions' as a challenging, text-free case, which further suggests that text is an uncontrolled and potentially dominant cue in the main retrieval evaluation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces WAFFLE, a dataset of 18,556 floorplan images scraped from Wikimedia Commons, with LLM-extracted pseudo-labels (building name, type, country, grounded architectural features) and OCR/detection metadata. The authors evaluate the dataset as a benchmark and as training data for four tasks: fine-tuned CLIP retrieval of building type (Section 4.1), open-vocabulary segmentation with CLIPSeg (Section 4.2), a semantic segmentation benchmark for existing models (Section 4.3), and text- and structure-conditioned floorplan generation with Stable Diffusion and ControlNet (Sections 4.4 and 4.5), including a user study. The central claim is that WAFFLE makes feasible new discriminative and generative building-understanding tasks and that fine-tuning on it improves performance relative to strong baselines.","tokens_in":21746,"tokens_out":6598,"duration_ms":58424,"significance":"If the central claims hold, WAFFLE would be a valuable community resource: it is substantially larger and more diverse in building type and country than existing floorplan datasets, it is built by an automatic pipeline, and it is intended for public release with code and models. The paper also takes useful precautions in the weakly supervised segmentation setup, notably deriving labels from Wikipedia text rather than from image pixels and inpainting OCR regions before training. The overall significance therefore hinges on whether the reported gains reflect visual understanding of floorplan layout rather than text leakage or unverified pseudo-labels; the experiments as reported do not yet establish that.","major_comments":[{"comment":"Table 2's building-type retrieval experiment feeds raw floorplan images to CLIP without masking or inpainting OCR text, even though Section 4.2 explicitly inpaints OCR regions \"to prevent leakage from the written text in the images.\" In-the-wild floorplans frequently contain titles, labels, and legends, and the dataset pipeline collects OCR detections precisely because this text is informative. The improvement from R@1=1.5% to 11.8% could therefore be driven by reading words such as \"cathedral\" or \"castle\" rather than by understanding floorplan layout, leaving the central claim of visual building-type understanding unsupported. The authors should repeat the evaluation on OCR-inpainted or masked images, or on a text-free subset such as the bottom images in Figure 1, to control for this confound.","section":"Section 4.1"},{"comment":"The pseudo-label validation is limited to 100 random samples from the whole dataset, with reported accuracies of 89% for building names, 85% for building types, and 96% for countries. The manual inspection of the test set only removes images that do not contain a valid floorplan and does not verify the building-type labels that are used as ground truth in the retrieval and generation evaluations (Tables 2, 5, 6). Since building-type pseudo-labels are both the training target and the evaluation target, systematic errors or biases in those labels could inflate the apparent improvements of CLIP_FT and SD_FT. The authors should report pseudo-label accuracy on the test set, with a per-type breakdown, or re-verify the test-set labels before drawing conclusions.","section":"Section 3.3"},{"comment":"The reported comparisons lack confidence intervals, significance tests, or multiple-seed variance, which is necessary to support the paper's \"significantly outperforms\" language. This is especially important for Table 3, which uses a 95-image manually selected evaluation set and an empirically chosen mIoU threshold of 0.25, and for Table 5, where the CLIP similarity differences (e.g., 24.9 vs 25.6) are small and FID is known to be unstable on small datasets. The user study in Section 4.4 reports a 70.42% preference without a confidence interval or participant-level variance. Adding standard errors, paired tests, or seed-level results would substantially strengthen the central claims.","section":"Tables 2, 3, 5"},{"comment":"The comparison against CC5K in the open-vocabulary segmentation experiment is not apples-to-apples: the footnote states that CC5K was evaluated only over a subset of residential buildings, while CLIPSeg and the fine-tuned model are evaluated on the full manually annotated 95-image set. The claim in Section 4.2 that the method \"outperform[s] the strongly-supervised CC5K\" is therefore not directly supported by the table as presented. The authors should evaluate all methods on the same image set or clearly report subset sizes and separate results for the common subset.","section":"Table 3"}],"minor_comments":[{"comment":"The phrase \"we leverage our data in an unsupervised manner\" is imprecise for the structure-conditioned experiment, because the ControlNet for internal structure is trained on CubiCasa5K annotations; only the boundary-conditioned variant uses automatically extracted edges from WAFFLE. The caption of Figure 7 also states that the condition is from CubiCasa5K. The text should clarify what is trained on WAFFLE versus external data.","section":"Section 4.5"},{"comment":"Fine-tuning actually hurts the Theatre category on FID (238.4 vs 189.5) and KMMD (0.17 vs 0.08) relative to the pretrained model; the claim that fine-tuning improves \"the vast majority of building types\" should be replaced with a precise statement of the number and identities of categories that improve.","section":"Table 6"},{"comment":"The semantic-similarity thresholds (0.7 positive, 0.4 negative) and the mIoU binarization threshold (0.25) are set empirically; reporting a short sensitivity analysis would help establish that the segmentation improvement does not depend on these choices.","section":"Appendix C.2"},{"comment":"The kernel used for KMMD is not specified; reproducible reporting of kernel maximum mean discrepancy requires the kernel choice and bandwidth.","section":"Section 4.4"},{"comment":"Minor typographical issues: \"W AF FLE\" and \"W AFFLE\" appear in the abstract, and \"manualy inpect\" appears in Appendix B.3.2.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The contribution is potentially valuable and fits the venue's scope; the major issues concern experimental controls and evaluation statistics rather than the dataset construction itself. If the authors can add OCR-masked retrieval results, verify the test-set pseudo-labels, and provide significance or error-bar information, I would be willing to reconsider favorably."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"WAFFLE is a genuinely new resource and probably the closest thing the floorplan literature has to a general-purpose corpus: roughly 19K real images, over 1K building types, more than 100 countries, and grounded legends and architectural features scraped from Wikimedia and structured with LLMs. That is the real contribution, and it is not in earlier datasets. The four experiments are sensible demonstrations that fine-tuning on WAFFLE helps, and the paper is honest about the historic/religious bias in the data and about pseudo-label noise.\n\nThe main soft spot is the building-type retrieval experiment. The authors know floorplans carry readable text—they OCR everything and inpaint OCR regions in the segmentation setup to prevent leakage—but in Section 4.1 CLIP is fine-tuned on raw images with no text masking. Many in-the-wild floorplans contain titles and labels that literally spell \"cathedral\" or \"castle.\" The jump from 1.5% to 11.8% R@1 could be partly reading words, not parsing layout. A masked or ablated version, where OCR boxes are inpainted as in the segmentation experiment, is needed before this result supports the \"layout understanding\" claim. This is not fatal to the dataset, but it is a load-bearing evaluation gap.\n\nOther soft spots are smaller. No confidence intervals or significance tests appear anywhere, and the reported gains in generation (CLIP similarity 24.9 to 25.6) are within noise. The pseudo-labels are validated on 100 samples; 85% accuracy on building type means roughly 15% of training labels are wrong, which is tolerable for training but matters when the same labels are ground truth in retrieval. The mIoU threshold is set empirically to 0.25, and the segmentation evaluation set is 95 images. These are fixable issues for a revised version. I do not see circularity: building-type labels come from Wikipedia text, not pixels, and the weakly supervised segmentation setup with inpainted text is standard.\n\nThe paper deserves a serious referee. If the dataset ships as promised, it will be a standard resource for diverse floorplan understanding. The main things to ask for in revision are a text-masked retrieval baseline, error bars or significance tests on the main tables, and a slightly larger pseudo-label audit. Send it out.","headline":"A genuinely new diverse floorplan dataset, but the retrieval evaluation likely overstates layout understanding by not masking text in the images.","tokens_in":22328,"tokens_out":2174,"would_cite":true,"duration_ms":22474,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"WAFFLE is a nearly 20K-image floorplan dataset from Wikimedia Commons with LLM-extracted metadata, and the paper argues it enables building-type retrieval, open-vocabulary segmentation, and type-conditioned floorplan generation that prior…","keywords":["floorplan understanding","in-the-wild dataset","multimodal learning","building type retrieval","open-vocabulary segmentation","text-to-image generation","pseudo-labeling","architectural features"],"falsifier":"Take a random sample of WAFFLE test images, have independent annotators label building type and grounded features without seeing the machine-generated labels, then evaluate the fine-tuned retrieval and segmentation models against those human labels; if performance collapses toward chance while the models still match the machine labels, the reported understanding is an artifact of label bias rather than building semantics.","tokens_in":21286,"feed_emoji":"🏛️","tokens_out":8066,"duration_ms":67501,"temperature":0.7,"pith_summary":"This paper introduces WAFFLE, a dataset of 18,556 floorplan images gathered from openly available Internet media, paired with structured metadata extracted by a large language model: building name, building type, country, and architectural features grounded to image regions. The central claim is that this automatically curated, noisy data supports building-understanding tasks that earlier floorplan datasets, which mostly cover apartments from a single country, cannot support. On this data, a fine-tuned contrastive vision-language model learns to retrieve a building's type from its floorplan, an open-vocabulary segmentation model learns to localize architectural terms, and a text-to-image diffusion model learns to generate floorplan images for specific building types. A sympathetic reader would take the paper as showing that in-the-wild schematic imagery plus LLM-generated labels can serve as a foundation for learning the semantics of buildings.","feed_headline":"20K wild floorplans teach models to read buildings","feed_subtitle":"New dataset links floorplan images to building type, country, and grounded architectural features, unlocking retrieval, segmentation, and…","key_machinery":"The machine behind the dataset is a fully automatic curation-and-labeling pipeline. A transformer object detector (DETR) fine-tuned on 200 annotated images locates layout components (floorplan, legend, compass, scale); the contrastive image-text model CLIP filters scraped images to floorplans; OCR is run on each image; and prompts to the large language model Llama-2 turn the noisy metadata and OCR text into structured pseudo-ground truth, including building name, type, country, and architectural features. Grounding is achieved by matching legend keys and feature names to OCR detections inside the detected floorplan and legend regions, giving the textual labels spatial coordinates. This pipeline is what converts unlabeled schematic images into paired image-text data usable for retrieval, segmentation, and generation.","core_discovery":"WAFFLE contains nearly 20K floorplan images spanning more than 100 countries, over 1K building types, and over 11K distinct grounded architectural features, with a train/test split of 18,259 and 297 images made by country to prevent building leakage. The paper's claim is that these data enable new discriminative and generative tasks: fine-tuning CLIP on paired images and building-type labels raises retrieval Recall@1 from 1.5% to 11.8%; fine-tuning CLIPSeg on grounded features raises open-vocabulary segmentation Average Precision from 0.157 to 0.226; and fine-tuning Stable Diffusion with prompts like \"a floor plan of a <building type>\" improves generation realism (FID from 194.8 to 145.3) and prompt alignment, with users preferring the fine-tuned generations 70.42% of the time. The same images also serve as a challenging benchmark: a supervised residential-floorplan segmentation model trained on CubiCasa5K attains only 0.488 IoU on walls in WAFFLE, suggesting that prior narrow datasets do not transfer to diverse real-world floorplans. The paper reports manual inspection of 100 random samples with 85-96% accuracy for the language-model-extracted labels.","pith_inferences":["Inference: because the test split is chosen by country, part of the retrieval and generation gains could come from country-specific visual style rather than building semantics; an audit that swaps or mixes countries would separate these factors.","Inference: the same LLM-curation pipeline could plausibly transfer to other schematic document families such as maps, engineering drawings, or historic blueprints, though the paper does not test this.","Inference: the reported label accuracy rests on 100 manual samples; verifying a larger random sample and the full test set would strengthen the case that the pseudo-labels, not metadata bias, drive the results.","Inference: the dataset's acknowledged skew toward historic and religious buildings implies that models trained on it may underperform on modern commercial and industrial building types."],"forward_implications":["A model fine-tuned on WAFFLE can identify the building type of a never-seen floorplan, including non-residential types such as castles, temples, and hospitals.","Open-vocabulary segmentation trained on WAFFLE can localize architectural terms such as nave, choir, and court inside floorplans, including terms absent from fixed residential label sets.","Text-conditioned generation on WAFFLE produces floorplan images whose structure reflects the prompt type, such as towers for castles and many small rooms for hospitals.","WAFFLE exposes the limited transfer of existing supervised residential models: their segmentation performance drops sharply on diverse, in-the-wild floorplans.","Structure-conditioned generation works for diverse building types even when the structural constraint is unusual for that type, because the text condition and layout condition are fused from separate sources."],"supporting_citations":[{"why":"A transformer object detector that, after fine-tuning on 200 images, supplies the bounding boxes for floorplan, legend, compass, and scale components.","marker":"[3]"},{"why":"CubiCasa5K provides the supervised residential-floorplan segmentation baseline and the annotated layout source used for structure-conditioned generation.","marker":"[18]"},{"why":"CLIPSeg is the open-vocabulary text-guided segmentation model that is fine-tuned on WAFFLE's grounded architectural features.","marker":"[23]"},{"why":"CLIP is the backbone used for image filtering and for the building-type retrieval model fine-tuned on WAFFLE.","marker":"[31]"},{"why":"Stable Diffusion is the text-to-image model fine-tuned for floorplan generation and used as the backbone for control-conditioned generation.","marker":"[32]"},{"why":"Llama-2 is the language model prompted to extract structured metadata, building names, building types, countries, and grounded feature lists from noisy text and OCR.","marker":"[36]"},{"why":"RPLAN is one of the prior apartment-only datasets whose limited scope WAFFLE is contrasted against.","marker":"[43]"},{"why":"ControlNet adds structural conditioning to the fine-tuned generation model, enabling layout-guided floorplan generation.","marker":"[49]"}],"fun_headline_variants":["20K floorplans from 100+ countries power new building AI tasks","WAFFLE: 20K diverse floorplans, 1K building types, 11K features","From 1 region to 100+: WAFFLE's 20K wild floorplans","20K real-world floorplans: WAFFLE dataset leaps from 1 region to 100+"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the building type, country, and architectural-feature labels produced by Llama-2 are accurate enough to serve both as training signal and as test ground truth; they are validated only by manual inspection of 100 random samples (85-96% accuracy), and the test-set labels themselves are not re-verified.","fun_headline_variants_meta":{"raw":{"variants":["20K floorplans from 100+ countries power new building AI tasks","WAFFLE: 20K diverse floorplans, 1K building types, 11K features","From 1 region to 100+: WAFFLE's 20K wild floorplans","20K real-world floorplans: WAFFLE dataset leaps from 1 region to 100+"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000986,"raw_usage":{"total_tokens":4207,"prompt_tokens":995,"completion_tokens":3212,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":611,"completion_tokens_details":{"reasoning_tokens":3112}},"tokens_in":611,"tokens_out":3212,"duration_ms":20545,"temperature":1.0,"reasoning_tokens":3112,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:49:38.295769+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of WAFFLE test images, have independent annotators label building type and grounded features without seeing the machine-generated labels, then evaluate the fine-tuned retrieval and segmentation models against those human labels; if performance collapses toward chance while the models still match the machine labels, the reported understanding is an artifact of label bias rather than building semantics.","supporting_citations":[{"cited_title":"End-to- end object detection with transformers","cited_arxiv_id":null,"evidence_quote":"A transformer object detector that, after fine-tuning on 200 images, supplies the bounding boxes for floorplan, legend, compass, and scale components."},{"cited_title":"Cubicasa5k: A dataset and an im- proved multi-task model for floorplan image analysis","cited_arxiv_id":null,"evidence_quote":"CubiCasa5K provides the supervised residential-floorplan segmentation baseline and the annotated layout source used for structure-conditioned generation."},{"cited_title":"Image segmenta- tion using text and image prompts","cited_arxiv_id":null,"evidence_quote":"CLIPSeg is the open-vocabulary text-guided segmentation model that is fine-tuned on WAFFLE's grounded architectural features."},{"cited_title":"Learning transferable visual models from natural language supervi- sion","cited_arxiv_id":null,"evidence_quote":"CLIP is the backbone used for image filtering and for the building-type retrieval model fine-tuned on WAFFLE."},{"cited_title":"High-resolution image syn- thesis with latent diffusion models, 2022","cited_arxiv_id":null,"evidence_quote":"Stable Diffusion is the text-to-image model fine-tuned for floorplan generation and used as the backbone for control-conditioned generation."},{"cited_title":"Data-driven interior plan genera- tion for residential buildings","cited_arxiv_id":null,"evidence_quote":"RPLAN is one of the prior apartment-only datasets whose limited scope WAFFLE is contrasted against."},{"cited_title":"Adding conditional control to text-to-image diffusion models, 2023","cited_arxiv_id":null,"evidence_quote":"ControlNet adds structural conditioning to the fine-tuned generation model, enabling layout-guided floorplan generation."}],"review_version":1}