Pith. sign in

REVIEW 3 cited by

LAD-Reasoner: Tiny Multimodal Models are Good Reasoners for Logical Anomaly Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.12749 v1 pith:HCRV7WQI submitted 2025-04-17 cs.CV

classification cs.CV
keywords anomalydetectionlogicalreasoninglad-reasoneraccuracydatainterpretable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in industrial anomaly detection have highlighted the need for deeper logical anomaly analysis, where unexpected relationships among objects, counts, and spatial configurations must be identified and explained. Existing approaches often rely on large-scale external reasoning modules or elaborate pipeline designs, hindering practical deployment and interpretability. To address these limitations, we introduce a new task, Reasoning Logical Anomaly Detection (RLAD), which extends traditional anomaly detection by incorporating logical reasoning. We propose a new framework, LAD-Reasoner, a customized tiny multimodal language model built on Qwen2.5-VL 3B. Our approach leverages a two-stage training paradigm that first employs Supervised Fine-Tuning (SFT) for fine-grained visual understanding, followed by Group Relative Policy Optimization (GRPO) to refine logical anomaly detection and enforce coherent, human-readable reasoning. Crucially, reward signals are derived from both the detection accuracy and the structural quality of the outputs, obviating the need for building chain of thought (CoT) reasoning data. Experiments on the MVTec LOCO AD dataset show that LAD-Reasoner, though significantly smaller, matches the performance of Qwen2.5-VL-72B in accuracy and F1 score, and further excels in producing concise and interpretable rationales. This unified design reduces reliance on large models and complex pipelines, while offering transparent and interpretable insights into logical anomaly detection. Code and data will be released.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Global Logic and Local Search: Dual-Stream Multimodal In-Context Learning for Verifiable Industrial Anomaly Detection

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A training-free dual-stream multimodal framework (PVLA + SAM 3 global logic + MCTS local search) improves verifiable industrial anomaly QA without defective training samples.

  2. EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A difficulty-aware GRPO training scheme with response resampling, advantage reweighting, GPT-generated text samples, and heatmap-guided contrastive embeddings improves InternVL3-8B by 7.77 percentage points on the MMA...

  3. OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning

    cs.CV 2025-05 reject novelty 5.0 of 10

    OmniAD unifies industrial anomaly detection and understanding in a single multimodal model using text-encoded masks and reinforcement learning, reporting 79.1 on MMAD and strong detection scores.

Pith tools