REVIEW 4 cited by
Fine-Tuning Vision-Language Model for Automated Engineering Drawing Information Extraction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Geometric Dimensioning and Tolerancing (GD&T) plays a critical role in manufacturing by defining acceptable variations in part features to ensure component quality and functionality. However, extracting GD&T information from 2D engineering drawings is a time-consuming and labor-intensive task, often relying on manual efforts or semi-automated tools. To address these challenges, this study proposes an automated and computationally efficient GD&T extraction method by fine-tuning Florence-2, an open-source vision-language model (VLM). The model is trained on a dataset of 400 drawings with ground truth annotations provided by domain experts. For comparison, two state-of-the-art closed-source VLMs, GPT-4o and Claude-3.5-Sonnet, are evaluated on the same dataset. All models are assessed using precision, recall, F1-score, and hallucination metrics. Due to the computational cost and impracticality of fine-tuning large closed-source VLMs for domain-specific tasks, GPT-4o and Claude-3.5-Sonnet are evaluated in a zero-shot setting. In contrast, Florence-2, a smaller model with 0.23 billion parameters, is optimized through full-parameter fine-tuning across three distinct experiments, each utilizing datasets augmented to different levels. The results show that Florence-2 achieves a 29.95% increase in precision, a 37.75% increase in recall, a 52.40% improvement in F1-score, and a 43.15% reduction in hallucination rate compared to the best-performing closed-source model. These findings highlight the effectiveness of fine-tuning smaller, open-source VLMs like Florence-2, offering a practical and efficient solution for automated GD&T extraction to support downstream manufacturing tasks.
Forward citations
Cited by 4 Pith papers
-
Benchmarking Deep Learning Approaches for AEC Engineering Drawing Layout Detection and Information Extraction
On a new 551-drawing facade dataset, RF-DETR achieved the best layout-detection accuracy (mAP50 0.949), Qwen3-VL the best F1 (0.911), and document-specific pre-training degraded DocLayout-YOLO relative to a COCO baseline.
-
VFEAgent: A Multimodal Agent Framework for End-to-End Automated Finite Element Analysis
VFEAgent is an end-to-end multi-agent framework that automates FEA modeling and simulation from multimodal inputs, achieving high success rates in generating physically valid simulations across engineering scenarios.
-
Linking heterogeneous microstructure informatics with expert characterization knowledge through customized and hybrid vision-language representations for industrial qualification
A hybrid CLIP/FLAVA similarity score with positive/negative text and image references classifies Ni-WC metal matrix composite micrographs against six expert criteria, but the in-sample evaluation does not demonstrate ...
-
Foundation Model Driven Robotics: A Comprehensive Review
A review of foundation-model-driven robotics that synthesizes recent work across perception, planning, control, HRI, simulation, and sim-to-real transfer, and highlights open challenges.
Discussion (0). Sign in to comment.