Pith. sign in

REVIEW 1 cited by

ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.06292 v1 pith:XBMYKIXX submitted 2024-12-09 cs.CV

classification cs.CV
keywords modelskeypointdetectionlanguagemllmsannotationsapproachdemonstrate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a novel zero-shot approach for keypoint detection on 3D shapes. Point-level reasoning on visual data is challenging as it requires precise localization capability, posing problems even for powerful models like DINO or CLIP. Traditional methods for 3D keypoint detection rely heavily on annotated 3D datasets and extensive supervised training, limiting their scalability and applicability to new categories or domains. In contrast, our method utilizes the rich knowledge embedded within Multi-Modal Large Language Models (MLLMs). Specifically, we demonstrate, for the first time, that pixel-level annotations used to train recent MLLMs can be exploited for both extracting and naming salient keypoints on 3D models without any ground truth labels or supervision. Experimental evaluations demonstrate that our approach achieves competitive performance on standard benchmarks compared to supervised methods, despite not requiring any 3D keypoint annotations during training. Our results highlight the potential of integrating language models for localized 3D shape understanding. This work opens new avenues for cross-modal learning and underscores the effectiveness of MLLMs in contributing to 3D computer vision challenges.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Estimating 2D Keypoints of Surgical Tools Using Vision-Language Models with Low-Rank Adaptation

    cs.CV 2025-08 conditional novelty 4.0 of 10

    Giving surgical-tool keypoint detection to Qwen2.5-VL via LoRA fine-tuning reaches MPJPE 0.0627 on SurgeoNet, comparable with or better than dedicated YOLOv8-Pose and SurgeoNet baselines.

Pith tools