Pith. sign in

REVIEW 3 cited by

Lenna: Language Enhanced Reasoning Detection Assistant

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.02433 v1 pith:37HXZ36E submitted 2023-12-05 cs.CV

classification cs.CV
keywords lennadetectionreasoninglanguageassistantlargemllmsmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the fast-paced development of multimodal large language models (MLLMs), we can now converse with AI systems in natural languages to understand images. However, the reasoning power and world knowledge embedded in the large language models have been much less investigated and exploited for image perception tasks. In this paper, we propose Lenna, a language-enhanced reasoning detection assistant, which utilizes the robust multimodal feature representation of MLLMs, while preserving location information for detection. This is achieved by incorporating an additional <DET> token in the MLLM vocabulary that is free of explicit semantic context but serves as a prompt for the detector to identify the corresponding position. To evaluate the reasoning capability of Lenna, we construct a ReasonDet dataset to measure its performance on reasoning-based detection. Remarkably, Lenna demonstrates outstanding performance on ReasonDet and comes with significantly low training costs. It also incurs minimal transferring overhead when extended to other tasks. Our code and model will be available at https://git.io/Lenna.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A recalibrated GRPO reinforcement learning method lets multimodal LLMs say 'None' for nonexistent referring expressions without sacrificing localization accuracy on objects that do exist.

  2. FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing

    cs.CV 2026-07 conditional novelty 5.0 of 10

    FAS-R1 combines long-CoT supervised fine-tuning with difficulty-aware GRPO and degradation-simulated augmentation to improve multi-task face anti-spoofing and explainable rationales.

  3. Language-guided Learning for Object Detection Tackling Multiple Variations in Aerial Images

    cs.CV 2025-05 conditional novelty 5.0 of 10

    LANGO adds an LLM-based visual semantic reasoner and a relation learning loss that aligns visual features with language representations, improving aerial detection AP on UAVDT and VisDrone.

Pith tools