Pith. sign in

REVIEW 5 cited by

TaskCLIP: Extend Large Vision-Language Model for Task Oriented Object Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.08108 v2 pith:BRBRTYTK submitted 2024-03-12 cs.CV

classification cs.CV
keywords objectdetectionmodelstasktaskclipvlmsdesignembeddings
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Task-oriented object detection aims to find objects suitable for accomplishing specific tasks. As a challenging task, it requires simultaneous visual data processing and reasoning under ambiguous semantics. Recent solutions are mainly all-in-one models. However, the object detection backbones are pre-trained without text supervision. Thus, to incorporate task requirements, their intricate models undergo extensive learning on a highly imbalanced and scarce dataset, resulting in capped performance, laborious training, and poor generalizability. In contrast, we propose TaskCLIP, a more natural two-stage design composed of general object detection and task-guided object selection. Particularly for the latter, we resort to the recently successful large Vision-Language Models (VLMs) as our backbone, which provides rich semantic knowledge and a uniform embedding space for images and texts. Nevertheless, the naive application of VLMs leads to sub-optimal quality, due to the misalignment between embeddings of object images and their visual attributes, which are mainly adjective phrases. To this end, we design a transformer-based aligner after the pre-trained VLMs to re-calibrate both embeddings. Finally, we employ a trainable score function to post-process the VLM matching results for object selection. Experimental results demonstrate that our TaskCLIP outperforms the state-of-the-art DETR-based model TOIST by 3.5% and only requires a single NVIDIA RTX 4090 for both training and inference.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TRINE: A Token-Aware, Runtime-Adaptive FPGA Inference Engine for Multimodal AI

    cs.AR 2026-03 conditional novelty 6.5 of 10

    TRINE unifies ViT/CNN/GNN/NLP as DDMM/SDDMM/SpMM on a runtime mode-switchable FPGA PE array with in-stream top-k pruning and DALO scheduling, cutting latency up to 22.57× vs RTX 4090 at ~21 W.

  2. Leverage Task Context for Object Affordance Ranking

    cs.CV 2024-11 conditional novelty 6.0 of 10

    The authors define task-context-conditioned object affordance ranking, release a 50,000-image benchmark, and report that their Context-embed Group Ranking model beats six saliency-ranking and multimodal-detection base...

  3. Continuous GNN-based Anomaly Detection on Edge using Efficient Adaptive Knowledge Graph Learning

    cs.LG 2024-11 conditional novelty 5.0 of 10

    A GNN-based video anomaly detector adapts its knowledge graph on-device through token-embedding updates, pruning, and node creation, avoiding cloud-based graph regeneration as anomaly types change.

  4. J-DDL: Surface Damage Detection and Localization System for Fighter Aircraft

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A rail-based 2D/3D aircraft inspection system with an improved YOLOv8 detector (AIR-YOLO) detects damage in images and localizes it in point clouds, plus a new 8,091-image damage dataset (AIRSD).

  5. Tell Me What to Track: Infusing Robust Language Guidance for Enhanced Referring Multi-Object Tracking

    cs.CV 2024-12 conditional novelty 4.0 of 10

    TellTrack improves referring multi-object tracking by adding collaborative query matching, direct query-level language infusion, and a reordered cross-modal encoder, achieving SOTA HOTA on Refer-KITTI and Refer-KITTI-V2.

Pith tools