REVIEW 19 cited by
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper introduces Grounding DINO 1.5, a suite of advanced open-set object detection models developed by IDEA Research, which aims to advance the "Edge" of open-set object detection. The suite encompasses two models: Grounding DINO 1.5 Pro, a high-performance model designed for stronger generalization capability across a wide range of scenarios, and Grounding DINO 1.5 Edge, an efficient model optimized for faster speed demanded in many applications requiring edge deployment. The Grounding DINO 1.5 Pro model advances its predecessor by scaling up the model architecture, integrating an enhanced vision backbone, and expanding the training dataset to over 20 million images with grounding annotations, thereby achieving a richer semantic understanding. The Grounding DINO 1.5 Edge model, while designed for efficiency with reduced feature scales, maintains robust detection capabilities by being trained on the same comprehensive dataset. Empirical results demonstrate the effectiveness of Grounding DINO 1.5, with the Grounding DINO 1.5 Pro model attaining a 54.3 AP on the COCO detection benchmark and a 55.7 AP on the LVIS-minival zero-shot transfer benchmark, setting new records for open-set object detection. Furthermore, the Grounding DINO 1.5 Edge model, when optimized with TensorRT, achieves a speed of 75.2 FPS while attaining a zero-shot performance of 36.2 AP on the LVIS-minival benchmark, making it more suitable for edge computing scenarios. Model examples and demos with API will be released at https://github.com/IDEA-Research/Grounding-DINO-1.5-API
Forward citations
Cited by 19 Pith papers
-
Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection
MoE fine-tuning with decomposed pre-trained FFN experts lets a real-time open-vocabulary detector beat a much larger-data baseline with similar active parameter count.
-
DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models
DINO-R1 trains visual-prompt detectors with group-relative query rewards and KL regularization, improving zero-shot and fine-tuned detection over supervised fine-tuning.
-
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
RynnBrain 1.1 reports state-of-the-art embodied cognition and localization scores with a 122B-A10B model and improved real-robot VLA policies via joint multi-embodiment training.
-
SceneBind: Binding What and Where Across Vision, Audio and Language
A scene is represented as a global semantic embedding plus object-centric semantic-spatial slots (azimuth, elevation, distance, confidence), which improves cross-modal retrieval and enables zero-shot audio-visual loca...
-
HERO: Learning Humanoid End-Effector Control for Visual Whole-Body Open-Vocabulary Object Grasping
HERO achieves 2.44 cm end-effector tracking error on a Unitree G1 humanoid and uses it, with open-vocabulary perception, to grasp novel objects at up to 90% success in diverse real scenes.
-
LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning
A VLM-based planner with task-oriented segmentation reranking and a clause-level condition retriever reports state-of-the-art scores on ActPlan-1K and ALFRED, though the reported ablation numbers are internally inconsistent.
-
EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow
EC-Flow predicts pixel trajectories on the robot body (embodiment-centric flow) from action-unlabeled videos and uses the robot's URDF model to convert those trajectories into executable actions, improving performance...
-
SAM4D: Segment Anything in Camera and LiDAR Streams
SAM4D is a promptable model that segments and tracks objects across camera and LiDAR streams with cross-modal prompts, trained on pseudo-labels generated by an automated data engine.
-
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
Training a multimodal LLM on GPT-4o-generated chain-of-thought referring traces, then optimizing with GRPO, improves referring accuracy and abstention on HumanRef.
-
Single Domain Generalization for Few-Shot Counting via Universal Representation Matching
URM distills CLIP vision-language representations into learnable prototypes for few-shot counting, improving single-domain generalization on unseen datasets.
-
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
A 0.2B-parameter V+L→A policy matches larger LLM-centric VLAs on LIBERO at 31 ms latency and 0.9 GB VRAM by fusing DINOv3 and BERT features with bidirectional cross-attention and ACT-style action chunks.
-
Embodied Operators and Benchmarking: Toward Reusable and Deployable Embodied Intelligence Systems
Embodied operators—deployable modules with task semantics and I/O contracts—should be the unit of optimization and multi-dimensional benchmarking for reusable robot intelligence systems.
-
Quantum orientation entanglement analysis of the interpolating helicity states between the instant form dynamics and the light-front dynamics
Interpolating helicity states expanded in Jacob–Wick helicity via Wigner d-matrix probabilities reveal a critical angle that bifurcates instant-form and light-front spin dynamics in contact-interaction pair production.
-
DatasetAgent: A Novel Multi-Agent System for Auto-Constructing Datasets from Real-World Images
DatasetAgent is an LLM-powered multi-agent pipeline that automatically constructs image classification, detection, and segmentation datasets from web images, with modest downstream gains shown but weak experimental controls.
-
3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model
A diffusion world model predicts 3D optical flow as an embodiment-agnostic action plan, and constrained optimization converts the flow into robot arm actions.
-
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
Argus adds explicit language-guided visual attention to multimodal LLMs by grounding questions to bounding boxes and re-engaging those regions, improving vision-centric reasoning and grounding accuracy.
-
One-Shot Crowd Counting With Density Guidance For Scene Adaptation
A one-shot crowd-counting method uses EM-clustered local and global density features from one labeled support image to adapt a fixed model to an unseen surveillance scene.
-
Model-Agnostic Open-Set Air-to-Air Visual Object Detection for Reliable UAV Perception
Fusing softmax confidence with embedding-space Gaussian mixture entropy improves open-set rejection in air-to-air UAV detection, reporting about 0.88 AUROC on real flight data.
-
Robust and Efficient 3D Gaussian Splatting for Urban Scene Reconstruction
A 3D Gaussian Splatting framework for urban scenes that combines visibility-based data partitioning, budgeted level-of-detail generation, and per-Gaussian appearance embeddings to enable efficient training and real-ti...
Discussion (0). Sign in to comment.