REVIEW 17 cited by
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recently, promising progress has been made by open-source vision-language models (VLMs) in bringing their capabilities closer to those of proprietary frontier models. However, most open-source models only publish their final model weights, leaving the critical details of data strategies and implementation largely opaque. In this work, we address VLM post-training from a data-centric perspective, showing the key role of data strategy in developing frontier VLMs. By studying and building our post-training data strategy from scratch, we share detailed insights into the development processes, aiming to benefit the development of competitive models for the open-source community. Our introduced data strategy, together with training recipes and model design, leads to a family of performant VLMs named Eagle2. Specifically, Eagle2-9B achieves state-of-the-art results across various multimodal benchmarks, matching certain competitive models with up to 70B parameters.
Forward citations
Cited by 17 Pith papers
-
Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
Foundation-model HOI work is organized into eight geometric, semantic, and visual sub-priors that enter six reconstruction/generation tasks and three robot-transfer routes.
-
RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation
Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.
-
Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack
A full learning stack—10K-hour UMI data, a flow-matching VLA, preference-optimization RL, and asynchronous deployment—reports SOTA RoboTwin results and cross-embodiment transfer to four real robots.
-
RhinoVLA Technical Report
RhinoVLA uses a token-efficient Qwen3-VL backbone, continuous Action Expert, and unified cross-robot interface to match π0.5 performance while hitting 11.69 Hz on Huixi R1 edge SoC.
-
Cambrian-P: Pose-Grounded Video Understanding
Cambrian-P adds per-frame camera pose tokens and a regression head to video MLLMs, delivering 4.5-6.5% gains on spatial benchmarks, generalization to other video QA tasks, and SOTA streaming pose estimation on ScanNet.
-
$M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills
M²-VLA shows that generalized VLMs can serve as direct backbones for robotic manipulation by selectively extracting task-critical features via Mixture of Layers and adding Meta Skill Modules for efficient trajectory learning.
-
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
RefBench-PRO organizes REC into attribute, position, interaction, relation, commonsense, and reject tasks; no tested MLLM exceeds 72%, and Ref-R1 raises Qwen2.5-VL-7B from 57.6 to 69.4 on it.
-
FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
A two-stage bilingual CLIP-style model with region-text supervision and a new text-side contrastive loss outperforms prior open models on fine-grained vision-language tasks in English and Chinese.
-
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
A clue-grounded audio-visual counting benchmark over 497 long videos and an RL-trained counting model, whose headline result is undermined by training on the DVD-Counting evaluation benchmark.
-
ReXVQA: A Large-scale Visual Question Answering Benchmark for Generalist Chest X-ray Understanding
ReXVQA is a large-scale chest X-ray VQA benchmark generated by an LLM pipeline, with a small reader study claiming MedGemma surpasses radiology residents.
-
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
Training a multimodal LLM on GPT-4o-generated chain-of-thought referring traces, then optimizing with GRPO, improves referring accuracy and abstention on HumanRef.
-
MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation
MEgoHand generates egocentric hand-object interaction motions from an RGB image, a text instruction, and an initial MANO hand pose using VLM-based semantics, monocular depth, and flow matching.
-
SigLIP-HD by Fine-to-Coarse Supervision
Fine-to-coarse L1 supervision lets a standard-resolution SigLIP 2 encoder produce better visual tokens for MLLMs without higher-resolution inference.
-
LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments
LeapBot-WA shows robot policies can be trained with latent world-model predictions instead of pixel video generation, hitting state-of-the-art for predictive action models and staying competitive with generative WAMs.
-
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models
A new resolution-focused benchmark and an open-source native-resolution training framework show that preserving original image resolution improves VLM performance on fine-grained visual tasks.
-
Llama Nemoretriever Colembed: Top-Performing Text-Image Retrieval Model
A new family of text-image retrieval models, built from Eagle2 with bidirectional attention and ColBERT-style late interaction, reports state-of-the-art NDCG@5 scores on ViDoRe V1 (91.0) and V2 (63.5).
-
Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision
A comprehensive review of cross-view video understanding that uses both first-person and third-person cameras, organized into a three-direction taxonomy with a dataset catalog and future research gaps.
Discussion (0). Continue with ORCID to comment.