Pith. sign in

REVIEW 17 cited by

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.14818 v1 pith:F5KIEAQK submitted 2025-01-20 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords modelsdatafrontieropen-sourcepost-trainingstrategyvlmsbuilding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, promising progress has been made by open-source vision-language models (VLMs) in bringing their capabilities closer to those of proprietary frontier models. However, most open-source models only publish their final model weights, leaving the critical details of data strategies and implementation largely opaque. In this work, we address VLM post-training from a data-centric perspective, showing the key role of data strategy in developing frontier VLMs. By studying and building our post-training data strategy from scratch, we share detailed insights into the development processes, aiming to benefit the development of competitive models for the open-source community. Our introduced data strategy, together with training recipes and model design, leads to a family of performant VLMs named Eagle2. Specifically, Eagle2-9B achieves state-of-the-art results across various multimodal benchmarks, matching certain competitive models with up to 70B parameters.

Discussion (0). Sign in to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Foundation-model HOI work is organized into eight geometric, semantic, and visual sub-priors that enter six reconstruction/generation tasks and three robot-transfer routes.

  2. RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.

  3. Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack

    cs.RO 2026-06 conditional novelty 6.0 of 10

    A full learning stack—10K-hour UMI data, a flow-matching VLA, preference-optimization RL, and asynchronous deployment—reports SOTA RoboTwin results and cross-embodiment transfer to four real robots.

  4. RhinoVLA Technical Report

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    RhinoVLA uses a token-efficient Qwen3-VL backbone, continuous Action Expert, and unified cross-robot interface to match π0.5 performance while hitting 11.69 Hz on Huixi R1 edge SoC.

  5. Cambrian-P: Pose-Grounded Video Understanding

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Cambrian-P adds per-frame camera pose tokens and a regression head to video MLLMs, delivering 4.5-6.5% gains on spatial benchmarks, generalization to other video QA tasks, and SOTA streaming pose estimation on ScanNet.

  6. $M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills

    cs.RO 2026-04 unverdicted novelty 6.0 of 10

    M²-VLA shows that generalized VLMs can serve as direct backbones for robotic manipulation by selectively extracting task-critical features via Mixture of Layers and adding Meta Skill Modules for efficient trajectory learning.

  7. RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension

    cs.CV 2025-12 conditional novelty 6.0 of 10

    RefBench-PRO organizes REC into attribute, position, interaction, relation, commonsense, and reject tasks; no tested MLLM exceeds 72%, and Ref-R1 raises Qwen2.5-VL-7B from 57.6 to 69.4 on it.

  8. FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A two-stage bilingual CLIP-style model with region-text supervision and a new text-side contrastive loss outperforms prior open models on fine-grained vision-language tasks in English and Chinese.

  9. AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

    cs.CV 2025-06 reject novelty 6.0 of 10

    A clue-grounded audio-visual counting benchmark over 497 long videos and an RL-trained counting model, whose headline result is undermined by training on the DVD-Counting evaluation benchmark.

  10. ReXVQA: A Large-scale Visual Question Answering Benchmark for Generalist Chest X-ray Understanding

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ReXVQA is a large-scale chest X-ray VQA benchmark generated by an LLM pipeline, with a small reader study claiming MedGemma surpasses radiology residents.

  11. Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Training a multimodal LLM on GPT-4o-generated chain-of-thought referring traces, then optimizing with GRPO, improves referring accuracy and abstention on HumanRef.

  12. MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MEgoHand generates egocentric hand-object interaction motions from an RGB image, a text instruction, and an initial MANO hand pose using VLM-based semantics, monocular depth, and flow matching.

  13. SigLIP-HD by Fine-to-Coarse Supervision

    cs.CV 2026-07 conditional novelty 5.5 of 10

    Fine-to-coarse L1 supervision lets a standard-resolution SigLIP 2 encoder produce better visual tokens for MLLMs without higher-resolution inference.

  14. LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments

    cs.RO 2026-07 conditional novelty 5.0 of 10

    LeapBot-WA shows robot policies can be trained with latent world-model predictions instead of pixel video generation, hitting state-of-the-art for predictive action models and staying competitive with generative WAMs.

  15. Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A new resolution-focused benchmark and an open-source native-resolution training framework show that preserving original image resolution improves VLM performance on fine-grained visual tasks.

  16. Llama Nemoretriever Colembed: Top-Performing Text-Image Retrieval Model

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A new family of text-image retrieval models, built from Eagle2 with bidirectional attention and ColBERT-style late interaction, reports state-of-the-art NDCG@5 scores on ViDoRe V1 (91.0) and V2 (63.5).

  17. Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision

    cs.CV 2025-06 accept novelty 3.0 of 10

    A comprehensive review of cross-view video understanding that uses both first-person and third-person cameras, organized into a three-direction taxonomy with a dataset catalog and future research gaps.

Pith tools