Pith. sign in

REVIEW 12 cited by

Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.14405 v2 pith:A6NO5W4H submitted 2024-11-21 cs.CL

classification cs.CL
keywords marco-o1reasoningmodelsopen-endedabsentaddressanswersbroader
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Currently OpenAI o1 sparks a surge of interest in the study of large reasoning models (LRM). Building on this momentum, Marco-o1 not only focuses on disciplines with standard answers, such as mathematics, physics, and coding -- which are well-suited for reinforcement learning (RL) -- but also places greater emphasis on open-ended resolutions. We aim to address the question: ''Can the o1 model effectively generalize to broader domains where clear standards are absent and rewards are challenging to quantify?'' Marco-o1 is powered by Chain-of-Thought (CoT) fine-tuning, Monte Carlo Tree Search (MCTS), reflection mechanisms, and innovative reasoning strategies -- optimized for complex real-world problem-solving tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What Were You Thinking? An LLM-Driven Large-Scale Study of Refactoring Motivations in Open-Source Projects

    cs.SE 2025-09 conditional novelty 6.0 of 10

    LLM-generated refactoring motivations agree with expert raters about 80% of the time, align with literature motivations in roughly half of cases, and correlate only weakly with software metrics.

  2. mSCoRe: a $M$ultilingual and Scalable Benchmark for $S$kill-based $Co$mmonsense $Re$asoning

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Introduces mSCoRe, a multilingual skill-annotated commonsense benchmark built from human-annotated seeds and LLM-based complexity scaling, and shows all eight tested LLMs degrade as complexity increases.

  3. InteChar: A Unified Oracle Bone Character List for Ancient Chinese Language Modeling

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    InteChar is a proposed standard character list for digitizing oracle bone script, paired with a corpus that reportedly improves ancient Chinese language modeling.

  4. Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new benchmark permutes CommonGen concept sets and asks LLMs to compose sentences in the specified order; even the best model achieves only about 75% ordered coverage.

  5. Com$^2$: A Causal-Guided Benchmark for Exploring Complex Commonsense Reasoning in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Com2 is a causal-graph-guided benchmark with 3,754 questions showing that LLMs struggle with complex commonsense reasoning, particularly on intervention and transition tasks.

  6. Chain of Methodologies: Scaling Test Time Computation without Training

    cs.CL 2025-06 reject novelty 6.0 of 10

    A prompt framework that interleaves methodology selection with reasoning steps improves LLM math and QA accuracy when combined with a Python interpreter, but the benefit of the methodology selection itself is small an...

  7. Table-r1: Self-supervised and Reinforcement Learning for Program-based Table Reasoning in Small Language Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Table-r1 combines a layout-transformation self-supervised task and a mix-paradigm GRPO stage so 7B/8B models outperform other small-model table reasoners and approach GPT-4o-level accuracy.

  8. Unlocking Recursive Thinking of LLMs: Alignment via Refinement

    cs.CL 2025-06 conditional novelty 6.0 of 10

    An offline alignment pipeline using reward-filtered self-refinement data and long chain-of-thought SFT raises an 8B model's AlpacaEval 2 win rate from 25.0% to 51.0% with roughly 14k training examples.

  9. MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos

    cs.CV 2025-06 conditional novelty 6.0 of 10

    MMR-V is a new video QA benchmark requiring long-range, multi-frame reasoning, on which the best AI model scores 52.5% versus 86% for humans.

  10. REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Asking a reasoning model several problems at once reveals large accuracy drops and exposes differences that single-question benchmarks miss.

  11. Psychology-Driven Enhancement of Humour Translation

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A decomposition-and-recomposition prompt method for humor translation reports large gains on LLM-based metrics, but the evaluation lacks human validation and statistical checks.

  12. Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models

    cs.CL 2025-07 reject novelty 4.0 of 10

    A step-level verifier-guided hybrid of Best-of-N sampling, Monte Carlo tree search, and conditional self-refinement improves reasoning in small instruction-tuned LLMs, claiming up to 28.6-point gains.

Pith tools