Pith. sign in

REVIEW 17 cited by

Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.14405 v2 pith:A6NO5W4H submitted 2024-11-21 cs.CL

classification cs.CL
keywords marco-o1reasoningmodelsopen-endedabsentaddressanswersbroader
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Currently OpenAI o1 sparks a surge of interest in the study of large reasoning models (LRM). Building on this momentum, Marco-o1 not only focuses on disciplines with standard answers, such as mathematics, physics, and coding -- which are well-suited for reinforcement learning (RL) -- but also places greater emphasis on open-ended resolutions. We aim to address the question: ''Can the o1 model effectively generalize to broader domains where clear standards are absent and rewards are challenging to quantify?'' Marco-o1 is powered by Chain-of-Thought (CoT) fine-tuning, Monte Carlo Tree Search (MCTS), reflection mechanisms, and innovative reasoning strategies -- optimized for complex real-world problem-solving tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning

    cs.CL 2025-05 conditional novelty 7.0 of 10

    A checkpoint-based search and candidate augmentation method improves small LLM mathematical reasoning accuracy over existing test-time scaling baselines.

  2. More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models

    cs.CL 2025-05 conditional novelty 7.0 of 10

    Reasoning models trade visual grounding for language-based inference, and this paper measures that trade-off with a new metric and benchmark.

  3. What Were You Thinking? An LLM-Driven Large-Scale Study of Refactoring Motivations in Open-Source Projects

    cs.SE 2025-09 conditional novelty 6.0 of 10

    LLM-generated refactoring motivations agree with expert raters about 80% of the time, align with literature motivations in roughly half of cases, and correlate only weakly with software metrics.

  4. mSCoRe: a $M$ultilingual and Scalable Benchmark for $S$kill-based $Co$mmonsense $Re$asoning

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Introduces mSCoRe, a multilingual skill-annotated commonsense benchmark built from human-annotated seeds and LLM-based complexity scaling, and shows all eight tested LLMs degrade as complexity increases.

  5. InteChar: A Unified Oracle Bone Character List for Ancient Chinese Language Modeling

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    InteChar is a proposed standard character list for digitizing oracle bone script, paired with a corpus that reportedly improves ancient Chinese language modeling.

  6. Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new benchmark permutes CommonGen concept sets and asks LLMs to compose sentences in the specified order; even the best model achieves only about 75% ordered coverage.

  7. Com$^2$: A Causal-Guided Benchmark for Exploring Complex Commonsense Reasoning in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Com2 is a causal-graph-guided benchmark with 3,754 questions showing that LLMs struggle with complex commonsense reasoning, particularly on intervention and transition tasks.

  8. Chain of Methodologies: Scaling Test Time Computation without Training

    cs.CL 2025-06 reject novelty 6.0 of 10

    A prompt framework that interleaves methodology selection with reasoning steps improves LLM math and QA accuracy when combined with a Python interpreter, but the benefit of the methodology selection itself is small an...

  9. Table-r1: Self-supervised and Reinforcement Learning for Program-based Table Reasoning in Small Language Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Table-r1 combines a layout-transformation self-supervised task and a mix-paradigm GRPO stage so 7B/8B models outperform other small-model table reasoners and approach GPT-4o-level accuracy.

  10. Unlocking Recursive Thinking of LLMs: Alignment via Refinement

    cs.CL 2025-06 conditional novelty 6.0 of 10

    An offline alignment pipeline using reward-filtered self-refinement data and long chain-of-thought SFT raises an 8B model's AlpacaEval 2 win rate from 25.0% to 51.0% with roughly 14k training examples.

  11. MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos

    cs.CV 2025-06 conditional novelty 6.0 of 10

    MMR-V is a new video QA benchmark requiring long-range, multi-frame reasoning, on which the best AI model scores 52.5% versus 86% for humans.

  12. TAT-R1: Terminology-Aware Translation with Reinforcement Learning and Word Alignment

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Word-alignment rewards for RL-trained translation raise terminology accuracy on RTT from 54.42 to 56.42 TA without hurting general translation quality.

  13. GraphGen: Enhancing Supervised Fine-Tuning for LLMs with Knowledge-Driven Synthetic Data Generation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A knowledge-graph-guided method that scores an LLM's knowledge gaps and generates atomic, aggregated, and multi-hop QA pairs, improving closed-book QA after fine-tuning.

  14. REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Asking a reasoning model several problems at once reveals large accuracy drops and exposes differences that single-question benchmarks miss.

  15. Psychology-Driven Enhancement of Humour Translation

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A decomposition-and-recomposition prompt method for humor translation reports large gains on LLM-based metrics, but the evaluation lacks human validation and statistical checks.

  16. Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes?

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A new 1,430-item multimodal benchmark shows that leading multimodal LLMs rarely notice small visual traps needed for commonsense safety reasoning, with top scores near 0.46 on a scale whose maximum is about 0.97.

  17. Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models

    cs.CL 2025-07 reject novelty 4.0 of 10

    A step-level verifier-guided hybrid of Best-of-N sampling, Monte Carlo tree search, and conditional self-refinement improves reasoning in small instruction-tuned LLMs, claiming up to 28.6-point gains.

Pith tools