REVIEW 17 cited by
Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Currently OpenAI o1 sparks a surge of interest in the study of large reasoning models (LRM). Building on this momentum, Marco-o1 not only focuses on disciplines with standard answers, such as mathematics, physics, and coding -- which are well-suited for reinforcement learning (RL) -- but also places greater emphasis on open-ended resolutions. We aim to address the question: ''Can the o1 model effectively generalize to broader domains where clear standards are absent and rewards are challenging to quantify?'' Marco-o1 is powered by Chain-of-Thought (CoT) fine-tuning, Monte Carlo Tree Search (MCTS), reflection mechanisms, and innovative reasoning strategies -- optimized for complex real-world problem-solving tasks.
Forward citations
Cited by 17 Pith papers
-
Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning
A checkpoint-based search and candidate augmentation method improves small LLM mathematical reasoning accuracy over existing test-time scaling baselines.
-
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
Reasoning models trade visual grounding for language-based inference, and this paper measures that trade-off with a new metric and benchmark.
-
What Were You Thinking? An LLM-Driven Large-Scale Study of Refactoring Motivations in Open-Source Projects
LLM-generated refactoring motivations agree with expert raters about 80% of the time, align with literature motivations in roughly half of cases, and correlate only weakly with software metrics.
-
mSCoRe: a $M$ultilingual and Scalable Benchmark for $S$kill-based $Co$mmonsense $Re$asoning
Introduces mSCoRe, a multilingual skill-annotated commonsense benchmark built from human-annotated seeds and LLM-based complexity scaling, and shows all eight tested LLMs degrade as complexity increases.
-
InteChar: A Unified Oracle Bone Character List for Ancient Chinese Language Modeling
InteChar is a proposed standard character list for digitizing oracle bone script, paired with a corpus that reportedly improves ancient Chinese language modeling.
-
Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability
A new benchmark permutes CommonGen concept sets and asks LLMs to compose sentences in the specified order; even the best model achieves only about 75% ordered coverage.
-
Com$^2$: A Causal-Guided Benchmark for Exploring Complex Commonsense Reasoning in Large Language Models
Com2 is a causal-graph-guided benchmark with 3,754 questions showing that LLMs struggle with complex commonsense reasoning, particularly on intervention and transition tasks.
-
Chain of Methodologies: Scaling Test Time Computation without Training
A prompt framework that interleaves methodology selection with reasoning steps improves LLM math and QA accuracy when combined with a Python interpreter, but the benefit of the methodology selection itself is small an...
-
Table-r1: Self-supervised and Reinforcement Learning for Program-based Table Reasoning in Small Language Models
Table-r1 combines a layout-transformation self-supervised task and a mix-paradigm GRPO stage so 7B/8B models outperform other small-model table reasoners and approach GPT-4o-level accuracy.
-
Unlocking Recursive Thinking of LLMs: Alignment via Refinement
An offline alignment pipeline using reward-filtered self-refinement data and long chain-of-thought SFT raises an 8B model's AlpacaEval 2 win rate from 25.0% to 51.0% with roughly 14k training examples.
-
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
MMR-V is a new video QA benchmark requiring long-range, multi-frame reasoning, on which the best AI model scores 52.5% versus 86% for humans.
-
TAT-R1: Terminology-Aware Translation with Reinforcement Learning and Word Alignment
Word-alignment rewards for RL-trained translation raise terminology accuracy on RTT from 54.42 to 56.42 TA without hurting general translation quality.
-
GraphGen: Enhancing Supervised Fine-Tuning for LLMs with Knowledge-Driven Synthetic Data Generation
A knowledge-graph-guided method that scores an LLM's knowledge gaps and generates atomic, aggregated, and multi-hop QA pairs, improving closed-book QA after fine-tuning.
-
REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once
Asking a reasoning model several problems at once reveals large accuracy drops and exposes differences that single-question benchmarks miss.
-
Psychology-Driven Enhancement of Humour Translation
A decomposition-and-recomposition prompt method for humor translation reports large gains on LLM-based metrics, but the evaluation lacks human validation and statistical checks.
-
Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes?
A new 1,430-item multimodal benchmark shows that leading multimodal LLMs rarely notice small visual traps needed for commonsense safety reasoning, with top scores near 0.46 on a scale whose maximum is about 0.97.
-
Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models
A step-level verifier-guided hybrid of Best-of-N sampling, Monte Carlo tree search, and conditional self-refinement improves reasoning in small instruction-tuned LLMs, claiming up to 28.6-point gains.
Discussion (0). Sign in to comment.