Pith. sign in

Omni-think: Scaling cross-domain generalization in llms via multi-task rl with hybrid rewards

7 Pith papers cite this work. Polarity classification is still indexing.

7 Pith papers citing it

years

2026 6 2025 1

representative citing papers

Prompt-Level Reward Specifications for Open-Ended Post-Training

cs.CL · 2026-05-28 · unverdicted · novelty 6.0

A prompt-level reward specification framework constructs reusable rubrics and executable checkers from prompts alone to deliver hybrid rewards combining requirement satisfaction, holistic quality, and deterministic constraints for LLM post-training.

Trust Region On-Policy Distillation

cs.LG · 2026-05-31 · unverdicted · novelty 5.0

TrOPD stabilizes on-policy distillation for LLMs with trust-region learning, outlier estimation, and off-policy guidance, outperforming prior OPD methods on reasoning and code benchmarks.

citing papers explorer

Showing 7 of 7 citing papers.