Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:19:33.840820Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 5 inbound Pith citation observations for arXiv:2505.24871.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:19:33.840820Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T08:15:52.772665Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T22:24:00.434533Z
71 of 71 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2dc90553-8ba9-43d5-a898-a69e774daed7 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Alayrac, J
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3e214dd-d14f-4ad1-af5b-4b898e03b140 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Efficient Online Data Mixing For Language Model Pre-Training
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00e694f5-71e2-49b2-aafa-6c45738c459f · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Constitutional AI: Harmlessness from AI Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b1e10be-3f0d-4fff-bd10-b63f1a4eb349 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0693db7d-7433-4d89-9fec-4b76efbc542e · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9fab8f8f-2d3a-4d91-b582-ce7504e130dc · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 09960a24-cca9-4205-b572-34c56db39c0a · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4f4c7406-6758-4888-825c-15e77db8b3cf · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2344a5a9-f500-4012-b701-1bcb8274e308 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning UniMax: Fairer and more Effective Language Sampling for Large-Scale Multilingual Pretraining
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b9990b6-8dc9-4bee-a015-bc7eb4ced88f · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.ArXiv Preprint, 2025
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 79346d66-c34d-4823-8d39-2c09636ce64d · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6ea1a21b-84b2-436f-ba79-57296eaa9c93 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Devlin, M.-W
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0789acc3-b5e4-4148-acb8-715928a1d3fb · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e135594d-1c46-4d77-8a61-9b2fe77ab4d0 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Teaching Large Language Models to Reason with Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37c2ea2a-7d24-4ba1-9296-5a561771ff43 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Optimizing Pretraining Data Mixtures with LLM-Estimated Utility
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca94893b-7245-4457-820f-027ef55c824e · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4778dbe-e24c-4b26-8d07-f77022815ca7 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Datamodels: Predicting Predictions from Training Data
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fec577c-30d3-4dca-bc9d-0105e7d12de0 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Kazemnejad, M
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cc549ef2-8849-4736-990d-a8cc6a122214 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 454816c9-254c-4c12-9dff-b963bfc79b26 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1cccf3fd-6ce4-423e-8ea0-56de82d54d1e · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c5086ab-e6f3-4a30-be7a-b50e5ced75d9 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0509d813-4b14-4f76-9936-5fb14ced841a · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c6b2d7ee-8903-4e5f-800a-4a7f606b1c26 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3306c2d2-a036-43da-9e03-02a335611353 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63641a10-1a9d-4b77-b9a8-2555350f9de6 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ebbcf531-7dbb-4c0c-ac47-90e97aff91e5 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1d85ef55-57af-4a92-a40e-93f2e3046e45 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fb0ce7a-3ae3-4735-a501-a108e128d0cb · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Visual Instruction Tuning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 102a6ad2-1187-4852-a81b-9834676a21ac · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning RegMix: Data Mixture as Regression for Language Model Pre-training
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd864dc0-3ecf-4ea1-9b0e-2e11bbfa8ceb · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0c60ca46-2971-49c5-948b-02858424b992 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b2c552c9-0fc2-461c-b13e-6cafda9e26f9 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 63d2c6ca-82ff-491b-9fac-58dc5be8544b · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 98ec5f18-7a3f-447c-8cad-a5c5a083d1b3 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 69ff9057-a129-4833-a766-311e787d8227 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4db49e7b-67f9-4fcc-924e-d1eeb347f82e · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d2b4e5c-f726-46dd-9e54-1c92dd1fc29a · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Mathew, V
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 07b05db5-4d96-4a24-883d-5afaf2c774c6 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning McKinzie, Z
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 506e6805-bec1-42ef-8741-9021d8de2142 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d96e84b5-90c7-418f-9cad-8895631444ce · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1668b6b2-9509-45c9-b2d5-cbb91bfdbbfd · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Cosmos-reason1: From physical common sense to embodied reasoning.ArXiv Preprint, 2025
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 36adb5f2-e95f-483b-895a-81d9efcd6413 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning GPT-4 Technical Report
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2b5241e-a48c-4c92-bda0-a3b0c03496aa · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Training language models to follow instructions with human feedback
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e7f3f0e-9d5a-45c9-8842-d72f6ed6f0ca · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Ouyang, J
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 05f165ac-e7f8-4847-bc0f-9120a054e969 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Peng, Chris, X
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1d5b4da2-473d-4f1a-90e3-d6c1dc18299a · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Radford, J
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbc401ff-1088-4f68-a918-586d04aee518 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd62f111-703f-48e8-8adc-cfc620dd31ba · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation befcf928-f64e-4cd1-a5d8-aeab00ba624a · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f92bc84-026a-4511-8e3a-6714df7829dd · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ef0fab3-046a-4843-8f4d-86c506c1ef34 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 367cc09a-05af-44dc-b842-ae928ee91792 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be5e165c-dd85-497e-bbf7-d2e9b81ce38d · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 230e477d-4ee5-4fb3-b729-95864a7ed00c · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Aligning Large Multimodal Models with Factually Augmented RLHF
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1fbe716-a90d-43ec-913a-c1820312abbb · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning mDPO: Conditional Preference Optimization for Multimodal Large Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be18043c-6a69-44e7-ab13-9d65bd318e31 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 78c93715-b0d7-4206-8442-5b4df4c41f82 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0a2e5b54-9d8c-44b1-8f15-2467a0288201 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bfee1f25-fabd-44c0-bbf0-cec72525f06d · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6984741f-9810-4566-8bd5-ca67811b28b8 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33f319fc-9290-4edf-ac9a-6f53d1eb7b31 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Self-rewarding correction for mathematical reasoning
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95606ce7-7304-4328-80ef-7b133816ede7 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9013f422-4ac0-4f27-b3a5-ea5697626f87 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 958484b5-336b-4ff0-83b4-8415090de637 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 191b6167-856d-46d8-9199-6fe42cb3ccd3 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37e4bc21-49cf-45a0-91ef-ecf3b2449ca6 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 180fd3ef-1eda-4c11-8770-d40abbb02e9f · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 19b5ab88-062e-4050-8a5e-838e1cfe7b40 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 743b0234-2b6f-4c89-a45c-ca3607796007 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5f4adcc-7c7e-4cf0-9684-045d4fbe4511 · outbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning aha moment
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bb4826de-a215-4f05-b0f8-2f7d09409a9e · inbound
Perception-Aware Policy Optimization for Multimodal Reasoning MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2917cdcd-d538-4c7a-851d-80e5c0bc0eed · inbound
Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed54ce13-893a-4ea7-b882-229b449251d0 · inbound
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2305ebed-33fd-4d4f-bbe1-e088f0d8af4c · inbound
Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5caa1d1d-90d8-4143-9426-0394621fd17e · inbound
Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.