Pith. sign in

Paper Citation Record · LEDGER

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 5 inbound Pith citation observations for arXiv:2505.24871.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24871 v2

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:19:33.840820Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T08:15:52.772665Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T22:24:00.434533Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved63
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2dc90553-8ba9-43d5-a898-a69e774daed7 · outbound

This paper cites Alayrac, J.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Alayrac, J

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:26.779861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:26.779861Z digest=sha256:cc589a3c532da790892e50dc5b8f2a0f346dadd9984214d355aa6f61bd22b756

Observation f3e214dd-d14f-4ad1-af5b-4b898e03b140 · outbound

This paper cites Efficient Online Data Mixing For Language Model Pre-Training.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Efficient Online Data Mixing For Language Model Pre-Training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:26.845188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:26.845188Z digest=sha256:ad33439768bdd918751fc3be4c7a270c0795401ab670ee46ae6ddfb44ccdfc37

Observation 00e694f5-71e2-49b2-aafa-6c45738c459f · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:26.919852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:26.919852Z digest=sha256:2ce5eb59301b7b7dba4ad100d8281295ebe7ee434905dd49fd88802342c9afde

Observation 1b1e10be-3f0d-4fff-bd10-b63f1a4eb349 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:40.339394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:27.126736Z digest=sha256:3c87363c3763d7737384a440ec8c5e5f3198060ec226a1d30ba11216fd727f10

Observation 0693db7d-7433-4d89-9fec-4b76efbc542e · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:40.188757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:27.222872Z digest=sha256:c74ef7349e1ad843eec052c82795fd5e4f6979738b1520a4c907982e69e51fc9

Observation 9fab8f8f-2d3a-4d91-b582-ce7504e130dc · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:40.080528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:27.317891Z digest=sha256:9eef2f4a47717d4e0b018f3b7e38fcda2d46f5f1014b4df2750cc923405dc6f5

Observation 09960a24-cca9-4205-b572-34c56db39c0a · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:39.936695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:27.412658Z digest=sha256:18839c88cf6aa61cac321a8ba183dd1e4920e4d40816116f845fee959f7c9a05

Observation 4f4c7406-6758-4888-825c-15e77db8b3cf · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:39.796438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:27.484727Z digest=sha256:3241c0c1d35f534b793d0d728bf62a93f95692287cec389acff7a5fe2ce42cd9

Observation 2344a5a9-f500-4012-b701-1bcb8274e308 · outbound

This paper cites UniMax: Fairer and more Effective Language Sampling for Large-Scale Multilingual Pretraining.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning UniMax: Fairer and more Effective Language Sampling for Large-Scale Multilingual Pretraining

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:27.565628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:27.565628Z digest=sha256:f1a6885d068ff8c36ccd4b277ac516c9297c1767a37b41a378ce8e27b31f61c8

Observation 2b9990b6-8dc9-4bee-a015-bc7eb4ced88f · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.ArXiv Preprint, 2025.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.ArXiv Preprint, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:19:39.655261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:27.638764Z digest=sha256:ab0a511871bca145783695459ffb34d1838c22ae14b1c2b1e424e719cf7e7ecf

Observation 79346d66-c34d-4823-8d39-2c09636ce64d · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:39.507566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:27.721544Z digest=sha256:402028d37bc23d29337705913d065e2f2387504cd621bcc8c405ae5cf249176c

Observation 6ea1a21b-84b2-436f-ba79-57296eaa9c93 · outbound

This paper cites Devlin, M.-W.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Devlin, M.-W

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:27.793435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:27.793435Z digest=sha256:2c728b2b2f92176adce038013cdc1705da15d84686435e9f74410e89aa0216b6

Observation 0789acc3-b5e4-4148-acb8-715928a1d3fb · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:39.359011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:27.940368Z digest=sha256:68bd3d730885b63b8fe429ee29a7fc33da0e908f31d9b16cd8687447cb14a2ca

Observation e135594d-1c46-4d77-8a61-9b2fe77ab4d0 · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:28.455709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:28.455709Z digest=sha256:9f000d62ee2eba604f36ce8f7225d56aa91219b8db9dbb11c95b0f22416e3c86

Observation 37c2ea2a-7d24-4ba1-9296-5a561771ff43 · outbound

This paper cites Optimizing Pretraining Data Mixtures with LLM-Estimated Utility.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Optimizing Pretraining Data Mixtures with LLM-Estimated Utility

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:28.566337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:28.566337Z digest=sha256:fba293dfe25f563b5473734aeb604b9bfc7c48b8173d010ecf99f67cb0bd553a

Observation ca94893b-7245-4457-820f-027ef55c824e · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:28.647433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:28.647433Z digest=sha256:e6541c20af853f14ae92dc2617b7a6c6ce6b76bcd8486ef87197d7c9eb97eee3

Observation e4778dbe-e24c-4b26-8d07-f77022815ca7 · outbound

This paper cites Datamodels: Predicting Predictions from Training Data.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Datamodels: Predicting Predictions from Training Data

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:28.727706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:28.727706Z digest=sha256:3659fedaab7c6fc21d26c37172944dc4126302fc2bb1deee2a933dffeb27ae5d

Observation 7fec577c-30d3-4dca-bc9d-0105e7d12de0 · outbound

This paper cites Kazemnejad, M.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Kazemnejad, M

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:19:39.270411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:28.796216Z digest=sha256:349396f937dfe263b1e83de2ee7192da72ecfe4945c3ba09bb3400939dfdd99a

Observation cc549ef2-8849-4736-990d-a8cc6a122214 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:28.867759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:28.867759Z digest=sha256:bb4fae2c2ac1e7f064e835cd29fc8cc671dbce93d4bb9ad116193ec70b24fad4

Observation 454816c9-254c-4c12-9dff-b963bfc79b26 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:39.139184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:28.981683Z digest=sha256:0e35effc5a8be0e5aab16e39993de588978a418037edd6d3ef15b97144cb1ede

Observation 1cccf3fd-6ce4-423e-8ea0-56de82d54d1e · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:29.080379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:29.080379Z digest=sha256:d237cb65c23e512554140a4efdcdf9ce60f4f409fc1f8a3084843060e3bc9f5c

Observation 4c5086ab-e6f3-4a30-be7a-b50e5ced75d9 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:39.023631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:29.172653Z digest=sha256:21b9e3b8ff87591677969537dde814aa4a3b4986d1aa59d3b184f6e237256fe8

Observation 0509d813-4b14-4f76-9936-5fb14ced841a · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:38.908662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:29.305826Z digest=sha256:c51f2f55f6d62fbdeeeb9c8eba8ba01b9f7fb9edbb60c1404d1c6cf94f514200

Observation c6b2d7ee-8903-4e5f-800a-4a7f606b1c26 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:38.768227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:29.375324Z digest=sha256:fa62278b603bbaef243b507b02c98109daea6b45f1653a785055da97448b17d6

Observation 3306c2d2-a036-43da-9e03-02a335611353 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:29.463877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:29.463877Z digest=sha256:f8f058dd9ddcb0d3c56b9553d6742c5b52a3273133937da8592b6663a89f7a2d

Observation 63641a10-1a9d-4b77-b9a8-2555350f9de6 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:38.662831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:29.585191Z digest=sha256:274d1566ca75fd57cd4cbbd0a5be64c7cf041bcb689880b7179d85bb059bea51

Observation ebbcf531-7dbb-4c0c-ac47-90e97aff91e5 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:38.470944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:29.678228Z digest=sha256:c2c0770d7205c9252fbfc2ddbad34f6e9a76579f8915b17b2e62695c44ef4b17

Observation 1d85ef55-57af-4a92-a40e-93f2e3046e45 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:29.779738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:29.779738Z digest=sha256:6f348ae57ba9cf778569b1a55c3bbb1615ee24bf2318d6eb79884e40c3cb7c15

Observation 7fb0ce7a-3ae3-4735-a501-a108e128d0cb · outbound

This paper cites Visual Instruction Tuning.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Visual Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:29.907802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:29.907802Z digest=sha256:61b7027e960e8ee53ae8353d6ffa1795bd08eb0ab0931cf09d670e1985ea4024

Observation 102a6ad2-1187-4852-a81b-9834676a21ac · outbound

This paper cites RegMix: Data Mixture as Regression for Language Model Pre-training.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning RegMix: Data Mixture as Regression for Language Model Pre-training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:30.003107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:30.003107Z digest=sha256:e0ae51ea078483551faca0aea429caaa7f15c8bb119d4691cd443ade108824fa

Observation dd864dc0-3ecf-4ea1-9b0e-2e11bbfa8ceb · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:38.256009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:30.076011Z digest=sha256:b9a893e88b2cfd302b45b14f1cbac3518182ba1fec683a5d8789a4c986b99241

Observation 0c60ca46-2971-49c5-948b-02858424b992 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:38.060622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:30.201326Z digest=sha256:57581a00565ed5bacfeceec743308242670496eb26e33aeacd74d8a149189fac

Observation b2c552c9-0fc2-461c-b13e-6cafda9e26f9 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:37.885730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:30.290991Z digest=sha256:4f01aeb53032532c8204febd64644218a5dfcf74773444dbbe35bff0fcc3b037

Observation 63d2c6ca-82ff-491b-9fac-58dc5be8544b · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:37.732944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:30.379806Z digest=sha256:9df0c3eb9568b7c7c77038fdad8f796b7fce431eeb01a133cc934a15694c703b

Observation 98ec5f18-7a3f-447c-8cad-a5c5a083d1b3 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:37.556322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:30.468138Z digest=sha256:2f3ac7ff283585fc0215b04e45c5cc0b7e33ed76119bb3fb2556145dba954bb7

Observation 69ff9057-a129-4833-a766-311e787d8227 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:37.336711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:30.566713Z digest=sha256:404dddeab728bd821fd919bf20619b276e5a66fc1cfe2896057a5510b6eb0255

Observation 4db49e7b-67f9-4fcc-924e-d1eeb347f82e · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:30.662556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:30.662556Z digest=sha256:518a9604bda59ab3665d681ab6cadb23102b881fb42bc65bb10e5c1384a02cb7

Observation 0d2b4e5c-f726-46dd-9e54-1c92dd1fc29a · outbound

This paper cites Mathew, V.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Mathew, V

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:19:37.114089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:30.770327Z digest=sha256:d1990511f86758d5160e7dcd5e1ba3f0eeedd91e87803a1ef7031e5b15e2a1f0

Observation 07b05db5-4d96-4a24-883d-5afaf2c774c6 · outbound

This paper cites McKinzie, Z.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning McKinzie, Z

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:19:36.939924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:30.866055Z digest=sha256:12d5a2e88fac7d891baaf78d62157bbab7042dc574fc79e1d677969a647b9f30

Observation 506e6805-bec1-42ef-8741-9021d8de2142 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:30.952675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:30.952675Z digest=sha256:1965a4457aa04146bcce761510119343f8b830138e7749906b414d0490fb8641

Observation d96e84b5-90c7-418f-9cad-8895631444ce · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:31.018935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:31.018935Z digest=sha256:bf9f496595796945cd7818966372a868a4d99880b15e8ebaa8d36e86a986e5b4

Observation 1668b6b2-9509-45c9-b2d5-cbb91bfdbbfd · outbound

This paper cites Cosmos-reason1: From physical common sense to embodied reasoning.ArXiv Preprint, 2025.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Cosmos-reason1: From physical common sense to embodied reasoning.ArXiv Preprint, 2025

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:19:36.771630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:31.104564Z digest=sha256:981920dee5fb4b857c56efee523e01d1a19e7afdfd105cc17c4cde26c3bbfa50

Observation 36adb5f2-e95f-483b-895a-81d9efcd6413 · outbound

This paper cites GPT-4 Technical Report.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning GPT-4 Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:31.205359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:31.205359Z digest=sha256:dfd825d78d58dd9fe80ea6fee3a483d14e035b8e39b1b534ae7090f0ccea4248

Observation b2b5241e-a48c-4c92-bda0-a3b0c03496aa · outbound

This paper cites Training language models to follow instructions with human feedback.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Training language models to follow instructions with human feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:31.310445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:31.310445Z digest=sha256:17460b9d271e51c2ee8d6f734819455f8f2ac189eabf13bfacfba882e9837ad0

Observation 4e7f3f0e-9d5a-45c9-8842-d72f6ed6f0ca · outbound

This paper cites Ouyang, J.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Ouyang, J

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:19:36.611155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:31.380667Z digest=sha256:67e0e8139f1e4d677bacc4ef9d884da2db5ea8716378bdca2ee3d005b652cfa8

Observation 05f165ac-e7f8-4847-bc0f-9120a054e969 · outbound

This paper cites Peng, Chris, X.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Peng, Chris, X

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:19:36.437611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:31.502640Z digest=sha256:59656f2a08d46c5e3f5b6f56b031f1edd6a76a59c5552a6faf14ab569351ec17

Observation 1d5b4da2-473d-4f1a-90e3-d6c1dc18299a · outbound

This paper cites Radford, J.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Radford, J

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:31.625386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:31.625386Z digest=sha256:83ca7955f4d2ff767a4df446b6fca1ccb183a066c2e6e0494cf4984e1cecfa24

Observation cbc401ff-1088-4f68-a918-586d04aee518 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:31.738904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:31.738904Z digest=sha256:fd112384c7f6a4d805ce5a40cea6a01f6f7aad40ed9b798882be33af13490740

Observation dd62f111-703f-48e8-8adc-cfc620dd31ba · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:36.227187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:31.793444Z digest=sha256:32cd0d07b3ea81566ed48844e48aed77f5c30b72465f429e003525a713ef0b05

Observation befcf928-f64e-4cd1-a5d8-aeab00ba624a · outbound

This paper cites Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:31.901276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:31.901276Z digest=sha256:f07c050fe47adc392b32c6852ff2d4de8243c299d35fc20f7ccd6d2f0e8928aa

Observation 9f92bc84-026a-4511-8e3a-6714df7829dd · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:32.007792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:32.007792Z digest=sha256:22e51f93096403d545226bf04467894f0d9c4a0b726f6e108deaef945d2d9682

Observation 5ef0fab3-046a-4843-8f4d-86c506c1ef34 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:32.059007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:32.059007Z digest=sha256:187fcb9838061cd5b2a8deeb55d9a82d22dec029d2e1af08f2d3244380ae50c7

Observation 367cc09a-05af-44dc-b842-ae928ee91792 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:32.177573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:32.177573Z digest=sha256:6371178f14f22293ce05a818bb6800f9ba9fcedec4b6f1af19d05eaaafdcfb5f

Observation be5e165c-dd85-497e-bbf7-d2e9b81ce38d · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:36.088926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:32.290607Z digest=sha256:3a79f7fc55d659b4979f5deff432a4bd43e3b63cbf4886d24b404b54307b6b82

Observation 230e477d-4ee5-4fb3-b729-95864a7ed00c · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:32.411773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:32.411773Z digest=sha256:a024d435b5751f9d060f180947e90a69d599f2150a38282013801a88a7a29b0f

Observation a1fbe716-a90d-43ec-913a-c1820312abbb · outbound

This paper cites mDPO: Conditional Preference Optimization for Multimodal Large Language Models.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning mDPO: Conditional Preference Optimization for Multimodal Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:32.483603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:32.483603Z digest=sha256:863ef5e8338ee62773f6f60c360646451a0f1176df2f7a8f8e01f732b28dd078

Observation be18043c-6a69-44e7-ab13-9d65bd318e31 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:35.910101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:32.601211Z digest=sha256:eb89a7067ef602a910957485d85ffb4d9586d451836caf36fe8744a5047de9af

Observation 78c93715-b0d7-4206-8442-5b4df4c41f82 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:35.741748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:32.703789Z digest=sha256:128352fcee9223ea5acfa2926703c3f8f720a28e7109c9c6a75ac08a40566893

Observation 0a2e5b54-9d8c-44b1-8f15-2467a0288201 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:35.548938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:32.784877Z digest=sha256:8a193b28707537ef9f8285b70c4bb6b568473e10a9a420c33eab5bc8334066f1

Observation bfee1f25-fabd-44c0-bbf0-cec72525f06d · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:35.363940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:32.887888Z digest=sha256:676ac2490e4151c693094bbcbba9cc05e011efd1b3ec6453e0942ab938ef9860

Observation 6984741f-9810-4566-8bd5-ca67811b28b8 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:33.002414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:33.002414Z digest=sha256:2dcd0e4adbc7a019ef478ad636f85479bff20ae795d0cbb604ce7f53a4ed319c

Observation 33f319fc-9290-4edf-ac9a-6f53d1eb7b31 · outbound

This paper cites Self-rewarding correction for mathematical reasoning.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Self-rewarding correction for mathematical reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:33.115481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:33.115481Z digest=sha256:e74cb5459ea9331388a2a5f572a8ffb98a451d57be105911e9e048313c1f3aec

Observation 95606ce7-7304-4328-80ef-7b133816ede7 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:35.209897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:33.199071Z digest=sha256:ff132fc64c9c2f5ae7088ec01b84346a9e586a827cdac367752dedbaaf089c77

Observation 9013f422-4ac0-4f27-b3a5-ea5697626f87 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:35.042683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:33.304892Z digest=sha256:26acf72ec688816b097f6b0696b1eee69e9d972a0ef994719086cdfc1fbcb3a5

Observation 958484b5-336b-4ff0-83b4-8415090de637 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:34.874529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:33.373216Z digest=sha256:3dad1ac3f2ad2b4f1c3c616a1a6acc6c8f5c24b7ee2460136f1fe348634607de

Observation 191b6167-856d-46d8-9199-6fe42cb3ccd3 · outbound

This paper cites Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:33.456884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:33.456884Z digest=sha256:87558abb1fd92107af225e2a670de1daefea66bc03f4a2d8103dd139beb8db6b

Observation 37e4bc21-49cf-45a0-91ef-ecf3b2449ca6 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:34.736096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:33.542724Z digest=sha256:d3d1a2497882728bd780b6341d6fcfca3408b90ced0681d6656085c3d338646a

Observation 180fd3ef-1eda-4c11-8770-d40abbb02e9f · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:34.617344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:33.618438Z digest=sha256:a2d6bc2013abcc26ae5e9923d1724418b489561f7f51a1283e0112817f8effc2

Observation 19b5ab88-062e-4050-8a5e-838e1cfe7b40 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:34.446814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:33.724474Z digest=sha256:47dd8e3b0d38a60659efb75aa9859b2dfced3e48ec979df371845b43c13296da

Observation 743b0234-2b6f-4c89-a45c-ca3607796007 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:33.786626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:33.786626Z digest=sha256:9f5e0bdf0781e62da6863eed0776f26215d910897d9de822fe3ac749bd2b7b5d

Observation d5f4adcc-7c7e-4cf0-9684-045d4fbe4511 · outbound

This paper cites aha moment.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning aha moment

Reference 74

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:19:34.067032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:19:33.840820Z digest=sha256:15c0acf5f8cfa7425730f01a5c9430e26902413953f8ec21feec8b56e379c052

Pith citing papers

Observation bb4826de-a215-4f05-b0f8-2f7d09409a9e · inbound

Perception-Aware Policy Optimization for Multimodal Reasoning cites this paper.

Perception-Aware Policy Optimization for Multimodal Reasoning MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:12:04.799535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:33673522e15e46dd0f4de8264c2b6fe9a0932184c9c7634f0cfe9011f4623ce5

Observation 2917cdcd-d538-4c7a-851d-80e5c0bc0eed · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:52.772665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:52.772665Z digest=sha256:612486dea863dd6c0a9873b5a0609b91cda6d725e3bc8b03b60863277dd51391

Observation ed54ce13-893a-4ea7-b882-229b449251d0 · inbound

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks cites this paper.

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T04:17:27.630062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:17:27.630062Z digest=sha256:d11fd1a6b55adcbf2425d05ceb65c79b394d5e3693a02380ab41935e777e23e5

Observation 2305ebed-33fd-4d4f-bbe1-e088f0d8af4c · inbound

Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models cites this paper.

Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:24:00.436153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T22:18:10.508034Z digest=sha256:fdb50e9acb3d8266591f8270f056cfa5a36c8015e05dac3981ffee31973a1bf4

Observation 5caa1d1d-90d8-4143-9426-0394621fd17e · inbound

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning cites this paper.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:24.668308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:24.668308Z digest=sha256:3187f22323a50c0fea13c90ba10656625677281da5d3bec8b7dc3ea6fac9a830