Pith. sign in

Paper Citation Record · LEDGER

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 5 inbound Pith citation observations for arXiv:2505.24871.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24871 v2

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:19:33.840820Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T08:15:52.772665Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T22:24:00.434533Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved63
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2dc90553-8ba9-43d5-a898-a69e774daed7 · outbound

This paper cites Alayrac, J.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Alayrac, J

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:26.779861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:26.779861Z digest=sha256:65ee7332068487ea9d94726785689e8ef6b7a2b30451c7d25f111f91544e95c4

Observation f3e214dd-d14f-4ad1-af5b-4b898e03b140 · outbound

This paper cites Efficient Online Data Mixing For Language Model Pre-Training.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Efficient Online Data Mixing For Language Model Pre-Training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:26.845188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:26.845188Z digest=sha256:41b2fb71e64d7eef38eb34f7141cc7903d54fd3a8db18d938935e0fadb8ce4d4

Observation 00e694f5-71e2-49b2-aafa-6c45738c459f · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:26.919852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:26.919852Z digest=sha256:4b11894148e43775c39bc22872ed8668c83a1e523751c669e159a1cb99b6adf5

Observation 1b1e10be-3f0d-4fff-bd10-b63f1a4eb349 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:40.339394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:27.126736Z digest=sha256:d62e6e147f205fa52d365ea104c742e41b77bcd8435e74d64d28053a61904f42

Observation 0693db7d-7433-4d89-9fec-4b76efbc542e · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:40.188757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:27.222872Z digest=sha256:f2fe67e13ab99026e9af5e8a1587fbe10108f96108aaffddf1e891f6bb50c6d3

Observation 9fab8f8f-2d3a-4d91-b582-ce7504e130dc · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:40.080528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:27.317891Z digest=sha256:dc374b50c3ff9441306979e8693b84d3268d784c7d0a768a5c0285078d993a0e

Observation 09960a24-cca9-4205-b572-34c56db39c0a · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:39.936695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:27.412658Z digest=sha256:6d4997634d7c01a7c1d88f47c15b30e7661a562a44548f11b43ef753c0e92237

Observation 4f4c7406-6758-4888-825c-15e77db8b3cf · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:39.796438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:27.484727Z digest=sha256:f17d70dd96f5b41cbc4f0904651efa94e88067385bc0e1ceb6af5da31d698a4c

Observation 2344a5a9-f500-4012-b701-1bcb8274e308 · outbound

This paper cites UniMax: Fairer and more Effective Language Sampling for Large-Scale Multilingual Pretraining.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning UniMax: Fairer and more Effective Language Sampling for Large-Scale Multilingual Pretraining

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:27.565628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:27.565628Z digest=sha256:c4aba75e6c15a62dcba51b6abe2f357ef158f78e5696d0d0a647f23e6ab1c646

Observation 2b9990b6-8dc9-4bee-a015-bc7eb4ced88f · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.ArXiv Preprint, 2025.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.ArXiv Preprint, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:19:39.655261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:27.638764Z digest=sha256:238f031e25f33d4db746c5de764629b83948da0ff8049a5814980e981b1b47ea

Observation 79346d66-c34d-4823-8d39-2c09636ce64d · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:39.507566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:27.721544Z digest=sha256:b4e8800b0145eea187d2cee2a090a006df889212b193bbba66045e5bda243540

Observation 6ea1a21b-84b2-436f-ba79-57296eaa9c93 · outbound

This paper cites Devlin, M.-W.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Devlin, M.-W

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:27.793435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:27.793435Z digest=sha256:ab2acf7a2d8fbd307a6ef95c120bfe8e2c9af826f52e07d71c5b44a677e547d3

Observation 0789acc3-b5e4-4148-acb8-715928a1d3fb · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:39.359011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:27.940368Z digest=sha256:f5f9ae4132fd69776dcff8907a9290d9165093670089f691ed7331a149b7f556

Observation e135594d-1c46-4d77-8a61-9b2fe77ab4d0 · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:28.455709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:28.455709Z digest=sha256:a6e73e70e9360149cf4c216d20457ebdb908ae139109d38407df05ce9b64bc95

Observation 37c2ea2a-7d24-4ba1-9296-5a561771ff43 · outbound

This paper cites Optimizing Pretraining Data Mixtures with LLM-Estimated Utility.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Optimizing Pretraining Data Mixtures with LLM-Estimated Utility

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:28.566337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:28.566337Z digest=sha256:de57853d5229d7ded13def983cd03ba860866efb547ee74a578d01f6a9e76b16

Observation ca94893b-7245-4457-820f-027ef55c824e · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:28.647433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:28.647433Z digest=sha256:d8cb42538a14445eeb4c42c2ded458356957532edff6cc43776e0d932a89630a

Observation e4778dbe-e24c-4b26-8d07-f77022815ca7 · outbound

This paper cites Datamodels: Predicting Predictions from Training Data.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Datamodels: Predicting Predictions from Training Data

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:28.727706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:28.727706Z digest=sha256:e2010781a426909840922871b60db063f48bc8864cf37f1869b77e365d98642c

Observation 7fec577c-30d3-4dca-bc9d-0105e7d12de0 · outbound

This paper cites Kazemnejad, M.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Kazemnejad, M

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:19:39.270411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:28.796216Z digest=sha256:39befe88ac64f840229459225ffbf979e7588864b50a640df0ce0a00b98d358e

Observation cc549ef2-8849-4736-990d-a8cc6a122214 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:28.867759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:28.867759Z digest=sha256:cc5a0ba037d118a6edaeccb4d69597c244c4d20b3556834835946f68b8c64ff2

Observation 454816c9-254c-4c12-9dff-b963bfc79b26 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:39.139184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:28.981683Z digest=sha256:cf881188bdf3785e4c321555ccc0dd9e1938a2676a0dd7bc521d8af9e97a238a

Observation 1cccf3fd-6ce4-423e-8ea0-56de82d54d1e · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:29.080379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:29.080379Z digest=sha256:cde05d202a5e7f8b536bfd634e040f8b95564de5de14081f6dc0e0adcc7fb914

Observation 4c5086ab-e6f3-4a30-be7a-b50e5ced75d9 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:39.023631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:29.172653Z digest=sha256:44e2ab68563be6269642881af1637b326aac09685057bf7cb22e2fea190dc7d1

Observation 0509d813-4b14-4f76-9936-5fb14ced841a · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:38.908662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:29.305826Z digest=sha256:bc39289f30f88326c995a8794a254e61f54aaa990151ed66e3cf2798bd44af00

Observation c6b2d7ee-8903-4e5f-800a-4a7f606b1c26 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:38.768227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:29.375324Z digest=sha256:62595b8523bdd899296dbd343905d8b751de794411d858a4a7c1fc16ad770128

Observation 3306c2d2-a036-43da-9e03-02a335611353 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:29.463877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:29.463877Z digest=sha256:4e172ba6639079713d1b021c3ffacc6d912195b2222d7289c57ace069f661ed9

Observation 63641a10-1a9d-4b77-b9a8-2555350f9de6 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:38.662831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:29.585191Z digest=sha256:4d8cc8977f07738da4798022ea114d8c1ff09f2404d3b41ab8df9dd80647783d

Observation ebbcf531-7dbb-4c0c-ac47-90e97aff91e5 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:38.470944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:29.678228Z digest=sha256:d6d476295a59c72c4ae861567951571d6ac264cabb32f4feaff9814b5c0ac60b

Observation 1d85ef55-57af-4a92-a40e-93f2e3046e45 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:29.779738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:29.779738Z digest=sha256:2029222c563b9f7011ef55e0a593851838c782f62c8d5efb566401f31dd83436

Observation 7fb0ce7a-3ae3-4735-a501-a108e128d0cb · outbound

This paper cites Visual Instruction Tuning.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Visual Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:29.907802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:29.907802Z digest=sha256:779e8b80aa345683ef05433aca1c0e44a930dc8c70d7f00c9dd8a072d825f77b

Observation 102a6ad2-1187-4852-a81b-9834676a21ac · outbound

This paper cites RegMix: Data Mixture as Regression for Language Model Pre-training.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning RegMix: Data Mixture as Regression for Language Model Pre-training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:30.003107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:30.003107Z digest=sha256:343ad9852ae7fffc9a96020258e54efb8e5751b41183994c970d1e916c3cac42

Observation dd864dc0-3ecf-4ea1-9b0e-2e11bbfa8ceb · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:38.256009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:30.076011Z digest=sha256:d686b10100a9c6c91586d60c5ae9ebde9c3ed74e6d82745f5b9cc1a98f7c4fcd

Observation 0c60ca46-2971-49c5-948b-02858424b992 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:38.060622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:30.201326Z digest=sha256:b34ad361ccd23fd0bf0db8953266662f8483d5582cff59e9de9538ed50a4eb32

Observation b2c552c9-0fc2-461c-b13e-6cafda9e26f9 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:37.885730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:30.290991Z digest=sha256:fbbcd85ffd993b948f87c61418156a168fbe5e975b927762a9917d564cd3473b

Observation 63d2c6ca-82ff-491b-9fac-58dc5be8544b · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:37.732944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:30.379806Z digest=sha256:76a9b6a822bb9b1b3c887e2baa7f78aa643e4e207a352ab7e88b898beaca9c12

Observation 98ec5f18-7a3f-447c-8cad-a5c5a083d1b3 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:37.556322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:30.468138Z digest=sha256:1fd001126e55b57b4b3c4b989dcfa58b382c72a15688a54509ffabf0f6b68d7e

Observation 69ff9057-a129-4833-a766-311e787d8227 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:37.336711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:30.566713Z digest=sha256:96377bf03d67337ac75ca665d32a952a046c97011f8caed103cca6a7b34d7caa

Observation 4db49e7b-67f9-4fcc-924e-d1eeb347f82e · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:30.662556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:30.662556Z digest=sha256:d5a50dfe1d1451e991fc6105e932aa8eacf4c1a1424ce093add9bc1a5206d855

Observation 0d2b4e5c-f726-46dd-9e54-1c92dd1fc29a · outbound

This paper cites Mathew, V.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Mathew, V

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:19:37.114089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:30.770327Z digest=sha256:4a7abf7fb85377fbe40535f66cb4809f9ec7d2e35efeb85b52bdf2f409139ee6

Observation 07b05db5-4d96-4a24-883d-5afaf2c774c6 · outbound

This paper cites McKinzie, Z.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning McKinzie, Z

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:19:36.939924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:30.866055Z digest=sha256:02e7fbfcf0173ca8dfdfe8813b7cdcda9498cd852523a093081d66fe699a7a19

Observation 506e6805-bec1-42ef-8741-9021d8de2142 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:30.952675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:30.952675Z digest=sha256:c1502099cc4b2a456da9955ef0edc68c24af1f65109f135673ed55b6115369d4

Observation d96e84b5-90c7-418f-9cad-8895631444ce · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:31.018935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:31.018935Z digest=sha256:3c03c46e657f92c76f24cea02f659269043f280dbed94cba0b8aa5016c7cce88

Observation 1668b6b2-9509-45c9-b2d5-cbb91bfdbbfd · outbound

This paper cites Cosmos-reason1: From physical common sense to embodied reasoning.ArXiv Preprint, 2025.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Cosmos-reason1: From physical common sense to embodied reasoning.ArXiv Preprint, 2025

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:19:36.771630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:31.104564Z digest=sha256:f5540548783b3dfd20727e2fd5456be1bdfc72cd6718786e09a6d77a4cd5a3a4

Observation 36adb5f2-e95f-483b-895a-81d9efcd6413 · outbound

This paper cites GPT-4 Technical Report.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning GPT-4 Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:31.205359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:31.205359Z digest=sha256:27f0b9ec82723fc3e607ac9ba021559244ce17ddef90d53cca6ccfa1dafa4cf9

Observation b2b5241e-a48c-4c92-bda0-a3b0c03496aa · outbound

This paper cites Training language models to follow instructions with human feedback.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Training language models to follow instructions with human feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:31.310445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:31.310445Z digest=sha256:04bd8382c0e1c1031dcd8039ec7e4eb6a5eda6e2d4a899c34cb3642a31650b3c

Observation 4e7f3f0e-9d5a-45c9-8842-d72f6ed6f0ca · outbound

This paper cites Ouyang, J.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Ouyang, J

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:19:36.611155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:31.380667Z digest=sha256:9562302437945eaafa7a6980b78bea4c95362f8069758e295e6bf2cddc41075a

Observation 05f165ac-e7f8-4847-bc0f-9120a054e969 · outbound

This paper cites Peng, Chris, X.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Peng, Chris, X

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:19:36.437611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:31.502640Z digest=sha256:46880297e9baf04a1714f9659e4f4ae930d90da3a40180c633d34a3edc41295e

Observation 1d5b4da2-473d-4f1a-90e3-d6c1dc18299a · outbound

This paper cites Radford, J.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Radford, J

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:31.625386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:31.625386Z digest=sha256:adf59b281284b0b106abf2649075b1f6ea30f36134c7244170f9adaea3e4d44d

Observation cbc401ff-1088-4f68-a918-586d04aee518 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:31.738904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:31.738904Z digest=sha256:658e8e089dcdf537cc07b127a6511430b37a7ea352e3c9f2d9bded8359284da9

Observation dd62f111-703f-48e8-8adc-cfc620dd31ba · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:36.227187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:31.793444Z digest=sha256:d5d1410650d96c19d5e4d0b5a3fb580c74f78641d4132945a08a2fdc2a47cb74

Observation befcf928-f64e-4cd1-a5d8-aeab00ba624a · outbound

This paper cites Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:31.901276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:31.901276Z digest=sha256:d56a3b063729d3fcced8599e070b47da96bae6a2460adc14ca963ee22c0a83d5

Observation 9f92bc84-026a-4511-8e3a-6714df7829dd · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:32.007792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:32.007792Z digest=sha256:7a650e8ca62fdaa1c95b4d2d1afaffa345990ea99d1b18fb70ac43b7efb62143

Observation 5ef0fab3-046a-4843-8f4d-86c506c1ef34 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:32.059007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:32.059007Z digest=sha256:c9c7f29a3b4cf2784628d257c83ce75181aa87025b340d32e10aa3841382772d

Observation 367cc09a-05af-44dc-b842-ae928ee91792 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:32.177573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:32.177573Z digest=sha256:74af5135ad1395494019d516832f02e4bb8221326c4e231832bdef37520a4867

Observation be5e165c-dd85-497e-bbf7-d2e9b81ce38d · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:36.088926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:32.290607Z digest=sha256:8115a1a93e7bd9d05dabcad43ce4aa8bce8323c6f83e0a5013df873d09d00a36

Observation 230e477d-4ee5-4fb3-b729-95864a7ed00c · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:32.411773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:32.411773Z digest=sha256:97ebf75a3d0008553a887522ca9b91393cd8d1c37a9e71a4b9a8a3d1e6a5aca2

Observation a1fbe716-a90d-43ec-913a-c1820312abbb · outbound

This paper cites mDPO: Conditional Preference Optimization for Multimodal Large Language Models.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning mDPO: Conditional Preference Optimization for Multimodal Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:32.483603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:32.483603Z digest=sha256:de67fd56b952b5b7eab4fbbdd0f9ddbf0ef1be094032054fa91ad8d7a861642b

Observation be18043c-6a69-44e7-ab13-9d65bd318e31 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:35.910101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:32.601211Z digest=sha256:6cd7a259f7adce4a61ea09f0c8a8e599de800a8abb183b4e66d12412d668ba5a

Observation 78c93715-b0d7-4206-8442-5b4df4c41f82 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:35.741748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:32.703789Z digest=sha256:790a86e2adb112f050abccb819ebb7b6a3d38d840223172094283d2e1a507c87

Observation 0a2e5b54-9d8c-44b1-8f15-2467a0288201 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:35.548938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:32.784877Z digest=sha256:16795ab0f7f0046f49cfbf459188d9dcb6ede8e770aa1811e0c8cadae7a8a4d4

Observation bfee1f25-fabd-44c0-bbf0-cec72525f06d · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:35.363940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:32.887888Z digest=sha256:062f6c5e2c45ed4cdd29cd2539292728458682ccb09e946b4563ae5e5cb752b8

Observation 6984741f-9810-4566-8bd5-ca67811b28b8 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:33.002414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:33.002414Z digest=sha256:12aaeb606acd205a35f00113bf05bf7a06c3c1754fc628b1c14cbd42ddc72db4

Observation 33f319fc-9290-4edf-ac9a-6f53d1eb7b31 · outbound

This paper cites Self-rewarding correction for mathematical reasoning.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Self-rewarding correction for mathematical reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:33.115481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:33.115481Z digest=sha256:11b64cf55bd97b257c5fc64062b9839b492782acfdf580848df83c44ca67c8b1

Observation 95606ce7-7304-4328-80ef-7b133816ede7 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:35.209897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:33.199071Z digest=sha256:6aebe8faa2a705eeb22534d2a00b844268d701e8dd6397904fcdea0f9df8566e

Observation 9013f422-4ac0-4f27-b3a5-ea5697626f87 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:35.042683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:33.304892Z digest=sha256:e793ed7c7b27dee738521ba60ae8c1d7ef24e25ace91dd4f27f13a70367a31f7

Observation 958484b5-336b-4ff0-83b4-8415090de637 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:34.874529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:33.373216Z digest=sha256:a901497c980722681f3dd53b01edecb82720d0135e74f09a684d66b4bbf4ad31

Observation 191b6167-856d-46d8-9199-6fe42cb3ccd3 · outbound

This paper cites Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:33.456884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:33.456884Z digest=sha256:495c2444f48a44013b5192f54ab897169387247cba379efbb41b02e190e06b0f

Observation 37e4bc21-49cf-45a0-91ef-ecf3b2449ca6 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:34.736096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:33.542724Z digest=sha256:98a66e8fb3fee5071b68fff4ae4d3b34cedd1117dc8f2b3f7a5e30272aea018e

Observation 180fd3ef-1eda-4c11-8770-d40abbb02e9f · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:34.617344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:33.618438Z digest=sha256:464502a40c1f2f1ef2523a521f0796894b247d25b26850c63fad1be6b890eb2c

Observation 19b5ab88-062e-4050-8a5e-838e1cfe7b40 · outbound

This paper cites an unresolved cited work.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:19:34.446814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:33.724474Z digest=sha256:22ecc369e1168617f49f6244d78c65ccd0a50708f20af727d05bcb1322194ec5

Observation 743b0234-2b6f-4c89-a45c-ca3607796007 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:33.786626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:33.786626Z digest=sha256:6ae3e2b6f57fef8be7284156e770d644279027772c373a2deb9a5d1b522d3417

Observation d5f4adcc-7c7e-4cf0-9684-045d4fbe4511 · outbound

This paper cites aha moment.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning aha moment

Reference 74

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:19:34.067032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:19:33.840820Z digest=sha256:0eebe6a9bff3153c15a21f4214cd656d3c388875229c7171a3e2fa1fd4eefc76

Pith citing papers

Observation bb4826de-a215-4f05-b0f8-2f7d09409a9e · inbound

Perception-Aware Policy Optimization for Multimodal Reasoning cites this paper.

Perception-Aware Policy Optimization for Multimodal Reasoning MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:12:04.799535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:b0d2ac5d5d2bb3d5f5b90437ea866288d89d775a939b71afe1a475029ec872a7

Observation 2917cdcd-d538-4c7a-851d-80e5c0bc0eed · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:52.772665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:52.772665Z digest=sha256:b181f118f4f5b82ef0e643b4a7943a960d529cc6a86c8e08c4a48d9365164c9d

Observation ed54ce13-893a-4ea7-b882-229b449251d0 · inbound

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks cites this paper.

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T04:17:27.630062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:17:27.630062Z digest=sha256:0a648535eba144c600a063740c31b2364993fcf8ed1a337336864b00e71d10e7

Observation 2305ebed-33fd-4d4f-bbe1-e088f0d8af4c · inbound

Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models cites this paper.

Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:24:00.436153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:18:10.508034Z digest=sha256:fc15a85429f86b8d3d6809efc7dc87a7a2ffc6a38078ecd26c926035549a6a58

Observation 5caa1d1d-90d8-4143-9426-0394621fd17e · inbound

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning cites this paper.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:24.668308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:24.668308Z digest=sha256:1720d10f137c4f2d6d054f562847a31edef94fa2caa38809545c3c43df1664af