Pith. sign in

Paper Citation Record · LEDGER

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing

As of 9 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 1 inbound Pith citation observation for arXiv:2508.18642.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.18642 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:24:52.802595Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T17:38:50.313305Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.225457Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved10
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f8721a2b-bd86-4f7e-a0d9-bb699f7d0df8 · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:55.969774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:50.825058Z digest=sha256:69e99d08e306b71224e76b389896a7c8c70fa3bc1b5e847e3c2fc757a5870fb1

Observation 0d22fe9f-74eb-4f4a-85a9-b96fdaf95e82 · outbound

This paper cites [System] You are an answer quality assessment expert.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing [System] You are an answer quality assessment expert

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:24:55.901772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:50.977092Z digest=sha256:95d0f94069857abb061d909b21770ef47a67a297d46b853cf9080b30991dd74c

Observation 296e50b2-f124-4d35-918a-7e3a8813e238 · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:55.484808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:51.389246Z digest=sha256:f739de51f42111ff108ef70ad0dce2d7f37443d397c9112f9b3f3a24e18083fa

Observation c15c6627-a112-4e83-b2d6-f5cbba1a9096 · outbound

This paper cites - Conclusion: Correct/Incorrect.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing - Conclusion: Correct/Incorrect

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:24:55.234743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:51.494733Z digest=sha256:353983685cf598ee2bce1539fe9a087ad4b13c08e41208fe1740aabc95a9531c

Observation 56238734-a64b-439c-bae8-20432c8c7ded · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:55.754742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:51.117105Z digest=sha256:fc649671ba1299030f9bcd07464e9792927deb112d5d368c1c9d5cc8ceb442df

Observation 3432a930-0580-48bc-a107-53702569a92b · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:55.649394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:51.232908Z digest=sha256:1b19d46f8ffd6a8e600294ef30d92418f7e7078d639c1a028fa50dd95f3f77bc

Observation 2f44dcde-9f93-41e9-9214-056e2e024254 · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 9

Resolution
parse uncertain
raw_fallback, observed 2026-08-05T16:24:55.080322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:51.594739Z digest=sha256:d6d9c65297b9c16a59255aef3c2b2d461d6fe677a14e2d4766072d7d6f27a7d3

Observation 678a2c23-aed3-4092-987a-c2bd8b9edb16 · outbound

This paper cites Write a script for a modern history video group assignment.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Write a script for a modern history video group assignment

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:24:54.910547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:51.776086Z digest=sha256:eae722c428fa8546489b59ada69002a3f5f28581d733b869388b6d9639869091

Observation 2e1cc6c4-06e7-493d-88d3-50f9322d3b87 · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:54.724737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:51.909221Z digest=sha256:b6903f0f385824782a0324aaf4f18942996f422cb139a4d643111db48a5f1256

Observation 6b7fd8a9-6f28-409b-99ff-b89d623f7ab7 · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:54.530654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:52.024731Z digest=sha256:404b5c6e1d6a4aa52cfb693d6ebfa77ce883250c319082f5f9d7363b666514d5

Observation 9212d3d3-1a8f-45a5-b5b9-3eb0f4a7bf66 · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:54.313759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:52.164758Z digest=sha256:8c33ec3a46183b4daa5f45840120654e1e9c189594c6d2f48640721d802f2977

Observation c485604b-9e36-4c4e-b4cc-fddbcdf531af · outbound

This paper cites an unresolved cited work.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-05T16:24:54.024945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:52.273448Z digest=sha256:476cee84cea06bbc0d06f227d5a5c021b604511b36da7bb45b256e85ed9fcdb4

Observation 8e0c6504-2b28-445b-9af7-f246677ecb54 · outbound

This paper cites My mother.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing My mother

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:24:53.644926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:52.574748Z digest=sha256:e600002cb0c96d8408dcad46b30b48397f96500a1751061e470779bd7142cbf7

Observation 56091e6d-dbc0-41f2-8462-4e7b70d157f9 · outbound

This paper cites come up with some three-character sword names.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing come up with some three-character sword names

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:24:53.314458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:24:52.802595Z digest=sha256:5017e4897b0ed9678ab962570563a8de6d10a81a1b0852e04c8ad11b1ce6caf6

Observation 9e8fb600-23f1-449c-88d1-27803be1d3aa · outbound

This paper cites Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:50.583302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:50.583302Z digest=sha256:f081e1a9b38cc251c72f2b7a6632d06c72861017f66a9eafcd188ed900dabfc8

Observation 87ace097-1856-44ad-a58a-9406b4e8e721 · outbound

This paper cites LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning.

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:50.674749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:50.674749Z digest=sha256:757d4b0867b0b1cc5575d9da0a641b216312293ce6ef447339d133f9a3dfec8d

Pith citing papers

Observation 96345c8f-51f8-4537-b756-e25287d8d992 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing

Reference 185

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.227008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:171952fb8115b65a47c64c46327241f72f0349d5918b714f03517ff3f4f9b92f