Pith. sign in

Paper Citation Record · LEDGER

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following

As of 11 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2508.02150.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.02150 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:14:36.480639Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved14
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 30ef7ae3-2cc4-46d8-8669-388d24235a35 · outbound

This paper cites To make the instructions more complex, I want you to identify and return five atomic constraints that can be added to the seed question.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following To make the instructions more complex, I want you to identify and return five atomic constraints that can be added to the seed question

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.131823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.131823Z digest=sha256:6c09debd339480772a3c8443a1281ecf90672a53a6f10456f639f76a0e0ed242

Observation 36fea6dd-fb2d-48b9-9067-968e99c5cdf4 · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.168195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.168195Z digest=sha256:5a5415d0cb52b5bd55b833810b4fd13da85b2c31cc13c5e44e5d478c732eb2ac

Observation 88453eb8-30b2-4c03-8bcd-bcf6610ad178 · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.226387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T05:14:35.964825Z digest=sha256:de73467bd44ea8e46a859f4174a929c36ff1dac3107cf5e904da76dd68fcc274

Observation 070800b4-ace1-497f-b440-6ef268daf33c · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.208220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T05:14:35.985981Z digest=sha256:ecc7b2781c7354d594522659e13975105dbbf11673f49d89a7a2c77bfd69ab97

Observation 9f26d81b-e226-42fa-809b-e6c91f2e289c · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.190858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T05:14:36.009417Z digest=sha256:27d9afa6c93043c3e7d9c715208cd3f7f72747ed216f7797e4b4958a962784f1

Observation a037e7a0-6edb-40ab-9a61-51cae0ebd1bf · outbound

This paper cites Step-by-Step Mastery: Enhancing Soft Constraint Following Ability of Large Language Models.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Step-by-Step Mastery: Enhancing Soft Constraint Following Ability of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.026653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.026653Z digest=sha256:f32b610888aef2aab6000121f7804560223e1471fb18c71e1cd772d13815e3fe

Observation 67020480-fe39-4fbd-813e-01d0dd2054e7 · outbound

This paper cites 5-thinking: Advancing superb rea- soning models with reinforcement learning.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following 5-thinking: Advancing superb rea- soning models with reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.169330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T05:14:36.032798Z digest=sha256:5688289b395be69d622550c0a0ad893b777b362a9a8218c314cb82c79f49d874

Observation 63de9e48-05d5-465a-9c44-e346864dee00 · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.149630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T05:14:36.044328Z digest=sha256:b5cf0070e8842e06174c5761456d318daec557e032e4618c4142ef68948f751f

Observation 2c1eee46-b55f-402a-bea0-bc133bc518f5 · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.127153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T05:14:36.060611Z digest=sha256:4d845f31bac34c5dd5f5870134c187d47c014348dca389a4ca50ffc183075f14

Observation c35bf93f-2f11-48ec-8320-bd251713d45a · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.100189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T05:14:36.086975Z digest=sha256:118ede511cebab544cef30f32a3037512c771f08ac8a313e4231bc7513b6ad15

Observation f10955ec-320c-4dcf-88b3-31dc5b3ad4b1 · outbound

This paper cites joy,” “anger,.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following joy,” “anger,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.053649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T05:14:36.104630Z digest=sha256:34a9a822eb947379eed935b6e0ee268f54d9d61e61116427a13034a8ba8d9e48

Observation 80807db4-1b09-4088-9b74-8b1f1dce4a6e · outbound

This paper cites You may choose one or more constraints from the list or propose new ones if needed.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following You may choose one or more constraints from the list or propose new ones if needed

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.185917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.185917Z digest=sha256:fb7ec6d7b1b05da51d1a143543ebf984b72cf7ed3ce5da38a3b7cf59af66130e

Observation f0fe6964-45f9-4668-b253-2224dcc20fbc · outbound

This paper cites Your task is only to generate new constraints that can be added to it.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Your task is only to generate new constraints that can be added to it

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.214321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.214321Z digest=sha256:6f6b3bd2164857f95fd281fc85b91887e37b1d176661bec36ec935a6b48a0a92

Observation f92f8690-123b-4b0e-be3b-87b7688341dd · outbound

This paper cites c1": "<first constraint>.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following c1": "<first constraint>

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.232092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.232092Z digest=sha256:7762c874938ef8c9c4cdd3ba63f021b5960d4e8dc3c7a0fb96b1329d1c98ec13

Observation 9361bf8c-c861-48cb-bb46-f8bf52b01c33 · outbound

This paper cites No explanation, no reformulated question, no analysis—only the JSON structure.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following No explanation, no reformulated question, no analysis—only the JSON structure

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.256221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.256221Z digest=sha256:78905bf4579f023b946f6a5406a3e5a478ee80b9de4a004fa61ea6438819017b

Observation 1e7f2873-a75e-42d1-8e88-bdb640d3c879 · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.274838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.274838Z digest=sha256:6147e6e9342d8b69b7d7469d20c8a159c7d34e559a2f9002efe6935605803779

Observation bf2fcc35-c469-47f5-824d-79fce6f97408 · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:14:36.865813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T05:14:36.314901Z digest=sha256:4f6bdc2abff871caa833bc94e7bcb8dcb88a2ff7c2745d84546a367fec1e2d7a

Observation b66d7156-b68c-4fe7-81ac-19e21775824d · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:14:36.759569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T05:14:36.366593Z digest=sha256:9cb0600a9ba75aa2710c2a9287b46f0c7109ed31b515d9fb24b1a31bfb1dbd3c

Observation e9c347f7-e45b-4334-83c9-30fd4bc1f17a · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:14:36.729126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T05:14:36.390194Z digest=sha256:8ddefefd119b187e71d23b85a19b8565d970b0c8fb708ea9b9a3e66295ed64b9

Observation b174ff10-11e9-43f2-ad0e-d195c6ee65df · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:14:36.706982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T05:14:36.417236Z digest=sha256:b23bca1b019e2803acd13b5cca8e27219a0e50ddaa311533cd63aa8eeda0dec7

Observation 2e865080-540f-4fbe-8c8a-f2986f6aa140 · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:14:36.671862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T05:14:36.455905Z digest=sha256:cece5dc7e98d2d3ca952926e1b4a81189714fa3666269bd943bf43c9a5c3cb65

Observation 3bae2fd9-3e61-43c5-9537-737eb96254dc · outbound

This paper cites You are a meticulous assistant who precisely adheres to all explicit and implicit constraints in user instructions.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following You are a meticulous assistant who precisely adheres to all explicit and implicit constraints in user instructions

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:36.788801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T05:14:36.344774Z digest=sha256:07e7b9481d8b118d00f3867e9a415bda2f8194615f0d57572d57640178dd2da9

Observation 26ba5337-378b-4f70-81c4-be680506db97 · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:14:36.632653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T05:14:36.465633Z digest=sha256:d3d7aa58745b9d9b2289e69e2928486b44e072179eaff58c60dd8ba04640cda4

Observation bbec9262-9820-4269-a73a-bdcc7c5e36f2 · outbound

This paper cites Whisker’s Quest.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Whisker’s Quest

Reference 27

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T05:14:36.609633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T05:14:36.480639Z digest=sha256:8e897844d6893d61e9431dfba8bba6adecab4cd708bf2633365e463e256cf098

Observation de110678-2512-4b55-bcfb-39ab272fa84a · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.243055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T05:14:35.823310Z digest=sha256:ac1c95d51de7b279e5d413f01143b31e87182d5034837d1299c337796afeb827

Observation 8919734b-a4e7-414d-a5ab-e7bf419868f0 · outbound

This paper cites In �������� �� ��� ����������� ��� ������������� ������������ ��� ����, pages 18632–18702.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following In �������� �� ��� ����������� ��� ������������� ������������ ��� ����, pages 18632–18702

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.259997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T05:14:35.700324Z digest=sha256:9c13876927c19abb5c9628e3e01d625f2cc3ecaf799e91986546e63af37ab9c5

Pith citing papers

No inbound Pith citation observations are available.