Pith. sign in

Paper Citation Record · LEDGER

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following

As of 12 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2508.02150.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.02150 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:14:36.480639Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved14
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 30ef7ae3-2cc4-46d8-8669-388d24235a35 · outbound

This paper cites To make the instructions more complex, I want you to identify and return five atomic constraints that can be added to the seed question.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following To make the instructions more complex, I want you to identify and return five atomic constraints that can be added to the seed question

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.131823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.131823Z digest=sha256:6c09debd339480772a3c8443a1281ecf90672a53a6f10456f639f76a0e0ed242

Observation 36fea6dd-fb2d-48b9-9067-968e99c5cdf4 · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.168195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.168195Z digest=sha256:5a5415d0cb52b5bd55b833810b4fd13da85b2c31cc13c5e44e5d478c732eb2ac

Observation 88453eb8-30b2-4c03-8bcd-bcf6610ad178 · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.226387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:14:35.964825Z digest=sha256:9ba605937643ce61c5e4f07f78fe7c5e38fcff7da9baccb3fc0ada3f8f6a3429

Observation 070800b4-ace1-497f-b440-6ef268daf33c · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.208220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:14:35.985981Z digest=sha256:383c23176da4f91a2a6e5dac63aaf0525260c0f623a1d720a87ee7d2f1b5eb3d

Observation 9f26d81b-e226-42fa-809b-e6c91f2e289c · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.190858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:14:36.009417Z digest=sha256:2097befed4bcf4d2e57a355866ae754577a31fc4b885ca15520ecd1926841653

Observation a037e7a0-6edb-40ab-9a61-51cae0ebd1bf · outbound

This paper cites Step-by-Step Mastery: Enhancing Soft Constraint Following Ability of Large Language Models.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Step-by-Step Mastery: Enhancing Soft Constraint Following Ability of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.026653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.026653Z digest=sha256:f32b610888aef2aab6000121f7804560223e1471fb18c71e1cd772d13815e3fe

Observation 67020480-fe39-4fbd-813e-01d0dd2054e7 · outbound

This paper cites 5-thinking: Advancing superb rea- soning models with reinforcement learning.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following 5-thinking: Advancing superb rea- soning models with reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.169330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:14:36.032798Z digest=sha256:b8b3873cfdf43ff7b29a003404933b03226c02ee7412cbb44b98b50f861189c4

Observation 63de9e48-05d5-465a-9c44-e346864dee00 · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.149630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:14:36.044328Z digest=sha256:b357e5576cc53b5e28d1882560ba3787d16ff444a21f6ba9fa1e9ab21880cabf

Observation 2c1eee46-b55f-402a-bea0-bc133bc518f5 · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.127153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:14:36.060611Z digest=sha256:fad1879c8fd4d1bccd4cee7d48cc7e9868a486cabfb0d8b2639f0a063f7e039a

Observation c35bf93f-2f11-48ec-8320-bd251713d45a · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.100189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:14:36.086975Z digest=sha256:a9dd2300b1c3cd19edaf13f02809d424b46162d5f333877e1e10c3ba36939049

Observation f10955ec-320c-4dcf-88b3-31dc5b3ad4b1 · outbound

This paper cites joy,” “anger,.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following joy,” “anger,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.053649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:14:36.104630Z digest=sha256:82e447dfce83bb2717361a4dd13bee65ceaacbf9bbde0f5bc9498b2ec33c3eba

Observation 80807db4-1b09-4088-9b74-8b1f1dce4a6e · outbound

This paper cites You may choose one or more constraints from the list or propose new ones if needed.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following You may choose one or more constraints from the list or propose new ones if needed

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.185917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.185917Z digest=sha256:fb7ec6d7b1b05da51d1a143543ebf984b72cf7ed3ce5da38a3b7cf59af66130e

Observation f0fe6964-45f9-4668-b253-2224dcc20fbc · outbound

This paper cites Your task is only to generate new constraints that can be added to it.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Your task is only to generate new constraints that can be added to it

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.214321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.214321Z digest=sha256:6f6b3bd2164857f95fd281fc85b91887e37b1d176661bec36ec935a6b48a0a92

Observation f92f8690-123b-4b0e-be3b-87b7688341dd · outbound

This paper cites c1": "<first constraint>.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following c1": "<first constraint>

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.232092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.232092Z digest=sha256:7762c874938ef8c9c4cdd3ba63f021b5960d4e8dc3c7a0fb96b1329d1c98ec13

Observation 9361bf8c-c861-48cb-bb46-f8bf52b01c33 · outbound

This paper cites No explanation, no reformulated question, no analysis—only the JSON structure.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following No explanation, no reformulated question, no analysis—only the JSON structure

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.256221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.256221Z digest=sha256:78905bf4579f023b946f6a5406a3e5a478ee80b9de4a004fa61ea6438819017b

Observation 1e7f2873-a75e-42d1-8e88-bdb640d3c879 · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:36.274838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:14:36.274838Z digest=sha256:6147e6e9342d8b69b7d7469d20c8a159c7d34e559a2f9002efe6935605803779

Observation bf2fcc35-c469-47f5-824d-79fce6f97408 · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:14:36.865813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:14:36.314901Z digest=sha256:766fdbb403073a1cb7d41867a59377a52965a00d1c3add6d93d2c487dd8ddad6

Observation b66d7156-b68c-4fe7-81ac-19e21775824d · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:14:36.759569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:14:36.366593Z digest=sha256:b81d5297ba0ab12201aaa37b509173b84445cc3e527c8e46457d258964f36fae

Observation e9c347f7-e45b-4334-83c9-30fd4bc1f17a · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:14:36.729126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:14:36.390194Z digest=sha256:4733ac7b3ffbc52ccce7d3ea3958127b022d7a1cbc9a1bb1a0d2f275092cade0

Observation b174ff10-11e9-43f2-ad0e-d195c6ee65df · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:14:36.706982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:14:36.417236Z digest=sha256:fad16a54bcf85d82bc6b15dff5ae2ec9a26d32c9a7063513d126e5552206d3e9

Observation 2e865080-540f-4fbe-8c8a-f2986f6aa140 · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:14:36.671862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:14:36.455905Z digest=sha256:ecdd899c5b8801531adb5197dcf14b3da465cd1c82fdd26ed205e6a303a5dcb2

Observation 3bae2fd9-3e61-43c5-9537-737eb96254dc · outbound

This paper cites You are a meticulous assistant who precisely adheres to all explicit and implicit constraints in user instructions.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following You are a meticulous assistant who precisely adheres to all explicit and implicit constraints in user instructions

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:36.788801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:14:36.344774Z digest=sha256:4a8dd3c29b97eb23738043aab0401c525b594cc8d54b1e02a36c3019d17a6842

Observation 26ba5337-378b-4f70-81c4-be680506db97 · outbound

This paper cites an unresolved cited work.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:14:36.632653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:14:36.465633Z digest=sha256:dc1890e71a9c406d6d7aae7cca808dc0da905136eec2065439d4c6136b3a8167

Observation bbec9262-9820-4269-a73a-bdcc7c5e36f2 · outbound

This paper cites Whisker’s Quest.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following Whisker’s Quest

Reference 27

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T05:14:36.609633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:14:36.480639Z digest=sha256:21ac1c75437eae288d734163223e9c66ac0e0d52ffc93f2881f039cc3bddbfbe

Observation de110678-2512-4b55-bcfb-39ab272fa84a · outbound

This paper cites ����� �������� ����������������.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following ����� �������� ����������������

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.243055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:14:35.823310Z digest=sha256:11dd177c7a8d762f17cdb4a98a0662e588f95e0c310f8796de13e9c84840ada4

Observation 8919734b-a4e7-414d-a5ab-e7bf419868f0 · outbound

This paper cites In �������� �� ��� ����������� ��� ������������� ������������ ��� ����, pages 18632–18702.

Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following In �������� �� ��� ����������� ��� ������������� ������������ ��� ����, pages 18632–18702

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:14:37.259997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:14:35.700324Z digest=sha256:25213f6cfa5b21d67d4b92a0ffb556e33d97f9d952ddb55372469ae23f4e4c7b

Pith citing papers

No inbound Pith citation observations are available.