Pith. sign in

Paper Citation Record · LEDGER

Length Penalties Make Chain-of-Thought Less Monitorable

As of 7 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2607.09786.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.09786 v3

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:30:30.646103Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T15:22:55.656489Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 97701c45-53b2-43a6-a8f8-5ae87c201e40 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Length Penalties Make Chain-of-Thought Less Monitorable L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:26.557823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:26.557823Z digest=sha256:59ced25cce748740cc5c6c17af0eee8aa2b73126b1a5d3042ee0e17d63fa52d6

Observation 28d8df4c-1a6e-480d-ac42-8561ae3cef0a · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

Length Penalties Make Chain-of-Thought Less Monitorable MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:26.666683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:26.666683Z digest=sha256:93452adf93fbdc5296d9f2886fe0d6f3df2bdbf9ccd73694c2e96a36dc8a003b

Observation 82ffa36e-827e-434b-853f-66e4ade443c8 · outbound

This paper cites Chain-of-Thought Reasoning In The Wild Is Not Always Faithful.

Length Penalties Make Chain-of-Thought Less Monitorable Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:26.776436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:26.776436Z digest=sha256:33b6d50410d9fb6a20f649be92eb2ec8cb176c7b236635c5d73ec3b9cc8125d2

Observation 633d15ed-7598-458d-a62d-015ff92eed21 · outbound

This paper cites CoT red-handed: Stress testing chain-of-thought monitoring.

Length Penalties Make Chain-of-Thought Less Monitorable CoT red-handed: Stress testing chain-of-thought monitoring

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:26.891885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:26.891885Z digest=sha256:7e1c6260b2d72348d982499fcfad51e9217a7b99ddee617977f3b16516ee6203

Observation 9efa272a-b3f8-49d5-94f7-f07f509e066e · outbound

This paper cites Training language models to reason efficiently.

Length Penalties Make Chain-of-Thought Less Monitorable Training language models to reason efficiently

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:27.005827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:27.005827Z digest=sha256:21098c6921f1fc10a69f0f9081dc112834cc33d2841410dae19bb0c4532116b4

Observation bd77c0cd-8eb6-41d3-916a-1de33f5ed9d8 · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

Length Penalties Make Chain-of-Thought Less Monitorable Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:27.115646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:27.115646Z digest=sha256:8907fa5f23882cb2704e3ad6a32abaf8cbb2abd7b388059f0b4df7becd1d645e

Observation 3efc6f17-67e3-4b1e-a489-4700e721431e · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Length Penalties Make Chain-of-Thought Less Monitorable Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:27.191039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:27.191039Z digest=sha256:f2f1e111cd97fb6d4eb770e92dd820a1ad0c540ecfc92c54defa610556cd318b

Observation a03e0ae6-eeb5-4032-b61f-e9d7eee66b59 · outbound

This paper cites Reasoning Models Don't Always Say What They Think.

Length Penalties Make Chain-of-Thought Less Monitorable Reasoning Models Don't Always Say What They Think

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:27.262689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:27.262689Z digest=sha256:b11e53dae45dc16f42e519430c365b337c80615b6be39712238c2bbf57348428

Observation 80149489-ac33-4a51-946b-fc884ff6ec92 · outbound

This paper cites Are DeepSeek R1 And Other Reasoning Models More Faithful?.

Length Penalties Make Chain-of-Thought Less Monitorable Are DeepSeek R1 And Other Reasoning Models More Faithful?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:27.374607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:27.374607Z digest=sha256:bec077ada78c57b518cd9e644ff279f8ff860158ebe6ece4abee77f912f6cf74

Observation ef30e1dc-c058-4e94-95c1-46f0948119e0 · outbound

This paper cites Stable Reinforcement Learning for Efficient Reasoning.

Length Penalties Make Chain-of-Thought Less Monitorable Stable Reinforcement Learning for Efficient Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:27.449308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:27.449308Z digest=sha256:becfa077003eaaf8881e9e1c767dc7df262754a9fcecfd0b2a0b57f11da6681e

Observation 41a9d687-c31d-4438-a8e6-4c06cbe6c0d5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Length Penalties Make Chain-of-Thought Less Monitorable DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:27.554149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:27.554149Z digest=sha256:7c20d8232aed5f25192c226cb60e868b650ef2576042c52980e2a4278744739d

Observation 03f629ac-98bd-4d01-845c-e27236735a97 · outbound

This paper cites When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors.

Length Penalties Make Chain-of-Thought Less Monitorable When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:27.658275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:27.658275Z digest=sha256:3b37e101f6be327fd3f699e44e1bbc4697d8d7faaaf747989fb5189ade0da533

Observation 5e5448be-75aa-41b4-bbf9-162edb9f3169 · outbound

This paper cites Guan, Miles Wang, Micah Carroll, Zehao Dou, Annie Y.

Length Penalties Make Chain-of-Thought Less Monitorable Guan, Miles Wang, Micah Carroll, Zehao Dou, Annie Y

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:27.763139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:27.763139Z digest=sha256:567c576d85dc739ddc2e9f4365ac03a99cb11522e68913f91b6f37222f289383

Observation d0f799cd-edf0-4487-9baa-cbd49d555b6b · outbound

This paper cites Verbalizable representations form a global workspace in language models.

Length Penalties Make Chain-of-Thought Less Monitorable Verbalizable representations form a global workspace in language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:27.863626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:27.863626Z digest=sha256:2b96f0db78aaf337105035b6f22a9600c05562230d3e44ac31efb09920cb562c

Observation 091de329-d6a2-4715-ac48-8e198bdf4189 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Length Penalties Make Chain-of-Thought Less Monitorable ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:27.962588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:27.962588Z digest=sha256:4b547de40b89e400970a5ae666cca169187932fde7d97085ad5e5fa8b46b7cea

Observation b59b549f-2dca-42ab-b805-c1971cb80f21 · outbound

This paper cites What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams.

Length Penalties Make Chain-of-Thought Less Monitorable What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:28.064643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:28.064643Z digest=sha256:d0d93be5b5554e8f722b7a3e91641fdc8386e240c8abab73c898fbed9fc9fc8f

Observation 38c6ebe6-faf7-4b68-b4ab-2d5137eefa4b · outbound

This paper cites Zimmermann, and Rohin Shah.

Length Penalties Make Chain-of-Thought Less Monitorable Zimmermann, and Rohin Shah

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:28.163403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:28.163403Z digest=sha256:5c4c5988ec58c2bea79a80f86a94f4f400f1bfd246fa23f2cfc161cba2fb57e7

Observation 9397e7b6-b61a-44b6-82e0-85d21f18e456 · outbound

This paper cites Large language models are zero-shot reasoners.

Length Penalties Make Chain-of-Thought Less Monitorable Large language models are zero-shot reasoners

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:28.291550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:28.291550Z digest=sha256:82e3b0ce213c944fc87b8fb9b1d129dec4c38e9c5f787a65f7206c6bd5a0fa63

Observation 4cd93958-929e-4098-bddf-200e0f8b80c0 · outbound

This paper cites Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety.

Length Penalties Make Chain-of-Thought Less Monitorable Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:28.364330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:28.364330Z digest=sha256:26ee343cdb8e79ede9aab48f6130975acb0ba1e6e6029e19fd8c3544db3f5269

Observation 4cd7ce18-472e-4807-936b-929d21c8f29b · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Length Penalties Make Chain-of-Thought Less Monitorable Gonzalez, Hao Zhang, and Ion Stoica

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:28.450839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:28.450839Z digest=sha256:458a69d686013c6cd547029a4e9460d7d84da60c2808c5ca6289121c526d54fb

Observation ec706780-0a3f-4c6b-ba0d-fe3d67fbab7e · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

Length Penalties Make Chain-of-Thought Less Monitorable Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:28.550839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:28.550839Z digest=sha256:77783c800d9ef00eb873fb5dd07dc2ff3e3befb14fac1b041fe4986db1c457c3

Observation 8ed3d504-ebaa-4019-a849-091634db5772 · outbound

This paper cites DeepCompress : A dual reward strategy for dynamically exploring and compressing reasoning chains.

Length Penalties Make Chain-of-Thought Less Monitorable DeepCompress : A dual reward strategy for dynamically exploring and compressing reasoning chains

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:28.679145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:28.679145Z digest=sha256:f2b47da310125a93f58655d2e1ee48febcf095c0a8e5c454f18dc37a2f8be869

Observation 508ab95b-bdf7-40d5-8508-8902496dfd05 · outbound

This paper cites an unresolved cited work.

Length Penalties Make Chain-of-Thought Less Monitorable Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:28.775772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:28.775772Z digest=sha256:ecc67c1f6cb27d887fe2fa41c84b8af565850b071b6f17e8a1c7bb0a03477c16

Observation 212f73b6-8632-4c75-9ee7-7723f0b457ce · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Length Penalties Make Chain-of-Thought Less Monitorable Understanding R1-Zero-Like Training: A Critical Perspective

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:28.871469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:28.871469Z digest=sha256:51b3a5f38bcad6263259681153c26d9a5b0ae89ce20e056b5119ae8c84dc3336

Observation 140526a6-46cb-4edf-a942-b65a7b27530a · outbound

This paper cites Faithful chain-of-thought reasoning.

Length Penalties Make Chain-of-Thought Less Monitorable Faithful chain-of-thought reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:29.011539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:29.011539Z digest=sha256:4ee8ae82d342b5add0f8659787800cb16a8a1e088be2802a125c3152999e959c

Observation 90fa1903-332b-4305-bfb6-29cebb2814ae · outbound

This paper cites Reasoning under pressure: How do training incentives influence chain-of-thought monitorability? arXiv preprint arXiv:2512.00218, 2025.

Length Penalties Make Chain-of-Thought Less Monitorable Reasoning under pressure: How do training incentives influence chain-of-thought monitorability? arXiv preprint arXiv:2512.00218, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:29.111462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:29.111462Z digest=sha256:9ef2992a77723eeaa030b74b6be1d5b86c341f50664cdfc13e41e3251c2ee239

Observation f6f18e25-42ec-49db-8372-e96bb50d6913 · outbound

This paper cites Chain-of-thought obfuscation learned from output supervision can generalise to unseen tasks.

Length Penalties Make Chain-of-Thought Less Monitorable Chain-of-thought obfuscation learned from output supervision can generalise to unseen tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:29.213030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:29.213030Z digest=sha256:82901bd8a863b78c1f75255a48487e9aa2666911075deed71f6ad995cd84e017

Observation 62774abb-7bfc-4014-a3cf-75eed839ff2d · outbound

This paper cites s1: Simple test-time scaling.

Length Penalties Make Chain-of-Thought Less Monitorable s1: Simple test-time scaling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:29.315426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:29.315426Z digest=sha256:6fad50f93363be5158de80fe40269c3d06fc363d5476f41436b7401f641c98f0

Observation 611fe834-32c6-4fbe-811d-fefe8563d370 · outbound

This paper cites NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model.

Length Penalties Make Chain-of-Thought Less Monitorable NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:29.444171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:29.444171Z digest=sha256:e12a84e2540d3c6bf5ae0b924fae6ac5f347a404d6050f38260ba67a5258a6f7

Observation 153b26d7-159a-4171-8320-f9bffd4cdde9 · outbound

This paper cites OpenAI o1 System Card.

Length Penalties Make Chain-of-Thought Less Monitorable OpenAI o1 System Card

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:29.549813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:29.549813Z digest=sha256:bdf3311fe9641b3fdc4288c3d396d0fff8ac187bbc28f297f7a3bd707d6f52c8

Observation 931c8414-328a-4505-96ac-e500acd05c13 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Length Penalties Make Chain-of-Thought Less Monitorable DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:29.616907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:29.616907Z digest=sha256:05b6464264a423a50ac7924dfe4655ddd866a1f88a27b45e06c2213e9eb52300

Observation 2bc338ce-45be-4bb1-9831-6981065e318a · outbound

This paper cites HybridFlow : A flexible and efficient RLHF framework.

Length Penalties Make Chain-of-Thought Less Monitorable HybridFlow : A flexible and efficient RLHF framework

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:29.701653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:29.701653Z digest=sha256:639a362d2dd4473a0ae43ed0c6515f961b68b3e75d5bd5b9fb54b81d1777e78c

Observation 026fc710-db73-4738-9a63-532022000159 · outbound

This paper cites an unresolved cited work.

Length Penalties Make Chain-of-Thought Less Monitorable Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:29.791096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:29.791096Z digest=sha256:6148ca8c9362262864265cd7dc776c7f7a8ae8c58932a3a3aab6197c27f3d46a

Observation a5384e5c-ef90-4c86-a496-3764e3e57b74 · outbound

This paper cites MonitorBench : A comprehensive benchmark for chain-of-thought monitorability in large language models.

Length Penalties Make Chain-of-Thought Less Monitorable MonitorBench : A comprehensive benchmark for chain-of-thought monitorability in large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:29.888965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:29.888965Z digest=sha256:5680a52bed70f2e119a61a5d4f76d9f7d00e9bb15316012af64e574d4710aa9f

Observation cc03df48-b946-4c2e-ab50-44787f31df68 · outbound

This paper cites MMLU-Pro : A more robust and challenging multi-task language understanding benchmark.

Length Penalties Make Chain-of-Thought Less Monitorable MMLU-Pro : A more robust and challenging multi-task language understanding benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:29.972402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:29.972402Z digest=sha256:a0a68f5c71f7ca5b921425793ffc188aaef6d8a036e9821351e9015f5bed64c1

Observation da070da8-6c13-4707-8003-e11d4dcbaa30 · outbound

This paper cites Le, and Denny Zhou.

Length Penalties Make Chain-of-Thought Less Monitorable Le, and Denny Zhou

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:30.064037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:30.064037Z digest=sha256:72fecd509c0d02b6b8b7fba7e5d52a665239ca60ed3a7f805808c85702bc55cf

Observation 1c6f5859-2199-4f0b-a129-13d2f68e69a2 · outbound

This paper cites Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning.

Length Penalties Make Chain-of-Thought Less Monitorable Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:30.198277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:30.198277Z digest=sha256:d813eb2de247e80b674726dd0ab81f54af6e6b69df1cbced6d54d11ef925bfa8

Observation fdf44ff4-68c5-438c-91eb-24e47b2b6d30 · outbound

This paper cites Qwen3 Technical Report.

Length Penalties Make Chain-of-Thought Less Monitorable Qwen3 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:30.288713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:30.288713Z digest=sha256:c829a4e82335cc8fcd943e4e7d4e2ac54f17104d8f21be53c3d843b516eff4a9

Observation a92e59f8-8c6d-4e93-9bd3-3c8af6441ff2 · outbound

This paper cites ShorterBetter : Guiding reasoning models to find optimal inference length for efficient reasoning.

Length Penalties Make Chain-of-Thought Less Monitorable ShorterBetter : Guiding reasoning models to find optimal inference length for efficient reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:30.356878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:30.356878Z digest=sha256:227fa047175db71c427d477701a9b8d84bf71b2c8343a31c72ed17f3fe12a200

Observation 89c28c4e-e42b-4702-8416-81ffd83d3c61 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Length Penalties Make Chain-of-Thought Less Monitorable DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:30.469902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:30.469902Z digest=sha256:0f639f3ab00edf93cf40e0f624fefd68e6c8c06920c2c71e8f8b6e584e3e4b2b

Observation adf9c2bf-5861-4f78-89d7-cd8ccaf75bb8 · outbound

This paper cites ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning.

Length Penalties Make Chain-of-Thought Less Monitorable ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:30.561818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:30.561818Z digest=sha256:51eb03a02ceac90b5e269ba9d78289816301d0ad35cf90ad713a34084a06225f

Observation 468f8bbb-9c8c-448c-accf-13bc3d193157 · outbound

This paper cites MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark.

Length Penalties Make Chain-of-Thought Less Monitorable MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:30.646103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:30.646103Z digest=sha256:eed314c0991697b982abb0470931957322e2ea13b1afbfd4c0ed83bf6d4ecb51

Pith citing papers

Observation 6c431f89-a143-491d-b303-ff92ff1eb2f6 · inbound

How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models cites this paper.

How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Length Penalties Make Chain-of-Thought Less Monitorable

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T15:22:55.656489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T15:22:55.656489Z digest=sha256:e9dd99861c4e5720ca5062a57a24daf2cd69d9955fba2c8f7c750fc9629f7da5