Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T04:30:30.646103Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2607.09786.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T04:30:30.646103Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T15:22:55.656489Z
A source-named dated measurement, never combined with another source.
Source: cited_works
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 97701c45-53b2-43a6-a8f8-5ae87c201e40 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28d8df4c-1a6e-480d-ac42-8561ae3cef0a · outbound
Length Penalties Make Chain-of-Thought Less Monitorable MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82ffa36e-827e-434b-853f-66e4ade443c8 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 633d15ed-7598-458d-a62d-015ff92eed21 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable CoT red-handed: Stress testing chain-of-thought monitoring
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9efa272a-b3f8-49d5-94f7-f07f509e066e · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Training language models to reason efficiently
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd77c0cd-8eb6-41d3-916a-1de33f5ed9d8 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3efc6f17-67e3-4b1e-a489-4700e721431e · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a03e0ae6-eeb5-4032-b61f-e9d7eee66b59 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Reasoning Models Don't Always Say What They Think
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80149489-ac33-4a51-946b-fc884ff6ec92 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Are DeepSeek R1 And Other Reasoning Models More Faithful?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef30e1dc-c058-4e94-95c1-46f0948119e0 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Stable Reinforcement Learning for Efficient Reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a9d687-c31d-4438-a8e6-4c06cbe6c0d5 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03f629ac-98bd-4d01-845c-e27236735a97 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e5448be-75aa-41b4-bbf9-162edb9f3169 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Guan, Miles Wang, Micah Carroll, Zehao Dou, Annie Y
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0f799cd-edf0-4487-9baa-cbd49d555b6b · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Verbalizable representations form a global workspace in language models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 091de329-d6a2-4715-ac48-8e198bdf4189 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b59b549f-2dca-42ab-b805-c1971cb80f21 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38c6ebe6-faf7-4b68-b4ab-2d5137eefa4b · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Zimmermann, and Rohin Shah
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9397e7b6-b61a-44b6-82e0-85d21f18e456 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Large language models are zero-shot reasoners
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cd93958-929e-4098-bddf-200e0f8b80c0 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cd7ce18-472e-4807-936b-929d21c8f29b · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Gonzalez, Hao Zhang, and Ion Stoica
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec706780-0a3f-4c6b-ba0d-fe3d67fbab7e · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Measuring Faithfulness in Chain-of-Thought Reasoning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ed3d504-ebaa-4019-a849-091634db5772 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable DeepCompress : A dual reward strategy for dynamically exploring and compressing reasoning chains
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 508ab95b-bdf7-40d5-8508-8902496dfd05 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 212f73b6-8632-4c75-9ee7-7723f0b457ce · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Understanding R1-Zero-Like Training: A Critical Perspective
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 140526a6-46cb-4edf-a942-b65a7b27530a · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Faithful chain-of-thought reasoning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90fa1903-332b-4305-bfb6-29cebb2814ae · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Reasoning under pressure: How do training incentives influence chain-of-thought monitorability? arXiv preprint arXiv:2512.00218, 2025
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6f18e25-42ec-49db-8372-e96bb50d6913 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Chain-of-thought obfuscation learned from output supervision can generalise to unseen tasks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62774abb-7bfc-4014-a3cf-75eed839ff2d · outbound
Length Penalties Make Chain-of-Thought Less Monitorable s1: Simple test-time scaling
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 611fe834-32c6-4fbe-811d-fefe8563d370 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 153b26d7-159a-4171-8320-f9bffd4cdde9 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable OpenAI o1 System Card
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 931c8414-328a-4505-96ac-e500acd05c13 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bc338ce-45be-4bb1-9831-6981065e318a · outbound
Length Penalties Make Chain-of-Thought Less Monitorable HybridFlow : A flexible and efficient RLHF framework
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 026fc710-db73-4738-9a63-532022000159 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5384e5c-ef90-4c86-a496-3764e3e57b74 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable MonitorBench : A comprehensive benchmark for chain-of-thought monitorability in large language models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc03df48-b946-4c2e-ab50-44787f31df68 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable MMLU-Pro : A more robust and challenging multi-task language understanding benchmark
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da070da8-6c13-4707-8003-e11d4dcbaa30 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Le, and Denny Zhou
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c6f5859-2199-4f0b-a129-13d2f68e69a2 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdf44ff4-68c5-438c-91eb-24e47b2b6d30 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable Qwen3 Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a92e59f8-8c6d-4e93-9bd3-3c8af6441ff2 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable ShorterBetter : Guiding reasoning models to find optimal inference length for efficient reasoning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89c28c4e-e42b-4702-8416-81ffd83d3c61 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adf9c2bf-5861-4f78-89d7-cd8ccaf75bb8 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 468f8bbb-9c8c-448c-accf-13bc3d193157 · outbound
Length Penalties Make Chain-of-Thought Less Monitorable MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c431f89-a143-491d-b303-ff92ff1eb2f6 · inbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Length Penalties Make Chain-of-Thought Less Monitorable
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.