Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:43:00.193690Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 0 inbound Pith citation observations for arXiv:2607.17823.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:43:00.193690Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
72 of 72 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 625e53fa-f725-4c3a-a903-19d02bb5b830 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Max k-armed bandit: On the extremehunter algorithm and beyond
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 07c00da6-1cb9-440b-9610-ffd8202ec4ab · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Model-based reinforcement learning with a generative model is minimax optimal
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47162232-2a07-4dee-a51e-1d65eeabdc70 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Rewarding behaviors
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f49fc2c4-5a51-461c-98c8-991e1fc12ad8 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning The best of n worlds: Aligning reinforcement learning with best-of-n sampling via max@ k optimisation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fd909a1-f7e5-4b00-84f3-c6bd2d5cb84b · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Optimal rates for feasible payoff set estimation in games
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation dc095108-3c1e-42c3-9fbc-567fd533d0f8 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Regret bounds for risk-sensitive reinforcement learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1b93c90e-f45a-484d-a499-3ff7ac0d225a · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Efficient Algorithms for Extreme Bandits
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1eb87b60-46ee-479d-82c1-a39110c44967 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning A distributional perspective on reinforcement learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 922f4659-6ee2-4e0e-beb0-e2f7211de686 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Distributional reinforcement learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 562fbd03-df68-4cdf-8a0e-3e55637ccec0 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Graph of thoughts: Solving elaborate problems with large language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 76f0cb61-4bc6-43ed-8015-1a75c354c82b · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Concentration inequalities
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bb58fe7f-850f-4322-bbcd-dc3597d8f089 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Ltlf/ldlf non-markovian rewards
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9fdfe351-46f4-495d-8ed8-4d466b2f21d0 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a285c5d-5d23-4918-b36b-cabf9dd18427 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Extreme bandits
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bfeca6e3-9d90-4053-8ae5-038570ccd3c0 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 40355f6a-ae50-479c-bbf6-c3036bc00e10 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Evaluating Large Language Models Trained on Code
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2671270e-a7bf-4526-b5e5-6280f75f798a · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b5455b1-d8ff-47a6-8215-6104926b509b · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model? Advances in Neural Information Processing Systems, 38: 0 57654--57689, 2026 b
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8554e31f-2321-4536-b3c6-5be9d4c27a01 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Robust reinforcement learning with general utility
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a2533bf9-107f-4c86-9d69-6d7ce6a91b78 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Algorithms for cvar optimization in mdps
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation aa071919-3714-43e4-b0ad-7743923b2799 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Inference-aware fine-tuning for best-of-n sampling in large language models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 629f9abf-1e80-4559-8c7a-5bf323b82482 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Goedel-Architect: Streamlining Formal Theorem Proving with Blueprint Generation and Refinement
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8c2f4e4c-eadd-4214-ace6-6e0ff1a60848 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Training Verifiers to Solve Math Word Problems
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e929eea7-8113-492b-9198-5cd1bc497a9d · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a57f932b-0639-4867-bf20-391465791a64 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Distributional reinforcement learning with quantile regression
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c1b5b99-2c8c-4610-8596-d6b86f17a714 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Sample complexity of episodic fixed-horizon reinforcement learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 71c59548-95e2-44ce-b301-6399e4e4a8ed · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Chain-of-verification reduces hallucination in large language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0b9aa258-404a-4eb5-9628-71baec40fa45 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Episodic reinforcement learning in finite mdps: Minimax lower bounds revisited
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6cdc281-69f9-4117-b890-a96d3ca60546 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Risk-sensitive reinforcement learning: Near-optimal risk-sample tradeoff in regret
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 59853193-3c51-4bf3-8d76-b54f8961256d · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Reinforcement learning with non-markovian rewards
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6f889ea7-c8bb-4e4e-9bbe-35fafa7a31b9 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Explore first, exploit next: The true shape of regret in bandit problems
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7c2ef413-981c-447a-876a-4ad6be31dd22 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 668552ad-4f7a-4bb7-9d75-c4183ec9e73e · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f55e9cc-c7e8-479d-9d3d-9372759007d4 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Provably efficient maximum entropy exploration
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 17730fc1-6c16-4a39-a6f1-762ade99ac0d · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Reward machines: Exploiting reward function structure in reinforcement learning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06732d28-280b-4403-8bbe-0417d21f2aa0 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Learning to Correct: Calibrated Reinforcement Learning for Multi-Attempt Chain-of-Thought
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ab01f87b-de69-421e-83cf-e32cf398b70b · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning OpenAI o1 System Card
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90a5d3d3-1c0b-4473-ac97-e58517198755 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Planning in markov decision processes with gap-dependent sample complexity
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22699b3b-7ddc-46f0-811d-a0fbd40a406e · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Policy Gradient for Reinforcement Learning with General Utilities
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d3e3b1d-0be1-4559-81ea-8cd12651222c · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcd66a85-f262-49b4-9429-9fe2ea430745 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Breaking the sample size barrier in model-based reinforcement learning with a generative model
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 81685593-ebe5-4b03-9595-da5efc84ff48 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Competition-level code generation with alphacode
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91a9539b-b2f6-4ff7-a710-2cd01fbf50ff · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Let's verify step by step
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation aeb5fe5c-079c-4c8b-92b6-50f6800bb1dc · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Self-refine: Iterative refinement with self-feedback
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2ebbdfa3-1951-496a-8bc0-acc6d8412474 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Challenging common assumptions in convex reinforcement learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8adae96a-6142-4b32-9a7b-1b874e16fdcb · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Convex reinforcement learning in finite trials
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation be78362e-7f03-4f6c-b83f-723801e46e75 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning No regret bound for extreme bandits
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 44e5bb33-067d-47de-b63c-0a47e274b126 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Markov decision processes
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e77ccf59-7b37-4a5f-a864-b189e2736d86 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Near-minimax-optimal distributional reinforcement learning with a generative model
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 16e479be-4ba3-421e-9332-4a85203e27d8 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Reflexion: Language agents with verbal reinforcement learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdf2dc89-efc2-4602-84db-0e7dfcdc4a67 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Near-Optimal Time and Sample Complexities for Solving Discounted Markov Decision Process with a Generative Model
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 194ae738-564d-4436-b827-29d982c6e385 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Scaling llm test-time compute optimally can be more effective than scaling parameters for reasoning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ce10b7e7-ec70-4dd4-b0da-719367611f40 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning On Advantage Estimates for Max@K Policy Gradients
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 29b53253-1846-4525-9ed1-32863b156d96 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Optimizing language models for inference time objectives using reinforcement learning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6c52c959-462b-42b7-8118-b8ff9fddab3f · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Finite-Time Regret Analysis of Retry-Aware Bandits
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3287bdd-5ede-4ab7-bf89-b864090bc05b · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Model-Based Reinforcement Learning in Discrete-Action Non-Markovian Reward Decision Processes
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f57f8433-a6f2-46bc-b670-d55b6d5b0f6f · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Advancing Mathematics Research with AI-Driven Formal Proof Search
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5459578-ee59-4c28-bdca-e4a58431d1c1 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Recursive self-aggregation unlocks deep thinking in large language models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d1235b0-03ef-4dcb-9484-78253068f15f · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Pass@ k policy optimization: Solving harder reinforcement learning problems
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1d36c955-e4b9-4b5a-8877-6cd7884a8c75 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Sample-efficient reinforcement learning for linearly-parameterized mdps with a generative model
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 08746916-6130-4183-a306-b7442a7ecd6f · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Near-minimax-optimal risk-sensitive reinforcement learning with cvar
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d44bf767-cedd-4e67-a821-4dd80d4ce061 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e438ab1-7139-48ec-891d-ff09a2f5fbb6 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Risk-sensitive Markov Decision Process and Learning under General Utility Functions
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaaaf2ff-7240-4b6f-b5d1-0bd167f1baa2 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dedb8a22-c4de-48e0-a547-c5498c0a8718 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fe2c7f4-2c1a-4375-91a7-89e3fa68afdd · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Tree of thoughts: Deliberate problem solving with large language models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6238607-8f5f-46ca-8713-11438ee98bd0 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Reward is enough for convex mdps
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 70cf53ee-3fe7-443e-bb61-d43a2682383d · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Star: Bootstrapping reasoning with reasoning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 515c704c-2b96-4e6b-bef0-c23004832a7e · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Variational policy gradient method for reinforcement learning with general utilities
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7d0718c-3cb6-4fc9-aea5-be94c0641cd8 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Estimation and inference in distributional reinforcement learning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9851e4c7-a1ef-4e60-8eaa-f776dfdf292e · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Beyond markovian: Reflective exploration via bayes-adaptive rl for llm reasoning
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0534ec8d-deb2-4ca9-84d1-a6cf374ac558 · outbound
Theoretical Foundations of $\max$@$k$ Reinforcement Learning Settling the sample complexity of online reinforcement learning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
No inbound Pith citation observations are available.