Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T05:29:30.210574Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2602.02150.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T05:29:30.210574Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-08T17:16:00.499779Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T17:46:07.008703Z
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ba16eefc-7a59-4215-85c1-4d67443a9cae · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Qwen Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbcb182c-7dc5-4836-a25d-eb073cf5cd7c · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Deep Think with Confidence
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f61ba3da-5e8d-4a0f-8c4f-55b8647b43e0 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning The Llama 3 Herd of Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8ccd421-5d78-4159-8c09-acdcfc0cf531 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bd47e3f-aed5-468e-9987-8a42d82fd06c · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Qwen2.5-Coder Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c513edd-9ded-498c-82f1-fb149569d5f3 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Tree search for llm agent reinforcement learning.arXiv preprint arXiv:2509.21240,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 534fd71e-c0aa-46e9-b074-44b0c6f12329 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning DeepSeek-V3 Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f38d4cf-570a-4a5a-a3c5-9710540765b5 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a7e2142-1ce1-41da-823f-11d16319bdfc · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d828de3e-b917-465e-92b1-fb455682e8fd · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Maximizing Confidence Alone Improves Reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56823190-2c6c-4186-b445-d26eaacfcf89 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70d81845-3612-492c-aa86-26df9a139765 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98f0c4b1-5a95-40a5-905c-070074d88b6e · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Qwen2 Technical Report
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be0ddac4-b876-4917-a912-8282b807067f · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Qwen3 Technical Report
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6515c923-53b5-4a8f-971c-50a0dc309ab1 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 767c0429-5729-4a53-80d7-008c6c71924e · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Learning to Reason without External Rewards
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bbb9519-929b-45cd-aadb-749ab98959b3 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Evolving language models without labels: Majority drives selection, novelty promotes variation.arXiv preprint arXiv:2509.15194,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1e7b6ed-41c7-4e17-9ee1-8df0b4d8bf23 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning TTRL: Test-Time Reinforcement Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e026cf04-701b-4003-829b-c6556f678359 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c298c9a4-1ef8-4d2f-a9ba-3318340809c1 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14a74240-5b58-45d8-9724-7e83b61f1252 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83e8a216-f516-40b9-8aad-a20078ad7f60 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Crafting papers on machine learning
Reference 1996
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf67b164-38b9-4af2-82bd-96ae7384a5f2 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Bapo: Stabilizing off-policy reinforcement learning for llms via balanced policy optimization with adaptive clipping.arXiv preprint arXiv:2510.18927,
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a4afd59-f7b6-4454-ba28-4157ea2649f4 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Can large reasoning models self-train? arXiv preprint arXiv:2505.21444,
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d58a01f-0df4-4816-a91b-31e1f344a4c5 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Unsupervised post-training for multi-modal llm reasoning via grpo.arXiv preprint arXiv:2505.22453,
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 223c20ca-91da-4c58-9739-b077b2209246 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da3fd933-8924-4d36-93e0-bd608cfddeef · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Qwen3-VL Technical Report
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aff2c1da-0680-408a-84bc-e783ab657f76 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17abc89d-d09e-4d49-9dfa-6d433e03d239 · outbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Agentic entropy-balanced policy optimization.arXiv preprint arXiv:2510.14545,
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 931ec558-fc4e-480d-909b-3cbad9cbb1ca · inbound
Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.