Pith. sign in

Paper Citation Record · LEDGER

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning

As of 4 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2602.02150.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.02150 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T05:29:30.210574Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T17:16:00.499779Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T17:46:07.008703Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ba16eefc-7a59-4215-85c1-4d67443a9cae · outbound

This paper cites Qwen Technical Report.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:27.148641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:27.148641Z digest=sha256:86046f89ccbe0a061486096c783e743c4ebccca778534ccc32832e8c89a17de9

Observation bbcb182c-7dc5-4836-a25d-eb073cf5cd7c · outbound

This paper cites Deep Think with Confidence.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Deep Think with Confidence

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:27.508814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:27.508814Z digest=sha256:24b94e9c6897466b1be11aaa04cd2b8554ba39d1e2f3a54515c8edf23479b170

Observation f61ba3da-5e8d-4a0f-8c4f-55b8647b43e0 · outbound

This paper cites The Llama 3 Herd of Models.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:27.641247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:27.641247Z digest=sha256:59996085b2a25d679e24baeb78f638b9f5f87442acd21854ac3631a653b5e0ca

Observation f8ccd421-5d78-4159-8c09-acdcfc0cf531 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:28.040915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:28.040915Z digest=sha256:85af6f76c1a59b6f9a920180ab141c9fe9c20c0a04464eed3808ddd95c0527ef

Observation 6bd47e3f-aed5-468e-9987-8a42d82fd06c · outbound

This paper cites Qwen2.5-Coder Technical Report.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Qwen2.5-Coder Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:28.159222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:28.159222Z digest=sha256:8b4320ffd129dcc1a6b8a7c7d9acd50dc6e5a11c01771caf9766d25f5f4e8906

Observation 1c513edd-9ded-498c-82f1-fb149569d5f3 · outbound

This paper cites Tree search for llm agent reinforcement learning.arXiv preprint arXiv:2509.21240,.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Tree search for llm agent reinforcement learning.arXiv preprint arXiv:2509.21240,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:28.259075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:28.259075Z digest=sha256:3403d98bb15f5cb4ff8bab5a156096fe9001795ff1de6086674d0b4e91e28657

Observation 534fd71e-c0aa-46e9-b074-44b0c6f12329 · outbound

This paper cites DeepSeek-V3 Technical Report.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning DeepSeek-V3 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:28.440217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:28.440217Z digest=sha256:2014ba364b3482c3121753502ab62beb2a8addabaeaed9df257bdf27e369e69c

Observation 3f38d4cf-570a-4a5a-a3c5-9710540765b5 · outbound

This paper cites ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:28.575603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:28.575603Z digest=sha256:89359980e23340affe1537bcd6af8d514ec9d90f6a30a6783b63e81521d3c4b4

Observation 4a7e2142-1ce1-41da-823f-11d16319bdfc · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:28.665263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:28.665263Z digest=sha256:adf73dab8edcb889eeaf5c96c5cdc439f821e37974489422dcdc89b58b048206

Observation d828de3e-b917-465e-92b1-fb455682e8fd · outbound

This paper cites Maximizing Confidence Alone Improves Reasoning.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Maximizing Confidence Alone Improves Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:28.713259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:28.713259Z digest=sha256:f7db2ebd63450924d7dbba53cda814cc24040a32d704e9f6a03e6f7905512fd2

Observation 56823190-2c6c-4186-b445-d26eaacfcf89 · outbound

This paper cites Proximal Policy Optimization Algorithms.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:28.789056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:28.789056Z digest=sha256:9f4a49662f7fcd1b8074eaddd366c43f1616ec49254a458f1cf9e538bafac023

Observation 70d81845-3612-492c-aa86-26df9a139765 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:29.019924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:29.019924Z digest=sha256:ef24b37c3e3b627ac482b60cf8ec9fefb5af219a702f4f3249a2571ff57ad234

Observation 98f0c4b1-5a95-40a5-905c-070074d88b6e · outbound

This paper cites Qwen2 Technical Report.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Qwen2 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:29.085109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:29.085109Z digest=sha256:7e03522db0343fe49e5e51a277c7e0a00cd86d99062bd9187a06fedd3572fc79

Observation be0ddac4-b876-4917-a912-8282b807067f · outbound

This paper cites Qwen3 Technical Report.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Qwen3 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:29.292079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:29.292079Z digest=sha256:f16f0d5c439a439adefe44d00da9a5fbfb1514fb502fa2402561ff0f11a15f6f

Observation 6515c923-53b5-4a8f-971c-50a0dc309ab1 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:29.393951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:29.393951Z digest=sha256:59eeae7fe9c921556e8ed8bec1207b9c7171a6fa443a70d5c862653e2b769b20

Observation 767c0429-5729-4a53-80d7-008c6c71924e · outbound

This paper cites Learning to Reason without External Rewards.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Learning to Reason without External Rewards

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:29.521368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:29.521368Z digest=sha256:d73f40a913342826db055f0b5b451514eee2f37fa16b38e9a20fcd71f2b714a6

Observation 5bbb9519-929b-45cd-aadb-749ab98959b3 · outbound

This paper cites Evolving language models without labels: Majority drives selection, novelty promotes variation.arXiv preprint arXiv:2509.15194,.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Evolving language models without labels: Majority drives selection, novelty promotes variation.arXiv preprint arXiv:2509.15194,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:29.670391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:29.670391Z digest=sha256:24ee287b4f0ac33bf6462965897f1be9d7b25ee772a74d7dae20ae28070cbe58

Observation d1e7b6ed-41c7-4e17-9ee1-8df0b4d8bf23 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning TTRL: Test-Time Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:29.795255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:29.795255Z digest=sha256:5f8bc20d200e243e706243c4cd8e47d5ec874ccdeabff6d974c464e7b9a21c2a

Observation e026cf04-701b-4003-829b-c6556f678359 · outbound

This paper cites an unresolved cited work.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:29.937356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:29.937356Z digest=sha256:f142fe64aebfe2db6ab2d9c124316106788c8a4076313fbc147b3e03a0f9b05c

Observation c298c9a4-1ef8-4d2f-a9ba-3318340809c1 · outbound

This paper cites an unresolved cited work.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:30.076673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:30.076673Z digest=sha256:3397fc01132055e45e6092b57b1a76546274f3e3144869b7348dee2ef1fa0b82

Observation 14a74240-5b58-45d8-9724-7e83b61f1252 · outbound

This paper cites an unresolved cited work.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:30.210574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:30.210574Z digest=sha256:11294a67e72c5dfffa4a1e5ca382c3265bc0a2274aa9e155b71c9dc210ac7f84

Observation 83e8a216-f516-40b9-8aad-a20078ad7f60 · outbound

This paper cites Crafting papers on machine learning.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Crafting papers on machine learning

Reference 1996

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:28.349671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:28.349671Z digest=sha256:36df9ab425ada587fff99fd21b1021dedbb98bf17181209054515ab064db4389

Observation bf67b164-38b9-4af2-82bd-96ae7384a5f2 · outbound

This paper cites Bapo: Stabilizing off-policy reinforcement learning for llms via balanced policy optimization with adaptive clipping.arXiv preprint arXiv:2510.18927,.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Bapo: Stabilizing off-policy reinforcement learning for llms via balanced policy optimization with adaptive clipping.arXiv preprint arXiv:2510.18927,

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:29.211496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:29.211496Z digest=sha256:e051e1c5546c6c80609ea2707f7f42a566aadb8cc8986622b6be595052215eb5

Observation 5a4afd59-f7b6-4454-ba28-4157ea2649f4 · outbound

This paper cites Can large reasoning models self-train? arXiv preprint arXiv:2505.21444,.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Can large reasoning models self-train? arXiv preprint arXiv:2505.21444,

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:28.878007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:28.878007Z digest=sha256:5fdc217e96508511784734299a15010097dd447d537c2b54a30d1f13e18ce5f6

Observation 2d58a01f-0df4-4816-a91b-31e1f344a4c5 · outbound

This paper cites Unsupervised post-training for multi-modal llm reasoning via grpo.arXiv preprint arXiv:2505.22453,.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Unsupervised post-training for multi-modal llm reasoning via grpo.arXiv preprint arXiv:2505.22453,

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:29.134691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:29.134691Z digest=sha256:babcea72d79aa0685c2d896be100636f6576488ec26549713d9adac768f7a274

Observation 223c20ca-91da-4c58-9739-b077b2209246 · outbound

This paper cites TreeRL: LLM Reinforcement Learning with On-Policy Tree Search.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:27.901205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:27.901205Z digest=sha256:7718c98d48aaf78d436b4fd9b368a7c4b416025e126ce4a9f0fe060a5e394d3d

Observation da3fd933-8924-4d36-93e0-bd608cfddeef · outbound

This paper cites Qwen3-VL Technical Report.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Qwen3-VL Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:27.189984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:27.189984Z digest=sha256:132a19a73b67960b313196b70318a75ca429e02a50ef99aee3f9e06039ffa9a5

Observation aff2c1da-0680-408a-84bc-e783ab657f76 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:27.789496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:27.789496Z digest=sha256:9996604eb853584cf9dcd35f0a2516898ea0ec632a86302a5a0c359dda6277e7

Observation 17abc89d-d09e-4d49-9dfa-6d433e03d239 · outbound

This paper cites Agentic entropy-balanced policy optimization.arXiv preprint arXiv:2510.14545,.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Agentic entropy-balanced policy optimization.arXiv preprint arXiv:2510.14545,

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:27.329790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:27.329790Z digest=sha256:422b6e9a0db614e411286ef871d2cb09d6cd5c1ad81cc5cb20b9dc95c475fdc3

Pith citing papers

Observation 931ec558-fc4e-480d-909b-3cbad9cbb1ca · inbound

Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning cites this paper.

Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-28T03:04:45.398688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T17:16:00.499779Z digest=sha256:85f7a1131d9c681ffa21130836dc4f025ae5447565640a89b0ea0ceeb7f51355