Pith. sign in

Paper Citation Record · LEDGER

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2608.03545.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03545 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:52:37.405393Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy42
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ed278b8-71eb-4459-9c40-e8e4f961a4ae · outbound

This paper cites an unresolved cited work.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:52:38.363913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.125672Z digest=sha256:df977b44d6e227fb942f124b16e2e312a49412c12cfe169c3ae72018250f22e1

Observation 35873716-33c1-4ee7-a5d7-687ed403528d · outbound

This paper cites , year = 1983, title =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning , year = 1983, title =

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.349715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.131388Z digest=sha256:c1b537c41092207133b889916e850466f4622182198fecb5f658ea365022c933

Observation 185ea821-dcd0-4fa2-a70b-1c8587f17197 · outbound

This paper cites , year = 1984, title =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning , year = 1984, title =

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.335520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.138291Z digest=sha256:e347e1dfb8fc0b474f7418d8bb2f6ac1121c1d86f6ed0d9e02ca280d60738d13

Observation b3d29bff-8b19-4eab-85af-8753c59bde22 · outbound

This paper cites , title =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning , title =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.143686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.143686Z digest=sha256:b3618dc408e562872ef31af3789216c435c7e7f3565a7ab4424b7b1e0672480a

Observation 3608bf1c-1911-4653-84ac-93be67822f81 · outbound

This paper cites , year = 1980, title =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning , year = 1980, title =

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.311285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.149323Z digest=sha256:c97a4b4d74c79fb314051989d7b5968802d12e12bb738fb538af1f502c92044f

Observation 7fe45bb2-cba3-4838-b941-ae90ab071f1d · outbound

This paper cites Clancey and Glenn Rennels , abstract =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Clancey and Glenn Rennels , abstract =

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.154451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.154451Z digest=sha256:0f398a510182cf8eb3aaf657ae0efd648c7b79c3d9b9bd72556874287907556e

Observation 0ac35a06-d4bf-4e1d-9431-1f99db176d23 · outbound

This paper cites and Rennels, Glenn R.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning and Rennels, Glenn R

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.296976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.160028Z digest=sha256:0d23bad1482706d7ba282833a56fbdd90640d80608f1975af47d5c92bc01d132

Observation cfc2ccef-ffbd-4879-bdb2-229b356bac7a · outbound

This paper cites an unresolved cited work.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:52:38.282460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.164832Z digest=sha256:40db0f6f80c53de4200b65c4728becb6f4d3dd77c6bd00cd2501cd3c7bfd2e30

Observation 6f41fcd9-91d7-4a90-85ab-5a1651d6b9e8 · outbound

This paper cites , year = 1979, title =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning , year = 1979, title =

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.268248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.169591Z digest=sha256:7f6a54078b1f63fdc5d1d8ffa1cd69e30793b7fc1f64442702efadcc6628a36e

Observation 8423e12b-d49c-4ca6-8743-803e28c5f2ff · outbound

This paper cites , title =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning , title =

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.253850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.174492Z digest=sha256:a101a42f49ecba87e4c6fa9dd1e2f5e19c8a3cb25893267b57b4c0357d16e5b8

Observation fef0cc84-d6bb-489c-bc68-e029c50a9b71 · outbound

This paper cites 2017 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2017 , eprint =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.179241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.179241Z digest=sha256:216fddb18c5a851f33690225051ede6ca5022f1156e824bf0e5f150a2224f690

Observation f57f9d02-2e29-4399-a2a4-e07407b6d9df · outbound

This paper cites an unresolved cited work.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.183822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.183822Z digest=sha256:b052e7b9848922c6ac4799d6a9c7088529a1b94d57c75bc6cb70b2c29593e73b

Observation 7230285f-38ad-475e-9636-945b7b7b6387 · outbound

This paper cites The Eleventh International Conference on Learning Representations , year =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning The Eleventh International Conference on Learning Representations , year =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.188675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.188675Z digest=sha256:cfb7dc956c5936679aad2a9ac9f281e57a267a143d5472a077effe5c5eee4234

Observation a192e1c2-29e3-421f-ba77-5439c7648c1b · outbound

This paper cites Advances in neural information processing systems , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Advances in neural information processing systems , volume =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.193820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.193820Z digest=sha256:fe77e2a5b1ba8ceb8ee0b0c27740e8fe74e2b6e1f8d9b331597f00e62073aab7

Observation 28f925f4-298b-49af-bc80-cafc8694a046 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning , url =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning TTRL: Test-Time Reinforcement Learning , url =

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.202867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.199002Z digest=sha256:1c19e0662aea090de8b6855a0094c90e75671032fba9ab15cea8bfd5f20fe7a2

Observation 7a9c20f4-41d6-4a78-8f6c-cc472feeaa1c · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.188738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.204174Z digest=sha256:c7d824d6601caa26d17351a6a2b23bbd7ce0aa9d2ba782c19603e9f296157cbb

Observation 7f668440-b608-4435-8fa4-1ee9e9964a04 · outbound

This paper cites 2025 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2025 , eprint =

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.174641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.209028Z digest=sha256:d0ae8141c4169fc6ecf12c5e9608a61e2584fb39b36ccc0505deaee7146186b2

Observation a27ffac8-fbe4-4c6c-930f-d5d2e403c0b3 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.159714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.214354Z digest=sha256:3425280aef1247b3d373de7a0852f31b5ad490599d71a50e0269717700715091

Observation adbaa938-1bcf-47fc-8526-8307834ead2d · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages =

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.144454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.218963Z digest=sha256:b95923b89e1b98de12b446a82024153d5771a6a6ee0c4c5af17afa76557fe4db

Observation 20e4baa6-33a1-486a-bb8e-f7f86972d041 · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.130677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.224796Z digest=sha256:91d632b981c55f4a3c7e3491a4a74edda1e163763a60795335c555338bc1ba2a

Observation 79e68d40-b81d-499d-968b-c94afc4010e3 · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.116799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.229542Z digest=sha256:457492fc8726ccc5d316037af00d4d7958491fd3088c733859dd566b59e8037d

Observation aad71683-7505-4632-bdae-ade1e461c9f4 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.103194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.234366Z digest=sha256:765487ee5363df13d94eed745c43568ef44bc64fb219590fb95a51e78e39ccb9

Observation 1718470a-a437-4913-a8fe-0b4324b23d6a · outbound

This paper cites Forty-third International Conference on Machine Learning , year =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Forty-third International Conference on Machine Learning , year =

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.089662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.239627Z digest=sha256:8533d73c75c0a9abe10a1054b09c9096e656528853d52d4ea7edd16bccbb3d9c

Observation 3c107409-0448-475b-91d3-245a5a6e1c69 · outbound

This paper cites DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning , volume =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.244689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.244689Z digest=sha256:6164e9ea3f5e5df61cfd9a57911af31ac7e23b4ff9f440844c98494c65910609

Observation 72850efe-d844-4d56-b9bf-396ea6340ca9 · outbound

This paper cites 2025 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2025 , eprint =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.249652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.249652Z digest=sha256:bfcdd5ad6a46dd25e24bb9d806618756a27290816d0b4bc0e4c3b2c32c25d8ea

Observation 7d739666-e02b-4c69-96ea-050634c320bc · outbound

This paper cites 2024 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2024 , eprint =

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.254564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.254564Z digest=sha256:f66ed2d52102909a138ce743f81e3c598054023a12db7b21979d95ab6a7781b2

Observation fb731896-e67e-4e71-bd9b-a1385c1ed121 · outbound

This paper cites ACM Computing Surveys , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning ACM Computing Surveys , volume =

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.058301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.259146Z digest=sha256:ae18432ab016f07e09ab0de36e614ac123444b454c60c2da46d78ea3d5e1e4da

Observation a2328cc5-d3fb-41e8-bc74-bead1257befe · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages =

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.043693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.264159Z digest=sha256:f8108c27e56e0cfe8d68c55ade9cfb6910f112aecc705e59bc2b332ea7ee2511

Observation b5eb781d-193a-4253-a3b9-c151273900a7 · outbound

This paper cites 2025 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2025 , eprint =

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.029749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.268766Z digest=sha256:394dcaca47a9af4fa763ee354fdfb09ebcec984a2c37df7281e1636de034a4c1

Observation 11539c81-0316-421a-b678-9166245911af · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.015451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.273663Z digest=sha256:ccccb0a5bf5783c9effc0492575979d6c7b57fdc1c3e259c1c49cbe6c4e9f47c

Observation f788e82d-6b0d-4c88-9236-6177cfd148ea · outbound

This paper cites Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization , url =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization , url =

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.000775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.278216Z digest=sha256:ad01eedae4a1999054fd6f99239f254b40cd7dfc1774da465840424a3b46c44b

Observation 813a7dde-7224-4df4-b079-24cc74824a68 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.985176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.282965Z digest=sha256:312c6f3bc274b9e3369fc5fa414c18fa0bb8c7b9f50c0a7b91f701fa7e7e6a92

Observation 19522351-5d6b-4f5a-8ea7-4483c3b21224 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.970491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.287476Z digest=sha256:1d849c6f7300b4871c677361fd12e3e5792b20e91d3b10480681267d4229ba88

Observation d29273ac-bbb1-4f0f-a13b-caf93ea42e2f · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , number =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the AAAI Conference on Artificial Intelligence , number =

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.955910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.292516Z digest=sha256:0f45e71d9f093e85bcc5c00c19c52fd86e4ccda2176cabc320729515069da115

Observation 613632e7-a01d-4dda-a218-95bcc6f1caae · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.940948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.297300Z digest=sha256:71809794db483fabf20851e8ac2f5a74e9ba5ee9f8168eed02698fe0dd4d9f81

Observation f8802d1c-dc37-4058-8c6c-ac8382597cb9 · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.926536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.301938Z digest=sha256:36238f1e7daeaef9d4bdc574aabc4fbc05125d677563d6c48bee559b6d5dc5b7

Observation d9c7f3dc-d4a5-49c1-91cc-cad23708d9e2 · outbound

This paper cites 2017 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2017 , eprint =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.306383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.306383Z digest=sha256:dd993380831a330afd3cc2288f7952ef4bb7899ae4a75b59bf35ba1430a8eb77

Observation ae6db8fd-e887-4b58-822f-50e922298525 · outbound

This paper cites Advances in neural information processing systems , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Advances in neural information processing systems , volume =

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.311151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.311151Z digest=sha256:3014361d5de1f38438c3a90b1beafcaa8cd8315f668b2daf9eab5307662288c3

Observation 6f09fccf-7350-4f5d-b300-9bb198b0639c · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Advances in Neural Information Processing Systems , volume =

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.894932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.315756Z digest=sha256:9fbf985dbfbc4bc8b3f61d352b391090926930042b3be71c145a2937664b28b1

Observation 7b77432a-8fd5-407a-99e6-c7719486bfa8 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.881012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.320316Z digest=sha256:9496bbf21b57b293783c919a3ed6843d470d7321a9ec207f052e20e6d60e0d24

Observation 685415de-be4c-495f-b9df-191984cf0c50 · outbound

This paper cites 2024 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2024 , eprint =

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.866652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.324949Z digest=sha256:daed8bf20f10d48dfa3889092963b3cfdeea475cb3ae93e4b9b466188101c642

Observation c9123113-4c72-4338-bad5-3c1160a8ae38 · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Advances in Neural Information Processing Systems , volume =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.852473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.330636Z digest=sha256:0caf140d1b50c20c89a4008b3834099ed0fb5d354e2532f1ae7d3e408732abce

Observation c83c0caf-7d57-436f-a0d7-282ad478b756 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.838805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.335195Z digest=sha256:88ab3b3d70eaeb7f0cb3cbb9405b732cf3485e7f2949a1010dcc07dc21da02c9

Observation 206a6009-f473-414a-b89e-7b37689ea0f4 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example , url =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example , url =

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.824996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.340347Z digest=sha256:e2911538418fd4612f709f56219697c7c224a38c6789596285ea15975407e6ad

Observation b5cfdda8-f7b2-4197-a10b-5932bf58afdc · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning , url =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning , url =

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.811006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.345104Z digest=sha256:f67993f770b37b86bfe9662c70e0886ef649342e2373cd5f66be10226b310af4

Observation 9f72fd16-6654-4ccb-8c74-5f5c6a2c1814 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.796660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.350015Z digest=sha256:ece49bb5ed3a2ab1ece3d90ca6325fd8f9e07a8e01efaf7821a8a9b2dd0ebbd5

Observation c4f1566f-16b2-421e-82dd-9063565dd468 · outbound

This paper cites 2025 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2025 , eprint =

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.781622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.354797Z digest=sha256:dd77a65346eeacdd36a0abdd32ee6ed9ac33deb3ff92c99c60094ea91187e2bf

Observation 4c46b567-af8e-4ca3-9634-9754c596e643 · outbound

This paper cites 2024 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2024 , eprint =

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.767694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.359492Z digest=sha256:d315166396f337bbb0ae517a3c09db0015e4055fc6cd19dc6c1a504c73fb1afa

Observation 1579bc71-7677-4a15-aa62-b16ae9304b4d · outbound

This paper cites 2021 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2021 , eprint =

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.753867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.363908Z digest=sha256:7de49a25516b3ca9cbab12e107fca9687b3119e49fe8d636551aa60753f5507e

Observation 19720dca-b1e2-48e0-9c2a-1b824bad68a0 · outbound

This paper cites Advances in neural information processing systems , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Advances in neural information processing systems , volume =

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.739927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.368739Z digest=sha256:a00218cbbe69822b9c47b5ac613df79093c20e6c8e0bd023ff0f055d8b9a2895

Observation 7e02f688-fc05-453e-81b9-93c2e71db1c8 · outbound

This paper cites Hugging Face repository , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Hugging Face repository , volume =

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.725796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.373259Z digest=sha256:6f893166f5607c24ffc70a0e611c496580a019122af6fb3d3bb388443a49f870

Observation 4988a4e9-2759-4038-844c-33085495760b · outbound

This paper cites an unresolved cited work.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.377789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.377789Z digest=sha256:34a329bd7aa45e78d81c42638f46f88fef19b86ea40ede23f0a556592dc44490

Observation 06de54c9-ba28-4e50-8f3b-0211714439d0 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Findings of the Association for Computational Linguistics: ACL 2024 , pages =

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.701090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.382398Z digest=sha256:ee36adbca8cd7bd9c37d1bf4a856c629c48b5fcb1256b3a350451757ede7e112

Observation cbb5181a-3311-4f3d-8d9e-152a201fe851 · outbound

This paper cites 2025 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2025 , eprint =

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.686100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.387143Z digest=sha256:ce4b2003d745c54f891124b4e3de8439c4f5683e2e0b8c3d4eaf292297b5e830

Observation 88758300-b891-42c6-ac45-1711f3e6a1f0 · outbound

This paper cites The Hidden Link Between.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning The Hidden Link Between

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.391538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.391538Z digest=sha256:09bdb81c01822234b3fcd1eb9b14569b28576f39b6332b8c8a0bd369eba3ba33

Observation 306e9743-db19-4fe3-9559-470fa2c9c306 · outbound

This paper cites The Fourteenth International Conference on Learning Representations , year =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning The Fourteenth International Conference on Learning Representations , year =

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.671530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.395884Z digest=sha256:2e43436ed8e70caf7d86f2d4a426837859f91fe994618910fd52182dbd84c8c8

Observation 35b06d4f-4564-49d9-83a5-ab510ec2c64d · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.656364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.400464Z digest=sha256:2bdbcab3f83d81cae6a79e628d39c032c942d1bce5750d1618be6ee75a05d0ec

Observation 7fe7ba90-5138-4f16-8f00-5575d2f34a96 · outbound

This paper cites A Survey on Human Preference Learning for Large Language Models.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning A Survey on Human Preference Learning for Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.405393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.405393Z digest=sha256:1048463ca0ea2d87f3aca194043e3024d490474d358dad9909c837c605e9ff2e

Pith citing papers

No inbound Pith citation observations are available.