Pith. sign in

Paper Citation Record · LEDGER

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator

As of 24 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 1 inbound Pith citation observation for arXiv:2511.16886.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.16886 v5

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T21:06:30.259347Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T14:03:01.974171Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ff9e43c3-9687-454e-ae99-7924ad2edec2 · outbound

This paper cites Universal Transformers.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Universal Transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:29.296424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:29.296424Z digest=sha256:5f0b64fe0543ac84e5dc5ec648ebfbb6aab5dbb88776ea77bb29357cd64f5608

Observation faa93361-0335-4d8e-aeb3-ccb0b08bc471 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Classifier-Free Diffusion Guidance

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:29.717017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:29.717017Z digest=sha256:69c6e839197ee42260d05fc0b0cb8d5a0de0c554f006a630709d6a1a6dc03181

Observation d9965a2c-7799-4554-9e43-271b2abce97e · outbound

This paper cites The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:30.009675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:30.009675Z digest=sha256:a85674dd594dac27394fac35ca07d7e4691c9fe9b7d38bc6fcb6fea51ca4ca9b

Observation f05a28e1-9b46-4a09-90ea-5f15886d9229 · outbound

This paper cites Looped Transformers are Better at Learning Learning Algorithms.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Looped Transformers are Better at Learning Learning Algorithms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:30.161963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:30.161963Z digest=sha256:57003a5478d574dbbd3951de5e50063d04fbab632d6371d6ab2b9f1ffed75887

Observation fd0c9e37-1de4-447d-8634-215537f7b208 · outbound

This paper cites Alternative Improvement Generators.As described in §4.2, there are several viable methods to generate intermediate steps.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Alternative Improvement Generators.As described in §4.2, there are several viable methods to generate intermediate steps

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:30.259347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:30.259347Z digest=sha256:b6e3f69534e726c543c143a048fff1384613183f940ddf6607268a237013cecc

Observation 21bcbd0d-a95d-42de-9a24-1c41b8f413f4 · outbound

This paper cites Hierarchical Reasoning Model.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Hierarchical Reasoning Model

Reference 1998

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:30.093930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:30.093930Z digest=sha256:482fe611f440174b884cf8b95a56a8a99fc6d2ead673b2b028eaa28d87bf0278

Observation 22b9ddee-1aa1-49cd-a886-ee384d2d7b4a · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Dream to Control: Learning Behaviors by Latent Imagination

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:29.597655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:29.597655Z digest=sha256:7e5e7ab57e0ef81e6c347eacab3b97813a5b51b64dbeec2c65eab119dc6b7d8a

Observation 21705f32-f9ce-49ff-a85b-0490e878683c · outbound

This paper cites Diffusion Guidance Is a Controllable Policy Improvement Operator.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:29.374560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:29.374560Z digest=sha256:a4f2b9e6e755852b8e10060eca742a8d7e6dd46799b4f89bfc1025062f716641

Observation 4f18f9f3-50c1-426c-954e-c5937e77cede · outbound

This paper cites On the Measure of Intelligence.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator On the Measure of Intelligence

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:29.227914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:29.227914Z digest=sha256:ea5e336662b115d9a80b7837206f1e90fc714c9af3640c3714f9df14cee60c1e

Observation 883b64f5-5f91-4300-9ae1-1a2d68106f17 · outbound

This paper cites Less is More: Recursive Reasoning with Tiny Networks.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Less is More: Recursive Reasoning with Tiny Networks

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:29.803009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:29.803009Z digest=sha256:1a39509bde4d38fc66b176e18056cd0c7e767a1f0a862a99d6fc0638e0d99853

Observation 53302236-baab-4592-86ae-5d7e6e2ac11a · outbound

This paper cites Adaptive Computation Time for Recurrent Neural Networks.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Adaptive Computation Time for Recurrent Neural Networks

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:29.486681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:29.486681Z digest=sha256:0d939185f583f7aba54a79a410df6dd61be6b61afcb1bbe0959c439bf8f90291

Observation f10b11f9-6161-40e2-b920-75d1698c4166 · outbound

This paper cites Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:29.881012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:29.881012Z digest=sha256:46ae45aa7b5db6a8d85af90b5df0d46122a78a6897e71f1b8118a3cafda49ad3

Pith citing papers

Observation fa662205-7a44-4221-88c9-4e704fca57ba · inbound

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook cites this paper.

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook Latent Reasoning in TRMs is Secretly a Policy Improvement Operator

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-13T14:09:38.535527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-13T14:03:01.974171Z digest=sha256:586aacd175091d64316ce17d31e774717f99f963a4cb9e6a7e4570443760921d