Pith. sign in

Paper Citation Record · LEDGER

Think before you speak: Training Language Models With Pause Tokens

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2310.02226.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.02226 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:22:32.565978Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ffa6a240-7238-4739-9f0e-2096a5dc696e · inbound

Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters cites this paper.

Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters Think before you speak: Training Language Models With Pause Tokens

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:16:23.868058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:16:23.842092Z digest=sha256:63cd11d3ff9d485f241bdab00330e08874b828add22a6baac43ac84fee7bb3a5

Observation bc82b9a3-74b9-40f7-b007-8126939e1953 · inbound

MeTHanol: Modularized Thinking Language Models with Intermediate Layer Thinking, Decoding and Bootstrapping Reasoning cites this paper.

MeTHanol: Modularized Thinking Language Models with Intermediate Layer Thinking, Decoding and Bootstrapping Reasoning Think before you speak: Training Language Models With Pause Tokens

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:03:26.316471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T21:02:36.873034Z digest=sha256:53fac3d8ecf8bbdd07a60f89030d02e99a3346a2bbbe65bf1e69ceb214895aa5

Observation 1f5f740c-529c-4a42-9587-4d242a2e1563 · inbound

Training Large Language Models to Reason in a Continuous Latent Space cites this paper.

Training Large Language Models to Reason in a Continuous Latent Space Think before you speak: Training Language Models With Pause Tokens

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:29:05.664360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T10:29:05.384381Z digest=sha256:2f12456adadfa8f13f3e3c745488278579197b887cce805740714750de42295c

Observation fd3878df-ac30-4218-b528-7128b9eccdb8 · inbound

Compressed Chain of Thought: Efficient Reasoning Through Dense Representations cites this paper.

Compressed Chain of Thought: Efficient Reasoning Through Dense Representations Think before you speak: Training Language Models With Pause Tokens

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T04:47:40.381114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T04:47:40.335477Z digest=sha256:5b5cf53344e86d2c5ab28c988d857f88d763b62e7fadc396e09fb9fae51dbe60

Observation ad26f8bf-777b-49bf-8912-6e541244df80 · inbound

Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning cites this paper.

Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning Think before you speak: Training Language Models With Pause Tokens

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T05:22:32.565978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:22:32.565978Z digest=sha256:4540d638bbea9cf0602b02685382662bb9be7f16e82b27faadf7d379ae3f7d71

Observation 4a505697-7948-4a32-933e-b00d2a86094f · inbound

Skip-Thinking: Chunk-wise Chain-of-Thought Distillation Enable Smaller Language Models to Reason Better and Faster cites this paper.

Skip-Thinking: Chunk-wise Chain-of-Thought Distillation Enable Smaller Language Models to Reason Better and Faster Think before you speak: Training Language Models With Pause Tokens

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:32.230744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:32.230744Z digest=sha256:0ec681e96611972914d68b6959b2848504d0aa4d27f076514f1d4e5a7c120651

Observation 8b8ef420-0d32-4d7a-a559-88dbb85524c3 · inbound

System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts cites this paper.

System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts Think before you speak: Training Language Models With Pause Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:38.221574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:38.221574Z digest=sha256:3af7e9c69e96d2fe45240b3f784b74ac97fae47290d64b3cff1b261d8d9f6c8e

Observation dd456392-e44e-41ce-bce8-39715e9adb07 · inbound

Knowing Before Saying: LLM Representations Encode Information About Chain-of-Thought Success Before Completion cites this paper.

Knowing Before Saying: LLM Representations Encode Information About Chain-of-Thought Success Before Completion Think before you speak: Training Language Models With Pause Tokens

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:19.638184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:19.638184Z digest=sha256:cdaad3dde80af6d48d52d27976dad11eb9363cbd79d2886198ef7797b50e9086

Observation 7a678ad2-fe4c-488a-b99d-81b9d4b172fe · inbound

Energy-Based Transformers are Scalable Learners and Thinkers cites this paper.

Energy-Based Transformers are Scalable Learners and Thinkers Think before you speak: Training Language Models With Pause Tokens

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:42:29.670598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:42:29.670598Z digest=sha256:983c08a34519c264dd5d8043df5dff08daaadeab84bf0fc35de78a36f76a36e3

Observation a06ac850-4280-4fab-9f07-12874da7d55f · inbound

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing cites this paper.

MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing Think before you speak: Training Language Models With Pause Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:30.625805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:30.625805Z digest=sha256:2fbcfbc2c9956f520aa80238f1a0c9fdf8848b33a0f14408e81162e10b5f19ce

Observation adae2c4c-54e3-43e1-877d-1c010e58537e · inbound

AnalogFed: Privacy-Preserving Discovery of Analog Circuits at Scale with Federated Generative AI cites this paper.

AnalogFed: Privacy-Preserving Discovery of Analog Circuits at Scale with Federated Generative AI Think before you speak: Training Language Models With Pause Tokens

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:56.781431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:56.781431Z digest=sha256:f5fe7f6b90e5479bd462b71834711308b19f13bbf93c254e17b625a875a26b74

Observation 02e29b84-6d12-4014-a3ca-f9e661c44304 · inbound

Learning Temporal Abstractions via Variational Homomorphisms in Option-Induced Abstract MDPs cites this paper.

Learning Temporal Abstractions via Variational Homomorphisms in Option-Induced Abstract MDPs Think before you speak: Training Language Models With Pause Tokens

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:34.903549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:17:34.903549Z digest=sha256:b21db15bc301b70c7b0e9cfb24d1432d7390eb781e14a5cf53fee97faade89a6

Observation 3ccb2d8f-3e9c-441b-a8e5-66eb0c32ca18 · inbound

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models cites this paper.

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models Think before you speak: Training Language Models With Pause Tokens

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:46:03.555461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T07:43:23.913399Z digest=sha256:9b3ef1e50b0535accc2c27da1ce4b9daf8461521f401e831be7f49add4518dab

Observation bb357db0-d87c-4e5b-92d5-44405ae7569f · inbound

Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thought cites this paper.

Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thought Think before you speak: Training Language Models With Pause Tokens

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:45:46.257095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T02:44:48.729794Z digest=sha256:df1cb5eb1567c12590d2e2fe24609ef899bb477081d04aceac4bc9566142fca3

Observation c905c2a2-5a3a-4962-b96b-8c13a3e000f9 · inbound

Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thought cites this paper.

Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thought Think before you speak: Training Language Models With Pause Tokens

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:09.543491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:09.543491Z digest=sha256:05f1ce8fd803c1d76e4901d6d03aea00561a8343a309c1aff04a47abd673dbd8

Observation 31af617b-0c5a-442d-aa30-a65efea124c8 · inbound

Next-Latent Prediction Transformers Learn Compact World Models cites this paper.

Next-Latent Prediction Transformers Learn Compact World Models Think before you speak: Training Language Models With Pause Tokens

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:25:29.419754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T07:21:04.684347Z digest=sha256:5b97e5b1df7daa61966c7217f8ba99cedd8c90837b76cfe7e1d0c1c4bc58b69d

Observation ff79b576-6d77-44a0-9ce4-9722e1e9c4d5 · inbound

Next-Latent Prediction Transformers Learn Compact World Models cites this paper.

Next-Latent Prediction Transformers Learn Compact World Models Think before you speak: Training Language Models With Pause Tokens

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T23:27:54.521220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:27:54.521220Z digest=sha256:0fc9e8d7f94cd0ca1f235e4af8e33aebdecfd33648491b4a1db1978f55fd6868

Observation a9780133-3271-4cc1-a15f-c7d98e65aa19 · inbound

ConFu: Contemplate the Future for Better Speculative Sampling cites this paper.

ConFu: Contemplate the Future for Better Speculative Sampling Think before you speak: Training Language Models With Pause Tokens

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:20:03.343507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T14:18:24.071422Z digest=sha256:d6e78e44c63a3291250d4b86c264d357380d53d8f9cc61de910924659107e26c

Observation ebdfda19-b095-4e6e-b79f-705874614819 · inbound

Representational Collapse in Multi-Agent LLM Committees: Measurement and Diversity-Aware Consensus cites this paper.

Representational Collapse in Multi-Agent LLM Committees: Measurement and Diversity-Aware Consensus Think before you speak: Training Language Models With Pause Tokens

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:03:04.471785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T17:59:52.185345Z digest=sha256:62058fff73cdfc9720069dff4ec8c0d2f6c4d9aa96c347d720e472d82a9ca4d3

Observation 8fddf0f9-438a-4a4c-9a01-99d6256b44e4 · inbound

SeLaR: Selective Latent Reasoning in Large Language Models cites this paper.

SeLaR: Selective Latent Reasoning in Large Language Models Think before you speak: Training Language Models With Pause Tokens

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:35:49.487904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T18:27:36.132030Z digest=sha256:fca9e141c7ec1586c0c7b4734bcafaa725038640a3d2150babb9519bf274b2d0

Observation 9456df79-ae04-4707-a066-ea9cbe63d7d3 · inbound

AI Achieves a Perfect LSAT Score cites this paper.

AI Achieves a Perfect LSAT Score Think before you speak: Training Language Models With Pause Tokens

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:40:57.923772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:35:01.079284Z digest=sha256:5222e5cd3c25d9b8fe77d42a35f9bfa286aed082c46e9b3d7e80684726de7e55

Observation 5adea2f9-71f0-43da-822f-a5f0292f3bdf · inbound

Latent Abstraction for Retrieval-Augmented Generation cites this paper.

Latent Abstraction for Retrieval-Augmented Generation Think before you speak: Training Language Models With Pause Tokens

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:01:04.901647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T04:22:05.341154Z digest=sha256:92789f6ee2711032b50d4255e01ace3333dc06eace48f6f22d6139f1d50eec5b

Observation e1cb126c-a791-4f63-8787-0cac22f0dd1c · inbound

HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering cites this paper.

HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering Think before you speak: Training Language Models With Pause Tokens

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:54:45.200043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T23:51:47.724033Z digest=sha256:bf7b4f38058873b4980a48842faf72db973772abb8d556f6dce410c712f3ef6c

Observation 95303fd2-2971-419e-a343-4aa1ba5f28f9 · inbound

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding cites this paper.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Think before you speak: Training Language Models With Pause Tokens

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:11.546187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T14:20:08.404090Z digest=sha256:56cb8bd8bb0dcfa5f9c9ba5641c3852aabb4ed8db0069285778afbc27aa1ae35

Observation 14890bca-94bc-42b8-a51e-7e46bef1126e · inbound

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding cites this paper.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Think before you speak: Training Language Models With Pause Tokens

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:16:39.972605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:e3b2a6025383c16bd0d8f4c9a527b527dc1061af3d742c437674e1b6d1053a89

Observation 2513bcb5-cb0e-4cb9-8484-adb152f34fef · inbound

NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning cites this paper.

NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning Think before you speak: Training Language Models With Pause Tokens

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:41:24.271480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T00:51:40.815981Z digest=sha256:c11a7fda107afabae2fc8ec3ab0112b26c45b3e9f83f960c8d2d6050e2977962

Observation 6d9cffc8-9a1c-4cdd-8fe1-de69bb944101 · inbound

Dynamic Latent Routing cites this paper.

Dynamic Latent Routing Think before you speak: Training Language Models With Pause Tokens

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:49:38.166124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T01:49:14.934586Z digest=sha256:104f98b9e9a3aba4db3d62d2fb5f7d21a6384cc2e4975a140c8621a3b64ea4b8

Observation e73dcb0c-bc57-4262-be13-37b3b997beb4 · inbound

TTE-Flash: Accelerating Reasoning-based Multimodal Representations via Think-Then-Embed Tokens cites this paper.

TTE-Flash: Accelerating Reasoning-based Multimodal Representations via Think-Then-Embed Tokens Think before you speak: Training Language Models With Pause Tokens

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:03:36.924513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T18:00:18.315737Z digest=sha256:dc3ecad216303152c1700e9de05e8c32f87cb70ea868b5a0f5537c699918aeb9

Observation 5ebad7b8-378d-46e4-8739-421ce11244bf · inbound

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning cites this paper.

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning Think before you speak: Training Language Models With Pause Tokens

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:36:10.403900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T06:34:57.483234Z digest=sha256:8140505654b9cb2b151d1b86dd643e1702d33c8136421dddff852ddac664fc19

Observation ea4364f7-53ee-4e8a-8c6b-a6a3648f7921 · inbound

Transformers Provably Learn to Internalize Chain-of-Thought cites this paper.

Transformers Provably Learn to Internalize Chain-of-Thought Think before you speak: Training Language Models With Pause Tokens

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:33:30.553941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T14:29:10.010212Z digest=sha256:2b8138c2c0adbd7e3d9c028993123a1081fbcc28fe010fc21a8af14a034b798c

Observation 3fe1e73e-24e8-4977-847e-97303368ee51 · inbound

MIRAGE: Mobile Agents with Implicit Reasoning and Generative World Models cites this paper.

MIRAGE: Mobile Agents with Implicit Reasoning and Generative World Models Think before you speak: Training Language Models With Pause Tokens

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:06:48.304321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T06:21:57.229717Z digest=sha256:84f687d71deb0eb330e5502749d89bc1ee2668d370a14c8a1fc2f9506b71a6e1

Observation 94387795-7d27-4bc0-b471-24a2fb05b2e4 · inbound

Test-Time Compute Scaling for ASR with Depth-Conditioned Looped Transformers cites this paper.

Test-Time Compute Scaling for ASR with Depth-Conditioned Looped Transformers Think before you speak: Training Language Models With Pause Tokens

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:26:45.923378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T06:56:57.423463Z digest=sha256:738f5939b545791471ceed445eaab3326ce2aa5b31d017af93c13e61e3ed7458

Observation 3915c544-a9a2-403f-bf09-b28b7487b39e · inbound

Why Limit the Residual Stream to Layers and Not Tokens? Persistent Memory for Continuous Latent Reasoning cites this paper.

Why Limit the Residual Stream to Layers and Not Tokens? Persistent Memory for Continuous Latent Reasoning Think before you speak: Training Language Models With Pause Tokens

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:47:17.690706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T21:54:19.244824Z digest=sha256:890ffdbfb05351ac852609c384d7c04c75f792318e1c9848c5d567d21b5192a6

Observation 49f8842d-9840-441c-ba72-947b02449465 · inbound

Forecasting Future Behavior as a Learning Task cites this paper.

Forecasting Future Behavior as a Learning Task Think before you speak: Training Language Models With Pause Tokens

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:41.342134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T12:55:39.494339Z digest=sha256:ea92a0bc68151c8e36f850e882f66c31a8b9d81eae042c6df46e9a2d95abc22a

Observation f024309d-532c-49f1-b1c4-c01e18d6225b · inbound

Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning cites this paper.

Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning Think before you speak: Training Language Models With Pause Tokens

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:58:22.055865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T07:20:38.534666Z digest=sha256:fbacde676d1ef94fa5502f2a5a2adb87d784240ec0c238f8db5279ac0819e155

Observation 3b16d7e7-ecc1-40f8-852d-752cca32ef64 · inbound

Does Verbose Chain-of-Thought Really Help? In-Distribution Evidence that Content, Not Length, Matters cites this paper.

Does Verbose Chain-of-Thought Really Help? In-Distribution Evidence that Content, Not Length, Matters Think before you speak: Training Language Models With Pause Tokens

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:54:21.082071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:45:56.613559Z digest=sha256:dc813abf955ad9dbaf35b93380c1cd43a64503ae484ca8704a13a4b678a47b5e

Observation 7d02fbc4-be96-4efa-8a32-adcc046593f5 · inbound

HALO: Hybrid Adaptive Latent Reasoning for Language Models cites this paper.

HALO: Hybrid Adaptive Latent Reasoning for Language Models Think before you speak: Training Language Models With Pause Tokens

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T08:04:15.548790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T08:04:15.548790Z digest=sha256:22ff6204fdc23fa7d91f316b46cfde9ba6edc60e88daa82776775b634cc8ec31

Observation 51fd1338-e877-4b76-9ef7-4fb9d4d55b48 · inbound

DeepLoop: Depth Scaling for Looped Transformers cites this paper.

DeepLoop: Depth Scaling for Looped Transformers Think before you speak: Training Language Models With Pause Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T05:05:25.293025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:05:25.293025Z digest=sha256:29bbd9bd1904c06a2d273c77381879dfea5549cd8e89c753fd73b2eb46c2d56a

Observation 40798c7b-bb3f-46c2-8587-b87bc6d2b770 · inbound

Is MoE Routing a Huffman Code? Discovering the Frequency-Diversity Law in Chain-of-Thought cites this paper.

Is MoE Routing a Huffman Code? Discovering the Frequency-Diversity Law in Chain-of-Thought Think before you speak: Training Language Models With Pause Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T14:39:31.073737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:39:31.073737Z digest=sha256:51e25c61c6f40bf234c638a11d09ef0a9b0a855cba2d734a9ed232dec0e3f874

Observation 4fbea894-1da0-4fe5-8aa6-8d0816cc42a4 · inbound

Not All LLM Reasoning is Visible in the Chain-of-Thought cites this paper.

Not All LLM Reasoning is Visible in the Chain-of-Thought Think before you speak: Training Language Models With Pause Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T04:14:43.722684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:14:43.722684Z digest=sha256:5f1fd1d8ebd7878064c2e888d0c58e962bf06af0d42ef9935e82130a62911d4a