Pith. sign in

Paper Citation Record · LEDGER

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning

As of 18 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2601.21909.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.21909 v2

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:52:57.142548Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 03949154-0246-4626-8627-8107394126c9 · outbound

This paper cites write newline.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:52.404126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:52.404126Z digest=sha256:6211a455d4947bf3aaef4ab7d0960e58c90572873e82c48d146feb93c5eb5623

Observation d299b128-0f29-4165-b051-41baea0c6acf · outbound

This paper cites Large Language Models for Mathematical Reasoning: Progresses and Challenges.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Large Language Models for Mathematical Reasoning: Progresses and Challenges

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:52.486546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:52.486546Z digest=sha256:f2fddec2d784e8573084c4d5368467fe565a64cc99a79eef6ec61c2f9e65781c

Observation 597b6f69-0fa5-4b2c-a2ba-c6d6f82b2314 · outbound

This paper cites T., Feltovich, P.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning T., Feltovich, P

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:52.557446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:52.557446Z digest=sha256:ce034a23ce5ec73bcd9b01557922843070c466cfb26427ba25a548c7498d6399

Observation 0adb5e2d-88aa-4787-bc8d-36d6c8ffda88 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:52.660916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:52.660916Z digest=sha256:ceb574f4c9b2a77bac394dbe5df00686d6df70ca3a9620e89ffe11095add18be

Observation 824797b2-2eab-49ae-887d-e0890fd45205 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:52.737461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:52.737461Z digest=sha256:bb9e98d86a599eeba6f790d570774b08ab0f651677a5b2a0b27c8be1a72a44b6

Observation fd41cb3a-75b2-4c37-9399-d6b65ca77288 · outbound

This paper cites Supervised reinforcement learning: From expert trajectories to step-wise reasoning.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Supervised reinforcement learning: From expert trajectories to step-wise reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:52.808175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:52.808175Z digest=sha256:3d09da72fb8a684cb9d5a9d68ea1ee2b054d41c0df94e14942f57815733e91f4

Observation 73c088e9-c972-4317-a839-67884c173e3d · outbound

This paper cites Theory-based causal transfer: Integrating instance-level induction and abstract-level structure learning.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Theory-based causal transfer: Integrating instance-level induction and abstract-level structure learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:52.870407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:52.870407Z digest=sha256:5fa000dc651d695f1856f50fbf0b44fc32faf89d3821c6a1beee1ac6dc6cc181

Observation df9b6a1c-a1b5-441b-9afe-7731a3d85953 · outbound

This paper cites PAL: Program-aided Language Models.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning PAL: Program-aided Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:52.931921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:52.931921Z digest=sha256:98f93302bcf04e45ad9f07a030efa1743a642f621c7205210aa6d50714d1ab46

Observation 881a0dfa-63aa-45f2-b32b-e6a75935e05e · outbound

This paper cites Structure-mapping: A theoretical framework for analogy.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Structure-mapping: A theoretical framework for analogy

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:52.993653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:52.993653Z digest=sha256:3f23e8d07dd680c8837915d54602d399fbbcaad4c6576c5f15408062f3f44701

Observation 1055b0a2-c601-4c8f-b63e-53dd459cef29 · outbound

This paper cites an unresolved cited work.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:53.043616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:53.043616Z digest=sha256:908a68084690e24563154f2fc1b5ecd21ab3070b818ea06c1cf4d6d7164d7ef5

Observation 38085235-490f-481d-967a-19f282a016ca · outbound

This paper cites The Llama 3 Herd of Models.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:53.099722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:53.099722Z digest=sha256:64959b73d2ef9bb494ba8abff1784a1a0d3bd2a5029e61121c1043a599f9bb84

Observation a1b5cdaa-abd4-4682-ac5e-a3ad8e8a9f89 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:53.153009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:53.153009Z digest=sha256:84a0e7e8564965dedb7cddc6a1e91fff3aeccf94e4be51d2b465a5d01a95e2bc

Observation df4db203-8bbe-46bf-8811-8df07b58f0c7 · outbound

This paper cites an unresolved cited work.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:53.218331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:53.218331Z digest=sha256:cf7a2ec4d6c87f9eae0328fb2908be3b07709a96b709e0b98090e14b150d70ae

Observation f3fd5b45-d18c-46ea-aba0-9bf7d3a6e83f · outbound

This paper cites Compositional generalization through abstract representations in human and artificial neural networks.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Compositional generalization through abstract representations in human and artificial neural networks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:53.265125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:53.265125Z digest=sha256:02559cbf17bff38b3ea05a9ef5d899a578e27a859430a8d20c8db528e85a299a

Observation 84ea9a23-a26b-43c0-8f05-4286be73c63a · outbound

This paper cites Adam: A Method for Stochastic Optimization.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Adam: A Method for Stochastic Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:53.321726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:53.321726Z digest=sha256:0e6130e0d97a74febc8158740ec0e70b73e85c5b235d1410260eea17443b67f8

Observation 2b567197-d207-410b-876c-2ca94fb4ada5 · outbound

This paper cites Mawps: A math word problem repository.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Mawps: A math word problem repository

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:53.380471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:53.380471Z digest=sha256:2e947b28bc1e8b09e4e533301281308142c8e72d198eae6067c824d225f3fbd0

Observation ad00f497-f171-44a3-ba37-c4661987a3df · outbound

This paper cites LLM Post-Training: A Deep Dive into Reasoning Large Language Models.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning LLM Post-Training: A Deep Dive into Reasoning Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:53.480561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:53.480561Z digest=sha256:5dec3fddd1508ca52644ee992fe1a7750391edc8c95e3ab28d1b29154fca740d

Observation 719c9849-4aec-4584-ade7-6c8d7cfd1fb0 · outbound

This paper cites D., Cohen, J.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning D., Cohen, J

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:53.644931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:53.644931Z digest=sha256:644e7c0e3ed16c12378ed8da737077725196b1743520f7a31d376f74fbc35d40

Observation b284252d-bd80-4853-b5c9-9039367201bd · outbound

This paper cites Let's verify step by step.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Let's verify step by step

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:53.794665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:53.794665Z digest=sha256:d3f542fbbf8874672593e9adfc77120751e5544438a526ce6c70a53a404a51ea

Observation 0f778b08-1513-4b56-9814-1927e1be69d3 · outbound

This paper cites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:53.945286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:53.945286Z digest=sha256:96c15d1ed3fafeb43c0cb8d37a6a30ed83dfb1413e9ee7ebf0899690a7cf6aa7

Observation 5234a885-8a0e-45ed-9088-f77498cf2e92 · outbound

This paper cites Towards a unified view of large language model post-training.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Towards a unified view of large language model post-training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:54.119238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:54.119238Z digest=sha256:1a3a0d2259f31e674d19fce11e3741c79f4e12b581ed32c00581d88c40eb63f9

Observation 991f4cc9-d33c-47b8-ba09-2ce868d5f25c · outbound

This paper cites W., and Behrens, T.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning W., and Behrens, T

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:54.246041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:54.246041Z digest=sha256:22c61866bb7b9acbdf8f998bbd89f5deab67f45dffe9ddbee4ee13c356c01515

Observation aed24ea8-d1cd-4731-b99f-36337b5a6877 · outbound

This paper cites A diverse corpus for evaluating and developing english math word problem solvers.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning A diverse corpus for evaluating and developing english math word problem solvers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:54.354117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:54.354117Z digest=sha256:485595fe51dc78255f7ce7887d92901a56547398e2a28a5b0d2a6be5306de656

Observation b9b6fbf3-2b7f-4a9c-9775-b33ea1aee5ba · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:54.465200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:54.465200Z digest=sha256:8834b90f9ca034c9c9e63e5aeaf68903de6b356ee53a683ac99cab98b97acfbc

Observation d77c5eee-0496-4cb1-bdee-1941e1b767d8 · outbound

This paper cites Abstraction and analogy-making in artificial intelligence.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Abstraction and analogy-making in artificial intelligence

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:54.554149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:54.554149Z digest=sha256:ce622fac58510c8e45fce44f75a1756429afcee913026a04332d52ba3621af95

Observation 4cf0a90d-4661-4c33-a30c-9e30a453e41f · outbound

This paper cites Training language models to follow instructions with human feedback.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Training language models to follow instructions with human feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:54.633374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:54.633374Z digest=sha256:c38692ce7d8e0c12a9c8e5d27f6925e37bdc7c0478fda141fd8cfb5ad25a1327

Observation 07e30e33-408e-4ecf-a777-7ff068f5df06 · outbound

This paper cites Plan-tuning: Post-training language models to learn step-by-step planning for complex problem solving.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Plan-tuning: Post-training language models to learn step-by-step planning for complex problem solving

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:54.696039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:54.696039Z digest=sha256:3166f3af216eab6569c4b7c29be531fa978a3b84b2d90fe8e16ed0a2a9755d92

Observation 47b807f3-2509-471b-bee4-cc3985adb1e6 · outbound

This paper cites Are NLP Models really able to Solve Simple Math Word Problems?.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Are NLP Models really able to Solve Simple Math Word Problems?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:54.747035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:54.747035Z digest=sha256:44726df9b649a26802e91be6a725c43b94251273bb830e9c41faecc0a013997b

Observation 5ada046a-e456-41ad-8116-866789fcc489 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:54.813802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:54.813802Z digest=sha256:9a6d8a57e4fe2d70b93c79157742ab7d5c30be24cd05551da902319e3b02fc32

Observation 6be17ceb-6d13-4a5d-98d0-563f1a98f517 · outbound

This paper cites Proximal Policy Optimization Algorithms.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Proximal Policy Optimization Algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:54.896415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:54.896415Z digest=sha256:0a9dbfd8f6a78fd45abfda60b3a79e8378ac001f2f04fe3bbc7cb9b5c98880d5

Observation 7f0ce054-c423-4a49-890b-4e73028b022d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:54.976326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:54.976326Z digest=sha256:8e505b66c0ab221e492a95f8d3632665f7c76728583d03c680badddc40c97988

Observation 1dfbcecf-f88a-41a9-917c-a0789484b3b8 · outbound

This paper cites an unresolved cited work.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:55.035510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:55.035510Z digest=sha256:5df6e974b0fa189d81b4fccf5b9f2e23641a3d3a327d97ba192f3407ebd86b1e

Observation d225e297-c130-4775-ae51-6de35305665d · outbound

This paper cites Cognitive load during problem solving: Effects on learning.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Cognitive load during problem solving: Effects on learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:55.120800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:55.120800Z digest=sha256:93f445d94508c5fa1e16ca2749dbb3e23737b96d6208170fcc2bae5c54cb1c64

Observation 02ef53d2-8ade-4ccb-a886-6e490e394334 · outbound

This paper cites Qwen2 Technical Report.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Qwen2 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:55.203049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:55.203049Z digest=sha256:b949a6b689912c97fb4a172197508de3785da08ef9eebe9a193bd7f0c6c2ccc4

Observation 25f6e867-6b54-4b22-9a09-7d0307f3f696 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Solving math word problems with process- and outcome-based feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:55.390388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:55.390388Z digest=sha256:83ace7ed34d980f91b766e2f0fc669375d1bcc35a0921fded6af9e6e6238492b

Observation 6b42ea3d-cb52-4bb2-a765-04e608af4a2d · outbound

This paper cites Post-Training Large Language Models via Reinforcement Learning from Self-Feedback.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Post-Training Large Language Models via Reinforcement Learning from Self-Feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:55.624948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:55.624948Z digest=sha256:95531ff7d21858323ec4e265924571da38c0e55cf384e637f1b25b9c0fc9edc8

Observation 57a344f1-34a7-4e3f-8772-3b5bc2852c5b · outbound

This paper cites X., Kurth-Nelson, Z., Kumaran, D., Tirumala, D., Soyer, H., Leibo, J.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning X., Kurth-Nelson, Z., Kumaran, D., Tirumala, D., Soyer, H., Leibo, J

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:55.778877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:55.778877Z digest=sha256:dbc94bbb270f73993658a7dcd252234172192900442f226ea1d97c5997042caa

Observation 42edb151-c342-42d8-acae-edab792b795b · outbound

This paper cites A Survey on Large Language Models for Mathematical Reasoning.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning A Survey on Large Language Models for Mathematical Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:56.005103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:56.005103Z digest=sha256:735db2c2f779d43c0648adbb331d2e1ebe14ad732d8e7322dcb61a90020ef829

Observation 3b9fa72b-c343-4503-a22c-74467f29fb60 · outbound

This paper cites UFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Function.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning UFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Function

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:56.188957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:56.188957Z digest=sha256:daf12ef77088ac628cdb941c6ef7b80961e73ef2986abe75fa841ca638f10333

Observation 8c00f3af-7000-416b-bbea-f5d2c84d4de0 · outbound

This paper cites V., Zhou, D., et al.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning V., Zhou, D., et al

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:56.343925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:56.343925Z digest=sha256:8a501186f203854f250895bf4df88066fb9eabac3e79d3857093a30911b08b72

Observation 98225bc7-fb84-4608-8e11-aab62dc43bc2 · outbound

This paper cites M., Meder, B., and Schulz, E.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning M., Meder, B., and Schulz, E

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:56.504633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:56.504633Z digest=sha256:50396fb5d0e68f3c08bf150bd5f547d0f1a6bfff812a36e55bc1797f42281f33

Observation 0f05880c-5c37-43ad-9651-56f7c529caa2 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:56.712409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:56.712409Z digest=sha256:f14eb56da046d0042b205310be324149c347a1aa3d514cb3749b070b8ce8e6a9

Observation 705591ad-5f76-428e-8bec-44a4f72c5fef · outbound

This paper cites Qwen3 Technical Report.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Qwen3 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:56.775626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:56.775626Z digest=sha256:e5e87987fe71f9fa2bc833d889c9a9d7b2bed29bceba2105d1b71676e0587142

Observation aa4efabc-a5cb-4037-a9a2-d13caf639587 · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Star: Bootstrapping reasoning with reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:56.923171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:56.923171Z digest=sha256:baa59d142bb14c42f734d6e1ec5dbfb7bab74049bf7dab51142a457c181c78dd

Observation 1f62c9ae-3988-48bd-8147-6254c80880f5 · outbound

This paper cites Groundedprm: Tree-guided and fidelity-aware process reward modeling for step-level reasoning.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Groundedprm: Tree-guided and fidelity-aware process reward modeling for step-level reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:57.031345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:57.031345Z digest=sha256:0570a7fed72580210b659903b5f310c41538dd4f374406aad4fe5f14b7f29379

Observation 1bbcb95d-0914-45b9-bdb9-a641c5ca952c · outbound

This paper cites Y., Garvert, M.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Y., Garvert, M

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:57.142548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:57.142548Z digest=sha256:fb557cceb1b7fe31330945f83dad3bdb030a8f23314b7577e3f719a3a535256d

Pith citing papers

No inbound Pith citation observations are available.