Pith. sign in

Paper Citation Record · LEDGER

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment

As of 23 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:2411.11681.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11681 v3

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:20:14.079797Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:06:21.444253Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T21:06:22.177325Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4064787b-f7f6-443e-8f16-7f71e4de827f · outbound

This paper cites G.; Guo, Z.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment G.; Guo, Z

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:20:14.762028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.855145Z digest=sha256:3ce66964126715a61b8e6d621fcd37459a5a37065d637d808efdee0b557d4f5e

Observation aa7c7899-4785-4ec0-a6aa-cfd9ab8ec6e0 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:20:14.743311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.862914Z digest=sha256:d7288d643fe78f4a8c3763436d2bc6aedbe33ee8d7cf95f3f691100815a5b166

Observation f80b4755-1396-47f1-9daa-b465b198ac61 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.869596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.869596Z digest=sha256:22a7d7e42096d4279f9403df2c5f39711ff5381c89a6591cd95790225d7e5220

Observation 06e1e333-0396-420c-ae1e-1cf16f1ec4c3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.876021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.876021Z digest=sha256:87ad8fe547a96f47598c7af91eef15ff333e9f0d2f3b65a3ef3e4b1881ec9f1a

Observation a0454d7f-5243-4c11-bf9b-2248ba093ae7 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.883203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.883203Z digest=sha256:301ff859e5f71f1d54ecb40f4153537e3e0fa7389a1fe029dfbada9fb04ec77a

Observation 8043d3c7-1bee-4bd9-afa5-387d339eff64 · outbound

This paper cites The Llama 3 Herd of Models.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.890056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.890056Z digest=sha256:aedb4dea5748aa1aadfb1a0cb3e10c96f531eaf313832aaee4ab60f775da2540

Observation eaccdb49-ae77-4a11-a633-333202d96141 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:20:14.699714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.896845Z digest=sha256:8c0054e83acec0c0635d59eebd2dfca482e57f9057f3a8ed492c312324318781

Observation 9497a753-89b6-4f82-9ce7-188687de12d1 · outbound

This paper cites J.; Shen, Y.; Wallis, P.; Allen - Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment J.; Shen, Y.; Wallis, P.; Allen - Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.902327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.902327Z digest=sha256:f2a2ed577a8d2e60247dede4375b9c94edcc949f00c93acdcf833c3d385a6d64

Observation 0f71042d-a3e2-432a-9b13-8d46bdd26366 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:20:14.666411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.907171Z digest=sha256:3ec90fa74c4f44c5d3a4590821a6a6743ef2ac647d07dd7736cadfa01b9e7db7

Observation 1ac934fb-3d04-40f6-91a5-53573afec9ee · outbound

This paper cites The Impact of Reasoning Step Length on Large Language Models.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment The Impact of Reasoning Step Length on Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.912332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.912332Z digest=sha256:994b02fe48632e49da0429d39756cd48871590275f7747a08257c13b242b886d

Observation 17919b9b-4c53-43bc-83b3-d6b809a27bc3 · outbound

This paper cites Large Language Models are Zero-Shot Reasoners.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Large Language Models are Zero-Shot Reasoners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.918672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.918672Z digest=sha256:b07c09b655ae77b60040753f25956eb5218d5dfbe833e905fc6eb56fd20ce5a4

Observation bf05dcc1-26d7-43c1-8876-d01f3cb28f33 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.925085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.925085Z digest=sha256:f4d597857c05fc44e9147338a97126b5541e3f6a61434aac4b9c36c2d2d018a5

Observation c4802d0c-712b-4a43-b151-360f2d971be9 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:20:14.648304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.931585Z digest=sha256:2653cafc4b2af4a8677710862101509ac1776f273ef47d7e61fe87a2d55d7d9b

Observation 80bd7a93-1d16-4484-a643-d5174506a6a4 · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:20:14.626816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.937717Z digest=sha256:a8d6eb1f678aebbd952f20b6d72989a4bce678ce792ccd05ad52abb1b0f2838d

Observation 288e386e-b59c-470a-8802-5dd726898d8c · outbound

This paper cites Let's Verify Step by Step.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Let's Verify Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.943239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.943239Z digest=sha256:edf2d1218f573faac2ed736e3515ffaebe1b4d27e6586f59ae1c646bf65fa22b

Observation 3637dcc3-b26e-4069-b076-782c3ec85f84 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.949599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.949599Z digest=sha256:4aad8ceac60d785496da7c62e1496ffb6c898cd537f09e939bb1b4c4c56b65e4

Observation 935a371f-ddf4-4ee1-b8bc-22d547689031 · outbound

This paper cites Let's reward step by step: Step-Level reward model as the Navigators for Reasoning.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Let's reward step by step: Step-Level reward model as the Navigators for Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.956530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.956530Z digest=sha256:5fa5cab68e4d5b32d1223abd78b268ad0d5a07ad402726754fb19dc148523433

Observation 69447051-366a-465c-8477-f75bb9e5cf60 · outbound

This paper cites L.; Bari, M.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment L.; Bari, M

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:20:14.605398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T18:20:13.964313Z digest=sha256:90872bf5941fe528b7270be6053afa87e25d64d519af94819f313065c16018ec

Observation 0c50e244-fe21-46ca-8f14-cea0f08fe4e8 · outbound

This paper cites L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.971334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.971334Z digest=sha256:ff086195ef3bc03226a03c485f88bb8ab8dd5bce7a7b287969a40d3f46efcb4d

Observation 9a045a4e-846e-477d-a9fe-30b6ebb88054 · outbound

This paper cites D.; Ermon, S.; and Finn, C.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment D.; Ermon, S.; and Finn, C

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.978109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.978109Z digest=sha256:293b054431292fecc4908d26042431cee9210721835b852645e362be9f122080

Observation 23db6937-3951-40e4-9360-c451fe2ac361 · outbound

This paper cites Proximal Policy Optimization Algorithms.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Proximal Policy Optimization Algorithms

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.984199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.984199Z digest=sha256:66a0bf099e089c6c0f8a5bf247309469c20bde6428cabeaf3ed8a2ea0d392c9b

Observation 9da5c1f5-4f36-48fc-b445-30169a388486 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:13.990185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:13.990185Z digest=sha256:b3e2a772bf309b0dbd6c280fff64d4fe1922e71adce9ef140639ad0006ea6b75

Observation 69932421-385d-4f3a-9afc-315608672f4b · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Solving math word problems with process- and outcome-based feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.001444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.001444Z digest=sha256:870a93f6a6d13cc4c64de5c2f7d243dee5c526f8ddce2f3bb0e81b18f4a9dc06

Observation 7ac11809-c912-46ac-a4e3-d6e3c29ad4a2 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.014172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.014172Z digest=sha256:8418759f84f111539cd7d1c9801f215fd54a2414173cfdc036154f89ece3d6b5

Observation 7a8d09c2-9393-499f-af82-283e414df92d · outbound

This paper cites V.; Chi, E.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment V.; Chi, E

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:20:14.560218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T18:20:14.020060Z digest=sha256:acf4e9b34d9bfcbcea7f1b8f97969629356e583dd3a8287e3127e21a6e617e01

Observation d195a577-dc7f-4dc9-bf04-853b4077e306 · outbound

This paper cites Aligning Large Language Models with Human: A Survey.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Aligning Large Language Models with Human: A Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.025705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.025705Z digest=sha256:204860d03d65134c5874d81db323f2cb629924739094c25aa303e2d6c218fcd6

Observation c7a366ea-27b8-4c33-909a-2a9edaa7432d · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.032820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.032820Z digest=sha256:8b42d98306251c9b47ebbd53b873d452f59d57946e985438b8daced8dd837e66

Observation 573e6d72-f60c-4d78-8631-f0b8bf4e563e · outbound

This paper cites H.; Le, Q.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment H.; Le, Q

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:20:14.542206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T18:20:14.038622Z digest=sha256:1c821cfd3cbafabc944ab454def1f7bda0e81ffb4a1688bf91c9ee0ba706ffc7

Observation 79d8f286-c094-44d7-87c7-74217c82891a · outbound

This paper cites Qwen2 Technical Report.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Qwen2 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.043754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.043754Z digest=sha256:1f6cb07244a04973da48cc0f376021bfbc5491b592aba52dab53a0957c9862cd

Observation df8da85d-3057-4c60-847e-8f10c225b51b · outbound

This paper cites an unresolved cited work.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:20:14.520712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-12T18:20:14.049470Z digest=sha256:79da35e99cce27b2630bcab790f2a53ca569ac262ef49328afc1f49df26f6518

Observation ec768db9-f861-464d-ae29-35ae60d00303 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.055687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.055687Z digest=sha256:156d08fff825ab1451521ce0da3bf64c549c2edb1db4268777ae16f054dc33c6

Observation e065f485-159f-40ea-b431-68dd1d3c3f96 · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.061179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.061179Z digest=sha256:d50300d7c393e793299462a3bcf60993d64b581fd59d3d27a31ab6a3713f556d

Observation 948fde31-2f2d-4ef5-8106-19293324a369 · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.066939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.066939Z digest=sha256:abefaf11b1d26f9f6565ea8dede843aa201a3f0ea53a61d058b4ae70c154e9f2

Observation 3e0b9875-2783-407a-a813-3cbad5e146a6 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment , " * write output.state after.block = add.period write newline

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.073328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.073328Z digest=sha256:3195a899afad713b65874e1a491b547bd74a017e24955cc7a8d78425070a0068

Observation de6d25ca-9a10-465f-b98d-d47e637beaef · outbound

This paper cites write newline.

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment write newline

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T18:20:14.079797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:20:14.079797Z digest=sha256:04a1945a1470b88dcd44a1de796768042bcff7c75757fecbf3bf072dfc205787

Pith citing papers

Observation 625b7171-e13e-4332-9566-8c078399fdcb · inbound

SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference Optimization cites this paper.

SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference Optimization PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:06:22.181532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:06:21.444253Z digest=sha256:0a75d86c001aa9c2054ac2fc44a9ea51dc6eab147ee5943918a49415cf162cb3