Pith. sign in

Paper Citation Record · LEDGER

Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2406.09136.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.09136 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:25:52.153198Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T09:42:04.341075Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5664a82c-edd2-4e6e-9d96-0f330c7ddd46 · inbound

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents cites this paper.

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:42:04.348991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T09:41:59.979595Z digest=sha256:f88a3080ae030c5c1496ff79427fc811a099f19b5262230e39e310688b298b28

Observation 27f697af-24bd-4abb-98f7-64b6c01245af · inbound

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information cites this paper.

LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T13:25:52.153198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:25:52.153198Z digest=sha256:a909dc9edbf22d0a8e222e996d7c2ab4acaa99ab732f94922347ffac3bb763d1

Observation 49cac612-8a2b-49e0-a987-f26a7b92acbb · inbound

CodeSteer: Symbolic-Augmented Language Models via Code/Text Guidance cites this paper.

CodeSteer: Symbolic-Augmented Language Models via Code/Text Guidance Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T12:16:27.445998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:16:27.445998Z digest=sha256:63d554a7d07bc1d2ccad3631526a4a39efad07d922a81347005de9a7600a5bd9

Observation a660e78a-0202-4906-8be5-b681e90e036b · inbound

CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning cites this paper.

CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:36:38.553844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:36:38.553844Z digest=sha256:dc21d12e649695f7123d904a77ef8df1fc77e98d7348c2bc54f794489fd6ee4e

Observation c554ab6f-cec7-4e27-9e0e-129f6f5858ed · inbound

NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning cites this paper.

NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:41:24.286587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T00:51:40.815981Z digest=sha256:43d1c9eca6a2abc6f2ec0266570b1eec8548b905335d3793467ac5576f487d5c