Pith. sign in

Paper Citation Record · LEDGER

Self-Consistency Preference Optimization

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2411.04109.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.04109 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:51:15.120818Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T13:33:19.400166Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 191a122f-9361-4dbe-a7cd-8653068959d5 · inbound

Self-Training Large Language Models with Confident Reasoning cites this paper.

Self-Training Large Language Models with Confident Reasoning Self-Consistency Preference Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:15.120818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:51:15.120818Z digest=sha256:14631902c71d6e0c2d20b01a552b6e64f5616273d844e54a466d713d7618c3a7

Observation 7c5d64e3-bd74-4310-87b3-46498df3ea7f · inbound

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning cites this paper.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Self-Consistency Preference Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.028303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.028303Z digest=sha256:a38095f4f3a2c01189c238a1045cc418841c569f5ee0a9ed975b4f935d22afe4

Observation 6ed41c72-6d8e-4c2f-b043-a25ab1fadf0e · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Self-Consistency Preference Optimization

Reference 293

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.327311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.327311Z digest=sha256:3fc73ebd716ae5da51f3bdd394b1ff657880aa2632105ae6e393779d40a54e19

Observation 6e28f834-ee8a-4d65-81b9-eb47024abda8 · inbound

PRInTS: Reward Modeling for Long-Horizon Information Seeking cites this paper.

PRInTS: Reward Modeling for Long-Horizon Information Seeking Self-Consistency Preference Optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T20:37:47.736132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:37:47.736132Z digest=sha256:f3e491ab9048a857af6307ee2598c576434e48d721eb355e15c8b85570d50613

Observation 09e4dea1-2b7c-4614-93da-5a41c3457393 · inbound

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model cites this paper.

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model Self-Consistency Preference Optimization

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:46:06.312666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T15:10:16.533927Z digest=sha256:2ffb08dfdb3f1eedb9e4828509e31ed287caa42c6c62c233ebaf1b5672cce4b7

Observation a25e04a3-5dd9-4280-bb24-9d877e87244c · inbound

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs cites this paper.

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs Self-Consistency Preference Optimization

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:45:59.758355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T16:58:10.013475Z digest=sha256:4ade7da8841d3073e7831b4404e88ffaa15493145d0372a8703efaf9667be9ff

Observation 0cc65ef4-4238-4b39-8c76-c2cdaf7592f5 · inbound

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination cites this paper.

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination Self-Consistency Preference Optimization

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T18:02:42.281607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T18:01:06.649723Z digest=sha256:3c2efa0f2d90a44fc29e2470382a3e149dbdd36417cad79401a4d27358a47a6a

Observation 6bd566af-f63d-4dde-b37d-7a3dc918f11b · inbound

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination cites this paper.

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination Self-Consistency Preference Optimization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T17:52:43.510864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:52:43.510864Z digest=sha256:a765caea713cd3a9a304994493b28c92ea1f63204c85e2898cab042495163581

Observation e4b6e54a-bb67-468a-a97d-a49592c36766 · inbound

Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era cites this paper.

Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era Self-Consistency Preference Optimization

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:33:19.402214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T13:29:11.509966Z digest=sha256:99426f002be2771df9275d1ccd761ebe4461a1046af63c1379e1a89fbb4f90ee

Observation 2d29a922-e717-4ef2-9cc0-0f2b45dcf30a · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Self-Consistency Preference Optimization

Reference 94

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:f85958e47825e3b3f71506a4501244a9f26ca5442367b1747bedc878b210fff4

Observation c55cb8ae-dc47-45c7-9e27-fe2c1bb069fc · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Self-Consistency Preference Optimization

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:42.168233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:42.168233Z digest=sha256:32bfb0161f953f9c788f35439af9d07a904ad4d460a2110ad84bf23ac91ae0f9

Observation 4caa0439-287f-4296-bd16-6aa4e452b899 · inbound

On-Policy Self-Distillation without Any Supervision cites this paper.

On-Policy Self-Distillation without Any Supervision Self-Consistency Preference Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.234006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.234006Z digest=sha256:0f92ae03e0ed89b928ac1735958826e5ec90eb0c9caf3c61009c2c259e800274