Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2411.04109.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:51:15.120818Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T13:33:19.400166Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 191a122f-9361-4dbe-a7cd-8653068959d5 · inbound
Self-Training Large Language Models with Confident Reasoning Self-Consistency Preference Optimization
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c5d64e3-bd74-4310-87b3-46498df3ea7f · inbound
Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Self-Consistency Preference Optimization
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ed41c72-6d8e-4c2f-b043-a25ab1fadf0e · inbound
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Self-Consistency Preference Optimization
Reference 293
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e28f834-ee8a-4d65-81b9-eb47024abda8 · inbound
PRInTS: Reward Modeling for Long-Horizon Information Seeking Self-Consistency Preference Optimization
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09e4dea1-2b7c-4614-93da-5a41c3457393 · inbound
VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model Self-Consistency Preference Optimization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a25e04a3-5dd9-4280-bb24-9d877e87244c · inbound
Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs Self-Consistency Preference Optimization
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0cc65ef4-4238-4b39-8c76-c2cdaf7592f5 · inbound
TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination Self-Consistency Preference Optimization
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6bd566af-f63d-4dde-b37d-7a3dc918f11b · inbound
TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination Self-Consistency Preference Optimization
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4b6e54a-bb67-468a-a97d-a49592c36766 · inbound
Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era Self-Consistency Preference Optimization
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2d29a922-e717-4ef2-9cc0-0f2b45dcf30a · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Self-Consistency Preference Optimization
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c55cb8ae-dc47-45c7-9e27-fe2c1bb069fc · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Self-Consistency Preference Optimization
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4caa0439-287f-4296-bd16-6aa4e452b899 · inbound
On-Policy Self-Distillation without Any Supervision Self-Consistency Preference Optimization
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.