Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T13:21:05.745640Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 9 inbound Pith citation observations for arXiv:2411.16345.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T13:21:05.745640Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T04:29:40.974895Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-13T22:16:16.564391Z
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 174dbf0b-60f1-45c0-a44f-29ce76466bd0 · outbound
Preference Optimization for Reasoning with Pseudo Feedback GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13c9d3e8-ae00-4b04-82ef-dc4e014b1d5a · outbound
Preference Optimization for Reasoning with Pseudo Feedback 480k, Iter
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 63f44db6-bc33-4bd8-8334-7573b55d8cf2 · outbound
Preference Optimization for Reasoning with Pseudo Feedback Program Synthesis with Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52744be6-3037-4c8f-8761-bbd77b8bdd44 · outbound
Preference Optimization for Reasoning with Pseudo Feedback Scaling Synthetic Data Creation with 1,000,000,000 Personas
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62556cbc-376b-4d06-a659-ee80689f64c4 · outbound
Preference Optimization for Reasoning with Pseudo Feedback Bootstrapping Language Models with DPO Implicit Rewards
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e07a0c6b-7b11-47dc-8d52-cf7a76238c08 · outbound
Preference Optimization for Reasoning with Pseudo Feedback (4) Estimate the expected returns of each prefix by checking the results approached by the completions
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1206815a-5c8f-4d7a-b4ae-d8e15f002a5f · outbound
Preference Optimization for Reasoning with Pseudo Feedback SAIL: Self-Improving Efficient Online Alignment of Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 742c295d-db99-4c47-8011-fd8e06a9928d · outbound
Preference Optimization for Reasoning with Pseudo Feedback RAFT: reward ranked finetuning for generative foundation model alignment
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6bbc4a7e-b8cc-4b9e-86d0-5d060bdf148d · outbound
Preference Optimization for Reasoning with Pseudo Feedback StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88cfdff8-8442-4637-94b5-ec8410ce3a47 · outbound
Preference Optimization for Reasoning with Pseudo Feedback The Llama 3 Herd of Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc02ae9a-4bb2-4e7e-a6bd-8032e0b7fe20 · outbound
Preference Optimization for Reasoning with Pseudo Feedback DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d74107e-23e0-44ae-ab34-1100b62bb662 · outbound
Preference Optimization for Reasoning with Pseudo Feedback Solving Math Word Problems by Combining Language Models With Symbolic Solvers
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a9d878a-4312-4a8d-acc5-514af8aca66f · outbound
Preference Optimization for Reasoning with Pseudo Feedback Large Language Models Can Self-Improve
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9973ac34-a7a1-49f4-b930-69c74e1919cb · outbound
Preference Optimization for Reasoning with Pseudo Feedback Qwen2.5-Coder Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a9df988-0678-4406-9028-f2e24909d756 · outbound
Preference Optimization for Reasoning with Pseudo Feedback LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5709a708-bbc7-4505-815a-6730cf786dff · outbound
Preference Optimization for Reasoning with Pseudo Feedback Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bd8fbcd-961d-4b06-b92d-302f4ed7e754 · outbound
Preference Optimization for Reasoning with Pseudo Feedback Prover-Verifier Games improve legibility of LLM outputs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6ee085e-1b7d-49ac-a12d-16e73d18f24b · outbound
Preference Optimization for Reasoning with Pseudo Feedback Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d807568-e2d1-415a-a1cc-302db3616210 · outbound
Preference Optimization for Reasoning with Pseudo Feedback Coderl: Mastering code generation through pretrained models and deep reinforcement learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7ebf23aa-7318-41e2-8901-da65faefeca4 · outbound
Preference Optimization for Reasoning with Pseudo Feedback Making Large Language Models Better Reasoners with Step-Aware Verifier
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7cdcc22-a9ac-466e-90fc-935bbc392553 · outbound
Preference Optimization for Reasoning with Pseudo Feedback RLTF: re- inforcement learning from unit test feedback
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a20168ce-c9bb-42fa-8891-296854d86799 · outbound
Preference Optimization for Reasoning with Pseudo Feedback WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c42a274a-8769-41e0-b0fc-7c518f53c694 · outbound
Preference Optimization for Reasoning with Pseudo Feedback Orca-Math: Unlocking the potential of SLMs in Grade School Math
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3be9682e-a04f-4a89-bf56-824d1c93ef78 · outbound
Preference Optimization for Reasoning with Pseudo Feedback Code Llama: Open Foundation Models for Code
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f48d79d-c892-4a5e-bcc2-a9e028bd1c30 · outbound
Preference Optimization for Reasoning with Pseudo Feedback Proximal Policy Optimization Algorithms
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fc22064-f5f7-43ef-b6f2-38f182566959 · outbound
Preference Optimization for Reasoning with Pseudo Feedback Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8ee0290-5a30-48a9-a85c-fdb7b78c804e · outbound
Preference Optimization for Reasoning with Pseudo Feedback MathScale: Scaling Instruction Tuning for Mathematical Reasoning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a46cb532-a946-4539-bdfa-2eae57a40fca · outbound
Preference Optimization for Reasoning with Pseudo Feedback A Survey on Self-Evolution of Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01e967de-159c-4f66-9986-428917799a3d · outbound
Preference Optimization for Reasoning with Pseudo Feedback Solving math word problems with process- and outcome-based feedback
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d87adb14-b5db-48e3-9332-d04f921936df · outbound
Preference Optimization for Reasoning with Pseudo Feedback Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34f12c37-2422-4376-8125-13d1e8e63036 · outbound
Preference Optimization for Reasoning with Pseudo Feedback Enabling Language Models to Implicitly Learn Self-Improvement
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c0632bf-d3c1-44f4-97b5-d612d366652c · outbound
Preference Optimization for Reasoning with Pseudo Feedback Large Language Models are Better Reasoners with Self-Verification
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 219e56c7-aafb-4429-97ca-55db9a689739 · outbound
Preference Optimization for Reasoning with Pseudo Feedback CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ee6d23c-8f26-4e25-a58a-0e9420d48bd2 · outbound
Preference Optimization for Reasoning with Pseudo Feedback Qwen2 Technical Report
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d30c269e-3703-499c-a810-2e225c7ff2d0 · outbound
Preference Optimization for Reasoning with Pseudo Feedback Self-Rewarding Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 757780c0-29b1-4677-992f-5d8bcace9bd7 · outbound
Preference Optimization for Reasoning with Pseudo Feedback MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 659c4b6a-699d-458e-93a6-3c57dca53e67 · outbound
Preference Optimization for Reasoning with Pseudo Feedback too easy
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 64c1ed25-ba48-4e3d-ad9e-a1f4ed295f28 · outbound
Preference Optimization for Reasoning with Pseudo Feedback 0001", "11
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d6b5d902-eb53-4160-b58c-3591936cac53 · outbound
Preference Optimization for Reasoning with Pseudo Feedback test_case_0
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b50a3e24-c244-41c5-9317-f53465fe7185 · outbound
Preference Optimization for Reasoning with Pseudo Feedback Return the least number of operators used
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 051bb9fa-5c7d-42cd-a092-722743d0964e · outbound
Preference Optimization for Reasoning with Pseudo Feedback I will show you a programming problem
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a553c65a-15e6-4899-899d-ff2bf26bccb6 · outbound
Preference Optimization for Reasoning with Pseudo Feedback test_case_0
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0b0f3bef-9296-43d0-923d-6a58d8aafcba · outbound
Preference Optimization for Reasoning with Pseudo Feedback Return the least number of operators used
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 94f8356a-8e62-4cdf-bc87-0aa32a1aa7a8 · outbound
Preference Optimization for Reasoning with Pseudo Feedback The Curse of Recursion: Training on Generated Data Makes Models Forget
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 294e1774-b989-4d44-b7fe-9e743c921d60 · outbound
Preference Optimization for Reasoning with Pseudo Feedback Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 042d3cc4-9493-4a1c-ab21-7ec0013f7a15 · outbound
Preference Optimization for Reasoning with Pseudo Feedback Self-Consuming Generative Models Go MAD
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4551775-ebc5-4bce-950f-c5d9e3dcf410 · outbound
Preference Optimization for Reasoning with Pseudo Feedback Measuring Progress on Scalable Oversight for Large Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a373ae0a-b483-4afb-96c9-5031545f29c5 · outbound
Preference Optimization for Reasoning with Pseudo Feedback Competition-Level Code Generation with AlphaCode
Reference 5333
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80cd2e01-5b30-42e1-98d0-4513221881c6 · inbound
ACECODER: Acing Coder RL via Automated Test-Case Synthesis Preference Optimization for Reasoning with Pseudo Feedback
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcd9bdc4-0142-4991-8947-eef73a0f7d43 · inbound
Visual-RFT: Visual Reinforcement Fine-Tuning Preference Optimization for Reasoning with Pseudo Feedback
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 26e99566-4cca-4a75-bc1b-f5fedb2a0b7a · inbound
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning Preference Optimization for Reasoning with Pseudo Feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8e3f768-a80e-443a-9303-c26bbbe2ec9a · inbound
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure Preference Optimization for Reasoning with Pseudo Feedback
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f452cb3-bc9a-4c92-8f82-d632e1321c79 · inbound
AdsQA: Towards Advertisement Video Understanding Preference Optimization for Reasoning with Pseudo Feedback
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 218d65f0-644c-4e71-90f8-e147acf3ad4f · inbound
Self-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM Reasoning Preference Optimization for Reasoning with Pseudo Feedback
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1f28dbfd-d9dd-415e-8e08-3788e98bfeb9 · inbound
Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization Preference Optimization for Reasoning with Pseudo Feedback
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 703e2227-665d-42d4-b3ba-11c08ed419cc · inbound
SocietyBench: Forecasting Counterfactual Social-World Evolution Preference Optimization for Reasoning with Pseudo Feedback
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b292d5d9-3f2f-4837-adf8-4d362e062637 · inbound
SocietyBench: Forecasting Counterfactual Social-World Evolution Preference Optimization for Reasoning with Pseudo Feedback
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.