Pith. sign in

Paper Citation Record · LEDGER

Preference Optimization for Reasoning with Pseudo Feedback

As of 15 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 9 inbound Pith citation observations for arXiv:2411.16345.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16345 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:21:05.745640Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:29:40.974895Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T22:16:16.564391Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved36
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 174dbf0b-60f1-45c0-a44f-29ce76466bd0 · outbound

This paper cites GPT-4 Technical Report.

Preference Optimization for Reasoning with Pseudo Feedback GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.514625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.514625Z digest=sha256:e313b8157cdbfae7dc8a2b5b872d9d02e76442503006ba4fa5335833eb72e324

Observation 13c9d3e8-ae00-4b04-82ef-dc4e014b1d5a · outbound

This paper cites 480k, Iter.

Preference Optimization for Reasoning with Pseudo Feedback 480k, Iter

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:06.424570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:21:05.723444Z digest=sha256:bfbea072ef98e94126ad15b03ebfcda7339fab00b29293fa438a6daad31d4abb

Observation 63f44db6-bc33-4bd8-8334-7573b55d8cf2 · outbound

This paper cites Program Synthesis with Large Language Models.

Preference Optimization for Reasoning with Pseudo Feedback Program Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.525112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.525112Z digest=sha256:82fce04bbf2c359fefdcdc665258ff94bd448a0a3b22dead2c55dc81357c9dce

Observation 52744be6-3037-4c8f-8761-bbd77b8bdd44 · outbound

This paper cites Scaling Synthetic Data Creation with 1,000,000,000 Personas.

Preference Optimization for Reasoning with Pseudo Feedback Scaling Synthetic Data Creation with 1,000,000,000 Personas

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.533533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.533533Z digest=sha256:3ebfaa05e5f00d66bb140328143d08b4866ec9ffc7706ec33c0834fda108236f

Observation 62556cbc-376b-4d06-a659-ee80689f64c4 · outbound

This paper cites Bootstrapping Language Models with DPO Implicit Rewards.

Preference Optimization for Reasoning with Pseudo Feedback Bootstrapping Language Models with DPO Implicit Rewards

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.539306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.539306Z digest=sha256:cea33d2f88b2f2ba26e7a2d6ec423026f1fbf8e4a2f77d62c618c5f9a6a671d5

Observation e07a0c6b-7b11-47dc-8d52-cf7a76238c08 · outbound

This paper cites (4) Estimate the expected returns of each prefix by checking the results approached by the completions.

Preference Optimization for Reasoning with Pseudo Feedback (4) Estimate the expected returns of each prefix by checking the results approached by the completions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:06.443302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:21:05.718597Z digest=sha256:9252b95ede73550d92e7d3cee6a299aa53dabbc36555687d3d9488774c7b1684

Observation 1206815a-5c8f-4d7a-b4ae-d8e15f002a5f · outbound

This paper cites SAIL: Self-Improving Efficient Online Alignment of Large Language Models.

Preference Optimization for Reasoning with Pseudo Feedback SAIL: Self-Improving Efficient Online Alignment of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.549512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.549512Z digest=sha256:008f7b4a75460344b6e90ea8476ad0dac078d15553756df01812bc011151a47d

Observation 742c295d-db99-4c47-8011-fd8e06a9928d · outbound

This paper cites RAFT: reward ranked finetuning for generative foundation model alignment.

Preference Optimization for Reasoning with Pseudo Feedback RAFT: reward ranked finetuning for generative foundation model alignment

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:06.501219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:21:05.554540Z digest=sha256:dae42b5bb630bf20c1f71bb561dffb13c3a3cba0963e2d2c8d35237e924b2c81

Observation 6bbc4a7e-b8cc-4b9e-86d0-5d060bdf148d · outbound

This paper cites StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback.

Preference Optimization for Reasoning with Pseudo Feedback StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.560207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.560207Z digest=sha256:b18df77cbecc9a462bef628d139ef47ca74478db8a3ffd225aed899aee909157

Observation 88cfdff8-8442-4637-94b5-ec8410ce3a47 · outbound

This paper cites The Llama 3 Herd of Models.

Preference Optimization for Reasoning with Pseudo Feedback The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.568833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.568833Z digest=sha256:9fc75418b4ab17dc744ec12081fd0681e8b60a20a9bc4946afb95f71bd8c37ce

Observation dc02ae9a-4bb2-4e7e-a6bd-8032e0b7fe20 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Preference Optimization for Reasoning with Pseudo Feedback DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.574501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.574501Z digest=sha256:d7199cc36a2961841b1f9633a851edb1005c11aacbf8851b167619c9f7b14981

Observation 9d74107e-23e0-44ae-ab34-1100b62bb662 · outbound

This paper cites Solving Math Word Problems by Combining Language Models With Symbolic Solvers.

Preference Optimization for Reasoning with Pseudo Feedback Solving Math Word Problems by Combining Language Models With Symbolic Solvers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.578977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.578977Z digest=sha256:b3189d93a7c52930ea14f2ec8f20492549986da588bd363217e54ed00bffdde9

Observation 2a9d878a-4312-4a8d-acc5-514af8aca66f · outbound

This paper cites Large Language Models Can Self-Improve.

Preference Optimization for Reasoning with Pseudo Feedback Large Language Models Can Self-Improve

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.583321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.583321Z digest=sha256:9d482ffaf213e79496f151435a48a9f71b7c1aec154fc2d2e884e6019ba76b81

Observation 9973ac34-a7a1-49f4-b930-69c74e1919cb · outbound

This paper cites Qwen2.5-Coder Technical Report.

Preference Optimization for Reasoning with Pseudo Feedback Qwen2.5-Coder Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.588551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.588551Z digest=sha256:3282b66e0234961610cd5d3bbf389481a0a0c75dde3fe024bf32ebae4fbedc39

Observation 8a9df988-0678-4406-9028-f2e24909d756 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Preference Optimization for Reasoning with Pseudo Feedback LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.592704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.592704Z digest=sha256:7fc3b4fe592b456605f8b6cfca0428d3daea237fc12b63d8bbb39c07ed568aef

Observation 5709a708-bbc7-4505-815a-6730cf786dff · outbound

This paper cites Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing.

Preference Optimization for Reasoning with Pseudo Feedback Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.597420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.597420Z digest=sha256:47ac8d2668544d86966aa0f96b0efcf5a78cd960e6e057b209b96def2d25b159

Observation 9bd8fbcd-961d-4b06-b92d-302f4ed7e754 · outbound

This paper cites Prover-Verifier Games improve legibility of LLM outputs.

Preference Optimization for Reasoning with Pseudo Feedback Prover-Verifier Games improve legibility of LLM outputs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.601952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.601952Z digest=sha256:dedc78780682db4fb3d300904db616cf739ada715e0c0be1bdb6703b1abf9d6d

Observation c6ee085e-1b7d-49ac-a12d-16e73d18f24b · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Preference Optimization for Reasoning with Pseudo Feedback Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.606869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.606869Z digest=sha256:d6d1c8bbbd1121d803729c707ad7eb789fc7d8e6d1393a5ecad306e0c5d561c9

Observation 8d807568-e2d1-415a-a1cc-302db3616210 · outbound

This paper cites Coderl: Mastering code generation through pretrained models and deep reinforcement learning.

Preference Optimization for Reasoning with Pseudo Feedback Coderl: Mastering code generation through pretrained models and deep reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:06.485975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:21:05.611468Z digest=sha256:759ecccc01381d987c9925548ec9a68223f60de552f3e746efe45acc78cf451a

Observation 7ebf23aa-7318-41e2-8901-da65faefeca4 · outbound

This paper cites Making Large Language Models Better Reasoners with Step-Aware Verifier.

Preference Optimization for Reasoning with Pseudo Feedback Making Large Language Models Better Reasoners with Step-Aware Verifier

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.617722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.617722Z digest=sha256:92f29b4a06abbe697bfaaa8eddc2a638f336d281aa1118fc7692b703447dcd22

Observation c7cdcc22-a9ac-466e-90fc-935bbc392553 · outbound

This paper cites RLTF: re- inforcement learning from unit test feedback.

Preference Optimization for Reasoning with Pseudo Feedback RLTF: re- inforcement learning from unit test feedback

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:06.472642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:21:05.628113Z digest=sha256:22d82f5b66507a6cac16ecb741330e54430cd629307524c0df3873a3cd5d9725

Observation a20168ce-c9bb-42fa-8891-296854d86799 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Preference Optimization for Reasoning with Pseudo Feedback WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.632235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.632235Z digest=sha256:d9ee2b241eb266ed5f3c53aaad24bebcc9867040609adc36cc7852c1e4b97d99

Observation c42a274a-8769-41e0-b0fc-7c518f53c694 · outbound

This paper cites Orca-Math: Unlocking the potential of SLMs in Grade School Math.

Preference Optimization for Reasoning with Pseudo Feedback Orca-Math: Unlocking the potential of SLMs in Grade School Math

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.636588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.636588Z digest=sha256:a69cf3ad33216c5eeaf07f5bacbf5a07a12828f0f3f037f10ad3362e144fa2a0

Observation 3be9682e-a04f-4a89-bf56-824d1c93ef78 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Preference Optimization for Reasoning with Pseudo Feedback Code Llama: Open Foundation Models for Code

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.641756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.641756Z digest=sha256:5e6b856e09f084d9f08acd2e91a030927c37764155867d5a2ce4a66ebbc4981c

Observation 7f48d79d-c892-4a5e-bcc2-a9e028bd1c30 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Preference Optimization for Reasoning with Pseudo Feedback Proximal Policy Optimization Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.646572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.646572Z digest=sha256:dfdc68973c1986dccd9183d750ae8a933d9ed2e46b78295cf6c4b57ce71d8fd3

Observation 9fc22064-f5f7-43ef-b6f2-38f182566959 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Preference Optimization for Reasoning with Pseudo Feedback Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.657171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.657171Z digest=sha256:74bad5dd52a39dfe0165db3772975c0cc884b907e9440d41f52cc5a4efc473b2

Observation e8ee0290-5a30-48a9-a85c-fdb7b78c804e · outbound

This paper cites MathScale: Scaling Instruction Tuning for Mathematical Reasoning.

Preference Optimization for Reasoning with Pseudo Feedback MathScale: Scaling Instruction Tuning for Mathematical Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.661293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.661293Z digest=sha256:4cf9f44e172d6261c2a03e17293a311e046a0646bd89bb15b1476924bb6b0abb

Observation a46cb532-a946-4539-bdfa-2eae57a40fca · outbound

This paper cites A Survey on Self-Evolution of Large Language Models.

Preference Optimization for Reasoning with Pseudo Feedback A Survey on Self-Evolution of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.665378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.665378Z digest=sha256:a58582195cb499b77f908942804bcd7e2490650b072ad7cc4543b2d43eb9a54d

Observation 01e967de-159c-4f66-9986-428917799a3d · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Preference Optimization for Reasoning with Pseudo Feedback Solving math word problems with process- and outcome-based feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.669950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.669950Z digest=sha256:d7884ebaec4e7a94866a831053158136a26b35b2bd3e3c20b8a4f810ad545059

Observation d87adb14-b5db-48e3-9332-d04f921936df · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Preference Optimization for Reasoning with Pseudo Feedback Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.674766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.674766Z digest=sha256:c8d03f257e9fe86ca659b594142954dca91138e9519fd8bead4d7d26b1ab17a6

Observation 34f12c37-2422-4376-8125-13d1e8e63036 · outbound

This paper cites Enabling Language Models to Implicitly Learn Self-Improvement.

Preference Optimization for Reasoning with Pseudo Feedback Enabling Language Models to Implicitly Learn Self-Improvement

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.681953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.681953Z digest=sha256:123cab1c4329baace1d94592461375545b02dfc8984fde2dd8ca18ad466f1550

Observation 2c0632bf-d3c1-44f4-97b5-d612d366652c · outbound

This paper cites Large Language Models are Better Reasoners with Self-Verification.

Preference Optimization for Reasoning with Pseudo Feedback Large Language Models are Better Reasoners with Self-Verification

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.689035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.689035Z digest=sha256:f41ae289d0ff78130dd2f9d81f8c57bfcaa8325c8e58f57d06af0ee6102da821

Observation 219e56c7-aafb-4429-97ca-55db9a689739 · outbound

This paper cites CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences.

Preference Optimization for Reasoning with Pseudo Feedback CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.693955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.693955Z digest=sha256:b56ae79617a50d70b191d813e218d3cac4b4237cfede6795a7834f96255fe141

Observation 1ee6d23c-8f26-4e25-a58a-0e9420d48bd2 · outbound

This paper cites Qwen2 Technical Report.

Preference Optimization for Reasoning with Pseudo Feedback Qwen2 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.698888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.698888Z digest=sha256:19b73c46f71b5e72681ab5fd2eda0451391399f278d9e12986803edb9f32b4ca

Observation d30c269e-3703-499c-a810-2e225c7ff2d0 · outbound

This paper cites Self-Rewarding Language Models.

Preference Optimization for Reasoning with Pseudo Feedback Self-Rewarding Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.703615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.703615Z digest=sha256:aaa413d9c8e8eb29f94e4398339ba8c229f994829d20e7e703c101217cdb33f1

Observation 757780c0-29b1-4677-992f-5d8bcace9bd7 · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

Preference Optimization for Reasoning with Pseudo Feedback MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.709656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.709656Z digest=sha256:9967edd0a1a9eea93d9fb924080de5e11874baf1f36b624951e9bee1f6de2bc4

Observation 659c4b6a-699d-458e-93a6-3c57dca53e67 · outbound

This paper cites too easy.

Preference Optimization for Reasoning with Pseudo Feedback too easy

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:06.457831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:21:05.713971Z digest=sha256:455a21a08d5d764e4669ad24542b9289b9f850a2f2d1c4efe346d87a3e2659bb

Observation 64c1ed25-ba48-4e3d-ad9e-a1f4ed295f28 · outbound

This paper cites 0001", "11.

Preference Optimization for Reasoning with Pseudo Feedback 0001", "11

Reference 43

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T13:21:06.402307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:21:05.727286Z digest=sha256:f7dd3ee4198dfea44a9dabf7bf570c4692bc4f508bacd486e9290628d68896ea

Observation d6b5d902-eb53-4160-b58c-3591936cac53 · outbound

This paper cites test_case_0.

Preference Optimization for Reasoning with Pseudo Feedback test_case_0

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:06.384029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:21:05.731031Z digest=sha256:fde793d1c44a6a110d54400e615d99848cf04c1620099c056240e8f0d6c2d9b9

Observation b50a3e24-c244-41c5-9317-f53465fe7185 · outbound

This paper cites Return the least number of operators used.

Preference Optimization for Reasoning with Pseudo Feedback Return the least number of operators used

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:06.369105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:21:05.734784Z digest=sha256:561e5eab4e9d3806238c96384571f466aad181f73e321ff1b3e6fe70ac1a935c

Observation 051bb9fa-5c7d-42cd-a092-722743d0964e · outbound

This paper cites I will show you a programming problem.

Preference Optimization for Reasoning with Pseudo Feedback I will show you a programming problem

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:06.354072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:21:05.738476Z digest=sha256:ebb226898930275316949eb946be0ee4ef4942a1d39ed04e75d8f1327b646d6d

Observation a553c65a-15e6-4899-899d-ff2bf26bccb6 · outbound

This paper cites test_case_0.

Preference Optimization for Reasoning with Pseudo Feedback test_case_0

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:06.340237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:21:05.742137Z digest=sha256:15169ae27ac3002d879b7389e4b1c697de332947d2d489c2d6297343387b4328

Observation 0b0f3bef-9296-43d0-923d-6a58d8aafcba · outbound

This paper cites Return the least number of operators used.

Preference Optimization for Reasoning with Pseudo Feedback Return the least number of operators used

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:06.325635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:21:05.745640Z digest=sha256:48cfc03c0d1fcad577b508e71810fe57faff9ffe3abf869d9e2b00ed41f6e339

Observation 94f8356a-8e62-4cdf-bc87-0aa32a1aa7a8 · outbound

This paper cites The Curse of Recursion: Training on Generated Data Makes Models Forget.

Preference Optimization for Reasoning with Pseudo Feedback The Curse of Recursion: Training on Generated Data Makes Models Forget

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.651965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.651965Z digest=sha256:2460907bb161b9464bf8c0c116b74936586281e2b3de566798dce469c93cea2a

Observation 294e1774-b989-4d44-b7fe-9e743c921d60 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Preference Optimization for Reasoning with Pseudo Feedback Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.544454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.544454Z digest=sha256:626c5f24db7709fbb117690e9af890e717566a8b7817bfa20b70fbbc3b3a1e72

Observation 042d3cc4-9493-4a1c-ab21-7ec0013f7a15 · outbound

This paper cites Self-Consuming Generative Models Go MAD.

Preference Optimization for Reasoning with Pseudo Feedback Self-Consuming Generative Models Go MAD

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.520773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.520773Z digest=sha256:7b11043d774c5134f44466942c7b8c9e15463340f04afbf42e5529494271a90f

Observation c4551775-ebc5-4bce-950f-c5d9e3dcf410 · outbound

This paper cites Measuring Progress on Scalable Oversight for Large Language Models.

Preference Optimization for Reasoning with Pseudo Feedback Measuring Progress on Scalable Oversight for Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.529270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.529270Z digest=sha256:c8625bc09e44ac5d3c3cc0fd502b09a1c09ccc04d6d34fd4284b5ac3892d90a2

Observation a373ae0a-b483-4afb-96c9-5031545f29c5 · outbound

This paper cites Competition-Level Code Generation with AlphaCode.

Preference Optimization for Reasoning with Pseudo Feedback Competition-Level Code Generation with AlphaCode

Reference 5333

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.623358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.623358Z digest=sha256:072a2e27e9ad32a1dd77b96b5915b45e894a2281b04dbb91acb0ca2e6d1fa666

Pith citing papers

Observation 80cd2e01-5b30-42e1-98d0-4513221881c6 · inbound

ACECODER: Acing Coder RL via Automated Test-Case Synthesis cites this paper.

ACECODER: Acing Coder RL via Automated Test-Case Synthesis Preference Optimization for Reasoning with Pseudo Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T14:53:49.750394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:53:49.750394Z digest=sha256:b8ce5128ae8c7b638cdc31dc326cfe9dbf276cce2b2f6551abbb22b9d4d0680c

Observation dcd9bdc4-0142-4991-8947-eef73a0f7d43 · inbound

Visual-RFT: Visual Reinforcement Fine-Tuning cites this paper.

Visual-RFT: Visual Reinforcement Fine-Tuning Preference Optimization for Reasoning with Pseudo Feedback

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:16:16.566145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T22:16:16.528682Z digest=sha256:1d9132d7abb6a06abc58ee15a32cff338d1c745b2c2991da7810d5f187e98ee4

Observation 26e99566-4cca-4a75-bc1b-f5fedb2a0b7a · inbound

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning cites this paper.

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning Preference Optimization for Reasoning with Pseudo Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:51.424112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:51.424112Z digest=sha256:f8cc6581cd3f9b5ea616142d2b7234e5ed217a89a249f58a25ac30a54f45ff23

Observation e8e3f768-a80e-443a-9303-c26bbbe2ec9a · inbound

Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure cites this paper.

Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure Preference Optimization for Reasoning with Pseudo Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:59:56.058442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:59:56.058442Z digest=sha256:33c30efccf8d051297fc7257a1bb128e63f5cdd4aca2ce0ed47b8fcd6c6cef44

Observation 7f452cb3-bc9a-4c92-8f82-d632e1321c79 · inbound

AdsQA: Towards Advertisement Video Understanding cites this paper.

AdsQA: Towards Advertisement Video Understanding Preference Optimization for Reasoning with Pseudo Feedback

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T20:20:36.762837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:20:36.762837Z digest=sha256:942619212f8a4383e59ce3eb2a9b20921a7d2d07942db506e94a0ed7becee787

Observation 218d65f0-644c-4e71-90f8-e147acf3ad4f · inbound

Self-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM Reasoning cites this paper.

Self-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM Reasoning Preference Optimization for Reasoning with Pseudo Feedback

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:41:01.874936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T05:40:25.414166Z digest=sha256:77bb891c17dc104d6f20962955648746dfef7a2166399034e9008c7c7def2cbc

Observation 1f28dbfd-d9dd-415e-8e08-3788e98bfeb9 · inbound

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization cites this paper.

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization Preference Optimization for Reasoning with Pseudo Feedback

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:37:29.881455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-13T07:32:58.404947Z digest=sha256:b76ed6fd94f7188e9a623022653914388c07dbc2382fb8d883f24830fbf2e272

Observation 703e2227-665d-42d4-b3ba-11c08ed419cc · inbound

SocietyBench: Forecasting Counterfactual Social-World Evolution cites this paper.

SocietyBench: Forecasting Counterfactual Social-World Evolution Preference Optimization for Reasoning with Pseudo Feedback

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T04:17:18.595854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:17:18.595854Z digest=sha256:cd58e6e7730139e835c89e5dd2d24cb2597e1ec22c0c6b2c4179fcf5f18bf60b

Observation b292d5d9-3f2f-4837-adf8-4d362e062637 · inbound

SocietyBench: Forecasting Counterfactual Social-World Evolution cites this paper.

SocietyBench: Forecasting Counterfactual Social-World Evolution Preference Optimization for Reasoning with Pseudo Feedback

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:40.974895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:29:40.974895Z digest=sha256:2e5e39824ac45f3a9b94d8540abca4ff37232346e71d55051435e96038126314