Pith. sign in

Paper Citation Record · LEDGER

Advancing LLM Reasoning Generalists with Preference Trees

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2404.02078.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.02078 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:17:46.033521Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T09:47:59.922599Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 751ebca6-97a9-4d89-845a-c0bddda875f0 · inbound

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs cites this paper.

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs Advancing LLM Reasoning Generalists with Preference Trees

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:18:01.688186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T16:18:01.560780Z digest=sha256:44960e2dfcbb69d5f52091cfa9d7e2ed1e52ace927106cd6092f74853cab6ebd

Observation fe2c7e9f-6eeb-4093-b116-adbf38238b95 · inbound

Training Software Engineering Agents and Verifiers with SWE-Gym cites this paper.

Training Software Engineering Agents and Verifiers with SWE-Gym Advancing LLM Reasoning Generalists with Preference Trees

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:20:40.532699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:20:40.483057Z digest=sha256:00eccfc331e1040f8fb6b80360d7ed67fda16bf509bbc362b3f2000efd795a49

Observation bbb07ed4-5c56-43c8-9149-8d4d764f1977 · inbound

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model cites this paper.

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model Advancing LLM Reasoning Generalists with Preference Trees

Reference 249

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:30:03.070601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T17:30:02.803757Z digest=sha256:5935952b0a76182c60735992d0b587dfab5ce0e1fdb2b2e369e29f0921af62c7

Observation 1403d4b0-07a5-4db9-9c9d-63e497ddc699 · inbound

DPO-Shift: Shifting the Distribution of Direct Preference Optimization cites this paper.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Advancing LLM Reasoning Generalists with Preference Trees

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:46.033521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:46.033521Z digest=sha256:2f2be06dc0fa6eab081dd5ee3db823ab6721c865eaf3c7696a5ff4dbe24d0ec1

Observation 165f613f-8b50-463b-87c7-d8edfa19b879 · inbound

Measuring Diversity in Synthetic Datasets cites this paper.

Measuring Diversity in Synthetic Datasets Advancing LLM Reasoning Generalists with Preference Trees

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:50.972531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T04:54:50.972531Z digest=sha256:57615be5b1db1aedf6b1191131a2649ad1287cdc21a6489cdb7a0beb487a1a67

Observation 97156b06-e60e-412f-bc12-c2c7951d7809 · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs Advancing LLM Reasoning Generalists with Preference Trees

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:43:00.303113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:13f6d4226223cf24f8c3f8b35128158e1eb4f86a9c43f22dabe5405a747b0bf4

Observation 5bab72b9-5519-4cfb-8f12-b78c64860908 · inbound

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization cites this paper.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Advancing LLM Reasoning Generalists with Preference Trees

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.431718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.431718Z digest=sha256:5c598761ef2e0d423118c35846c9b78e1d0d0f14d8b1f60ce3055e36753966fc

Observation a0780e28-bddc-4838-9500-e3b1140c06a7 · inbound

Towards Reliable, Uncertainty-Aware Alignment cites this paper.

Towards Reliable, Uncertainty-Aware Alignment Advancing LLM Reasoning Generalists with Preference Trees

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.967384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.967384Z digest=sha256:ee0137cb4f4c0c561bba9c852a8d60cef9d202238cdbe2f47b8bf8a43e2ed7a5

Observation 7053f5d5-dea0-40bc-9503-8717edc46fbd · inbound

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks cites this paper.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Advancing LLM Reasoning Generalists with Preference Trees

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.162000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.162000Z digest=sha256:6f7faa1bee3617b1b0a531d7b7e52bb3abff7f638bc3d046753719542c5e6597

Observation de9fd7fd-02f4-48a6-90e0-a0ae86389acd · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Advancing LLM Reasoning Generalists with Preference Trees

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:21.344039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:21.344039Z digest=sha256:ce1eb88b57618611a61414cbb7e00f7cd192da0876cc3787333d1aeb285ead0d

Observation 66f4bc15-23dc-4d73-aba2-5d4f2540df54 · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Advancing LLM Reasoning Generalists with Preference Trees

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.578345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.578345Z digest=sha256:383a977f3f84b3f6d655c940caf9d260cbdb48869506c97189f9d7d3a5937668

Observation b03affbf-8748-48a8-b01b-0021ec3c3e58 · inbound

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling cites this paper.

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling Advancing LLM Reasoning Generalists with Preference Trees

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:20:49.262191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T03:15:57.744706Z digest=sha256:d0204371b229f7aa99bb95ea3024b0aab489f56018f22253eec41cd36bd273b6

Observation c93977e0-94bb-4c22-9f53-e65977d49b3d · inbound

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems cites this paper.

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems Advancing LLM Reasoning Generalists with Preference Trees

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:26:02.393606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:56:00.449776Z digest=sha256:f95229853855d44ad4e7ee7e7c1877260ffd07ff1dce85baf1dfb6a913a2ac04

Observation af10dbf6-ac5e-4141-b505-c0e394271f2b · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Advancing LLM Reasoning Generalists with Preference Trees

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:51:30.256032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:51:52.375703Z digest=sha256:b5c7f385a223d0fe7416a3cecd84de448cf9519d572e0b5344e021c1c8137380

Observation afbbf399-40d4-47e1-8cc6-6f4a444e99ce · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Advancing LLM Reasoning Generalists with Preference Trees

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:15:03.285300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T05:11:32.053440Z digest=sha256:fd54d47d13037ff2f1f8305d11a26c9fe9d5c0f0694a6c887b3cc9cdfe4e9e69

Observation c197a95d-51d7-4132-993a-48c8947a0adb · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation Advancing LLM Reasoning Generalists with Preference Trees

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:12:22.771833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T06:08:28.855671Z digest=sha256:d99dcdc630fc23e6960c4dc9338f5b790655f152e5bae9dca87426b4e3e70dde

Observation cdafa392-70f2-41ba-b055-d46aef9d1d2e · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation Advancing LLM Reasoning Generalists with Preference Trees

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.055182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:c8562bd216a72265fe9baa435998aa02fd4d0a7ed2fefd4dec96ff90e8a0cd81

Observation 0d72bc2b-36a2-4861-a354-1bc2b1e43cf5 · inbound

Selective-Advantage Entropy-Adaptive Horizon GRPO: Asymmetric Token-Level Discounting for Efficient Reinforcement Learning of Language Models cites this paper.

Selective-Advantage Entropy-Adaptive Horizon GRPO: Asymmetric Token-Level Discounting for Efficient Reinforcement Learning of Language Models Advancing LLM Reasoning Generalists with Preference Trees

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:26:46.066822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T06:56:01.601701Z digest=sha256:a21bc0269257a2cba3acf433b0cc96ffc55e7c9da16be4c641fa4496af236482

Observation 0a1dc478-abb7-42a0-b754-eaed8449afc6 · inbound

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output cites this paper.

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output Advancing LLM Reasoning Generalists with Preference Trees

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:47:38.255909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T13:39:17.701196Z digest=sha256:5347066c4285582445c1301c86a53633f1e06463bcadafe9dea5e353a6df5054

Observation 35613fd2-39ca-4567-9df6-1dd09537c8c0 · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning Advancing LLM Reasoning Generalists with Preference Trees

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:47:59.923933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:54198bdcff694bc494fdbb2d4b6421c5b4f832e6a3cce3d8513c5391c5cf6247

Observation cf321595-60bd-4b5a-a8f3-66f03fa2802c · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Advancing LLM Reasoning Generalists with Preference Trees

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:3df6b751ab6d3c92b6b00ad90b1408ea8fe7c4188c5317cdfb27af11871c4de5

Observation 350e3b83-e95c-497e-96fa-f4c9f3db9a24 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Advancing LLM Reasoning Generalists with Preference Trees

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:39.023948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:39.023948Z digest=sha256:22fc84c796c3ca8f612fff945f15d12c6f3302a8406723b11cc9dd347e051f60