Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:59:10.184366Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2501.03486.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:59:10.184366Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:44:47.323075Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T14:44:51.507749Z
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f3dbc39f-d98b-4946-8524-69f0e91ce6c1 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Secrets of rlhf in large language models part ii: Reward modeling
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e7bed5ab-415b-4c58-92e4-1b5364fee949 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment A survey of reinforcement learning from human feedback
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6116827e-cf89-4023-b118-7c0d6f88fa36 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment More RLHF, More Trust? On The Impact of Human Preference Alignment On Language Model Trustworthiness
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 18713587-ed6e-41fc-8679-dd676ae9625e · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Safe rlhf: Safe reinforcement learning from human feedback
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a7166d65-becc-46a7-8972-7f215e8d80c2 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Principled Reinforcement Learning with Human Feedback from Pairwise orK-wise Comparisons
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 29f680c8-bac2-490b-9edb-a496cfe2ec85 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment A general theoretical paradigm to understand learning from human preferences
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a3f8eccf-19a6-40a1-9468-0771028071ab · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Fine-Tuning Language Models from Human Preferences
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 391c5cb3-d8b5-49ac-b31f-1a8abcb5dcb4 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Open problems and fundamental limitations of reinforcement learning from human feedback
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation faf43128-af30-4199-8203-5c34d79c08ed · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Training language models to follow instructions with human feedback
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 451a0237-393a-411f-b747-9d97cdec58a4 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Black-Box Prompt Learning for Pre-trained Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e44f31cb-560f-413d-9c4e-1c080c84cfb7 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation de5a6c1d-ba17-4318-ab0f-05d7be0252b3 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Prompt Optimization with Human Feedback
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0b766e39-0900-4563-9c1e-d51fb9106020 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Learning overparameterized neural networks via stochastic gradient descent on structured data
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fcfe79d1-167c-4dde-9fa7-18861793afd6 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment The Power of Scale for Parameter-Efficient Prompt Tuning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation acdeefd7-43be-4106-81c9-fddbc78f261d · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment PRewrite: Prompt Rewriting with Reinforcement Learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 56b98ff1-d3ff-4bfc-8255-9e838e19fab7 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment PromptAgent: Strategic planning with language models enables expert-level prompt optimization
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fcb8f4c1-3cfe-4ea0-8e9f-cd5ffca179f3 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Alpacafarm: A simulation framework for methods that learn from human feedback
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 562926ee-1217-4ae1-a667-cfaad434bce6 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Fine-tuning language models from human preferences
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d18022bb-15a5-4864-aec3-69662755b378 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7becd67c-580b-4559-a3da-63f2671a61e4 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Direct preference optimization: Your language model is secretly a reward model
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 276465e7-0c69-41ef-8326-2cf77d19cd43 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Slic-hf: Sequence likelihood calibration with human feedback
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 312fa8a8-7906-4cf3-94f6-60cd8e4841f8 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Direct Preference Optimization with an Offset
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bda6fc92-a6e3-4b8e-801f-0cc62d25afd6 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment A general theoretical paradigm to understand learning from human preferences
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fd32cd9f-1ff1-4766-950d-7bc08cf6402f · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Mixed Preference Optimization: Reinforcement Learning with Data Selection and Better Reference Model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ef2e944e-cf69-4612-bbce-a07125d2fd51 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment LiPO: Listwise Preference Optimization through Learning-to-Rank
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation acccd9f2-e803-4e7d-a598-21b33f1a4190 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Filtered Direct Preference Optimization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9f50da3e-b0af-4bf3-b854-06da05888027 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Generalized Preference Optimization: A Unified Approach to Offline Alignment
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 79c9d017-81a4-49ba-9a5b-e82bd0174600 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 831659da-f0f9-432b-a8c9-bfb3568592b4 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Efficient Exploration for LLMs
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6cb06d3e-57b7-46aa-9b64-7996b7e80d80 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment ORPO: Monolithic Preference Optimization without Reference Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71467c5b-ca64-498d-ad56-06f457e4537a · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Intuitive Fine-Tuning: Towards Simplifying Alignment into a Single Process
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64b6647a-c6a1-451d-a077-a3130b34f3b8 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Toward Human Readable Prompt Tuning: Kubrick’s The Shining is a good movie, and a good prompt too? In: Proc
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d5c25949-cf2d-42fb-810d-944cd447bffe · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Prefix-Tuning: Optimizing Continuous Prompts for Generation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d345a886-0366-4586-b950-576a41d19a03 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Factual Probing Is [MASK]: Learning vs
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c0d3c806-2aca-48ac-94a8-b8ca6bc2bdf0 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Clip-Tuning: Towards Derivative-free Prompt Learning with a Mixture of Rewards
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3622ab08-5ec3-4f66-bfcc-072b8be0ad4f · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Black-box tuning for language-model-as-a-service
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 46e78e99-c8c2-4041-879d-de6b1e959055 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment BBTv2: Pure Black-Box Optimization Can Be Comparable to Gradient Descent for Few-Shot Learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e68c2763-ce7e-42e5-a131-0e2e0a852c4a · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment MultiPrompter: Cooperative Prompt Optimization with Multi-Agent Reinforcement Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a5f0095c-972c-4273-9dc7-ad532465c38c · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Curiosity-driven red-teaming for large language models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b092b334-e32b-462b-b4cc-09b113439480 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Discovering Language Model Behaviors with Model-Written Evaluations
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dcfc53a-4602-4135-942d-6df245602b2a · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Gradient-based language model red teaming
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f4541b2d-78e4-4bec-8e95-1c95d142d157 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Learning diverse attacks on large language models for robust red-teaming and safety tuning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 82a672a9-7e88-4cc4-bfab-1f566211638a · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment LIAR: Leveraging Alignment (Best-of-N) to Jailbreak LLMs in Seconds
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c3661996-7ddb-4a79-9589-8da45f9648bc · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Rank analysis of incomplete block designs: I
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 40faa462-52eb-4e62-8fe3-ec3fc011815f · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Large Language Models Are Human-Level Prompt Engineers
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation af3e1967-17e4-4c35-9a13-0fec56182b5c · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8372c7cc-3d3d-4eb5-8ad2-29ac7363d3c6 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Reinforcement learning by reward-weighted regression for operational space control
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 014d6aa8-19ef-41d3-9101-89fce8150344 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Direct preference optimization: Your language model is secretly a reward model
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 13b2585d-fd8d-41ca-94d8-326791ea42ef · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment ULTRAFEEDBACK: Boosting Language Models with Scaled AI Feedback
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5370850f-8f2f-4e0c-b091-5a7c8ecdd6c0 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Helpsteer: Multi-attribute helpfulness dataset for steerlm
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5dc24dbd-7e44-4abe-b802-6d683215fcbc · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Orca: Progressive learning from complex explanation traces of gpt-4
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d3f032c6-9ece-4a74-b5af-1af73d25d73e · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 444e9694-5020-48a5-bc32-58c3ba765577 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment While this aspect certainly played a crucial role, it oversimplifies the broader economic and structural issues
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e67b434d-cc16-47dc-8e30-aac1facce29e · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment These tools allowed financial institutions to shift risk off their balance sheets and increase leverage, ultimately contributing to instability
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 24e09e9e-8087-4c34-ab23-dada236ba7d0 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Birdhouses and Animal Habitats:Smaller bottles can serve as habitats for birds or insects
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 73192f2e-d815-474c-9e2a-c1bcd1a10d26 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Garden Tools: Convert old bottles into garden markers, plant markers, or simple tools like a mini watering sprayer
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b7e26d24-e29b-48f9-a72e-00c3c684fb64 · outbound
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment Covers and Protectors:Use them as covers for plants during winters or protect delicate surfaces in transit
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 476f0a0c-491b-4a7c-9159-8d27bfe9a932 · inbound
SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.