Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:34:25.301260Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 117 outbound references and 3 inbound Pith citation observations for arXiv:2507.13158.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:34:25.301260Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T11:17:36.742733Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T07:59:40.152268Z
100 of 117 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 77084717-f2be-497f-8d93-392cf620e010 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a422fa13-9e47-499a-a907-8db37dd31b07 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1257466-8376-4141-82e0-9e481af8b394 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Language models are few-shot learners
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77be57c1-9825-4171-a4c3-dca895a4b655 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities PAD: Personalized Alignment of LLMs at Decoding-Time
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12d4acc6-c2fe-47e3-8ed9-3c0a9e4d5f3e · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Reward Model Ensembles Help Mitigate Overoptimization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91495224-4271-46dd-b6f0-d9345a56b743 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities SEE: Strategic Exploration and Exploitation for Cohesive In-Context Prompt Optimization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6761f935-4f6b-4abd-97c8-6a9572c0cc0d · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities A Survey of Automatic Prompt Optimization with Instruction-focused Heuristic-based Search Algorithm
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02831b20-4cac-4fee-97a0-fbe7e43d64c8 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Bert: Pre-training of deep bidirectional transformers for language understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e79d17c-964f-441e-b61a-0d03c59a8ba3 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 234ef8e0-e046-41dc-8a62-bd88d24ab84a · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c99d0c59-2b0a-4352-94b3-00b75f366909 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities PaLM-E: An Embodied Multimodal Language Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c01026a-8ea6-45ac-b721-2b1569ef96ae · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fa477cb-68ae-4a5e-a4f9-a9649ee2ea25 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities KTO: Model Alignment as Prospect Theoretic Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e90adc2d-d0b0-4c82-816c-7f3fad3051e9 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Self-Attentional Credit Assignment for Transfer in Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f8a58f3-a49d-42a4-b821-d9b72e930c01 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Learning Robust Rewards with Adversarial Inverse Reinforcement Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93ddae34-6421-4d2e-a4a4-6e5dfa256901 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cd38935-20b2-47eb-a5ba-8fa46c2c118a · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b868f90-ea30-4833-9110-e319d1f4a64d · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e45fe2a-0649-4a2e-88a4-cf1875fbf5d2 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Direct Language Model Alignment from Online AI Feedback
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad89a21f-894c-43aa-beaa-cc1ad8ea8a61 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f633d741-1147-4b73-8843-2e5ae5178b08 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Soft Actor-Critic Algorithms and Applications
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db3d3eb6-22a5-47aa-a7d5-f9f1b50f5ccd · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Prompt-to-Prompt Image Editing with Cross Attention Control
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ab5c693-073c-41f4-b2ef-6d36a1654708 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Explaining Length Bias in LLM-Based Preference Evaluations
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da8b9743-caaa-4d03-9a5d-143f401474c0 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Large Language Models Cannot Self-Correct Reasoning Yet
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 028f2908-3e8d-4bfa-a97f-5e506b9f1667 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f457f343-d5d0-4591-86e8-36c7a6fa49e7 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities OpenAI o1 System Card
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93a3a016-32f8-40a2-a29d-41d011bc36ec · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Towards Efficient Exact Optimization of Language Model Alignment
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7045ef3e-baf2-4297-88d2-97e11e689a8d · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Scaling Laws for Neural Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50e73708-3b7d-4fc0-90bc-a99fa3c9ceb4 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities ARGS: Alignment as Reward-Guided Search
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d1a173a-0a9d-4c7c-90be-1b997d674bd9 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Decomposed Prompting: A Modular Approach for Solving Complex Tasks
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e12ffca-9b85-4c25-ba3e-61cd0bbcb210 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0251ed79-ed5f-4d75-a2e3-93e6133dc99c · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities DiffWave: A Versatile Diffusion Model for Audio Synthesis
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87e7605d-5449-4237-87ed-abfbf8fe873e · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Privacy in Large Language Models: Attacks, Defenses and Future Directions
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cd7afb4-fe5b-495e-9338-84609153062c · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Personalized Language Modeling from Personalized Human Feedback
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbf519a6-9680-4607-8070-a3251649d9d9 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Reward-Guided Speculative Decoding for Efficient LLM Reasoning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a006a67-de87-4587-9da6-2604c988b059 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities AI Transparency in the Age of LLMs: A Human-Centered Research Roadmap
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddd944c0-7925-4e6a-9efe-88cafb060228 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities 24 (AAAI 2025 and ACL
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8345db02-524c-42cc-9e93-99dfba1bf963 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Statistical Rejection Sampling Improves Preference Optimization
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9236ba57-eb05-42d2-a5e3-a1d44d921573 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities RRM: Robust Reward Model Training Mitigates Reward Hacking
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7399c96-4795-48db-842d-c9751bf8efca · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b49c1be-5347-4cca-825b-e4b3fba7e550 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Inference-time scaling for generalist reward modeling.arXiv preprint arXiv:2504.02495,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28372cfb-437c-4cf6-9e3c-4a6219d7a720 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Rethinking Diverse Human Preference Learning through Principal Component Analysis
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f60977e9-086c-499f-b5b7-6954f51060d3 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Generative Reward Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4186af3e-c5b1-429b-aa9a-2686c2ec0536 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities A Survey of Explainable Reinforcement Learning
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1710d386-8bfe-4cb3-a838-d648fcb7b511 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Playing Atari with Deep Reinforcement Learning
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac4fc5df-b4c9-4000-9c7b-a74f276d7b1e · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Active Preference Learning for Large Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10f4b0a1-0457-4552-9c92-2537507d433a · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Unresolved cited work
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6516752-655b-4a0b-9193-adbdc6190678 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Deep Research is a new agentic AI capability integrated within ChatGPT that autonomously conducts multi-step web research and synthesizes comprehensive reports
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87571df1-77b6-48be-be6f-6d30f6d46119 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2993deae-6611-47ef-b968-555e2a0ce527 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78ea5e32-c17c-4513-a199-c79c0c6fdb00 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Automatic Prompt Optimization with "Gradient Descent" and Beam Search
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05ccdabb-33dd-469e-9c2c-2d05c4d0ea58 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities A Survey on Explainable Reinforcement Learning: Concepts, Algorithms, Challenges
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcec6494-167d-4b74-a35b-9e7ab6db4d5e · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ceb9cd7-bbbc-49c6-a586-76b6bd52e915 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Unresolved cited work
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87501923-c62b-4545-b5ef-5c4fd60de35e · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd2fe72f-dacb-4593-a527-847122627c44 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities A critical look at tokenwise reward-guided text generation.arXiv preprint arXiv:2406.07780,
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e80c9d6-ee00-4950-b41e-170613f0611b · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities A Generalist Agent
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2502d7e-e280-42d3-b361-6ba68091870a · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Learning Long-Term Reward Redistribution via Randomized Return Decomposition
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 21ae2890-0204-44ba-b4bf-2b3cff27a916 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c615db8a-b98d-4bb0-aa2a-419d1f96996b · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Spurious Rewards: Rethinking Training Signals in RLVR
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86484ea3-4055-4753-a933-a2e3a05de87c · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f256eb6-64ed-4697-934f-42d9e1153fb3 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Reviving The Classics: Active Reward Modeling in Large Language Model Alignment
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08498943-cdf1-4c8e-a21a-b4efbf1bb866 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b622793b-3d91-4a6e-b0db-24460636a79a · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities When Life Gives You Lemons, Make Cherryade: Converting Feedback from Bad Responses into Good Labels
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae7730dd-9eb7-48e4-a4de-b85e157906e8 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Unresolved cited work
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c36e4d87-d47b-44e8-8317-feb967e0fa88 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 129db947-a33a-4513-93e2-67f981934587 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e9d55e6-b139-4cc8-a47c-3cd574ed1282 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49f8f937-525f-4198-9dfc-d182860057d2 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities A Roadmap to Pluralistic Alignment
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 299da6d6-f1b3-43da-9533-97f5226241fe · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Inverse-RLignment: Large Language Model Alignment from Demonstrations through Inverse Reinforcement Learning
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b8a5cf1-704d-4ee7-9611-ac9b10123401 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36944cbd-6b9a-4553-9207-488a235ed765 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Learning to Repair: Repairing model output errors after deployment using a dynamic memory of feedback
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbeeeba4-80c0-4693-9d85-79b85eef1422 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities DeepMind Control Suite
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c3029af-049b-429f-a029-722c15b9e226 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Grpo: Generalized reinforcement preference optimization.arXiv preprint arXiv:2405.00000,
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fef2af8-216f-45a4-9f61-02b0e6243bcd · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83ad950e-a994-42d4-a596-dd77890634e7 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Gram: A generative foundation reward model for reward generalization.arXiv preprint arXiv:2506.14175, 2025a
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02dac0dd-db99-4c97-ae6b-d9eee97d771c · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0519097-48ef-4274-8731-66dc94072c1e · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b81d8143-0d29-4284-9d02-6ba928991711 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation caa7a167-a786-40b8-ad11-8e33e6c466dd · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11c90913-72bb-4836-a257-2e2879251b27 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34aa152a-be09-43ba-852b-232695cee304 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5809d262-c85c-43f6-ae1f-2ecea71cf6bc · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Reference 104
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 738a22bb-d8db-4848-92be-c26a1f7ba4ff · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Large Language Models as Optimizers
Reference 105
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a357dcd6-bad0-41da-8319-1a491f2a2c77 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Rethinking Goal-conditioned Supervised Learning and Its Connection to Offline RL
Reference 106
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b248e687-6a40-4976-a20b-c4b0a2463933 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs
Reference 107
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac69e7f0-10ed-4d9b-8ad0-9b06eba55910 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts
Reference 108
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0daaf2f3-714b-4697-b4c1-8ef1564e3f8b · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 109
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26a849c7-7e62-4a1a-9237-5698fd209cb9 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 110
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c79daecb-c563-4a67-b965-4f7e470b53a2 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Diversifying AI: Towards Creative Chess with AlphaZero
Reference 111
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b131a45-5c49-4f61-b78d-06947941fe55 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities When scaling meets LLM finetuning: The effect of data, model and finetuning method
Reference 112
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 963c5887-8ece-4b8d-8c6c-f0a2e45def27 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 113
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3396fdf-285c-4f06-9ab9-8ce1b7348ab2 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Overcoming Reward Overoptimization via Adversarial Policy Optimization with Lightweight Uncertainty Estimation
Reference 114
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4730a3b9-3f73-48fe-bbaa-275bc9827592 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Secrets of RLHF in Large Language Models Part I: PPO
Reference 115
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99c591d9-b88d-4fbb-b681-f880534ff6fc · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 116
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccac6194-08ff-4523-a40f-e40eac577756 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Reference 117
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90b10786-5de4-4378-88bc-07dde36cb0ba · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 1952
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f63be04-b7ef-4d04-826a-b31d7e63bd06 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 1973
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3889c48b-10e2-4ab4-bbfb-46d8a41f0f78 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Retrieval Augmented Thought Process for Private Data Handling in Healthcare
Reference 1991
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8c77f3b-3746-4371-abd6-e3bb2c761ef8 · outbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Style Over Substance: Evaluation Biases for Large Language Models
Reference 1992
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58ce233a-7689-4a88-977a-8ab27b2358d0 · inbound
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2654dec6-ac02-4a04-9249-fd87fe4db4ce · inbound
The Alignment Flywheel: A Governance-Centric Hybrid MAS for Architecture-Agnostic Safety Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 72c87fdd-80a0-4329-9873-91904335d9f2 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
Reference 188
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.