REVIEW 6 cited by
Training Socially Aligned Language Models on Simulated Social Interactions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Social alignment in AI systems aims to ensure that these models behave according to established societal values. However, unlike humans, who derive consensus on value judgments through social interaction, current language models (LMs) are trained to rigidly replicate their training corpus in isolation, leading to subpar generalization in unfamiliar scenarios and vulnerability to adversarial attacks. This work presents a novel training paradigm that permits LMs to learn from simulated social interactions. In comparison to existing methodologies, our approach is considerably more scalable and efficient, demonstrating superior performance in alignment benchmarks and human evaluations. This paradigm shift in the training of LMs brings us a step closer to developing AI systems that can robustly and accurately reflect societal norms and values.
Forward citations
Cited by 6 Pith papers
-
Step-Level Preference Learning for Generative Agents in Social Simulations
Step-level human preference data collected via SimPref, then SFT+DPO, improves long-horizon social-simulation behavior of open-weight LLM agents on held-out events.
-
Structural transparency of societal AI alignment through Institutional Logics
Introduces a five-component analytical framework, grounded in Institutional Logics, for making visible the organizational and institutional decisions that shape AI alignment.
-
Modeling Earth-Scale Human-Like Societies with One Billion Agents
Light Society scales LLM-agent social simulations to one billion agents by substituting most LLM interactions with a distilled surrogate model.
-
Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks
CoopGuard's defer-tempt-analyze-coordinate agents cut reported jailbreak success and raise attacker token costs on the new EMRA benchmark, but the deceptive-rate metric is partly defined by the paper's own scoring rubric.
-
Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives
A survey proposes a macro-meso-micro value framework for agentic AI alignment and maps applications, methods, and benchmarks onto it.
-
Bridging the Gap: In-Context Learning for Modeling Human Disagreement
Across four open-source LLMs and three subjective-task datasets, in-context learning with multi-perspective prompts improves aggregated-label predictions in zero-shot, but disaggregated hard and soft label predictions...
Discussion (0). Sign in to comment.