Pith. sign in

REVIEW 6 cited by

Training Socially Aligned Language Models on Simulated Social Interactions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.16960 v3 pith:OJKCLLS5 submitted 2023-05-26 cs.CL cs.AIcs.CYcs.HC

classification cs.CLcs.AIcs.CYcs.HC
keywords socialtrainingmodelsalignmentinteractionslanguageparadigmsimulated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Social alignment in AI systems aims to ensure that these models behave according to established societal values. However, unlike humans, who derive consensus on value judgments through social interaction, current language models (LMs) are trained to rigidly replicate their training corpus in isolation, leading to subpar generalization in unfamiliar scenarios and vulnerability to adversarial attacks. This work presents a novel training paradigm that permits LMs to learn from simulated social interactions. In comparison to existing methodologies, our approach is considerably more scalable and efficient, demonstrating superior performance in alignment benchmarks and human evaluations. This paradigm shift in the training of LMs brings us a step closer to developing AI systems that can robustly and accurately reflect societal norms and values.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Step-Level Preference Learning for Generative Agents in Social Simulations

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Step-level human preference data collected via SimPref, then SFT+DPO, improves long-horizon social-simulation behavior of open-weight LLM agents on held-out events.

  2. Structural transparency of societal AI alignment through Institutional Logics

    cs.CY 2026-02 conditional novelty 6.0 of 10

    Introduces a five-component analytical framework, grounded in Institutional Logics, for making visible the organizational and institutional decisions that shape AI alignment.

  3. Modeling Earth-Scale Human-Like Societies with One Billion Agents

    cs.MA 2025-06 conditional novelty 6.0 of 10

    Light Society scales LLM-agent social simulations to one billion agents by substituting most LLM interactions with a distilled surrogate model.

  4. Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks

    cs.CR 2026-07 reject novelty 4.0 of 10

    CoopGuard's defer-tempt-analyze-coordinate agents cut reported jailbreak success and raise attacker token costs on the new EMRA benchmark, but the deceptive-rate metric is partly defined by the paper's own scoring rubric.

  5. Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A survey proposes a macro-meso-micro value framework for agentic AI alignment and maps applications, methods, and benchmarks onto it.

  6. Bridging the Gap: In-Context Learning for Modeling Human Disagreement

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Across four open-source LLMs and three subjective-task datasets, in-context learning with multi-perspective prompts improves aggregated-label predictions in zero-shot, but disaggregated hard and soft label predictions...

Pith tools