Pith. sign in

REVIEW 3 cited by

Anticipating Safety Issues in E2E Conversational AI: Framework and Tooling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.03451 v3 pith:ZXKE7XTB submitted 2021-07-07 cs.CL cs.AI

classification cs.CLcs.AI
keywords conversationalmodelsend-to-enddecisionsframeworkpotentialreleaseresearchers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Over the last several years, end-to-end neural conversational agents have vastly improved in their ability to carry a chit-chat conversation with humans. However, these models are often trained on large datasets from the internet, and as a result, may learn undesirable behaviors from this data, such as toxic or otherwise harmful language. Researchers must thus wrestle with the issue of how and when to release these models. In this paper, we survey the problem landscape for safety for end-to-end conversational AI and discuss recent and related work. We highlight tensions between values, potential positive impact and potential harms, and provide a framework for making decisions about whether and how to release these models, following the tenets of value-sensitive design. We additionally provide a suite of tools to enable researchers to make better-informed decisions about training and releasing end-to-end conversational AI models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models

    cs.LG 2025-07 conditional novelty 7.0 of 10

    A demographically diverse annotation dataset shows that safety perceptions for text-to-image outputs vary by rater identity and that conventional safety classifiers under-detect bias harms flagged by minority-group raters.

  2. Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints

    cs.LG 2025-06 conditional novelty 6.0 of 10

    HC-RLHF returns an aligned language model only after a held-out safety test certifies, with probability at least 1-delta, that expected harm (as judged by a learned cost model) is below a chosen threshold.

  3. Safe Inference-Time Alignment via Lagrangian Reward Augmentation

    cs.LG 2026-07 conditional novelty 5.5 of 10

    Dualizing Safe RLHF yields a one-dimensional convex calibration of λ that defines a drop-in safety-aware reward for Best-of-N and token-level inference-time decoders.

Pith tools