Pith. sign in

REVIEW 7 cited by

BlenderBot 3: a deployed conversational agent that continually learns to responsibly engage

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.03188 v3 pith:2VZRFUUD submitted 2022-08-05 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelagentsblenderbotdeployeddeploymentdialogueincludingopen-domain
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present BlenderBot 3, a 175B parameter dialogue model capable of open-domain conversation with access to the internet and a long-term memory, and having been trained on a large number of user defined tasks. We release both the model weights and code, and have also deployed the model on a public web page to interact with organic users. This technical report describes how the model was built (architecture, model and training scheme), and details of its deployment, including safety mechanisms. Human evaluations show its superiority to existing open-domain dialogue agents, including its predecessors (Roller et al., 2021; Komeili et al., 2022). Finally, we detail our plan for continual learning using the data collected from deployment, which will also be publicly released. The goal of this research program is thus to enable the community to study ever-improving responsible agents that learn through interaction.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 98 citations worldwide. Full citation record

  1. Momentum Based Reward Design for Low Emission Traffic Signal Control

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    A progressive multi-turn text-to-vis agent with rule-guided ReAct validation beats one-shot baselines by large execution-accuracy margins on a new reverse-constructed benchmark.

  2. How Stylistic Similarity Shapes Preferences in Dialogue Dataset with User and Third Party Evaluations

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new open-domain dialogue dataset shows that users' own judgments of stylistic similarity correlate with their preference (Spearman r=0.67-0.75), but third-party stylistic similarity judgments do not, indicating a ga...

  3. Entriever: Energy-based Retriever for Knowledge-Grounded Dialog Systems

    cs.CL 2025-05 conditional novelty 6.0 of 10

    An energy-based retriever that jointly scores sets of knowledge items improves retrieval accuracy and semi-supervised dialog performance over independently-scoring baselines.

  4. Agentic AI and Retrieval-Augmented Models in Straight-Through Underwriting

    cs.AI 2026-07 conditional novelty 5.0 of 10

    On 635 synthetic BOP applications, multi-agent Agentic RAG reaches 86.5% decision accuracy versus 77.6% single-LLM and 76.9% naive RAG, with largest gains on multi-step and missing-information cases.

  5. Exchange of Perspective Prompting Enhances Reasoning in Large Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A two-branch prompting method that exchanges answers between an original math question and a paraphrased version improves accuracy on several math benchmarks, but the gain is not separated from the extra compute or ru...

  6. From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI

    cs.AI 2026-06 conditional novelty 4.0 of 10

    Autonomous AI becomes dependable when tool use is embedded in persistent workspaces with reusable skills, shifting evaluation from answers to task closure.

  7. Embodied AI Agents: Modeling the World

    cs.AI 2025-06 conditional novelty 4.0 of 10

    Embodied AI agents should be built around physical world models plus a mental world model of the user, with virtual, wearable, and robotic agents sharing this core.

Pith tools