Pith. sign in

REVIEW 12 cited by

Towards a Human-like Open-Domain Chatbot

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2001.09977 v3 pith:L3VWXRRY submitted 2020-01-27 cs.CL cs.LGcs.NEstat.ML

classification cs.CLcs.LGcs.NEstat.ML
keywords perplexitymeenamulti-turntrainedchatbotend-to-endevaluationhuman-like
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Meena, a multi-turn open-domain chatbot trained end-to-end on data mined and filtered from public domain social media conversations. This 2.6B parameter neural network is simply trained to minimize perplexity of the next token. We also propose a human evaluation metric called Sensibleness and Specificity Average (SSA), which captures key elements of a human-like multi-turn conversation. Our experiments show strong correlation between perplexity and SSA. The fact that the best perplexity end-to-end trained Meena scores high on SSA (72% on multi-turn evaluation) suggests that a human-level SSA of 86% is potentially within reach if we can better optimize perplexity. Additionally, the full version of Meena (with a filtering mechanism and tuned decoding) scores 79% SSA, 23% higher in absolute SSA than the existing chatbots we evaluated.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Momentum Based Reward Design for Low Emission Traffic Signal Control

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    A progressive multi-turn text-to-vis agent with rule-guided ReAct validation beats one-shot baselines by large execution-accuracy margins on a new reverse-constructed benchmark.

  2. When Should LLMs Be Less Specific? Selective Abstraction for Reliable Long-Form Text Generation

    cs.AI 2026-02 conditional novelty 6.0 of 10

    Atom-wise selective abstraction—replacing low-confidence factual claims with higher-confidence, less specific versions—improves the risk-coverage trade-off in long-form generation by up to 27.73% AURC over claim removal.

  3. Dr.Copilot: A Multi-Agent Prompt Optimized Assistant for Improving Patient-Doctor Communication in Romanian

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A multi-agent LLM assistant with DSPy-optimized prompts improved Romanian doctors' written communication quality and patient satisfaction in a live telemedicine deployment, but the evaluation is confounded by self-sel...

  4. BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A pipeline that combines contextual embeddings with two LMs' per-word probabilities and sparse autoencoders to automatically find interpretable slices where one model outperforms another.

  5. InFact: Informativeness Alignment for Improved LLM Factuality

    cs.CL 2025-05 conditional novelty 6.0 of 10

    InFACT trains LLMs with hierarchical informativeness rewards plus abstention, improving factual precision on QA benchmarks while largely preserving recall.

  6. Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications

    cs.AI 2026-07 conditional novelty 5.0 of 10

    Kaleidoscope combines persona-based test generation, contextual rubrics, and human-reliability-gated LLM judging into a practical, inspectable evaluation workflow.

  7. Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Simple tries and n-gram models beat large neural models for chat autocompletion on seen prefixes, while fine-tuned transformers and conversational context lead on unseen ones.

  8. EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices

    cs.DC 2025-07 conditional novelty 5.0 of 10

    EdgeLoRA combines automatic adapter routing, LRU caching with a memory pool, and grouped LoRA batching to serve thousands of LoRA adapters on edge devices with up to 4x higher throughput than llama.cpp.

  9. Format-Adapter: Improving Reasoning Capability of LLMs by Adapting Suitable Format

    cs.CL 2025-06 conditional novelty 5.0 of 10

    FORMAT-ADAPTER automatically generates and selects per-question reasoning formats for LLMs, improving vote-based accuracy by around 4.3% over prior multi-format methods.

  10. From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI

    cs.AI 2026-06 conditional novelty 4.0 of 10

    Autonomous AI becomes dependable when tool use is embedded in persistent workspaces with reusable skills, shifting evaluation from answers to task closure.

  11. From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR

    cs.CL 2025-08 unverdicted novelty 4.0 of 10

    A systematic analysis of LLM exploration in RLVR, introducing capability-boundary metrics and examining entropy-performance exchange across training stages and token levels.

  12. Optimizing Conversational Product Recommendation via Reinforcement Learning

    cs.IR 2025-06 reject novelty 1.0 of 10

    A position paper sketching how RL (DQN, PPO, RLHF) could optimize conversational product recommendation, without any validation.

Pith tools