REVIEW 12 cited by
Towards a Human-like Open-Domain Chatbot
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present Meena, a multi-turn open-domain chatbot trained end-to-end on data mined and filtered from public domain social media conversations. This 2.6B parameter neural network is simply trained to minimize perplexity of the next token. We also propose a human evaluation metric called Sensibleness and Specificity Average (SSA), which captures key elements of a human-like multi-turn conversation. Our experiments show strong correlation between perplexity and SSA. The fact that the best perplexity end-to-end trained Meena scores high on SSA (72% on multi-turn evaluation) suggests that a human-level SSA of 86% is potentially within reach if we can better optimize perplexity. Additionally, the full version of Meena (with a filtering mechanism and tuned decoding) scores 79% SSA, 23% higher in absolute SSA than the existing chatbots we evaluated.
Forward citations
Cited by 12 Pith papers
-
Momentum Based Reward Design for Low Emission Traffic Signal Control
A progressive multi-turn text-to-vis agent with rule-guided ReAct validation beats one-shot baselines by large execution-accuracy margins on a new reverse-constructed benchmark.
-
When Should LLMs Be Less Specific? Selective Abstraction for Reliable Long-Form Text Generation
Atom-wise selective abstraction—replacing low-confidence factual claims with higher-confidence, less specific versions—improves the risk-coverage trade-off in long-form generation by up to 27.73% AURC over claim removal.
-
Dr.Copilot: A Multi-Agent Prompt Optimized Assistant for Improving Patient-Doctor Communication in Romanian
A multi-agent LLM assistant with DSPy-optimized prompts improved Romanian doctors' written communication quality and patient satisfaction in a live telemedicine deployment, but the evaluation is confounded by self-sel...
-
BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language Models
A pipeline that combines contextual embeddings with two LMs' per-word probabilities and sparse autoencoders to automatically find interpretable slices where one model outperforms another.
-
InFact: Informativeness Alignment for Improved LLM Factuality
InFACT trains LLMs with hierarchical informativeness rewards plus abstention, improving factual precision on QA benchmarks while largely preserving recall.
-
Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications
Kaleidoscope combines persona-based test generation, contextual rubrics, and human-reliability-gated LLM judging into a practical, inspectable evaluation workflow.
-
Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems
Simple tries and n-gram models beat large neural models for chat autocompletion on seen prefixes, while fine-tuned transformers and conversational context lead on unseen ones.
-
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
EdgeLoRA combines automatic adapter routing, LRU caching with a memory pool, and grouped LoRA batching to serve thousands of LoRA adapters on edge devices with up to 4x higher throughput than llama.cpp.
-
Format-Adapter: Improving Reasoning Capability of LLMs by Adapting Suitable Format
FORMAT-ADAPTER automatically generates and selects per-question reasoning formats for LLMs, improving vote-based accuracy by around 4.3% over prior multi-format methods.
-
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
Autonomous AI becomes dependable when tool use is embedded in persistent workspaces with reusable skills, shifting evaluation from answers to task closure.
-
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR
A systematic analysis of LLM exploration in RLVR, introducing capability-boundary metrics and examining entropy-performance exchange across training stages and token levels.
-
Optimizing Conversational Product Recommendation via Reinforcement Learning
A position paper sketching how RL (DQN, PPO, RLHF) could optimize conversational product recommendation, without any validation.
Discussion (0). Sign in to comment.