Pith. sign in

REVIEW 21 cited by

LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.11998 v4 pith:MNAPQPW2 submitted 2023-09-21 cs.CL cs.AI

classification cs.CLcs.AI
keywords datasetlmsys-chat-1mmodelsreal-worldbenchmarkcontentlarge-scalellms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Studying how people interact with large language models (LLMs) in real-world scenarios is increasingly important due to their widespread use in various applications. In this paper, we introduce LMSYS-Chat-1M, a large-scale dataset containing one million real-world conversations with 25 state-of-the-art LLMs. This dataset is collected from 210K unique IP addresses in the wild on our Vicuna demo and Chatbot Arena website. We offer an overview of the dataset's content, including its curation process, basic statistics, and topic distribution, highlighting its diversity, originality, and scale. We demonstrate its versatility through four use cases: developing content moderation models that perform similarly to GPT-4, building a safety benchmark, training instruction-following models that perform similarly to Vicuna, and creating challenging benchmark questions. We believe that this dataset will serve as a valuable resource for understanding and advancing LLM capabilities. The dataset is publicly available at https://huggingface.co/datasets/lmsys/lmsys-chat-1m.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 26 citations worldwide. Full citation record

  1. Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling

    cs.DC 2026-08 conditional novelty 7.0 of 10

    A resource-fair batching policy (ISJL) that keeps co-batched LLM requests within a token-progress window is proved 3/4-competitive in an offline model and empirically outperforms FCFS, SJF, and LJF on throughput and latency.

  2. K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

    cs.CL 2026-05 conditional novelty 7.0 of 10

    K12-KGraph is a textbook-derived knowledge graph that powers a new benchmark revealing LLMs' poor curriculum cognition and a small training corpus that outperforms general instruction data on educational tasks.

  3. Understanding Refusal in Language Models with Sparse Autoencoders

    cs.CL 2025-05 conditional novelty 7.0 of 10

    Refusal in Gemma-2-2B and Llama-3.1-8B is mediated by a small set of SAE features, harm features causally activate refusal features, and adversarial jailbreaks suppress those refusal features.

  4. CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A chest X-ray VLM co-trained with classification and grounding heads, tuned with DAPO reinforcement learning, and augmented with deterministic measurement tools outperforms prior radiology VLMs on report generation, V...

  5. RH-RAG: Trustworthy Long-Form Generation for Privacy-Constrained Settings

    cs.CL 2026-08 reject novelty 6.0 of 10

    A multi-agent RAG framework that adds planning, bounded memory, and NLI-based revision to local 7-8B models, reported to improve faithfulness and coherence in long-form generation.

  6. The One-Word Census: Answer-Choice Conformity Across 44 Language Models

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Forty-four language models asked to name one thing per category converge on the same modal answers far more than people do, with newest flagships most conformist and persona-tuned models most divergent.

  7. Economic Evaluations of Language Models

    cs.CY 2026-06 conditional novelty 6.0 of 10

    Using O*NET and 4.5M chatbot conversations plus synthetic prompts, EconEvals measures LM performance on U.S. work activities and predicts substantial time savings in 47% of occupations, with usage lagging.

  8. Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders

    cs.CL 2026-02 conditional novelty 6.0 of 10

    Coverage of sparse-autoencoder-identified task features predicts post-training performance and can guide synthesis of small, high-impact datasets (2,000 vs. 300,000 samples).

  9. After Talking with 1,000 Personas: Learning Preference-Aligned Proactive Assistants From Large-Scale Persona Interactions

    cs.HC 2026-02 conditional novelty 6.0 of 10

    A two-stage framework — category-structured fine-tuning on LLM-simulated personas plus on-device activation steering — improves proactive-assistant timing and perceived quality, though the biggest gains are measured w...

  10. Auditing LLM Editorial Bias in News Media Exposure

    cs.CY 2025-10 conditional novelty 6.0 of 10

    Compared with Google News, GPT-4o-Mini, Claude-3.7-Sonnet, and Gemini-2.0-Flash surface fewer unique news outlets, distribute attention more unevenly, and lean ideologically in system-specific ways.

  11. A global log for medical AI

    cs.AI 2025-10 conditional novelty 6.0 of 10

    MedLog defines a nine-field, syslog-style event log for clinical AI, intended to support real-world surveillance and auditing; the four-deployment validation claimed in the abstract is absent from the body.

  12. Developer-LLM Conversations: An Empirical Study of Interactions and Generated Code Quality

    cs.SE 2025-09 conditional novelty 6.0 of 10

    Analyzing 82,845 real ChatGPT coding chats shows generated code frequently has language-specific issues, with some quality problems persisting or worsening over multiple turns.

  13. TweakLLM: A Routing Architecture for Dynamic Tailoring of Cached Responses

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A small language model can rewrite cached large-model responses to fit similar new queries, preserving quality while cutting inference cost.

  14. Mass-Scale Analysis of In-the-Wild Conversations Reveals Complexity Bounds on LLM Jailbreaking

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Across 2M+ in-the-wild LLM conversations, jailbreak attempts show no higher complexity than normal chats, and assistant toxicity has declined over time, suggesting bounded attack sophistication.

  15. Evaluating the Sensitivity of LLMs to Prior Context

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Prior conversational context, especially from a different knowledge domain, can sharply reduce LLM multiple-choice accuracy, and repeating the task near the query mitigates the drop.

  16. Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models?

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Large reasoning models default to English or Chinese as internal 'reasoning hubs', and forcing non-hub reasoning lowers math accuracy, especially for low-resource languages, while sometimes improving safety and cultur...

  17. Improving Aviation Safety Analysis: Automated HFACS Classification Using Reinforcement Learning with Group Relative Policy Optimization

    cs.CL 2025-08 conditional novelty 4.0 of 10

    GRPO fine-tuning of Llama 3.1 8B improves multi-label HFACS classification of aviation narratives, reaching 18% exact match and 88% partial match on a 100-sample test set.

  18. DialogueForge: LLM Simulation of Human-Chatbot Dialogue

    cs.CL 2025-07 conditional novelty 4.0 of 10

    DialogueForge generates synthetic human-chatbot dialogues by pitting an inquirer LLM against a responder LLM, and finds that fine-tuned small models can approach GPT-4o-level realism on LLM-judged metrics.

  19. Toward Real-World Chinese Psychological Support Dialogues: CPsDD Dataset and a Co-Evolving Multi-Agent System

    cs.CL 2025-07 conditional novelty 4.0 of 10

    CPsDD is a 68K-dialogue Chinese psychological support dataset with strategy annotations, and CADSS is a multi-agent system reporting state-of-the-art results on Chinese and English emotional support tasks.

  20. Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A GRPO decoding framework that cuts memory via micro-batched KV-cache reuse and improves decoding-round efficiency with predicted-length scheduling, at the cost of serialization.

  21. Is There a Case for Conversation Optimized Tokenizers in Large Language Models?

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Retraining LLM tokenizers on chatbot conversation data reduces token counts by 5-10% on conversational text with minimal impact on general text.

Pith tools