Pith. sign in

REVIEW 6 cited by

Clio: Privacy-Preserving Insights into Real-World AI Use

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.13678 v1 pith:4SSTXQAM submitted 2024-12-18 cs.CY cs.AIcs.CLcs.CRcs.LG

classification cs.CYcs.AIcs.CLcs.CRcs.LG
keywords clioconversationsclaudeinsightssystemsworldacrossassistants
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

How are AI assistants being used in the real world? While model providers in theory have a window into this impact via their users' data, both privacy concerns and practical challenges have made analyzing this data difficult. To address these issues, we present Clio (Claude insights and observations), a privacy-preserving platform that uses AI assistants themselves to analyze and surface aggregated usage patterns across millions of conversations, without the need for human reviewers to read raw conversations. We validate this can be done with a high degree of accuracy and privacy by conducting extensive evaluations. We demonstrate Clio's usefulness in two broad ways. First, we share insights about how models are being used in the real world from one million Claude.ai Free and Pro conversations, ranging from providing advice on hairstyles to providing guidance on Git operations and concepts. We also identify the most common high-level use cases on Claude.ai (coding, writing, and research tasks) as well as patterns that differ across languages (e.g., conversations in Japanese discuss elder care and aging populations at higher-than-typical rates). Second, we use Clio to make our systems safer by identifying coordinated attempts to abuse our systems, monitoring for unknown unknowns during critical periods like launches of new capabilities or major world events, and improving our existing monitoring systems. We also discuss the limitations of our approach, as well as risks and ethical concerns. By enabling analysis of real-world AI usage, Clio provides a scalable platform for empirically grounded AI safety and governance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 8 citations worldwide. Full citation record

  1. When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration

    cs.AI 2025-06 conditional novelty 7.0 of 10

    Model benchmark performance only weakly predicts how well people learn from AI explanations, with notable outliers across code and math.

  2. The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    LLMs align with human moral judgments only under high consensus, concentrate on a narrow set of moral values, and the profile-based prompting method's reported improvement is evaluated in-sample.

  3. The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It

    cs.CL 2025-05 accept novelty 6.0 of 10

    LLM safety research at ACL venues from 2020 to 2024 is predominantly English-only, and the language gap is growing over time.

  4. CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data

    cs.AI 2026-07 conditional novelty 5.0 of 10

    Fine-tuning on weaknesses diagnosed at individual rubric-criterion level beats prompt-level and random data selection on most of 8 model×domain settings.

  5. HERCULES: Hierarchical Embedding-based Recursive Clustering Using LLMs for Efficient Summarization

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A recursive k-means clustering pipeline that asks an LLM to summarize each cluster at every level, with a demonstration on 20 Newsgroups.

  6. Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects

    cs.CY 2025-05 conditional novelty 4.0 of 10

    A position paper argues that understanding AI's second-order effects requires moving from static benchmarks to an ecosystem of field testing, red teaming, and contextual evaluation.

Pith tools