Pith. sign in

REVIEW 5 cited by

Goals, Process, and Challenges of Exploratory Data Analysis: An Interview Study

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.00568 v1 pith:HYN6IN3R submitted 2019-11-01 cs.HC

classification cs.HC
keywords dataanalysisanalystsdiscoveryexplorationexploratorygoalsanalyses
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

How do analysis goals and context affect exploratory data analysis (EDA)? To investigate this question, we conducted semi-structured interviews with 18 data analysts. We characterize common exploration goals: profiling (assessing data quality) and discovery (gaining new insights). Though the EDA literature primarily emphasizes discovery, we observe that discovery only reliably occurs in the context of open-ended analyses, whereas all participants engage in profiling across all of their analyses. We describe the process and challenges of EDA highlighted by our interviews. We find that analysts must perform repetitive tasks (e.g., examine numerous variables), yet they may have limited time or lack domain knowledge to explore data. Analysts also often have to consult other stakeholders and oscillate between exploration and other tasks, such as acquiring and wrangling additional data. Based on these observations, we identify design opportunities for exploratory analysis tools, such as augmenting exploration with automation and guidance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 45 citations worldwide. Full citation record

  1. Dirty Data in the Newsroom: Comparing Data Preparation in Journalism and Data Science

    cs.HC 2025-07 conditional novelty 7.0 of 10

    Data journalists' data preparation overlaps with data scientists' but has distinct challenges, captured in a taxonomy of 60 dirty data issues and four integration scenarios.

  2. Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks

    cs.HC 2025-07 conditional novelty 6.0 of 10

    Data scientists construct prediction targets through bricolage, applying five reformulation strategies (piggybacking, composing, swapping, bridging, refining) to balance five criteria: validity, simplicity, predictabi...

  3. NoteFlow: Recommending Charts as Sight Glasses for Tracing Data Flow in Computational Notebooks

    cs.HC 2025-02 conditional novelty 6.0 of 10

    NoteFlow automatically tracks the flow of data tables in notebooks, recommends charts, and lets users trace a chart across transformations to locate anomalies and understand analysis.

  4. Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline

    cs.HC 2025-01 conditional novelty 6.0 of 10

    Twenty-nine interviews show AI practitioners rely on synthetic data across nearly every pipeline stage while validation remains mostly manual spot-checking.

  5. Jupybara: Operationalizing a Design Space for Actionable Data Analysis and Storytelling with LLMs

    cs.HC 2025-01 conditional novelty 5.0 of 10

    Jupybara is an LLM-powered Jupyter extension that operationalizes a semantic, rhetorical, and pragmatic design space for actionable data analysis and storytelling.

Pith tools