REVIEW 5 cited by
Goals, Process, and Challenges of Exploratory Data Analysis: An Interview Study
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
How do analysis goals and context affect exploratory data analysis (EDA)? To investigate this question, we conducted semi-structured interviews with 18 data analysts. We characterize common exploration goals: profiling (assessing data quality) and discovery (gaining new insights). Though the EDA literature primarily emphasizes discovery, we observe that discovery only reliably occurs in the context of open-ended analyses, whereas all participants engage in profiling across all of their analyses. We describe the process and challenges of EDA highlighted by our interviews. We find that analysts must perform repetitive tasks (e.g., examine numerous variables), yet they may have limited time or lack domain knowledge to explore data. Analysts also often have to consult other stakeholders and oscillate between exploration and other tasks, such as acquiring and wrangling additional data. Based on these observations, we identify design opportunities for exploratory analysis tools, such as augmenting exploration with automation and guidance.
Forward citations
Cited by 5 Pith papers
-
Dirty Data in the Newsroom: Comparing Data Preparation in Journalism and Data Science
Data journalists' data preparation overlaps with data scientists' but has distinct challenges, captured in a taxonomy of 60 dirty data issues and four integration scenarios.
-
Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks
Data scientists construct prediction targets through bricolage, applying five reformulation strategies (piggybacking, composing, swapping, bridging, refining) to balance five criteria: validity, simplicity, predictabi...
-
NoteFlow: Recommending Charts as Sight Glasses for Tracing Data Flow in Computational Notebooks
NoteFlow automatically tracks the flow of data tables in notebooks, recommends charts, and lets users trace a chart across transformations to locate anomalies and understand analysis.
-
Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline
Twenty-nine interviews show AI practitioners rely on synthetic data across nearly every pipeline stage while validation remains mostly manual spot-checking.
-
Jupybara: Operationalizing a Design Space for Actionable Data Analysis and Storytelling with LLMs
Jupybara is an LLM-powered Jupyter extension that operationalizes a semantic, rhetorical, and pragmatic design space for actionable data analysis and storytelling.
Discussion (0). Continue with ORCID to comment.