Pith. sign in

REVIEW 4 major objections 6 minor 89 references

Dirty Data in the Newsroom: Comparing Data Preparation in Journalism and Data Science

T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Interviews with 36 data journalists show all 60 identified dirty-data issues fit on one two-axis taxonomy, and that table integration follows four recurring patterns.

desk verdict Solid qualitative synthesis with a genuinely new taxonomy; the comparison to data science is shakier than the rest, but the paper's own transparency about that asymmetry keeps it worth engaging. read the letter →

arxiv 2507.07238 v1 pith:MOIRTNJB submitted 2025-07-09 cs.HC cs.CY

classification cs.HCcs.CY
keywords datajournalismsciencewranglingcleaningdirtyqualitymulti-tableintegrationmentalmodels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Data preparation eats up huge amounts of time in any data-driven project, yet almost all empirical studies of it have focused on data scientists, not on the journalists who assemble, clean, and check data for news stories. This paper reports 36 interviews with professional data journalists and compares what they do with what the data-science literature says data scientists do. It claims the two groups share a common process, which the paper extends to 23 concrete preparation activities, and that all 60 dirty-data issues raised by journalists or by 16 earlier taxonomies can be classified on one new grid: four data objects (table, item, attribute, value) by six data qualities (completeness, accuracy, form, granularity, relation, semantics). The same analysis yields four recurring difficulties in combining multiple tables: regional, diachronic, fragmented, and disparate datasets. If the taxonomy and process model are right, tool builders and researchers gain a common language for studying and supporting all data workers, not just scientists.

What carries the argument

The load-bearing object is the model-discrepancy taxonomy: a two-axis design space with four data objects (table, item, attribute, value) crossed with six data qualities (completeness, accuracy, form, granularity, relation, semantics). Its job is to give every one of the 60 synthesized dirty-data issues a cell, reconciling the bottom-up, domain-oriented language journalists use with the top-down taxonomy language of database research. A second mechanism is the extended process model, which slots 23 preparation activities into the preparation subprocesses (initiate, gather, create, profile, wrangle) and communication subprocesses (disseminate, document) of the prior model [16]. Together the two mechanisms let the authors compare evidence bases and locate divergences such as the journalists' verify-transformations activity and absence of label-data activity.

What would settle it

Run the same semi-structured interview protocol with a matched group of 36 data scientists: if they too report verify-transformation checks and no label-data work, the paper's claimed journalist-specific divergences would vanish. Alternatively, find a dirty-data issue among the 60 that cannot be assigned to exactly one of the 24 object-by-quality cells without an arbitrary choice.

Watch

Extended reading notes

Core claim

The central claim is that dirty data is best understood as a discrepancy between the mental model a data worker holds about a dataset and the data model that actually exists, because datasets are design artifacts made by people. The paper operationalizes that claim with a two-dimensional taxonomy: four data objects (table, item, attribute, value) on one axis and six data qualities (completeness, accuracy, form, granularity, relation, semantics) on the other. The taxonomy was built to cover all 60 dirty-data issues that emerged from the union of the authors' journalism interviews, prior data-science workflow papers, and 16 database and statistics taxonomies of dirty data. Alongside the taxonomy, the paper claims that data journalists' preparation can be mapped onto the same high-level process model used for data scientists, extended to 23 fine-grained activities; notable divergences are that journalists verify their transformations, avoid imputing or synthesizing data, and never label data items. The paper also identifies four challenges in combining multiple tables: regional inconsistencies, evolving diachronic schema, fragmented entities that must be reassembled, and topically disparate tables linked only by shared entities or geography.

Load-bearing premise

The comparison of data scientists and journalists rests on two differently collected evidence bases, published workflow studies for scientists and new interviews for journalists, so a group difference could be an artifact of the difference in method rather than a real difference in practice.

Editorial extensions

If this is right

  • A tool designer can use the 24 cells to decide which classes of dirty data are and are not addressed by a given data preparation tool.
  • The four integration challenges give data journalism support systems a concrete checklist: handle independently collected regional tables, schemas that drift across releases, fragments of a single logical dataset, and entity resolution across unrelated topics.
  • Since no journalist in the study reported labeling data for machine learning, the paper implies that data journalists prepare data for descriptive reporting rather than predictive modeling, so tools that emphasize label-data workflows will miss their needs.
  • The extended process model gives future studies a common activity vocabulary for comparing data work across professions.
  • The taxonomy's mental-model framing suggests that cleaning is not an objective fix but an act of conforming a dataset to a particular worker's expectations, which may explain why the same table can be considered clean by one organization and dirty by another.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • I infer that the model-discrepancy framing is the paper's most portable idea: any population that consumes data without storing it, such as policy analysts, auditors, or citizens, could in principle be described with the same object-by-quality grid, and the paper itself conjectures this for domain-oriented data workers.
  • A testable extension would be to run the same interview protocol on a matched group of data scientists; if verify-transformation checks and the absence of label-data work still appear, the reported journalist-scientist divergences are population differences rather than protocol artifacts.
  • The four integration challenges, regional, diachronic, fragmented, and disparate table combinations, could plausibly generalize beyond journalism to any multi-source integration problem and would make a useful scenario catalog for schema-matching and entity-resolution benchmarks.
  • A consequence the paper leaves implicit is that treating dirty data as a discrepancy between mental models shifts responsibility from the dataset to undocumented expectations, which strengthens the case for making data dictionaries and provenance first-class artifacts in preparation tools.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This paper reports a four-phase qualitative study of data preparation in data journalism. Phase 1 derives a priori codes from 16 data science workflow papers; Phase 2 analyzes 36 semi-structured interviews with data journalists, yielding 566 coded passages and 23 consolidated activity codes; Phase 3 synthesizes 60 dirty data issues from 16 taxonomies and the interview data; Phase 4 identifies four multi-table integration challenges (regional, diachronic, fragmented, disparate) from 69 'nightmare stories.' The contributions are an extended process model, a two-axis model-discrepancy taxonomy, and the four integration challenges.

Significance. The interview corpus is a genuine empirical contribution, and the paper's systematic accounting—150 excerpts, 315 taxonomy instances, 566 coded passages, 60 synthesized issues—is transparently documented via OSF supplements. The proposed taxonomy is concrete and testable in principle: its distributional claims about which issues are reported by database researchers versus domain-oriented users can be checked by coding a new sample. The four integration challenges are a useful, design-relevant vocabulary. The central risk is not in the data collection but in the comparative framing: the paper contrasts interview-based codes with literature-based codes as if they were commensurable measurements of the two populations.

major comments (4)
  1. [§3.1–§3.2, Figure 2] The central comparison between data scientists and journalists is built on two non-commensurable evidence bases: Phase 1 codes come from 16 published workflow papers (150 excerpts), while Phase 2 codes come from 36 new interviews using a different protocol, including pre-interview artifact sharing and prompts for 'nightmare stories.' The paper treats an activity's absence from the literature as a divergence (blue/green coding in Figure 2), but silence in a paper is not evidence of absence in practice; the authors concede this for 'identify items' in §4.5, calling it 'under-reported in data science workflows.' The same logic affects 'verify transformations' (§4.4) and the claimed lack of 'label data items' (§4.5). Because no commensurability test is offered (e.g., applying the interview protocol to a matched sample of data scientists, or recoding the 16 papers using interview-derived codes and reporting coverage), the title's 'comparing' claim is stronger than the data support. I recommend reframing all divergences as 'reported in the data science literature' versus 'reported in journalist interviews,' and adding an explicit limitations paragraph, or providing a supplementary coding check on a subset of the literature corpus.
  2. [§3.2.3, §3.2.5, §5.3] All coding was performed by the first author, with no second coder, inter-rater reliability, member checking, or audit trail beyond the OSF materials. The paper nevertheless reports precise counts—43 initial activity codes, 26 issue codes, 13 unique to interviews, 16 unique to previous work, 31 overlapping (Fig. 3a)—as though they were stable measurements. Since the synthesis of 60 issues and the four integration challenges depend on these codings, the lack of any reliability or reflexivity check leaves open the possibility that the taxonomy is idiosyncratic. Adding a second coder on a sample of passages or at least a clear statement of why reliability is not appropriate for this design would strengthen the central claims.
  3. [§3.2.4] The saturation claim is stated but not evidenced: 'growth of our codebook's cardinality' is a reasonable proxy, but no saturation curve or stopping criterion is reported, and the reader cannot tell how many new codes appeared in the last five interviews. Given that the paper's later counts depend on codebook stability, the saturation argument needs at least a supplemental plot or a statement of the threshold used.
  4. [§6] The fourth contribution is supported by a single aggregate number: 63 of 69 nightmare stories exhibit at least one of the four challenges, and 11 involve more than one. But the section does not report how many stories fall into each of the four categories (counts are given only for fragmented and disparate, 17/36 and 14/36), nor how the remaining six stories were classified. Because the 'nightmare story' prompt specifically elicits difficulties, the prevalence figures may overstate how common these challenges are in routine journalistic work; a limitations note is needed.
minor comments (6)
  1. [Table 1] Table 1 misspells 'Oliveira' as 'Oliveria' and 'Manssour' as 'Mannssour'; these should be corrected.
  2. [References] Reference [17] lists the author as 'Theordore Dasu, Tamraparni & Johnson'; the first author's name should be 'Theodore' and the attribution to Dasu & Johnson corrected.
  3. [Figure 3] Figure 3(a) is very dense; a larger or interactive version of the 60-issue mapping in the supplement would help readers verify the object/quality assignments referenced in Supp. Section 4.
  4. [§5.2] The definitions of 'relation' and 'semantics' overlap for duplicate items, which are classified under Semantics; clarify the boundary with an example of a relational duplicate.
  5. [§4.2] The acronym FOI is used without expansion at first use in §4.2; spell out 'freedom of information.'
  6. [§7.2] The term 'MacGyvering' should be introduced with a neutral definition of its scope before the pop-culture reference, since the paper uses it as a category of practice.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's contributions are an inductive empirical synthesis, and its self-citations are peripheral to the central argument.

full rationale

The paper's derivation chain is an inductive qualitative synthesis. Phase 1 derives a priori codes from 16 data-science workflow studies; Phase 2 codes 36 journalist interviews with those codes plus new inductive codes; Phase 3 merges interview issues with 16 dirty-data taxonomies; Phase 4 re-analyzes nightmare stories to induce four integration challenges. No equation or fitted parameter is used, and no claim is derived by definition from its own input. The taxonomy is explicitly built from the 60 issue codes (Sections 3.3 and 5.3), so its coverage statement is a description of the construction rather than a hidden prediction; this is normal taxonomic synthesis, not circularity. The comparison between journalists and data scientists rests on two different evidence corpora, and absence of a practice in published workflow studies may reflect reporting granularity; but this is a methodological limitation of the comparison, not a circular reduction. Self-citations are peripheral: Table Scraps [39] is cited as complementary and for the 'nerd box' concept, and Munzner [54] supplies terminology for the four data objects; neither carries the central argument. No uniqueness theorem or prior author-derived constraint is invoked. The paper is self-contained as an empirical and taxonomic contribution, so the circularity score is 0.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The study rests on qualitative assumptions about sample representativeness, single-coder coding reliability, and corpus comparability. There are no fitted parameters or invented entities; the taxonomy is a conceptual organization, not a postulated mechanism.

assumptions (5)
  • domain assumption Self-reported interview accounts accurately reflect data journalists' real preparation practices.
    Section 3.2.2-3.2.3: All activity and issue codes derive from retrospective self-reports, which may be incomplete, motivated, or post-hoc rationalized.
  • domain assumption The 16 data science workflow papers and the 16 dirty data taxonomies constitute a representative and sufficient corpus for comparison and synthesis.
    Sections 3.1 and 3.3: The baseline for data scientists and the source of prior dirty data issues are curated via a systematic review subset and snowball sampling; a different corpus would shift the 13/16/31 overlap counts.
  • domain assumption Single-coder thematic analysis with reflective synthesis yields reliable categories.
    Section 3.2.3: The first author applied a priori and a posteriori codes without reported inter-rater reliability or an independent audit, so coding drift or bias is not measured.
  • domain assumption Theoretical saturation after 36 interviews, proxied by codebook cardinality growth, justifies stopping.
    Section 3.2.4: Saturation is operationalized as code-count growth; a larger or different sample might surface additional issues not in the taxonomy.
  • domain assumption The Crisan et al. four-process model is the correct scaffold for mapping journalist activities.
    Section 4: The 23 activities are categorized within this external model; if the model is ill-suited to journalism, the 'extension' contribution weakens.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Dirty Data in the Newsroom: Comparing Data Preparation in Journalism and Data Science." pith.science (2026). https://pith.science/paper/MOIRTNJB

@misc{pith2026250707238,
  author       = {Pith},
  title        = {Pith review of: Dirty Data in the Newsroom: Comparing Data Preparation in Journalism and Data Science},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MOIRTNJB}},
  note         = {Machine review of arXiv:2507.07238}
}
read the original abstract

The work involved in gathering, wrangling, cleaning, and otherwise preparing data for analysis is often the most time consuming and tedious aspect of data work. Although many studies describe data preparation within the context of data science workflows, there has been little research on data preparation in data journalism. We address this gap with a hybrid form of thematic analysis that combines deductive codes derived from existing accounts of data science workflows and inductive codes arising from an interview study with 36 professional data journalists. We extend a previous model of data science work to incorporate detailed activities of data preparation. We synthesize 60 dirty data issues from 16 taxonomies on dirty data and our interview data, and we provide a novel taxonomy to characterize these dirty data issues as discrepancies between mental models. We also identify four challenges faced by journalists: diachronic, regional, fragmented, and disparate data sources.

Figures

Figures reproduced from arXiv: 2507.07238 by the authors.

Figure 1
Figure 1. Process, products, and contributions: Our hybrid deductive-inductive thematic analysis [ [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Data preparation activities: From our thematic analysis, we identify 23 activities that data scientists and data journalists [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. (a) Sixty data issues and which source of data they occur in (data science workflows, data journalism interviews, or [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

89 extracted references · 52 canonical work pages

  1. [1]

    Timo Aho, Outi Sievi-Korte, Terhi Kilamo, Sezin Yaman, and Tommi Mikkonen

  2. [2]

    Sara Alspaugh, Nava Zokaei, Andrea Liu, Cindy Jin, and Marti A. Hearst. 2019. Futzing and Moseying: Interviews with Professional Data Analysts on Exploration Practices. Transactions Visualization and Computer Graphics 25, 1 (Jan. 2019), 22–31. https://doi.org/10.1109/TVCG.2018.2865040

  3. [3]

    In the Beginning Were the Data

    Ángel Arrese. 2022. "In the Beginning Were the Data": Economic Journalism as/and Data Journalism. Journalism Studies 23, 4 (Feb. 2022), 487–505. https: //doi.org/10.1080/1461670X.2022.2032803

  4. [4]

    José Barateiro and Helena Galhardas. 2005. A Survey of Data Quality Tools. Datenbank-Spektrum 4, 14 (Aug. 2005), 15–21. http://dc-pubs.dbs.uni-leipzig.de/ files/Barateiro2005ASurveyofDataQuality.pdf

  5. [5]

    Andrea Batch and Niklas Elmqvist. 2018. The Interactive Visualization Gap in Initial Exploratory Data Analysis. Transactions Visualization and Computer Graphics 24, 1 (Jan. 2018), 278–287. https://doi.org/10.1109/TVCG.2017.2743990

  6. [6]

    Leilani Battle and Jeffrey Heer. 2019. Characterizing Exploratory Visual Analysis: A Literature Review and Evaluation of Analytic Provenance in Tableau.Computer Graphics Forum 38, 3 (July 2019), 145–159. https://doi.org/10.1111/cgf.13678

  7. [7]

    Charles Berret and Cheryl Phillips. 2016. Teaching Data and Computational Journalism. Columbia Journalism School, New York, NY, USA

  8. [8]

    Eddy Borges-Rey. 2021. Journalism with Machines? From Computational Think- ing to Distributed Cognition. In The Data Journalism Handbook 2: Towards a Critical Data Practice , Jonathan Gray and Liliana Bounegru (Eds.). Amster- dam University Press, Amsterdam, Netherlands, 92–95. https://doi.org/10.1515/ 9789048542079-023

Show all 89 references
  1. [9]

    1998.Transforming Qualitative Information: Thematic analysis and code development

    Richard E Boyatzis. 1998.Transforming Qualitative Information: Thematic analysis and code development. SAGE, Thousand Oaks, California

  2. [10]

    Paul Bradshaw. 2011. The Inverted Pyramid of Data Journalism. Retrieved Aug. 13, 2021 from https://onlinejournalismblog.com/2011/07/07/the-inverted- pyramid-of-data-journalism

  3. [11]

    Suzana Guedes Cardoso. 2022. The Practice of Data Journalism and Changes in the Professional Profile of Journalists in Newsrooms in the United States, United Kingdom, and Brazil. In Digital Convergence in Contemporary Newsrooms: Media Innovation, Content Adaptation, Digital Tr...

  4. [12]

    Abhirup Chatterjee and Arie Segev. 1991. Data Manipulation in Heterogeneous Databases. ACM SIGMOD Record 20, 4 (Dec. 1991), 64–68. https://doi.org/10. 1145/141356.141385

  5. [13]

    Fanny Chevalier, Melanie Tory, Bongshin Lee, Jarke van Wijk, Giuseppe Santucci, Marian Dörk, and Jessica Hullman. 2018. From Analysis to Communication: Sup- porting the Lifecycle of a Story. InData-Driven Storytelling, Nathalie Henry Riche, Dirty Data in the Newsroom CHI ’23, ...

  6. [14]

    Hamilton, and Fred Turner

    Sarah Cohen, James T. Hamilton, and Fred Turner. 2011. Computational Journal- ism. Commun. ACM 54, 10 (Oct. 2011), 66–71. https://doi.org/10.1145/2001269. 2001288

  7. [15]

    Anamaria Crisan and Brittany Fiore-Gartland. 2021. Fits and Starts: Enterprise Use of AutoML and the Role of Humans in the Loop. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21). ACM, New York, NY, USA, 1–15. https://doi.or...

  8. [16]

    Anamaria Crisan, Brittany Fiore-Gartland, and Melanie Tory. 2020. Passing the Data Baton: A Retrospective Analysis on Data Science Work and Workers. Transactions Visualization and Computer Graphics 27, 2 (Oct. 2020), 1860–1870. https://doi.org/10.1109/TVCG.2020.3030340

  9. [17]

    Theordore Dasu, Tamraparni & Johnson. 2003. Exploratory Data Mining and Data Cleaning. John Wiley & Sons, Hoboken, NJ, USA

  10. [18]

    Wesley Gongora de Almeida, Rafael Timóteo de Sousa, Flávio Elias de Deus, Georges Daniel Amvame Nze, and Fábio Lúcio Lopes de Mendonça. 2013. Tax- onomy of Data Quality Problems in Multidimensional Data Warehouse Mod- els. In Proceedings 8th Iberian Conference on Information S...

  11. [19]

    Catherine D’Ignazio and Lauren F. Klein. 2020. Data Feminism . MIT Press, Cambridge, Massachusetts

  12. [20]

    David Donoho. 2017. 50 Years of Data Science. Journal of Computational and Graphical Statistics 26, 4 (Oct. 2017), 745–766. https://doi.org/10.1080/10618600. 2017.1384734

  13. [21]

    Shari L. Dworkin. 2012. Sample Size Policy for Qualitative Studies Using In- Depth Interviews. Archives Sexual Behavior 41, 6 (Sept. 2012), 1319–1320. https: //doi.org/10.1007/s10508-012-0016-6

  14. [22]

    Marc A Feldman. 2020. Data Quality: Dimensions, Measurement, Strategy, Man- agement, and Governance. Quality Progress 53, 2 (2020), 54–54

  15. [23]

    U.M. Feyyad. 1996. Data Mining and Knowledge Discovery: Making Sense Out of Data. Expert 11, 5 (Oct. 1996), 20–25. https://doi.org/10.1109/64.539013

  16. [24]

    Katherine Fink and C. W. Anderson. 2015. Data Journalism in the United States. Journalism Studies 16, 4 (July 2015), 467–481. https://doi.org/10.1080/1461670X. 2014.939852

  17. [25]

    Bruce Frey (Ed.). 2018. The SAGE Encyclopedia of Educational Research, Measure- ment, and Evaluation . SAGE, Thousand Oaks, California. https://doi.org/10. 4135/9781506326139

  18. [26]

    Garrett Grolemund and Hadley Wickham. 2014. A Cognitive Interpretation of Data Analysis: A Cognitive Interpretation of Data Analysis. International Statistical Review 82, 2 (Aug. 2014), 184–204. https://doi.org/10.1111/insr.12028

  19. [27]

    Theresia Gschwandtner, Johannes Gärtner, Wolfgang Aigner, and Silvia Miksch

  20. [28]

    Joseph M Hellerstein. 2008. Quantitative Data Cleaning for Large Databases. United Nations Economic Commission for Europe. https://dsf.berkeley.edu/jmh/ papers/cleaning-unece.pdf

  21. [29]

    Bahareh Heravi, Kathryn Cassidy, Edie Davis, and Natalie Harrower. 2022. Pre- serving Data Journalism: A Systematic Literature Review. Journalism Practice 16, 10 (March 2022), 2083–2105. https://doi.org/10.1080/17512786.2021.1903972

  22. [30]

    Karen Holtzblatt and Hugh Beyer. 2015. Contextual Design: Evolved . Spinger Nature, Berlin, Germany

  23. [31]

    Sarah Hutchins. 2020. Data Dive: School’s Out. The Investigative Reporters & Editors Journal 43, 1 (Feb. 2020), 6–7

  24. [32]

    Kaggle. 2019. State of Data Science and Machine Learning . Kaggle. Retrieved May 15, 2022 from https://www.kaggle.com/kaggle-survey-2019

  25. [33]

    Sean Kandel, Jeffrey Heer, Catherine Plaisant, Jessie Kennedy, Frank van Ham, Nathalie Henry Riche, Chris Weaver, Bongshin Lee, Dominique Brodbeck, and Paolo Puono. 2011. Research Directions in Data Wrangling: Visualizations and Transformations for Usable and Credible Data. In...

  26. [34]

    Sean Kandel, Andreas Paepcke, Joseph Hellerstein, and Jeffrey Heer. 2011. Wran- gler: Interactive Visual Specification of Data Transformation Scripts. In Pro- ceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’11). ACM, Vancouver, Canada, 3363–3372. htt...

  27. [35]

    Hellerstein, and Jeffrey Heer

    Sean Kandel, Andreas Paepcke, Joseph M. Hellerstein, and Jeffrey Heer. 2012. Enterprise Data Analysis and Visualization: An Interview Study. Transactions Visualization and Computer Graphics 18, 12 (Dec. 2012), 2917–2926. https://doi. org/10.1109/TVCG.2012.219

  28. [36]

    Hellerstein, and Jeffrey Heer

    Sean Kandel, Ravi Parikh, Andreas Paepcke, Joseph M. Hellerstein, and Jeffrey Heer. 2012. Profiler: Integrated Statistical Analysis and Visualization for Data Quality Assessment. In Proceedings of the International Working Conference on Advanced Visual Interfaces (Capri Island...

  29. [37]

    Haber, and Jeffrey S

    Eser Kandogan, Aruna Balakrishnan, Eben M. Haber, and Jeffrey S. Pierce. 2014. From Data to Insight: Work Practices of Analysts in the Enterprise. Computer Graphics and Applications 34, 5 (Sept. 2014), 42–50. https://doi.org/10.1109/MCG. 2014.62

  30. [38]

    H. Kang, L. Getoor, B. Shneiderman, M. Bilgic, and L. Licamele. 2008. Interactive Entity Resolution in Relational Data: A Visual Analytic Tool and Its Evaluation. Transactions Visualization and Computer Graphics 14, 5 (Sept. 2008), 999–1014. https://doi.org/10.1109/TVCG.2008.55

  31. [39]

    Stephen Kasica, Charles Berret, and Tamara Munzner. 2020. Table Scraps: An Actionable Framework for Multi-Table Data Wrangling From An Artifact Study of Computational Journalism. Transactions Visualization and Computer Graphics 27, 2 (2020), 957–966. https://doi.org/10.1109/TV...

  32. [40]

    Miryung Kim, Thomas Zimmermann, Robert DeLine, and Andrew Begel. 2016. The Emerging Role of Data Scientists on Software Development Teams. In Pro- ceedings of the 38th International Conference on Software Engineering (Austin, Texas) (ICSE ’16). ACM, New York, NY, USA, 96–107. ...

  33. [41]

    Miryung Kim, Thomas Zimmermann, Robert DeLine, and Andrew Begel. 2018. Data Scientists in Software Teams: State of the Art and Challenges. Transactions Software Engineering 44, 11 (Nov. 2018), 1024–1038. https://doi.org/10.1109/TSE. 2017.2754374

  34. [42]

    Won Kim, Byoung-Ju Choi, Eui-Kyeong Hong, Soo-Kyung Kim, and Doheon Lee

  35. [43]

    Won Kim and Jungyun Seo. 1991. Classifying Schematic and Data Heterogeneity in Multidatabase Systems. Computer 24 (Dec. 1991), 12–18. https://doi.org/10. 1109/2.116884

  36. [44]

    Lin Li, Taoxin Peng, and Jessie Kennedy. 2011. A Rule Based Taxonomy of Dirty Data. International Journal of Computing 1, 2 (Feb. 2011), 140–148

  37. [45]

    Yang Liu, Tim Althoff, and Jeffrey Heer. 2020. Paths Explored, Paths Omitted, Paths Obscured: Decision Points & Selective Reporting in End-to-End Data Analysis. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). ACM, New Y...

  38. [46]

    Varshney, Ioana Baldini, Casey Dugan, and Aleksandra Mojsilović

    Yaoli Mao, Dakuo Wang, Michael Muller, Kush R. Varshney, Ioana Baldini, Casey Dugan, and Aleksandra Mojsilović. 2019. How Data Scientists Work Together With Domain Experts in Scientific Collaborations: To Find The Right Answer Or To Ask The Right Question? Proceedings of the A...

  39. [47]

    Hilary Mason and Chris Wiggins. 2010. A Taxonomy of Data Science. Retrieved February 9, 2021 from https://sites.google.com/a/isim.net.in/datascience_isim/ taxonomy

  40. [48]

    Dirty Data

    Marcus Messner and Bruce Garrison. 2017. Journalism’s "Dirty Data" Below Researchers’ Radar. Newspaper Research Journal 28, 4 (Aug. 2017), 88–100. https: //doi.org/10.1177/073953290702800408

  41. [49]

    Philip Meyer. 2002. Precision Journalism: A Reporter’s Introduction to Social Science Methods (4th ed.). Rowman & Littlefield, Lanham, MD, USA

  42. [50]

    Microsoft. 2017. What is the Team Data Science Process? https://docs.microsoft. com/en-us/azure/architecture/data-science-process/overview

  43. [51]

    Microsoft. 2022. Schema Drift in Mapping Data Flow. https://learn.microsoft. com/en-us/azure/data-factory/concepts-data-flow-schema-drift

  44. [52]

    Alessandra Maciel Pax Milani, Fernando V Paulovich, and Isabel Herb Manssour

  45. [53]

    Vera Liao, Casey Dugan, and Thomas Erickson

    Michael Muller, Lange Ingrid, Dakuo Wang, David Piorkowski, Jason Tsay, Q. Vera Liao, Casey Dugan, and Thomas Erickson. 2019. How Data Science Workers Work with Data: Discovery, Capture, Curation, Design, Creation. In Proceedings of the CHI Conference on Human Factors in Compu...

  46. [54]

    Tamara Munzner. 2014. Visualization Analysis and Design . CRC Press, Boca Raton, FL, USA

  47. [55]

    Heiko Müller and Johann-Christoph Freytag. 2003. Problems, Methods, and Chal- lenges in Comprehensive Data Cleansing. Technical Report. Humboldt Universität, Berlin, Germany

  48. [56]

    Information Visualization 19, 4 (Jan

    Visualization in the Preprocessing Phase: Getting Insights From Enterprise Professionals. Information Visualization 19, 4 (Jan. 2020), 273–287. https://doi. org/10.1177/1473871619896101

  49. [57]

    Adegboyega Ojo and Bahareh Heravi. 2018. Patterns in Award Winning Data Storytelling. Digital Journalism 6, 6 (Nov. 2018), 693–718. https://doi.org/10. 1080/21670811.2017.1403291

  50. [58]

    Paulo Oliveira, Fátima Rodrigues, and Pedro Henriques. 2005. A Formal Definition of Data Quality Problems. In Proceedings of the International Conference Infor- mation Quality (Cambridge, MA, USA) (ICIQ ’05). MIT, Cambridge, MA, USA,

  51. [59]

    Paulo Oliveira, Fátima Rodrigues, Pedro Henriques, and Helena Galhardas. 2005. A Taxonomy of Data Quality Problems . Instituto de Engenharia de Sistemas e Computadores - Investigação e Desenvolvimento. Retrieved April 6, 2022 from CHI ’23, April 23–28, 2023, Hamburg, Germany S...

  52. [60]

    Donald A. Norman. 2013. The Design of Everyday Things . Basic Books, New York, NY, USA

  53. [61]

    Ryan Pitts and Lindsay Muscato. 2021. Open-Source Coding Practices in Data Journalism. In The Data Journalism Handbook 2: Towards a Critical Data Prac- tice, Jonathan Gray and Liliana Bounegru (Eds.). Amsterdam University Press, Amsterdam, Netherlands, 191–193. https://doi.org...

  54. [62]

    Bernstein

    Erhard Rahm and Philip A. Bernstein. 2001. A Survey of Approaches to Automatic Schema Matching. The International Journal on Very Large Data Bases 10 (Dec. 2001), 334–350. https://doi.org/10.1007/s007780100057

  55. [63]

    http://mitiq.mit.edu/ICIQ/Documents/IQ%20Conference%202005/Papers/ AFormalDefinitionofDQProblems.pdf

  56. [64]

    Sabbir M Rashid, James P McCusker, Paulo Pinheiro, Marcello P Bax, Henrique O Santos, Jeanette A Stingone, Amar K Das, and Deborah L McGuinness. 2020. The Semantic Data Dictionary: An Approach for Describing and Annotating Data. Data Intelligence 2, 4 (Oct. 2020), 443–486. htt...

  57. [65]

    Sylvain Parasie. 2022. Computing the News: Data Journalism and the Search for Objectivity. Columbia University Press, New York, New York, USA

  58. [66]

    Jan Roeder, Jan Muntermann, and Thomas Kneib. 2020. Towards a Taxonomy of Data Heterogeneity. In 15th International Conference on Wirtschaftsinformatik (Potsdam, Germany). GITO Verlag, Berlin, Germany, 293–308. https://doi.org/10. 30844/wi_2020_c6-roeder

  59. [67]

    Simon Rogers. 2013. Data Journalism Broken Down: What We Do to the Data Before You See It. The Guardian. Retrieved August 16, 2022 from https://www. theguardian.com/news/datablog/2011/apr/07/data-journalism-workflow

  60. [68]

    Erhard Rahm and Hong Hai Do. 2000. Data Cleaning: Problems and Current Approaches. IEEE Data Eng. Bull. 23, 4 (2000), 3–13

  61. [69]

    Adam Rule, Aurélien Tabard, and James D. Hollan. 2018. Exploration and Ex- planation in Computational Notebooks. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Montréal, Canada) (CHI ’18). ACM, New York, NY, USA, 1–12. https://doi.org/10.1145/31735...

  62. [70]

    Robinson

    Rebecca S. Robinson. 2014. Purposive Sampling. In Encyclopedia of Quality of Life and Well-Being Research , Alex C. Michalos (Ed.). Springer, New York, NY, USA, 5243–5245. https://doi.org/10.1007/978-94-007-0753-5_2337

  63. [71]

    Dilruba Showkat and Eric PS Baumer. 2021. Where do Stories Come From? Examining the Exploration Process in Investigative Data Journalism. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (2021), 1–31. https: //doi.org/10.1145/3479534

  64. [72]

    Tableau Software. 2018. Tableau Prep Builder. https://www.tableau.com/ products/prep

  65. [73]

    Simon Rogers, Jonathan Schwabish, and Dainelle Bowers. 2017. Data Journalism in 2017: The Current State and Challenges Facing the Field Today. Technical Report. Google News Lab, Mountain View, CA, USA

  66. [74]

    Jonathan Stray. 2017. Making NLP Work for Investigative Journalism. Retrieved November 13, 2021 from https://www.youtube.com/watch?v=yRP9DL8E36A

  67. [75]

    Everyone Wants to do the Model Work, not the data work

    Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo. 2021. "Everyone Wants to do the Model Work, not the data work": Data Cascades in High-Stakes AI. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Y...

  68. [76]

    Jonathan Stray. 2021. Making Algorithms Work for Reporting. In The Data Journalism Handbook 2: Towards a Critical Data Practice , Jonathan Gray and Liliana Bounegru (Eds.). Amsterdam University Press, Amsterdam, Netherlands, 90–91

  69. [77]

    Jon Swain. 2018. A Hybrid Approach to Thematic Analysis in Qualitative Research: Using a Practical Example . SAGE, Thousand Oaks, FA, United States. https: //doi.org/10.4135/9781526435477

  70. [78]

    Daniele R de Souza, Lorenzo P Leuck, Caroline Q Santos, Milene S Silveira, Isabel H Manssour, and Roberto Tietzmann. 2018. Interacting with Data to Create Journalistic Stories: A Systematic review. In International Conference on Human Interface and the Management of Informatio...

  71. [79]

    Weisz, Michael Muller, Parikshit Ram, Werner Geyer, Casey Dugan, Yla Tausczik, Horst Samulowitz, and Alexander Gray

    Dakuo Wang, Justin D. Weisz, Michael Muller, Parikshit Ram, Werner Geyer, Casey Dugan, Yla Tausczik, Horst Samulowitz, and Alexander Gray. 2019. Human- AI Collaboration in Data Science: Exploring Data Scientists’ Perceptions of Au- tomated AI. Proceedings of the ACM on Human-C...

  72. [80]

    Jonathan Stray. 2019. Making Artificial Intelligence Work for Investigative Journalism. Digital Journalism 7, 8 (July 2019), 1076–1097. https://doi.org/10. 1080/21670811.2019.1630289

  73. [81]

    Rüdiger Wirth and Jochen Hipp. 2000. CRISP-DM: Towards a Standard Process Model for Data Mining. In Proceedings of the Fourth International Conference on the Practical Application of Knowledge Discovery and Data Mining (Manchester, UK) (PAKDDM ’00). Practical Application Compa...

  74. [82]

    Kanit Wongsuphasawat, Yang Liu, and Jeffrey Heer. 2019. Goals, Process, and Challenges of Exploratory Data Analysis: An Interview Study . https://doi.org/10. 48550/arXiv.1911.00568 arXiv:1911.00568

  75. [83]

    April Yi Wang, Anant Mittal, Christopher Brooks, and Steve Oney. 2019. How Data Scientists Use Computational Notebooks for Real-Time Collaboration. Pro- ceedings of the ACM on Human-Computer Interaction 3, CSCW (Nov. 2019), 1–30. https://doi.org/10.1145/3359141

  76. [84]

    Zhang, Michael Muller, and Dakuo Wang

    Amy X. Zhang, Michael Muller, and Dakuo Wang. 2020. How do Data Science Workers Collaborate? Roles, Workflows, and Tools. Proceedings of the ACM on Human-Computer Interaction 4, CSCW1 (May 2020), 1–23. https://doi.org/10. 1145/3392826

  77. [85]

    Hadley Wickham. 2014. Tidy Data. Journal of Statistical Software 59, 10 (Sept. 2014), 1–23. https://doi.org/10.18637/jss.v059.i10

  78. [88]

    Mary S Woodley. 2008. Crosswalks, Metadata Harvesting, Federated Searching, Metasearching: Using Metadata to Connect Users and Information. InIntroduction to Metadata (3rd ed.), Murtha Baca (Ed.). Getty Research Institute, Los Angeles, CA, USA

  79. [2003]

    Data Mining and Knowledge Discovery 7, 1 (Jan

    A Taxonomy of Dirty Data. Data Mining and Knowledge Discovery 7, 1 (Jan. 2003), 81–99. https://doi.org/10.1023/A:1021564703268

  80. [2012]

    InMultidisciplinary Research and Practice for Information Systems (Prague, Czech Republic)(CD-ARES ’12)

    A Taxonomy of Dirty Time-Oriented Data. InMultidisciplinary Research and Practice for Information Systems (Prague, Czech Republic)(CD-ARES ’12). Springer, New York, NY, USA, 58–72. https://doi.org/10.1007/978-3-642-32498-7_5

  81. [2020]

    In International Conference on Product-Focused Software Process Improvement (Turin, Italy) (PROFES ’20)

    Demystifying Data Science Projects: A Look on the People and Process of Data Science Today. In International Conference on Product-Focused Software Process Improvement (Turin, Italy) (PROFES ’20). Springer, New York, NY, USA, 153–167. https://doi.org/10.1007/978-3-030-64148-1_10

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.