REVIEW 4 major objections 6 minor 89 references
Dirty Data in the Newsroom: Comparing Data Preparation in Journalism and Data Science
T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Interviews with 36 data journalists show all 60 identified dirty-data issues fit on one two-axis taxonomy, and that table integration follows four recurring patterns.
desk verdict Solid qualitative synthesis with a genuinely new taxonomy; the comparison to data science is shakier than the rest, but the paper's own transparency about that asymmetry keeps it worth engaging. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the model-discrepancy taxonomy: a two-axis design space with four data objects (table, item, attribute, value) crossed with six data qualities (completeness, accuracy, form, granularity, relation, semantics). Its job is to give every one of the 60 synthesized dirty-data issues a cell, reconciling the bottom-up, domain-oriented language journalists use with the top-down taxonomy language of database research. A second mechanism is the extended process model, which slots 23 preparation activities into the preparation subprocesses (initiate, gather, create, profile, wrangle) and communication subprocesses (disseminate, document) of the prior model [16]. Together the two mechanisms let the authors compare evidence bases and locate divergences such as the journalists' verify-transformations activity and absence of label-data activity.
What would settle it
Run the same semi-structured interview protocol with a matched group of 36 data scientists: if they too report verify-transformation checks and no label-data work, the paper's claimed journalist-specific divergences would vanish. Alternatively, find a dirty-data issue among the 60 that cannot be assigned to exactly one of the 24 object-by-quality cells without an arbitrary choice.
Extended reading notes
Core claim
The central claim is that dirty data is best understood as a discrepancy between the mental model a data worker holds about a dataset and the data model that actually exists, because datasets are design artifacts made by people. The paper operationalizes that claim with a two-dimensional taxonomy: four data objects (table, item, attribute, value) on one axis and six data qualities (completeness, accuracy, form, granularity, relation, semantics) on the other. The taxonomy was built to cover all 60 dirty-data issues that emerged from the union of the authors' journalism interviews, prior data-science workflow papers, and 16 database and statistics taxonomies of dirty data. Alongside the taxonomy, the paper claims that data journalists' preparation can be mapped onto the same high-level process model used for data scientists, extended to 23 fine-grained activities; notable divergences are that journalists verify their transformations, avoid imputing or synthesizing data, and never label data items. The paper also identifies four challenges in combining multiple tables: regional inconsistencies, evolving diachronic schema, fragmented entities that must be reassembled, and topically disparate tables linked only by shared entities or geography.
Load-bearing premise
The comparison of data scientists and journalists rests on two differently collected evidence bases, published workflow studies for scientists and new interviews for journalists, so a group difference could be an artifact of the difference in method rather than a real difference in practice.
Editorial extensions
If this is right
- A tool designer can use the 24 cells to decide which classes of dirty data are and are not addressed by a given data preparation tool.
- The four integration challenges give data journalism support systems a concrete checklist: handle independently collected regional tables, schemas that drift across releases, fragments of a single logical dataset, and entity resolution across unrelated topics.
- Since no journalist in the study reported labeling data for machine learning, the paper implies that data journalists prepare data for descriptive reporting rather than predictive modeling, so tools that emphasize label-data workflows will miss their needs.
- The extended process model gives future studies a common activity vocabulary for comparing data work across professions.
- The taxonomy's mental-model framing suggests that cleaning is not an objective fix but an act of conforming a dataset to a particular worker's expectations, which may explain why the same table can be considered clean by one organization and dirty by another.
Reading between the lines
- I infer that the model-discrepancy framing is the paper's most portable idea: any population that consumes data without storing it, such as policy analysts, auditors, or citizens, could in principle be described with the same object-by-quality grid, and the paper itself conjectures this for domain-oriented data workers.
- A testable extension would be to run the same interview protocol on a matched group of data scientists; if verify-transformation checks and the absence of label-data work still appear, the reported journalist-scientist divergences are population differences rather than protocol artifacts.
- The four integration challenges, regional, diachronic, fragmented, and disparate table combinations, could plausibly generalize beyond journalism to any multi-source integration problem and would make a useful scenario catalog for schema-matching and entity-resolution benchmarks.
- A consequence the paper leaves implicit is that treating dirty data as a discrepancy between mental models shifts responsibility from the dataset to undocumented expectations, which strengthens the case for making data dictionaries and provenance first-class artifacts in preparation tools.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a four-phase qualitative study of data preparation in data journalism. Phase 1 derives a priori codes from 16 data science workflow papers; Phase 2 analyzes 36 semi-structured interviews with data journalists, yielding 566 coded passages and 23 consolidated activity codes; Phase 3 synthesizes 60 dirty data issues from 16 taxonomies and the interview data; Phase 4 identifies four multi-table integration challenges (regional, diachronic, fragmented, disparate) from 69 'nightmare stories.' The contributions are an extended process model, a two-axis model-discrepancy taxonomy, and the four integration challenges.
Significance. The interview corpus is a genuine empirical contribution, and the paper's systematic accounting—150 excerpts, 315 taxonomy instances, 566 coded passages, 60 synthesized issues—is transparently documented via OSF supplements. The proposed taxonomy is concrete and testable in principle: its distributional claims about which issues are reported by database researchers versus domain-oriented users can be checked by coding a new sample. The four integration challenges are a useful, design-relevant vocabulary. The central risk is not in the data collection but in the comparative framing: the paper contrasts interview-based codes with literature-based codes as if they were commensurable measurements of the two populations.
major comments (4)
- [§3.1–§3.2, Figure 2] The central comparison between data scientists and journalists is built on two non-commensurable evidence bases: Phase 1 codes come from 16 published workflow papers (150 excerpts), while Phase 2 codes come from 36 new interviews using a different protocol, including pre-interview artifact sharing and prompts for 'nightmare stories.' The paper treats an activity's absence from the literature as a divergence (blue/green coding in Figure 2), but silence in a paper is not evidence of absence in practice; the authors concede this for 'identify items' in §4.5, calling it 'under-reported in data science workflows.' The same logic affects 'verify transformations' (§4.4) and the claimed lack of 'label data items' (§4.5). Because no commensurability test is offered (e.g., applying the interview protocol to a matched sample of data scientists, or recoding the 16 papers using interview-derived codes and reporting coverage), the title's 'comparing' claim is stronger than the data support. I recommend reframing all divergences as 'reported in the data science literature' versus 'reported in journalist interviews,' and adding an explicit limitations paragraph, or providing a supplementary coding check on a subset of the literature corpus.
- [§3.2.3, §3.2.5, §5.3] All coding was performed by the first author, with no second coder, inter-rater reliability, member checking, or audit trail beyond the OSF materials. The paper nevertheless reports precise counts—43 initial activity codes, 26 issue codes, 13 unique to interviews, 16 unique to previous work, 31 overlapping (Fig. 3a)—as though they were stable measurements. Since the synthesis of 60 issues and the four integration challenges depend on these codings, the lack of any reliability or reflexivity check leaves open the possibility that the taxonomy is idiosyncratic. Adding a second coder on a sample of passages or at least a clear statement of why reliability is not appropriate for this design would strengthen the central claims.
- [§3.2.4] The saturation claim is stated but not evidenced: 'growth of our codebook's cardinality' is a reasonable proxy, but no saturation curve or stopping criterion is reported, and the reader cannot tell how many new codes appeared in the last five interviews. Given that the paper's later counts depend on codebook stability, the saturation argument needs at least a supplemental plot or a statement of the threshold used.
- [§6] The fourth contribution is supported by a single aggregate number: 63 of 69 nightmare stories exhibit at least one of the four challenges, and 11 involve more than one. But the section does not report how many stories fall into each of the four categories (counts are given only for fragmented and disparate, 17/36 and 14/36), nor how the remaining six stories were classified. Because the 'nightmare story' prompt specifically elicits difficulties, the prevalence figures may overstate how common these challenges are in routine journalistic work; a limitations note is needed.
minor comments (6)
- [Table 1] Table 1 misspells 'Oliveira' as 'Oliveria' and 'Manssour' as 'Mannssour'; these should be corrected.
- [References] Reference [17] lists the author as 'Theordore Dasu, Tamraparni & Johnson'; the first author's name should be 'Theodore' and the attribution to Dasu & Johnson corrected.
- [Figure 3] Figure 3(a) is very dense; a larger or interactive version of the 60-issue mapping in the supplement would help readers verify the object/quality assignments referenced in Supp. Section 4.
- [§5.2] The definitions of 'relation' and 'semantics' overlap for duplicate items, which are classified under Semantics; clarify the boundary with an example of a relational duplicate.
- [§4.2] The acronym FOI is used without expansion at first use in §4.2; spell out 'freedom of information.'
- [§7.2] The term 'MacGyvering' should be introduced with a neutral definition of its scope before the pop-culture reference, since the paper uses it as a category of practice.
Circularity Check
No significant circularity: the paper's contributions are an inductive empirical synthesis, and its self-citations are peripheral to the central argument.
full rationale
The paper's derivation chain is an inductive qualitative synthesis. Phase 1 derives a priori codes from 16 data-science workflow studies; Phase 2 codes 36 journalist interviews with those codes plus new inductive codes; Phase 3 merges interview issues with 16 dirty-data taxonomies; Phase 4 re-analyzes nightmare stories to induce four integration challenges. No equation or fitted parameter is used, and no claim is derived by definition from its own input. The taxonomy is explicitly built from the 60 issue codes (Sections 3.3 and 5.3), so its coverage statement is a description of the construction rather than a hidden prediction; this is normal taxonomic synthesis, not circularity. The comparison between journalists and data scientists rests on two different evidence corpora, and absence of a practice in published workflow studies may reflect reporting granularity; but this is a methodological limitation of the comparison, not a circular reduction. Self-citations are peripheral: Table Scraps [39] is cited as complementary and for the 'nerd box' concept, and Munzner [54] supplies terminology for the four data objects; neither carries the central argument. No uniqueness theorem or prior author-derived constraint is invoked. The paper is self-contained as an empirical and taxonomic contribution, so the circularity score is 0.
Assumptions & free parameters
assumptions (5)
- domain assumption Self-reported interview accounts accurately reflect data journalists' real preparation practices.
- domain assumption The 16 data science workflow papers and the 16 dirty data taxonomies constitute a representative and sufficient corpus for comparison and synthesis.
- domain assumption Single-coder thematic analysis with reflective synthesis yields reliable categories.
- domain assumption Theoretical saturation after 36 interviews, proxied by codebook cardinality growth, justifies stopping.
- domain assumption The Crisan et al. four-process model is the correct scaffold for mapping journalist activities.
Cite this review
Pith. "Pith review of Dirty Data in the Newsroom: Comparing Data Preparation in Journalism and Data Science." pith.science (2026). https://pith.science/paper/MOIRTNJB
@misc{pith2026250707238,
author = {Pith},
title = {Pith review of: Dirty Data in the Newsroom: Comparing Data Preparation in Journalism and Data Science},
year = {2026},
howpublished = {\url{https://pith.science/paper/MOIRTNJB}},
note = {Machine review of arXiv:2507.07238}
}
read the original abstract
The work involved in gathering, wrangling, cleaning, and otherwise preparing data for analysis is often the most time consuming and tedious aspect of data work. Although many studies describe data preparation within the context of data science workflows, there has been little research on data preparation in data journalism. We address this gap with a hybrid form of thematic analysis that combines deductive codes derived from existing accounts of data science workflows and inductive codes arising from an interview study with 36 professional data journalists. We extend a previous model of data science work to incorporate detailed activities of data preparation. We synthesize 60 dirty data issues from 16 taxonomies on dirty data and our interview data, and we provide a novel taxonomy to characterize these dirty data issues as discrepancies between mental models. We also identify four challenges faced by journalists: diachronic, regional, fragmented, and disparate data sources.
Figures
Reference graph
Works this paper leans on
-
[1]
Timo Aho, Outi Sievi-Korte, Terhi Kilamo, Sezin Yaman, and Tommi Mikkonen
-
[2]
Sara Alspaugh, Nava Zokaei, Andrea Liu, Cindy Jin, and Marti A. Hearst. 2019. Futzing and Moseying: Interviews with Professional Data Analysts on Exploration Practices. Transactions Visualization and Computer Graphics 25, 1 (Jan. 2019), 22–31. https://doi.org/10.1109/TVCG.2018.2865040
arXiv 2019
-
[3]
In the Beginning Were the Data
Ángel Arrese. 2022. "In the Beginning Were the Data": Economic Journalism as/and Data Journalism. Journalism Studies 23, 4 (Feb. 2022), 487–505. https: //doi.org/10.1080/1461670X.2022.2032803
-
[4]
José Barateiro and Helena Galhardas. 2005. A Survey of Data Quality Tools. Datenbank-Spektrum 4, 14 (Aug. 2005), 15–21. http://dc-pubs.dbs.uni-leipzig.de/ files/Barateiro2005ASurveyofDataQuality.pdf
2005
-
[5]
Andrea Batch and Niklas Elmqvist. 2018. The Interactive Visualization Gap in Initial Exploratory Data Analysis. Transactions Visualization and Computer Graphics 24, 1 (Jan. 2018), 278–287. https://doi.org/10.1109/TVCG.2017.2743990
arXiv 2018
-
[6]
Leilani Battle and Jeffrey Heer. 2019. Characterizing Exploratory Visual Analysis: A Literature Review and Evaluation of Analytic Provenance in Tableau.Computer Graphics Forum 38, 3 (July 2019), 145–159. https://doi.org/10.1111/cgf.13678
-
[7]
Charles Berret and Cheryl Phillips. 2016. Teaching Data and Computational Journalism. Columbia Journalism School, New York, NY, USA
2016
-
[8]
Eddy Borges-Rey. 2021. Journalism with Machines? From Computational Think- ing to Distributed Cognition. In The Data Journalism Handbook 2: Towards a Critical Data Practice , Jonathan Gray and Liliana Bounegru (Eds.). Amster- dam University Press, Amsterdam, Netherlands, 92–95. https://doi.org/10.1515/ 9789048542079-023
2021
Show all 89 references
-
[9]
1998.Transforming Qualitative Information: Thematic analysis and code development
Richard E Boyatzis. 1998.Transforming Qualitative Information: Thematic analysis and code development. SAGE, Thousand Oaks, California
1998
-
[10]
Paul Bradshaw. 2011. The Inverted Pyramid of Data Journalism. Retrieved Aug. 13, 2021 from https://onlinejournalismblog.com/2011/07/07/the-inverted- pyramid-of-data-journalism
2011
-
[11]
Suzana Guedes Cardoso. 2022. The Practice of Data Journalism and Changes in the Professional Profile of Journalists in Newsrooms in the United States, United Kingdom, and Brazil. In Digital Convergence in Contemporary Newsrooms: Media Innovation, Content Adaptation, Digital Tr...
2022 doi
-
[12]
Abhirup Chatterjee and Arie Segev. 1991. Data Manipulation in Heterogeneous Databases. ACM SIGMOD Record 20, 4 (Dec. 1991), 64–68. https://doi.org/10. 1145/141356.141385
1991
-
[13]
Fanny Chevalier, Melanie Tory, Bongshin Lee, Jarke van Wijk, Giuseppe Santucci, Marian Dörk, and Jessica Hullman. 2018. From Analysis to Communication: Sup- porting the Lifecycle of a Story. InData-Driven Storytelling, Nathalie Henry Riche, Dirty Data in the Newsroom CHI ’23, ...
2018 doi
-
[14]
Hamilton, and Fred Turner
Sarah Cohen, James T. Hamilton, and Fred Turner. 2011. Computational Journal- ism. Commun. ACM 54, 10 (Oct. 2011), 66–71. https://doi.org/10.1145/2001269. 2001288
2011 doi
-
[15]
Anamaria Crisan and Brittany Fiore-Gartland. 2021. Fits and Starts: Enterprise Use of AutoML and the Role of Humans in the Loop. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21). ACM, New York, NY, USA, 1–15. https://doi.or...
2021
-
[16]
Anamaria Crisan, Brittany Fiore-Gartland, and Melanie Tory. 2020. Passing the Data Baton: A Retrospective Analysis on Data Science Work and Workers. Transactions Visualization and Computer Graphics 27, 2 (Oct. 2020), 1860–1870. https://doi.org/10.1109/TVCG.2020.3030340
2020
-
[17]
Theordore Dasu, Tamraparni & Johnson. 2003. Exploratory Data Mining and Data Cleaning. John Wiley & Sons, Hoboken, NJ, USA
2003
-
[18]
Wesley Gongora de Almeida, Rafael Timóteo de Sousa, Flávio Elias de Deus, Georges Daniel Amvame Nze, and Fábio Lúcio Lopes de Mendonça. 2013. Tax- onomy of Data Quality Problems in Multidimensional Data Warehouse Mod- els. In Proceedings 8th Iberian Conference on Information S...
2013
-
[19]
Catherine D’Ignazio and Lauren F. Klein. 2020. Data Feminism . MIT Press, Cambridge, Massachusetts
2020
-
[20]
David Donoho. 2017. 50 Years of Data Science. Journal of Computational and Graphical Statistics 26, 4 (Oct. 2017), 745–766. https://doi.org/10.1080/10618600. 2017.1384734
2017
-
[21]
Shari L. Dworkin. 2012. Sample Size Policy for Qualitative Studies Using In- Depth Interviews. Archives Sexual Behavior 41, 6 (Sept. 2012), 1319–1320. https: //doi.org/10.1007/s10508-012-0016-6
2012 doi
-
[22]
Marc A Feldman. 2020. Data Quality: Dimensions, Measurement, Strategy, Man- agement, and Governance. Quality Progress 53, 2 (2020), 54–54
2020
-
[23]
U.M. Feyyad. 1996. Data Mining and Knowledge Discovery: Making Sense Out of Data. Expert 11, 5 (Oct. 1996), 20–25. https://doi.org/10.1109/64.539013
1996 doi
-
[24]
Katherine Fink and C. W. Anderson. 2015. Data Journalism in the United States. Journalism Studies 16, 4 (July 2015), 467–481. https://doi.org/10.1080/1461670X. 2014.939852
2015
-
[25]
Bruce Frey (Ed.). 2018. The SAGE Encyclopedia of Educational Research, Measure- ment, and Evaluation . SAGE, Thousand Oaks, California. https://doi.org/10. 4135/9781506326139
2018
-
[26]
Garrett Grolemund and Hadley Wickham. 2014. A Cognitive Interpretation of Data Analysis: A Cognitive Interpretation of Data Analysis. International Statistical Review 82, 2 (Aug. 2014), 184–204. https://doi.org/10.1111/insr.12028
2014 doi
-
[27]
Theresia Gschwandtner, Johannes Gärtner, Wolfgang Aigner, and Silvia Miksch
-
[28]
Joseph M Hellerstein. 2008. Quantitative Data Cleaning for Large Databases. United Nations Economic Commission for Europe. https://dsf.berkeley.edu/jmh/ papers/cleaning-unece.pdf
2008
-
[29]
Bahareh Heravi, Kathryn Cassidy, Edie Davis, and Natalie Harrower. 2022. Pre- serving Data Journalism: A Systematic Literature Review. Journalism Practice 16, 10 (March 2022), 2083–2105. https://doi.org/10.1080/17512786.2021.1903972
2022
-
[30]
Karen Holtzblatt and Hugh Beyer. 2015. Contextual Design: Evolved . Spinger Nature, Berlin, Germany
2015
-
[31]
Sarah Hutchins. 2020. Data Dive: School’s Out. The Investigative Reporters & Editors Journal 43, 1 (Feb. 2020), 6–7
2020
-
[32]
Kaggle. 2019. State of Data Science and Machine Learning . Kaggle. Retrieved May 15, 2022 from https://www.kaggle.com/kaggle-survey-2019
2019
-
[33]
Sean Kandel, Jeffrey Heer, Catherine Plaisant, Jessie Kennedy, Frank van Ham, Nathalie Henry Riche, Chris Weaver, Bongshin Lee, Dominique Brodbeck, and Paolo Puono. 2011. Research Directions in Data Wrangling: Visualizations and Transformations for Usable and Credible Data. In...
2011 doi
-
[34]
Sean Kandel, Andreas Paepcke, Joseph Hellerstein, and Jeffrey Heer. 2011. Wran- gler: Interactive Visual Specification of Data Transformation Scripts. In Pro- ceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’11). ACM, Vancouver, Canada, 3363–3372. htt...
2011
-
[35]
Hellerstein, and Jeffrey Heer
Sean Kandel, Andreas Paepcke, Joseph M. Hellerstein, and Jeffrey Heer. 2012. Enterprise Data Analysis and Visualization: An Interview Study. Transactions Visualization and Computer Graphics 18, 12 (Dec. 2012), 2917–2926. https://doi. org/10.1109/TVCG.2012.219
2012 doi
-
[36]
Hellerstein, and Jeffrey Heer
Sean Kandel, Ravi Parikh, Andreas Paepcke, Joseph M. Hellerstein, and Jeffrey Heer. 2012. Profiler: Integrated Statistical Analysis and Visualization for Data Quality Assessment. In Proceedings of the International Working Conference on Advanced Visual Interfaces (Capri Island...
2012
-
[37]
Haber, and Jeffrey S
Eser Kandogan, Aruna Balakrishnan, Eben M. Haber, and Jeffrey S. Pierce. 2014. From Data to Insight: Work Practices of Analysts in the Enterprise. Computer Graphics and Applications 34, 5 (Sept. 2014), 42–50. https://doi.org/10.1109/MCG. 2014.62
2014 doi
-
[38]
H. Kang, L. Getoor, B. Shneiderman, M. Bilgic, and L. Licamele. 2008. Interactive Entity Resolution in Relational Data: A Visual Analytic Tool and Its Evaluation. Transactions Visualization and Computer Graphics 14, 5 (Sept. 2008), 999–1014. https://doi.org/10.1109/TVCG.2008.55
2008 doi
-
[39]
Stephen Kasica, Charles Berret, and Tamara Munzner. 2020. Table Scraps: An Actionable Framework for Multi-Table Data Wrangling From An Artifact Study of Computational Journalism. Transactions Visualization and Computer Graphics 27, 2 (2020), 957–966. https://doi.org/10.1109/TV...
2020
-
[40]
Miryung Kim, Thomas Zimmermann, Robert DeLine, and Andrew Begel. 2016. The Emerging Role of Data Scientists on Software Development Teams. In Pro- ceedings of the 38th International Conference on Software Engineering (Austin, Texas) (ICSE ’16). ACM, New York, NY, USA, 96–107. ...
2016
-
[41]
Miryung Kim, Thomas Zimmermann, Robert DeLine, and Andrew Begel. 2018. Data Scientists in Software Teams: State of the Art and Challenges. Transactions Software Engineering 44, 11 (Nov. 2018), 1024–1038. https://doi.org/10.1109/TSE. 2017.2754374
2018
-
[42]
Won Kim, Byoung-Ju Choi, Eui-Kyeong Hong, Soo-Kyung Kim, and Doheon Lee
-
[43]
Won Kim and Jungyun Seo. 1991. Classifying Schematic and Data Heterogeneity in Multidatabase Systems. Computer 24 (Dec. 1991), 12–18. https://doi.org/10. 1109/2.116884
1991
-
[44]
Lin Li, Taoxin Peng, and Jessie Kennedy. 2011. A Rule Based Taxonomy of Dirty Data. International Journal of Computing 1, 2 (Feb. 2011), 140–148
2011
-
[45]
Yang Liu, Tim Althoff, and Jeffrey Heer. 2020. Paths Explored, Paths Omitted, Paths Obscured: Decision Points & Selective Reporting in End-to-End Data Analysis. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). ACM, New Y...
2020
-
[46]
Varshney, Ioana Baldini, Casey Dugan, and Aleksandra Mojsilović
Yaoli Mao, Dakuo Wang, Michael Muller, Kush R. Varshney, Ioana Baldini, Casey Dugan, and Aleksandra Mojsilović. 2019. How Data Scientists Work Together With Domain Experts in Scientific Collaborations: To Find The Right Answer Or To Ask The Right Question? Proceedings of the A...
2019 doi
-
[47]
Hilary Mason and Chris Wiggins. 2010. A Taxonomy of Data Science. Retrieved February 9, 2021 from https://sites.google.com/a/isim.net.in/datascience_isim/ taxonomy
2010
-
[48]
Dirty Data
Marcus Messner and Bruce Garrison. 2017. Journalism’s "Dirty Data" Below Researchers’ Radar. Newspaper Research Journal 28, 4 (Aug. 2017), 88–100. https: //doi.org/10.1177/073953290702800408
2017 doi
-
[49]
Philip Meyer. 2002. Precision Journalism: A Reporter’s Introduction to Social Science Methods (4th ed.). Rowman & Littlefield, Lanham, MD, USA
2002
-
[50]
Microsoft. 2017. What is the Team Data Science Process? https://docs.microsoft. com/en-us/azure/architecture/data-science-process/overview
2017
-
[51]
Microsoft. 2022. Schema Drift in Mapping Data Flow. https://learn.microsoft. com/en-us/azure/data-factory/concepts-data-flow-schema-drift
2022
-
[52]
Alessandra Maciel Pax Milani, Fernando V Paulovich, and Isabel Herb Manssour
-
[53]
Vera Liao, Casey Dugan, and Thomas Erickson
Michael Muller, Lange Ingrid, Dakuo Wang, David Piorkowski, Jason Tsay, Q. Vera Liao, Casey Dugan, and Thomas Erickson. 2019. How Data Science Workers Work with Data: Discovery, Capture, Curation, Design, Creation. In Proceedings of the CHI Conference on Human Factors in Compu...
2019
-
[54]
Tamara Munzner. 2014. Visualization Analysis and Design . CRC Press, Boca Raton, FL, USA
2014
-
[55]
Heiko Müller and Johann-Christoph Freytag. 2003. Problems, Methods, and Chal- lenges in Comprehensive Data Cleansing. Technical Report. Humboldt Universität, Berlin, Germany
2003
-
[56]
Information Visualization 19, 4 (Jan
Visualization in the Preprocessing Phase: Getting Insights From Enterprise Professionals. Information Visualization 19, 4 (Jan. 2020), 273–287. https://doi. org/10.1177/1473871619896101
2020 doi
-
[57]
Adegboyega Ojo and Bahareh Heravi. 2018. Patterns in Award Winning Data Storytelling. Digital Journalism 6, 6 (Nov. 2018), 693–718. https://doi.org/10. 1080/21670811.2017.1403291
2018 arXiv
-
[58]
Paulo Oliveira, Fátima Rodrigues, and Pedro Henriques. 2005. A Formal Definition of Data Quality Problems. In Proceedings of the International Conference Infor- mation Quality (Cambridge, MA, USA) (ICIQ ’05). MIT, Cambridge, MA, USA,
2005
-
[59]
Paulo Oliveira, Fátima Rodrigues, Pedro Henriques, and Helena Galhardas. 2005. A Taxonomy of Data Quality Problems . Instituto de Engenharia de Sistemas e Computadores - Investigação e Desenvolvimento. Retrieved April 6, 2022 from CHI ’23, April 23–28, 2023, Hamburg, Germany S...
2005
-
[60]
Donald A. Norman. 2013. The Design of Everyday Things . Basic Books, New York, NY, USA
2013
-
[61]
Ryan Pitts and Lindsay Muscato. 2021. Open-Source Coding Practices in Data Journalism. In The Data Journalism Handbook 2: Towards a Critical Data Prac- tice, Jonathan Gray and Liliana Bounegru (Eds.). Amsterdam University Press, Amsterdam, Netherlands, 191–193. https://doi.org...
2021 doi
-
[62]
Bernstein
Erhard Rahm and Philip A. Bernstein. 2001. A Survey of Approaches to Automatic Schema Matching. The International Journal on Very Large Data Bases 10 (Dec. 2001), 334–350. https://doi.org/10.1007/s007780100057
2001 doi
-
[63]
http://mitiq.mit.edu/ICIQ/Documents/IQ%20Conference%202005/Papers/ AFormalDefinitionofDQProblems.pdf
-
[64]
Sabbir M Rashid, James P McCusker, Paulo Pinheiro, Marcello P Bax, Henrique O Santos, Jeanette A Stingone, Amar K Das, and Deborah L McGuinness. 2020. The Semantic Data Dictionary: An Approach for Describing and Annotating Data. Data Intelligence 2, 4 (Oct. 2020), 443–486. htt...
2020 doi
-
[65]
Sylvain Parasie. 2022. Computing the News: Data Journalism and the Search for Objectivity. Columbia University Press, New York, New York, USA
2022
-
[66]
Jan Roeder, Jan Muntermann, and Thomas Kneib. 2020. Towards a Taxonomy of Data Heterogeneity. In 15th International Conference on Wirtschaftsinformatik (Potsdam, Germany). GITO Verlag, Berlin, Germany, 293–308. https://doi.org/10. 30844/wi_2020_c6-roeder
2020
-
[67]
Simon Rogers. 2013. Data Journalism Broken Down: What We Do to the Data Before You See It. The Guardian. Retrieved August 16, 2022 from https://www. theguardian.com/news/datablog/2011/apr/07/data-journalism-workflow
2013
-
[68]
Erhard Rahm and Hong Hai Do. 2000. Data Cleaning: Problems and Current Approaches. IEEE Data Eng. Bull. 23, 4 (2000), 3–13
2000
-
[69]
Adam Rule, Aurélien Tabard, and James D. Hollan. 2018. Exploration and Ex- planation in Computational Notebooks. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Montréal, Canada) (CHI ’18). ACM, New York, NY, USA, 1–12. https://doi.org/10.1145/31735...
2018
-
[70]
Robinson
Rebecca S. Robinson. 2014. Purposive Sampling. In Encyclopedia of Quality of Life and Well-Being Research , Alex C. Michalos (Ed.). Springer, New York, NY, USA, 5243–5245. https://doi.org/10.1007/978-94-007-0753-5_2337
2014 doi
-
[71]
Dilruba Showkat and Eric PS Baumer. 2021. Where do Stories Come From? Examining the Exploration Process in Investigative Data Journalism. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (2021), 1–31. https: //doi.org/10.1145/3479534
2021 doi
-
[72]
Tableau Software. 2018. Tableau Prep Builder. https://www.tableau.com/ products/prep
2018
-
[73]
Simon Rogers, Jonathan Schwabish, and Dainelle Bowers. 2017. Data Journalism in 2017: The Current State and Challenges Facing the Field Today. Technical Report. Google News Lab, Mountain View, CA, USA
2017
-
[74]
Jonathan Stray. 2017. Making NLP Work for Investigative Journalism. Retrieved November 13, 2021 from https://www.youtube.com/watch?v=yRP9DL8E36A
2017
-
[75]
Everyone Wants to do the Model Work, not the data work
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo. 2021. "Everyone Wants to do the Model Work, not the data work": Data Cascades in High-Stakes AI. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Y...
2021
-
[76]
Jonathan Stray. 2021. Making Algorithms Work for Reporting. In The Data Journalism Handbook 2: Towards a Critical Data Practice , Jonathan Gray and Liliana Bounegru (Eds.). Amsterdam University Press, Amsterdam, Netherlands, 90–91
2021
-
[77]
Jon Swain. 2018. A Hybrid Approach to Thematic Analysis in Qualitative Research: Using a Practical Example . SAGE, Thousand Oaks, FA, United States. https: //doi.org/10.4135/9781526435477
2018 doi
-
[78]
Daniele R de Souza, Lorenzo P Leuck, Caroline Q Santos, Milene S Silveira, Isabel H Manssour, and Roberto Tietzmann. 2018. Interacting with Data to Create Journalistic Stories: A Systematic review. In International Conference on Human Interface and the Management of Informatio...
2018 doi
-
[79]
Weisz, Michael Muller, Parikshit Ram, Werner Geyer, Casey Dugan, Yla Tausczik, Horst Samulowitz, and Alexander Gray
Dakuo Wang, Justin D. Weisz, Michael Muller, Parikshit Ram, Werner Geyer, Casey Dugan, Yla Tausczik, Horst Samulowitz, and Alexander Gray. 2019. Human- AI Collaboration in Data Science: Exploring Data Scientists’ Perceptions of Au- tomated AI. Proceedings of the ACM on Human-C...
2019 doi
-
[80]
Jonathan Stray. 2019. Making Artificial Intelligence Work for Investigative Journalism. Digital Journalism 7, 8 (July 2019), 1076–1097. https://doi.org/10. 1080/21670811.2019.1630289
2019 arXiv
-
[81]
Rüdiger Wirth and Jochen Hipp. 2000. CRISP-DM: Towards a Standard Process Model for Data Mining. In Proceedings of the Fourth International Conference on the Practical Application of Knowledge Discovery and Data Mining (Manchester, UK) (PAKDDM ’00). Practical Application Compa...
2000
- [82]
-
[83]
April Yi Wang, Anant Mittal, Christopher Brooks, and Steve Oney. 2019. How Data Scientists Use Computational Notebooks for Real-Time Collaboration. Pro- ceedings of the ACM on Human-Computer Interaction 3, CSCW (Nov. 2019), 1–30. https://doi.org/10.1145/3359141
2019 doi
-
[84]
Zhang, Michael Muller, and Dakuo Wang
Amy X. Zhang, Michael Muller, and Dakuo Wang. 2020. How do Data Science Workers Collaborate? Roles, Workflows, and Tools. Proceedings of the ACM on Human-Computer Interaction 4, CSCW1 (May 2020), 1–23. https://doi.org/10. 1145/3392826
2020
-
[85]
Hadley Wickham. 2014. Tidy Data. Journal of Statistical Software 59, 10 (Sept. 2014), 1–23. https://doi.org/10.18637/jss.v059.i10
2014 doi
-
[88]
Mary S Woodley. 2008. Crosswalks, Metadata Harvesting, Federated Searching, Metasearching: Using Metadata to Connect Users and Information. InIntroduction to Metadata (3rd ed.), Murtha Baca (Ed.). Getty Research Institute, Los Angeles, CA, USA
2008
-
[2003]
Data Mining and Knowledge Discovery 7, 1 (Jan
A Taxonomy of Dirty Data. Data Mining and Knowledge Discovery 7, 1 (Jan. 2003), 81–99. https://doi.org/10.1023/A:1021564703268
2003 doi
-
[2012]
InMultidisciplinary Research and Practice for Information Systems (Prague, Czech Republic)(CD-ARES ’12)
A Taxonomy of Dirty Time-Oriented Data. InMultidisciplinary Research and Practice for Information Systems (Prague, Czech Republic)(CD-ARES ’12). Springer, New York, NY, USA, 58–72. https://doi.org/10.1007/978-3-642-32498-7_5
-
[2020]
In International Conference on Product-Focused Software Process Improvement (Turin, Italy) (PROFES ’20)
Demystifying Data Science Projects: A Look on the People and Process of Data Science Today. In International Conference on Product-Focused Software Process Improvement (Turin, Italy) (PROFES ’20). Springer, New York, NY, USA, 153–167. https://doi.org/10.1007/978-3-030-64148-1_10
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.