{"id":"a87f75ad-db33-4a2a-8131-e544c8b27c61","arxiv_id":"2507.11976","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A six-dimensional taxonomy of 102 conformance checking tasks, derived from 33 case studies, provides a structured vocabulary for evaluating and designing conformance checking visualizations.","lead":"Conformance checking compares what actually happened in a business process with what should happen, but researchers have lacked a systematic way to say what analytical questions such analyses answer. This paper derives a taxonomy of 102 such tasks from 33 real-world case studies to help match visualizations to the questions they are supposed to support.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The taxonomy's external validity rests on 33 academic case studies, but the central claim that it captures real conformance checking analysis purposes is only validated for 7 of 102 tasks; a broader mapping test against commercial tool visualizations is needed.","rationale":"The reader's weakest_assumption identifies the representativeness of tasks extracted from academic case studies as the load-bearing assumption. I agree with that identification and sharpen it: the paper's own external validation covers only 7 of 102 tasks, so the gap between the academic case-study origin and real-world analytical purposes is not yet closed. This is the most load-bearing concern because the central claim is about the taxonomy's ability to determine analytical purpose of visualizations; if coverage fails, the taxonomy cannot anchor evaluation and design. However, the paper is a design-science contribution that explicitly acknowledges incompleteness, provides a systematic and transparent method, double-codes all tasks, and releases the data. The internal counts are consistent (task goals sum to 102, data targets sum to 102, etc.), and the commercial tool examples offer some evidence of real-world relevance. The concern is therefore a call for further validation, not a reason to reject the taxonomy as a first structured artifact. The reader's ACCEPT verdict with moderate confidence remains appropriate, so no verdict change is warranted.","tokens_in":21591,"tokens_out":6996,"duration_ms":76593,"concrete_test":"Take a stratified sample of roughly 50 conformance checking visualizations from commercial process mining tools (e.g., Celonis, UiPath, ARIS, myInvenio, SAP Signavio) that were not part of the case-study corpus. Have two independent coders map each visualization to the 102 tasks and their six dimensions. If more than 20% of the visualizations cannot be mapped to any existing task or require a new data characteristic or task means, the taxonomy's coverage of real analytical purposes is insufficient and the central claim weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the taxonomy gives researchers a way to determine the analytical purpose of conformance checking visualizations, which requires the 102 tasks to cover the tasks that real analysts perform. The task list is generated in Section 3.1 from 33 academic case studies, not from practitioner observation, and the authors explicitly state in Section 7.2 that the list is most likely not complete. The only external validation is Section 5, where 7 of the 102 tasks are matched to commercial tool visualizations. If a substantial fraction of real-world visualizations cannot be mapped to any of the 102 tasks, the taxonomy's usefulness for visualization evaluation and design is not established. This is a load-bearing concern because it directly affects whether the taxonomy can serve its stated purpose, even though the authors appropriately acknowledge the limitation and present the work as a first structured artifact.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a task taxonomy for conformance checking, motivated by the need to give visualization evaluation and design a task-level anchor. The authors generate tasks from 33 conformance checking case studies (screened from 209 publications), code 102 tasks into six dimensions (task goal, task means, data characteristics, constraint type, data target, data cardinality), refine the taxonomy over five iterations using the Nickerson et al. method, analyze the frequency of and dependencies between tasks, and illustrate selected tasks with visualizations from commercial process mining tools. The central contribution is the taxonomy itself, together with the claim that it provides a structured way to determine the analytical purpose of conformance checking visualizations.","tokens_in":21679,"tokens_out":7685,"duration_ms":86753,"significance":"If the taxonomy is accepted, it gives the conformance checking community a structured vocabulary for analysis tasks and provides a bridge between visual analytics and process mining. The paper's method is a genuine strength: the case-study selection uses explicit inclusion criteria, the coding is double-coded with consensus, the taxonomy development is documented across five iterations, and the data are made available online. The dependency analysis via process discovery is a useful addition that goes beyond a static classification. The external-validity concern raised by the reliance on 33 academic case studies is real but appropriately scoped: the paper explicitly frames the task list as a first observation and states in Section 7.2 that it is most likely not complete, and Section 5 presents the commercial-tool mapping as illustrative rather than exhaustive. I therefore do not regard this as a load-bearing flaw for the paper's stated contribution, although a broader mapping of all 102 tasks to tool visualizations would strengthen future work.","major_comments":[],"minor_comments":[{"comment":"The text says 'As a validation of all 101 tasks is out of scope,' but the taxonomy contains 102 tasks everywhere else in the paper (Section 3.2.3, Table 4, Section 4.1). Please correct the number.","section":"Section 5, first paragraph"},{"comment":"The text says 'Despite occurring only two times in our literature review,' but Table 3 lists 'Process conformance over time' with frequency 1 and Table 5 lists 'Describe: Derive Process conformance over time' with frequency 1. Please reconcile the count or explain which two occurrences are meant.","section":"Section 5.6"},{"comment":"The taxonomy requirements state that each dimension has mutually exclusive characteristics and that each task has exactly one characteristic per dimension, but the constraint type is later described as a 'subset choice' and as allowing tasks to be associated with multiple realizations. Please clarify that the subset itself is the single characteristic value, or revise the wording to remove the apparent contradiction.","section":"Section 3.2 and Section 4.1"},{"comment":"The supplementary literature search is described as limiting results to the '50 most relevant papers' based on a gradual decline in relevance after scanning titles and abstracts. This selection heuristic should be documented more precisely, including the search date and the decision rule, so that the screening is reproducible.","section":"Section 3.1.1"},{"comment":"The double-coding procedure is described, but no inter-rater reliability statistic (for example, Cohen's kappa) is reported. Reporting agreement would strengthen the reliability of the coding procedure that underpins the taxonomy.","section":"Section 3.2.2"},{"comment":"The caption contains a stray template fragment, '19.12.2017Beispiel-Fußzeile 2', which appears unrelated to the figure and should be removed.","section":"Figure 2 caption"},{"comment":"The dependency analysis uses an encoded task order as the timestamp in the event log. It would be helpful to state how tasks that are described in parallel or without a clear ordering within a case study were handled, since this affects the discovered process models.","section":"Section 4.3"}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it delivers something the field has been calling for: a task taxonomy specifically for conformance checking, where none existed. Second, the authors are unusually candid about its limits—they explicitly say the task list is most likely not complete and that completeness was never the goal. That candor is earned, because the method is solid for what it claims to do.\n\nThe core artifact is a six-dimensional classification of 102 tasks extracted from 33 case studies, using double coding with consensus and five iterations of taxonomy refinement following Nickerson et al. The dimensions are adapted from visualization task design spaces, but the domain-specific content—constraint type, data targets like log/trace/event, and the long tail of data characteristics—is new. The dependencies analysis via process discovery is a nice touch, giving evidence that analysts move from log-level overviews to finer-grained event-level analysis. The section matching seven tasks to visualizations in commercial tools is real external grounding, even if thin.\n\nNow the soft spots, in proportion. The stress-test concern about external validity is legitimate but not fatal. The taxonomy was built from academic case studies, not practitioner observation, and only 7 of 102 tasks are matched to commercial tool visualizations. If the taxonomy is meant to be exhaustive for real-world analyst tasks, that evidence is weak. But the authors do not claim exhaustive coverage; Section 7.2 says the list is most likely not complete and offers it as a first observation. The stated purpose—helping researchers determine the analytical purpose of visualizations—does not require a complete task list to be useful, though a substantially incomplete list would limit its power. That is a real limitation, but it is honestly disclosed and does not undermine the contribution as a first structured artifact.\n\nA second soft spot: the coding outcomes are not independently verifiable from the paper. The figshare link is provided, but without a permanent DOI or independent inspection, the reliability of the 102-task extraction rests on the authors' process description. That is typical for this kind of design-science work, and the process description is detailed enough to be credible.\n\nWho is this for? Researchers working at the intersection of process mining and visualization, especially those designing or evaluating conformance checking visualizations. It gives them a vocabulary and a structured basis for task-based evaluation. Practitioners will find it less directly useful. The paper deserves a serious referee—it is well-documented, methodologically transparent, and fills a real gap. I would send it to review, with the expectation that the completeness limitation be discussed but not treated as a fatal flaw.","headline":"A genuine first artifact: a systematically built 102-task taxonomy for conformance checking, with honest limitations that are acknowledged rather than hidden.","tokens_in":22235,"tokens_out":1104,"would_cite":true,"duration_ms":16421,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Conformance checking analyses reduce to 102 distinct tasks, each expressible as a six-dimensional tuple, giving visualization designers and evaluators a concrete target to anchor their work.","keywords":["Process Mining","Conformance Checking","Visualization","Visual Analytics","Task Taxonomy","event log analysis"],"falsifier":"Interview or shadow professional process analysts in industry and collect the conformance-related questions they ask; if a substantial share of those questions cannot be expressed as one of the 102 six-tuples without adding a new characteristic to a dimension, the taxonomy's completeness claim is refuted.","tokens_in":21373,"feed_emoji":"📊","tokens_out":5702,"duration_ms":57333,"temperature":0.7,"pith_summary":"The paper claims that conformance checking—the comparison of recorded process executions against a prescribed process model—can be systematically structured as 102 distinct analysis tasks. Each task is expressible as a six-tuple spanning task goal, task means, data characteristics, constraint type, data target, and data cardinality. The taxonomy is built from 33 academic case studies that applied conformance checking to real-world event data, and the tasks are deliberately independent of any specific conformance checking technique. If the taxonomy holds, visualization researchers and tool vendors gain a yardstick that currently does not exist: instead of asking whether a visualization is generally good, they can ask which of the 102 tasks it supports and how well.","feed_headline":"102 tasks now define the purpose of conformance-checking views","feed_subtitle":"The six-tuple scheme anchors visualization design and evaluation to concrete analytical purposes.","key_machinery":"The carrying object is the six-dimensional taxonomy itself, with each task coded as a six-tuple; the constraint type dimension—a subset choice over control-flow, data, resource, and time perspectives—was introduced to capture tasks that differ only in process perspective. The construction method follows the iterative taxonomy development method with ending conditions, seeded by the generic design space of visualization tasks as a meta-characteristic. Task means are taken from a standard classification of visualization actions, while data characteristics and data targets were derived inductively from the case studies. To find dependencies among tasks, the authors treat each case study as a trace in an event log whose activities are the tasks, then apply process discovery to the resulting log.","core_discovery":"The central claim is that the analytical purposes of conformance checking form a finite, structured space that can be captured as 102 tasks, each described by a six-tuple along the dimensions task goal (describe, explore, explain, confirm, present), task means (derive, identify, summarize, compare, present, discover, annotate, explore), data characteristics (such as guideline violations, process conformance, reasons for violations), constraint type (control-flow, data, resource, time, or a subset of these), data target (log, trace, event), and data cardinality (single, multiple, all). The taxonomy was derived through an iterative coding of tasks reported in 33 case studies and harmonized over five coding rounds. It also reports frequency patterns and dependencies: conformance analyses are predominantly exploratory rather than confirmatory, and they typically progress from a log-level overview to trace- and event-level detail, with explanatory analyses occurring last. The authors' purpose is to give visualization evaluation and design a task-based reference point, so that a visualization's usefulness can be assessed against concrete analytical purposes rather than generic visual idioms.","pith_inferences":["The same case-study-to-taxonomy pipeline could be applied to other process mining subfields, such as process discovery or predictive monitoring, yielding comparable task taxonomies that the visual analytics community could integrate into a unified process mining task model.","Adopting the six-tuple as a machine-readable task descriptor would allow a visualization to be annotated with the tasks it supports, enabling automated matching between user questions and tool features, and gap analysis across a tool's visualizations.","The taxonomy predicts testable differences in visualization effectiveness: for example, tasks with data cardinality 'all' (e.g., conformance distributions) plausibly benefit from aggregated idioms, whereas 'single'-trace tasks plausibly benefit from per-trace chart idioms such as chevron diagrams; this can be checked in user experiments.","The observed workflow pattern—log to trace to event, describe and present early, explain late—could be turned into a prescriptive interaction design pattern for conformance checking tools, including default drill-down paths and placement of explanation views."],"forward_implications":["A visualization can now be mapped to the specific tasks it supports, such as 'Describe: Derive Process conformance' or 'Explore: Identify Guideline violations', making tool evaluations comparable.","Empirical user studies can measure effectiveness and efficiency per task rather than judging a visualization in the abstract; Section 7.1 lays out such a study design for identifying guideline violations.","Taxonomy tasks can guide the design of new visualizations by making explicit which analytical purposes a design must serve, potentially reducing ad-hoc vendor-driven choices.","The dependency patterns suggest tool workflows should start with log-level overviews, support drilling through traces to events, and place explanatory analysis near the end.","The taxonomy points to underexplored tasks—confirmatory analyses appear in only two tasks—as candidate areas for future conformance checking research and tooling."],"supporting_citations":[{"why":"Supplies the initial corpus of 161 process mining case studies, extended by the authors' own review.","marker":"[4]"},{"why":"Provides the iterative taxonomy development method with ending conditions used to build and finalize the taxonomy.","marker":"[31]"},{"why":"Defines the design space of visualization tasks used as the meta-characteristic for the initial five dimensions.","marker":"[33]"},{"why":"Supplies the classification of task means (analyze, search, query) adopted for the task means dimension.","marker":"[27]"},{"why":"Provides the criteria for constructing and evaluating visualization task classifications that structure the method.","marker":"[20]"},{"why":"Defines the conformance checking techniques and terminology that delimit the domain the taxonomy describes.","marker":"[2]"}],"fun_headline_variants":["102 tasks map the purpose of conformance-checking visuals","Six-tuple scheme: 102 conformance-checking task types","Task taxonomy defines 102 conformance-checking purposes","Six dimensions, 102 tasks: conformance-checking view purposes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The list of 102 tasks is assumed to cover the questions that real conformance-checking analysts actually ask, but it is generated from 33 academic case studies that had to literally contain the term 'conformance checking' and use real-world event data; the authors explicitly state the task list is most likely not complete.","fun_headline_variants_meta":{"raw":{"variants":["102 tasks map the purpose of conformance-checking visuals","Six-tuple scheme: 102 conformance-checking task types","Task taxonomy defines 102 conformance-checking purposes","Six dimensions, 102 tasks: conformance-checking view purposes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00054,"raw_usage":{"total_tokens":2606,"prompt_tokens":978,"completion_tokens":1628,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":594,"completion_tokens_details":{"reasoning_tokens":1558}},"tokens_in":594,"tokens_out":1628,"duration_ms":14768,"temperature":1.0,"reasoning_tokens":1558,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:56:34.037482+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Interview or shadow professional process analysts in industry and collect the conformance-related questions they ask; if a substantial share of those questions cannot be expressed as one of the 102 six-tuples without adding a new characteristic to a dimension, the taxonomy's completeness claim is refuted.","supporting_citations":[{"cited_title":"Emamjome, R","cited_arxiv_id":null,"evidence_quote":"Supplies the initial corpus of 161 process mining case studies, extended by the authors' own review."},{"cited_title":"Nickerson, U","cited_arxiv_id":null,"evidence_quote":"Provides the iterative taxonomy development method with ending conditions used to build and finalize the taxonomy."},{"cited_title":"Schulz, T","cited_arxiv_id":null,"evidence_quote":"Defines the design space of visualization tasks used as the meta-characteristic for the initial five dimensions."},{"cited_title":"Munzner, Visualization analysis and design, CRC press, 2014","cited_arxiv_id":null,"evidence_quote":"Supplies the classification of task means (analyze, search, query) adopted for the task means dimension."},{"cited_title":"Kerracher, J","cited_arxiv_id":null,"evidence_quote":"Provides the criteria for constructing and evaluating visualization task classifications that structure the method."},{"cited_title":"Carmona, B","cited_arxiv_id":null,"evidence_quote":"Defines the conformance checking techniques and terminology that delimit the domain the taxonomy describes."}],"review_version":1}