{"id":"adbb766b-5206-48a9-9e7a-49f7e231e31c","arxiv_id":"2412.03991","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper maps quality control techniques for crowdsourcing tasks with large or infinite answer spaces into a two-tiered framework.","lead":"This survey reviews how researchers measure and improve answer quality in open-ended crowdsourcing tasks such as translation, image segmentation and creative writing, where many different answers can be acceptable. It organizes the literature into a two-tiered framework covering task, worker, answer and system. A smart generalist might read it to see how open-ended crowd data is vetted before it becomes training data or published content.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The two-tier framework is not applied to the System aspect: Section 6 lacks the quality-dimension and evaluation-metric structure promised in Section 2.3, so the central claim overstates what the survey delivers.","rationale":"The reader's weakest assumption concerns literature selection in Section 2.4, which is a real threat to the survey's map and gap analysis. I agree that the corpus could be incomplete or biased. However, the more direct threat to the central claim is internal: the two-tiered framework, which is the paper's headline contribution, is not actually applied to the System aspect. Sections 3, 4, and 5 each have explicit Quality Dimensions, Quality Evaluation, and Quality Control Methods subsections; Section 6 instead organizes material around task assignment, answer aggregation, and workflow design, with no parallel structure and no System-level evaluation metrics. This is independent of corpus coverage and is verifiable from the text alone. The reader's conditional verdict remains appropriate because the gap is structural and fixable: either Section 6 should be restructured to match the promised second tier, or the paper should be explicit that System is a cross-cutting integration layer rather than a fourth aspect with its own dimensions and metrics. My concern does not push the verdict to rejection; it strengthens the case for revision before acceptance, so I leave the reader's conditional verdict unchanged.","tokens_in":31456,"tokens_out":11341,"duration_ms":111361,"concrete_test":"Conduct a section-by-section structural audit of the paper's promised taxonomy. For each of the four aspects (Task, Worker, Answer, System), list the specific subsections that define its quality dimensions, its evaluation metrics, and its design decisions. If Section 6 has no subsection that can be tagged as 'quality evaluation' and no dimension definitions beyond the three process topics in 6.1-6.3, then the claim that the second tier applies 'in each aspect' is unsupported. The minimal revision would be to add a short System-model subsection defining System-level dimensions and metrics, or to explicitly redefine System as a cross-aspect integration layer and adjust the abstract and Section 2.3 accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's core contribution is the two-tiered framework, so the central claim requires every first-tier aspect to receive the promised second-tier treatment. Section 2.3 states that 'each section corresponds to one aspect in the quality model, with quality dimensions, evaluation metrics and design decisions reviewed,' and the abstract repeats that the second tier applies 'in each aspect.' Sections 3, 4, and 5 follow this structure explicitly, with subsections like 3.1 Quality Dimensions, 3.2 Quality Evaluation, and 3.3 Quality Control Methods. Section 6, which is the 'System' aspect, does not. Its subsections are 6.1 Task Assignment, 6.2 Answer Aggregation, and 6.3 Workflow Design, and the section frames them as 'the three core tasks in the crowdsourcing execution process' rather than as quality dimensions. No System-level evaluation metrics are reviewed at all. This is an internal mismatch with the stated taxonomy, not a matter of field consensus: as written, the framework is not actually two-tiered across all four aspects. The issue is load-bearing because the framework itself is the central contribution; if the System aspect is simply a cross-cutting integration layer, then the abstract and Section 2.3 need to say so, or Section 6 needs a genuine second-tier structure.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript surveys quality control in open-ended crowdsourcing and proposes a two-tiered framework: the first tier identifies four aspects (task, worker, answer, system), and the second tier classifies works within each aspect into quality dimensions, evaluation metrics, and design decisions. The survey describes the literature selection method, reviews representative works for the task, worker, and answer aspects, discusses system-level concerns as cross-aspect issues, and outlines challenges and future directions, including implications of large language models for crowdsourcing.","tokens_in":31721,"tokens_out":5308,"duration_ms":45981,"significance":"If the proposed framework is applied consistently, the survey could serve as a useful organizing map for a research area that is increasingly important as open-ended annotation tasks and LLM-assisted crowdsourcing grow. The paper does several things well: it gives concrete examples of open-ended tasks and their answer-space properties (Table 1), it provides a clear three-part second-tier structure for the task, worker, and answer aspects, and it includes an explicit, reproducible-looking literature search procedure in Section 2.4. The discussion of LLMs in Section 7.2 is timely and connects crowdsourcing quality control to current practice. The main caveat is that the System aspect does not receive the same second-tier treatment, which creates an internal inconsistency in the central contribution.","major_comments":[{"comment":"Section 2.3 states that \"each section corresponds to one aspect in the quality model, with quality dimensions, evaluation metrics and design decisions reviewed,\" and the abstract promises the second tier \"in each aspect.\" Sections 3, 4, and 5 indeed follow this structure (e.g., 3.1 Quality Dimensions, 3.2 Quality Evaluation, 3.3 Quality Control Methods). Section 6, however, is organized around \"the three core tasks in the crowdsourcing execution process\" (6.1 Task Assignment, 6.2 Answer Aggregation, 6.3 Workflow Design) and provides no quality-dimension or evaluation-metric subsections for the System aspect. This is an internal inconsistency in the paper's central contribution. The authors should either (a) restructure Section 6 so that the System aspect also receives the promised second-tier treatment, identifying System-level quality dimensions and evaluation metrics, or (b) revise the framework description to state explicitly that System is a cross-cutting integration layer to which the second tier does not apply in full, and adjust the abstract and Section 2.3 accordingly.","section":"Section 2.3 and Section 6"},{"comment":"The literature selection is described as a keyword search across 22 conferences and 14 journals for 2012-2023, followed by one-pass title/abstract screening by the authors, plus \"additional papers derived from the authors' prior knowledge.\" However, Section 1.5 claims a \"systematic review of all related works.\" The selection process has no inter-rater reliability, no citation chaining or snowballing, and no explicit inclusion/exclusion criteria beyond excluding simple crowdsourcing tasks, so the completeness and representativeness of the corpus are not established. Since the proposed taxonomy and the gap analysis in Section 7.1 depend on which papers are included, the authors should either strengthen the methodology (e.g., dual screening, inter-rater agreement, snowballing) or temper the claim of a systematic review and add a limitations discussion in the conclusion.","section":"Section 2.4 and Section 1.5"}],"minor_comments":[{"comment":"The text attributes reference [94] to \"Li et al.,\" but reference [94] is Zheng et al., \"Truth inference in crowdsourcing: Is the problem solved?\" (PVLDB 2017). The citation should be corrected.","section":"Section 1.4"},{"comment":"The fourth aspect is called \"context\" in the contributions list (Section 1.5), \"workflow\" in Section 7, and \"system\" elsewhere (Section 2.1 and the abstract). Unify the naming to avoid confusion about the framework's first tier.","section":"Section 1.5 and Section 7"},{"comment":"The first bullet reads \"Quality model. in each aspect\" — the period should be removed so the sentence reads \"Quality model in each aspect refers to a collection of quality dimensions...\"","section":"Section 2.3"},{"comment":"There is a typo in \"froms open-ended answers\" — this should be \"from open-ended answers.\" Other minor spacing/ligature artifacts appear in the abstract and several places (e.g., \"su ffi ciently\"), likely from LaTeX rendering; these should be cleaned up.","section":"Section 5.2.1"}],"recommendation":"major_revision","confidential_remarks":"The self-citation [160] (Chai et al.) appears once in Section 1.1 as an example of open-ended question answering and does not prop up the taxonomy; I do not see a conflict-of-interest concern. The main issue is the structural mismatch between the promised two-tier framework and the actual organization of Section 6. This is fixable by either restructuring Section 6 or revising the framework claim. The literature-selection concern is well-taken but also fixable through additional validation or a more cautious statement of completeness. The paper is within scope for a survey-oriented venue and should be reconsidered after these revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This survey is a genuinely useful reference for the crowdsourcing quality-control community. The open-ended focus fills a real gap: Daniel et al. and Jin et al. organize general or Boolean quality control, and this paper's two-tiered framework — aspects (task, worker, answer, system) broken down into quality dimensions, evaluation metrics, and design decisions — is a reasonable way to structure the space. The task, worker, and answer sections (3–5) follow the framework consistently and cover the literature broadly, including a structured search of 147 papers plus the authors' prior knowledge.\n\nThe most serious problem is that Section 6, the System aspect, does not deliver the promised second tier. The abstract and Section 2.3 say each aspect gets quality dimensions, evaluation metrics, and design decisions. Sections 3, 4, and 5 do that. Section 6 is organized by task assignment, answer aggregation, and workflow design, and there are no System-level evaluation metrics reviewed. The section frames these as 'the three core tasks in the crowdsourcing execution process' rather than as quality dimensions. That is an internal mismatch with the stated framework. It is fixable — either restructure Section 6 or soften the claim that the framework applies 'in each aspect' — but as written it overstates what the survey delivers.\n\nThere are also smaller issues. Section 1.4 attributes the truth-inference survey [94] to 'Li et al.'; the reference is Zheng et al. (VLDB 2017). The literature selection in Section 2.4 is described, but there is no inter-rater reliability, the screening is one pass by title and abstract, and 'papers derived from the authors' prior knowledge' is vague. That means coverage could have blind spots, though nothing I saw suggests major omissions.\n\nThese flaws are addressable and do not invalidate the framework. The open-ended focus is underserved, the two-tier structure is useful, and the LLM discussion in Section 7.2 is timely.\n\nI'd send this to peer review with a request for revision. The referee should push on the System section and on the transparency of the literature search. Anyone working on crowdsourcing quality control, especially for open-ended tasks, would want to cite this after those fixes.","headline":"A useful survey with a promising two-tier framework for open-ended crowdsourcing quality control, undermined by a System section that doesn't follow the promised taxonomy and some citation and screening issues.","tokens_in":32212,"tokens_out":3848,"would_cite":true,"duration_ms":30819,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey proposes a two-tiered framework that organizes quality control research for open-ended crowdsourcing into task, worker, answer, and system aspects.","keywords":["crowdsourcing","open-ended tasks","quality control","survey","two-tiered framework","task model","worker model","answer aggregation"],"falsifier":"Re-run the literature search with a broader keyword set (for example, 'free-form annotation', 'subjective annotation', 'generative tasks', 'LLM annotation') and with two independent screeners; if this finds many relevant quality control papers that do not fit the two-tiered framework, or finds substantial work in venues outside the chosen list, the survey's coverage claim would be refuted.","tokens_in":31287,"feed_emoji":"🗺️","tokens_out":6263,"duration_ms":50188,"temperature":0.7,"pith_summary":"Crowdsourcing tasks with open-ended answers—such as translation, image segmentation, and story writing—accept a large or infinite space of correct responses, so standard majority-vote and probabilistic-model quality control designed for Boolean tasks does not transfer. This survey argues that quality control for these tasks is best understood through a two-tiered framework: a first tier covering the aspects of task, worker, answer, and system, and a second tier breaking each aspect into quality dimensions, evaluation metrics, and design decisions. The paper reviews the resulting literature, organizes it into this framework, and identifies which areas are mature and which are missing. If the framework is right, it gives requesters, workers, and platform developers a usable map for choosing and designing quality control methods, and it sharpens the research agenda for open-ended crowdsourcing.","feed_headline":"New framework maps quality control for open-ended crowdsourcing","feed_subtitle":"Survey organizes 147 papers on translation, segmentation, and writing to show what works and what gaps remain.","key_machinery":"The central object is the two-tiered quality control framework (the taxonomy). The first tier partitions the literature by the aspect of the crowdsourcing process being optimized: task, worker, answer, and system. The second tier further classifies each aspect into quality dimensions (the attributes that are modeled, such as task design, worker expertise, answer reliability), evaluation metrics (how quality is measured, including automatic metrics, peer feedback, and expert ratings), and design decisions (the methods used to optimize quality, such as task mapping, workflow design, teaching, incentives, and aggregation algorithms). The framework's work is to make scattered papers comparable and to expose what is under-studied.","core_discovery":"The paper's central claim is that a two-tiered framework capturing quality dimensions, evaluation metrics, and design decisions across the aspects of task, worker, answer, and system accounts for the state of quality control research in open-ended crowdsourcing. The first tier provides a holistic view of what determines answer quality, while the second tier exposes the internal structure of each aspect: which quality attributes are modeled, how quality is measured, and what design choices are made to improve it. Surveying papers from 2012 to 2023 across major venues, the survey finds that most proposed methods are task-specific, that a few cross-task approaches exist for answer aggregation and evaluation, and that system-level joint optimization of task assignment, aggregation, and workflow is comparatively rare. The survey also positions open-ended crowdsourcing relative to Boolean crowdsourcing and to the emerging role of large language models as both annotators and sources of quality problems.","pith_inferences":["The boundary between Boolean and open-ended tasks is likely a spectrum rather than a dichotomy, so a graded version of the framework might better predict when Boolean methods such as majority voting or probabilistic graphical models can be adapted.","The 'system' aspect is the thinnest in the survey's account, suggesting that a formal definition of system-level quality metrics and their interactions would be a natural next step, though the paper does not develop one.","As LLMs increasingly generate crowd answers, the 'worker' aspect may need to be reinterpreted: worker modeling could become model behavior modeling, shifting quality control toward prompt design, consistency checks, and output filtering—an extension the paper gestures at but leaves open."],"forward_implications":["Researchers can position new quality control methods within the framework and immediately see which combinations of aspect, dimension, and metric are already populated.","The survey's gap analysis implies that general or cross-task quality control methods, applicable across data types, are a priority because only a few such approaches exist.","Because the framework treats quality control as a system-level problem, it points toward joint optimization of task assignment, answer aggregation, and workflow design rather than optimizing each step in isolation.","In the era of large language models, the framework can be used to design quality control for hybrid human-AI annotation pipelines, where answers may originate from crowd workers or from LLMs."],"supporting_citations":[{"why":"Supplies the definition of quality as 'conformance to requirements' that sets the survey's scope for what counts as quality control.","marker":"[216]"},{"why":"The prior survey of crowdsourcing quality control from statistical and mechanism-design perspectives that this survey extends to open-ended tasks.","marker":"[215]"},{"why":"A prior survey organized by quality attributes, assessment techniques, and assurance actions, which the paper contrasts with its own framework.","marker":"[213]"},{"why":"A prior summary of truth-inference algorithms for crowdsourcing, used as the baseline for the answer-aggregation discussion.","marker":"[94]"},{"why":"Dawid-Skene, the canonical probabilistic model for Boolean answers, used to illustrate why standard aggregation methods do not transfer to open-ended answers.","marker":"[99]"},{"why":"ZenCrowd, an example of a Boolean-oriented quality control method that relies on a finite answer space and unique ground truth.","marker":"[47]"},{"why":"VisualGenome, a large open-ended annotation dataset that motivates the need for quality control in tasks with infinite answer spaces.","marker":"[153]"},{"why":"One of the few cross-task answer-aggregation methods via annotation distances, used as evidence for the survey's claim that general methods are rare.","marker":"[78]"},{"why":"Merging-and-matching aggregation for complex annotations, another rare cross-task approach that supports the gap analysis.","marker":"[9]"}],"fun_headline_variants":["Two-tier framework maps quality control for open-ended crowdsourcing","Open-ended crowdsourcing quality: survey reveals task-specific methods","System-level quality control rare in open-ended crowdsourcing survey","New survey: LLMs reshape open-ended crowdsourcing quality control"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's taxonomy and gap analysis rest on its literature selection: a keyword search of 22 conferences and 14 journals for 2012 to 2023, screened once by title and abstract, with no inter-rater reliability check and with extra papers added from the authors' prior knowledge.","fun_headline_variants_meta":{"raw":{"variants":["Two-tier framework maps quality control for open-ended crowdsourcing","Open-ended crowdsourcing quality: survey reveals task-specific methods","System-level quality control rare in open-ended crowdsourcing survey","New survey: LLMs reshape open-ended crowdsourcing quality control"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000295,"raw_usage":{"total_tokens":1709,"prompt_tokens":934,"completion_tokens":775,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":550,"completion_tokens_details":{"reasoning_tokens":707}},"tokens_in":550,"tokens_out":775,"duration_ms":7483,"temperature":1.0,"reasoning_tokens":707,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:51:50.738352+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the literature search with a broader keyword set (for example, 'free-form annotation', 'subjective annotation', 'generative tasks', 'LLM annotation') and with two independent screeners; if this finds many relevant quality control papers that do not fit the two-tiered framework, or finds substantial work in venues outside the chosen list, the survey's coverage claim would be refuted.","supporting_citations":[{"cited_title":"Quality is free: The art of making quality certain","cited_arxiv_id":null,"evidence_quote":"Supplies the definition of quality as 'conformance to requirements' that sets the survey's scope for what counts as quality control."},{"cited_title":"A technical survey on statisti- cal modelling and design methods for crowdsourcing quality control","cited_arxiv_id":null,"evidence_quote":"The prior survey of crowdsourcing quality control from statistical and mechanism-design perspectives that this survey extends to open-ended tasks."},{"cited_title":"Quality control in crowdsourcing: A survey of quality attributes, as- sessment techniques, and assurance actions","cited_arxiv_id":null,"evidence_quote":"A prior survey organized by quality attributes, assessment techniques, and assurance actions, which the paper contrasts with its own framework."}],"review_version":1}