{"id":"abbcb961-a8e8-4c87-9f36-1356bebbbb83","arxiv_id":"2501.01738","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A taxonomy with nine new annotation categories maps EU political advertising rules (DSA, TTPA, G-E-DSA) to concrete content-analysis codes.","lead":"This paper builds a codebook of nine new content categories for analyzing political ads and user posts under three EU digital election rules. It is a methodological tool for researchers and auditors, but it has not yet been tested on real content.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Several Ro codes (Ro-4, Ro-8, Ro-9) target platform-level obligations, not content features, so annotating individual ads with this codebook cannot by itself measure compliance as claimed.","rationale":"The reader's weakest-assumption analysis is correct but under-specified: the legal mapping's weakness is not only that the codebook is untested, but that several categories are not content-level codes at all. This is an internal construct-validity issue, visible from the codebook's own definitions and examples, not merely an absence of inter-coder reliability. The paper's stated method of applying annotation categories to ad content cannot produce scores for platform-level obligations unless additional data sources and a separate coding unit are introduced. This does not invalidate the taxonomy as a risk-signal framework, so the reader's CONDITIONAL verdict remains appropriate. A revised version could add a platform-level compliance module, specify data linkage (e.g., ad-repository metadata, silence-period calendars), and pilot the instrument. The reader and I both identify the same broad area; I sharpen the concern to unit-of-analysis mismatch.","tokens_in":14099,"tokens_out":3130,"duration_ms":33548,"concrete_test":"Operationalization test: select 100 political ads from Meta Ad Library during a national election period; two coders independently apply the full codebook using only the ad creative and visible text, as the paper's content-analysis method implies. Record the number of items where Ro-2, Ro-3, Ro-4, Ro-5, Ro-8, or Ro-9 cannot be assigned because the needed information (sponsor, targeting criteria, repository status, silence-period context, fact-check label) is absent from the content item. If this exceeds a small threshold (e.g., more than 10%), the codebook is not self-contained and the central claim requires revision to combine content coding with platform-data extraction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the taxonomy/codebook enables systematic content analysis of user-generated and political ad content to assess compliance with DSA/TTPA/G-E-DSA. The weakest link is a unit-of-analysis mismatch: the codebook mixes content-level categories with platform-level compliance obligations. In the Annex, Ro-4 'Ad Repository Requirements' (Art 13 TTPA, Art 39 DSA) requires judging whether a platform maintains a public repository; Ro-8 'Silence Period Enforcement' requires knowing whether an ad ran during a national silence period; Ro-9 'Monitoring and Fact-checking' requires knowing whether the platform attached a fact-check label. Even Ro-2 (sponsor identification), Ro-3 (targeting criteria disclosure), and Ro-5 (political ad labelling) depend on ad-repository metadata or platform display state, not on the observable content artifact. Section 5 says coders assess compliance by applying annotation categories to a sample of content, but no separate instrument or data linkage is specified for these platform-level codes. Therefore, even if each legal citation is doctrinally correct, the taxonomy as presented cannot deliver the claimed compliance measurement on individual content items; at best it measures content-level risk signals, with compliance requiring an additional platform-audit layer.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a taxonomy/codebook for content analysis of political content and political ads under the EU's Digital Services Act (DSA), the Regulation on Transparency and Targeting of Political Advertising (TTPA), and the Commission's Guidelines for electoral processes (G-E-DSA). It introduces nine new annotation categories (Ro-1 to Ro-9), maps each to legal provisions, and augments them with previously published categories for electoral-rights risks and disinformation. The authors argue that the resulting taxonomy enables systematic content analysis of user-generated and ad content to assess compliance with regulatory mandates and to support systemic risk assessments under Art. 34 and Art. 37 DSA. The paper contains a legal doctrinal review, a table of the nine new categories, and an annex codebook, but no empirical validation of the codebook and no specification of how platform-level codes can be applied to individual content items.","tokens_in":14310,"tokens_out":4542,"duration_ms":42424,"significance":"If validated, this taxonomy would provide a useful bridge between EU regulatory texts and empirical content analysis, especially for external audits under Art. 37 DSA and for research on Art. 34(1)(c) DSA systemic electoral risks. The legal mapping is transparent and the codebook structure (code, source, definition, example) is a sensible, replicable format. The paper is best understood as a well-organized doctrinal proposal rather than a completed empirical instrument; its main value at this stage is as a starting point for annotation studies, provided the proposed categories are pretested and platform-level data are integrated.","major_comments":[{"comment":"The central claim that the taxonomy 'enables systematic content analysis ... to assess compliance' is not supported by the proposed annotation procedure because several codes apply to platform-level obligations rather than to individual content items. Ro-4 ('Ad Repository Requirements') requires a judgment about whether a platform maintains a public repository, Ro-8 ('Silence Period Enforcement') requires knowing whether an ad ran during a national silence period, and Ro-9 ('Monitoring and Fact-checking') requires knowing whether the platform attached a fact-check label; Ro-2, Ro-3, and Ro-5 similarly depend on ad-repository metadata or platform display state. Section 4 describes coders applying each annotation category to a sample of content, but no separate instrument or data linkage is specified for these platform-level codes, so the codebook as presented cannot deliver compliance measurement on individual content items without an additional platform-audit layer.","section":"§4 and Annex 7.2 (Ro-4, Ro-8, Ro-9)"},{"comment":"The claim that the taxonomy enables systematic content analysis is not empirically tested. The methodology states that coding consistency is measured through Krippendorff's Alpha and pre-tests, but the paper reports no pretest results, no inter-coder reliability coefficient, and no annotated sample; the conclusion explicitly defers validation to 'future work.' As it stands, the contribution is a proposed codebook, not a validated instrument, and the abstract's wording should be revised accordingly or supplemented with a pilot annotation study.","section":"§4 Methodology and §6 Conclusion"},{"comment":"The legal mapping for Ro-8 ('Silence Period Enforcement') is incomplete: the source column cites only 'Rec. 14,' while the body text in Section 3.3 says platforms are 'encouraged to respect' silence periods, reflecting that silence periods are primarily imposed by Member State law rather than by a directly binding EU obligation. Since the paper's central claim rests on the correctness and completeness of the code-to-provision mapping, this category needs a fuller legal basis or a re-framing as a Member-State-dependent indicator.","section":"Annex 7.2, Ro-8"},{"comment":"Ro-1 is tied to 'Rec 81 DSA, Art 34(1)(d) DSA' and to the Directive on combating violence against women, whereas the paper's stated focus is the electoral-process risk under Art. 34(1)(c) DSA; the relationship between gender-based violence and electoral-process risk is asserted rather than argued, so the inclusion of this category within an electoral-risk codebook needs a clearer justification.","section":"Table 1 and §5"}],"minor_comments":[{"comment":"There is a typo in the definition of 'qualification': it reads 'e.g.,hoxes' and should be 'e.g., hoaxes.'","section":"§4"},{"comment":"The abbreviation 'TTP' is used interchangeably with 'TTPA' (e.g., RQ2 and Section 3.2); please use a single abbreviation consistently.","section":"Throughout"},{"comment":"The source list mixes article ranges and individual articles in a way that is ambiguous ('Art. 3-9 Directive ... Art. 3 Directive ... Art. 4 Forced marriage'); please format each source as a complete, unambiguous citation.","section":"Annex 7.2, Ro-1 source list"},{"comment":"The cited reference (Zapf et al.) is a comparison of inter-rater coefficients, not the primary source for Krippendorff's Alpha; please cite Krippendorff's 'Content Analysis' directly or a methodology text that defines the coefficient.","section":"§4, reference to Krippendorff's Alpha"},{"comment":"The text states that the taxonomy includes 'misrepresentation in political advertisements' as a category, but no such standalone category appears in Table 1 or the annex; please align the narrative with the actual codebook.","section":"§5"}],"recommendation":"major_revision","confidential_remarks":"The paper sits at the boundary of legal-doctrinal work and empirical social science. For a cs.CY venue, the lack of any empirical validation is a significant gap, and the unit-of-analysis problem with the Ro codes is the main technical obstacle. The novel contribution of the nine Ro categories relative to the authors' prior codebooks should also be stated more explicitly, since the annex relies heavily on earlier published categories."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a method paper, not an empirical result. The new artifact is a nine-category extension (Ro-1 to Ro-9) of the authors' earlier German-election codebook, meant to annotate political content under the DSA, TTPA, and the Commission's electoral guidelines. The legal mapping is careful and the annex is usable; that is the real value. If you need a starting vocabulary for ad-compliance audits, this is a fair baseline.\n\nThe soft spot is bigger than the reader's report suggests, though not fatal to the taxonomy's existence. The paper's core sentence says the codebook 'enables systematic content analysis... to assess compliance.' But several categories are not content features. Ro-4 (ad repository), Ro-8 (silence periods), and Ro-9 (fact-check labeling) require platform-level or contextual data. Even Ro-2, Ro-3, and Ro-5 depend on ad-repository metadata or the display state of the ad, not on the ad's text or image alone. Section 5 says coders apply categories to a sample of content, but the paper doesn't specify a second instrument or data linkage for platform state. So as written, the taxonomy can generate content-level risk signals, not compliance measurements. The authors should either narrow the claim or add a platform-audit layer.\n\nThe other gap is validation. The paper explicitly defers pre-tests and Krippendorff's alpha to future work, and the conclusion says future work must validate and operationalize the taxonomy. That is honest, but it means the load-bearing premise—that the legal reading is correct and annotatable—is untested. The reader's conditional verdict is right. There is no fitted-parameter circularity here; the main debt is to the authors' own earlier codebooks, and that is legitimate since the nine Ro categories are genuinely new.\n\nWho should read it: people designing transparency audits or coding political ads in the EU. It is a serious referee candidate because the taxonomy is grounded, reproducible in its annex, and likely to be used regardless of validation. I'd recommend peer review with a required revision: either reframe as a content-level risk taxonomy or specify the platform-level data needed for compliance codes, and report at least a small pilot.\n\nMy recommendation: send it to peer review, but expect the authors to cut the compliance claim down to size.","headline":"A transparent codebook for EU political-ad compliance whose central claim outruns its evidence: the legal mapping is careful, but several codes target platform-level obligations that a content sample cannot see, and the paper explicitly defers validation.","tokens_in":14852,"tokens_out":1952,"would_cite":false,"duration_ms":20697,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper presents a nine-category codebook that maps the EU's DSA, TTPA, and electoral guidelines onto annotation categories for measuring political-content compliance.","keywords":["Digital Services Act","political advertising","content analysis","codebook","electoral integrity","systemic risk","transparency","European elections"],"falsifier":"A direct test would be to take a sample of political ads that regulators have already found non-compliant, such as ads with missing sponsor information or synthetic media without labels, have independent coders apply the Ro-1 to Ro-9 codes, and check whether the codebook flags all known violations without excessive false positives. A second check is legal: if a code's source provisions were mis-cited or a mandatory duty such as the Article 39 DSA repository fields is not represented, the mapping fails.","tokens_in":13874,"feed_emoji":"🗳️","tokens_out":4580,"duration_ms":43636,"temperature":0.7,"pith_summary":"The paper argues that the EU's new digital electoral rules — the Digital Services Act, the Transparency and Targeting of Political Advertising Regulation, and the Commission's electoral guidelines — can be rendered as a concrete content-analysis taxonomy. It proposes a codebook with nine new annotation categories (Ro-1 to Ro-9) covering gender-based violence, sponsor identification, targeting disclosure, ad repositories, ad labelling, influencer content, synthetic media, silence periods, and fact-checking, each tied to specific legal provisions. If the taxonomy is right, external auditors and platform risk teams could sample user-generated and paid political content and systematically score compliance with electoral-integrity obligations. The paper does not yet test the codebook; its stated next step is empirical validation.","feed_headline":"Nine codes map EU election laws onto online political content","feed_subtitle":"A codebook turns DSA and TTPA duties into annotation categories for compliance checks.","key_machinery":"The load-bearing mechanism is the legal-doctrinal mapping from each annotation code to named provisions in the three instruments. The codebook's unit is a code: a short alphanumeric label such as Ro-5 for political ad labelling, with a source line citing provisions like Articles 26 and 39 DSA or Article 11 TTPA, a qualification that determines what counts, a definition, and an example to anchor coder judgment. The nine new regulatory codes are designed to capture the transparency and risk-mitigation duties that older electoral and disinformation taxonomies did not cover. The reliability machinery is inter-coder agreement measured by Krippendorff's Alpha on sample content drawn by cluster or random sampling.","core_discovery":"The paper's central claim is that a legally grounded codebook can operationalize the compliance duties in the DSA, TTPA, and G-E-DSA for empirical content analysis. By doing doctrinal close reading of the three instruments and merging them with existing disinformation and election-risk category sets, the authors construct a three-part annotation system: electoral-rights codes, regulatory codes Ro-1 through Ro-9, and disinformation codes. Each code carries a source provision, a qualification, a definition, and examples, so that annotators can apply it to text, image, video, and audio content. The intended use is to test systemic risks under Article 34(1)(c) DSA, specifically negative effects on civic discourse and electoral processes, and to support external audit of platform claims under Article 37 DSA. The authors are explicit that the taxonomy is a tool for future validation, not a validated instrument.","pith_inferences":["Editorial inference: the paper's legal mapping is untested; the most decisive next study would be an inter-coder reliability test on real ad-repository data from a very large online platform, with the codes applied by coders who have no stake in the taxonomy.","Editorial inference: the taxonomy may need extension for non-ad political content and for organic influencer posts whose status under the TTPA definition of political advertising is unsettled; the paper itself notes the definitional complexity of political content.","Editorial inference: if validated, the codebook could become a bridge between the DSA audit regime and the AI Act's transparency duties for synthetic content, since it already codes deepfakes and AI-generated political media.","Editorial inference: the absence of a test means the taxonomy should be read as a hypothesis about which legal provisions are observable in content, not as evidence of what platforms currently fail to do."],"forward_implications":["Auditors and civil-society monitors can apply the codebook to a sample of ads and user posts to produce a structured compliance report across the three regulatory domains.","The taxonomy gives Article 34(1)(c) DSA a measurable empirical meaning: negative effects on civic discourse and electoral processes become observable annotation categories rather than an open-ended legal phrase.","Platforms can use the same categories when building ad repositories, labeling synthetic media, enforcing silence periods, and responding to fact-checker labels, aligning content moderation with the G-E-DSA guidance.","Because each code cites specific provisions, the framework can support automated classifiers or language-model detection pipelines that need a labeled dataset aligned with EU law.","Comparative studies across member states and elections become possible using a shared codebook keyed to the same legal texts."],"supporting_citations":[{"why":"The Digital Services Act is the core legal source; its provisions on ad transparency, systemic risk, and ad repositories are mapped to the Ro codes.","marker":"3"},{"why":"The Transparency and Targeting of Political Advertising Regulation supplies sponsor identification, targeting disclosure, labelling, and repository duties used in Ro-2 through Ro-5.","marker":"4"},{"why":"The Commission's electoral guidelines provide the basis for influencer content, synthetic media labelling, silence periods, and fact-checking duties in Ro-6 through Ro-9.","marker":"5"},{"why":"The prior German election study supplies the electoral-risk categories and the approach of defining systemic risks to the electoral process that the codebook adapts.","marker":"11"},{"why":"The disinformation taxonomy provides the D-1 to D-10 definitions adopted for the disinformation section of the codebook.","marker":"38"},{"why":"The prior content-moderation codebook demonstrates the method of mapping legal provisions to annotation categories and is cited as the basis for the coding guideline.","marker":"41"},{"why":"Krippendorff's Alpha is selected as the inter-coder reliability measure for pretesting and refining the taxonomy.","marker":"40"}],"fun_headline_variants":["Codebook turns EU digital election laws into annotator categories","New taxonomy maps DSA and TTPA duties to content codes","Legal codebook operationalizes EU campaign compliance checks","EU election rules coded for content analysis: a taxonomy","From law to labels: EU digital election compliance codes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the paper's reading of the DSA, TTPA, and G-E-DSA is legally correct and complete enough that applying the nine codes to content actually measures compliance.","fun_headline_variants_meta":{"raw":{"variants":["Codebook turns EU digital election laws into annotator categories","New taxonomy maps DSA and TTPA duties to content codes","Legal codebook operationalizes EU campaign compliance checks","EU election rules coded for content analysis: a taxonomy","From law to labels: EU digital election compliance codes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000593,"raw_usage":{"total_tokens":2695,"prompt_tokens":777,"completion_tokens":1918,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":393,"completion_tokens_details":{"reasoning_tokens":1839}},"tokens_in":393,"tokens_out":1918,"duration_ms":12752,"temperature":1.0,"reasoning_tokens":1839,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:20:44.692197+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would be to take a sample of political ads that regulators have already found non-compliant, such as ads with missing sponsor information or synthetic media without labels, have independent coders apply the Ro-1 to Ro-9 codes, and check whether the codebook flags all known violations without excessive false positives. A second check is legal: if a code's source provisions were mis-cited or a mandatory duty such as the Article 39 DSA repository fields is not represented, the mapping fails.","supporting_citations":[{"cited_title":"Federal Elections, 2004-2020, Journal of Quantitative Description: Digital Media 2,","cited_arxiv_id":null,"evidence_quote":"The prior content-moderation codebook demonstrates the method of mapping legal provisions to annotation categories and is cited as the basis for the coding guideline."},{"cited_title":"This provision demands the identifiable and consistent indication that the content is labeled as advertising content (Art","cited_arxiv_id":null,"evidence_quote":"Krippendorff's Alpha is selected as the inter-coder reliability measure for pretesting and refining the taxonomy."}],"review_version":1}