Pith. sign in

REVIEW 3 major objections 6 minor 33 references

Behind Closed Words: Creating and Investigating the forePLay Annotated Dataset for Polish Erotic Discourse

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper introduces forePLay, the first Polish manually annotated dataset of erotic content—24,768 sentences labeled across five categories—and reports that Polish-specific models beat multilingual alternatives on every label…

desk verdict A real Polish erotic-content dataset with honest documentation, but the benchmark claims rest on labels that may be too unstable to trust until the outlier annotator is removed and the results re-checked. read the letter →

arxiv 2412.17533 v3 pith:JIJFJTMD submitted 2024-12-23 cs.CL

classification cs.CL
keywords forePLayPolishlanguageeroticcontentdetectionannotateddatasetmoderationinter-annotatoragreementlanguage-specificmodelsmorphologicallycomplexlanguages
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is trying to establish that Polish-language erotic content can be detected reliably by models trained on a new resource, forePLay, the first manually annotated Polish dataset of erotic discourse. The dataset uses a five-way taxonomy—erotic, ambiguous, violence-related, socially unacceptable, and neutral—that the authors argue captures the context-dependence and moral dimensions that binary schemes miss. Their evaluations show that specialized Polish models (HerBERT, Polish RoBERTa, PLLuM, Bielik) outperform multilingual and general-purpose systems on this benchmark, which they take as evidence that content moderation for morphologically complex languages needs language-specific resources. The paper also documents substantial disagreement among annotators, so part of its contribution is a candid account of how unstable the ground truth is for this task.

What carries the argument

The load-bearing object is the annotation scheme: a five-way exclusive taxonomy (erotic, ambiguous, violence-related, socially unacceptable, neutral) with a fixed priority order for overlapping categories and a deliberately separated ambiguous class for context-dependent erotic connotations. This scheme produces the final labels through majority voting with a superannotator resolving total disagreements, and every model score in the paper is a measurement of that aggregated ground truth.

What would settle it

Re-annotate a random sample of roughly 500 forePLay sentences with a fresh annotator pool using the same guidelines, then measure agreement between the new labels and the published majority-vote labels; if agreement for the ambiguous category falls near chance, the ground truth that all model comparisons rest on is not stable. Alternatively, retrain the same models on per-annotator labels rather than majority votes—if the ranking of Polish-specific versus multilingual models flips, the paper's central comparison is an artifact of label aggregation.

Watch

Extended reading notes

Core claim

The central discovery is forePLay itself: a dataset of 24,768 Polish sentences drawn from online fiction repositories and published literature, annotated by six raters with majority vote and a superannotator breaking three-way ties. The five exclusive labels follow a fixed priority order (socially unacceptable outranks violence-related, which outranks erotic), and the separate 'ambiguous' label is designed for sentences whose erotic reading depends on context. On this benchmark the paper reports that Polish-specialized encoder models reach macro-F1 scores of 0.929–0.944 in binary classification, with Polish RoBERTa leading HerBERT, and that fine-tuned PLLuM-Mistral-12B reaches 0.946. Polish-specific models consistently beat multilingual baselines such as GPT-4o, Llama 3.1, Mixtral, and Command-R, although all models lose accuracy as the number of classes grows, dropping to roughly 0.66 on the five-class task.

Load-bearing premise

The load-bearing premise is that the aggregated majority-vote labels, with a superannotator breaking ties, are reliable enough to serve as ground truth—but the paper's own numbers show a Krippendorff's alpha of only 0.387 before removing the most divergent annotator, so if the labels are unstable every reported model score is a measurement of that instability.

Editorial extensions

If this is right

  • Content moderation for Polish should use Polish-specific models rather than English-centric multilingual tools, since the paper's comparisons show a consistent advantage for the specialized models on every label configuration.
  • The taxonomy is a reusable template for erotic-content detection in other morphologically complex languages, provided each language receives its own manually annotated dataset.
  • The strong performance drop with more classes (from roughly 0.94 in binary to 0.66 in five-class macro-F1) implies that fine-grained moderation needs either more training data for rare categories or a different evaluation strategy such as learning with disagreements.
  • The public release of 3,704 erotic and ambiguous sentences gives researchers a reproducible benchmark, although the rare violence-related and socially unacceptable classes are excluded from the release for ethical reasons.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's own agreement statistics suggest the majority-vote labels may be too noisy to serve as a stable ground truth; a natural test is to re-annotate a random sample with a fresh annotator pool and measure agreement with the published labels.
  • If the ranking of Polish-specific versus multilingual models were computed on per-annotator labels or soft labels instead of majority votes, the reported advantage might change, since the aggregated labels are where the noise is concentrated.
  • Because 69% of the corpus comes from amateur online fiction, the models' edge on this benchmark may not transfer to other Polish registers such as chat, social media, or professional prose, which the evaluation does not directly test.
  • A direct extension would be to train the same models on disagreement-weighted objectives and test whether the Polish-model advantage survives when the target is an individual reader's judgment rather than an aggregated label.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces forePLay, a Polish-language dataset of 24,768 sentences annotated for erotic content detection, with a five-class taxonomy (erotic, ambiguous, violence-related, socially unacceptable, neutral). The authors document the data collection from online fiction repositories and literary works, the annotation process with six annotators, and the aggregation via majority vote with a superannotator for three-way ties. They report inter-annotator agreement (Krippendorff's alpha 0.387 overall, 0.716 after excluding one outlier annotator, Fem1) and benchmark a wide range of Polish-specific and multilingual models on binary, three-class, four-class, and five-class classification tasks. The central claims are that forePLay is the first Polish manually annotated dataset of erotic content and that specialized Polish language models outperform multilingual alternatives on this benchmark.

Significance. If the dataset labels are reliable, forePLay is a valuable resource: it addresses a genuine gap for a morphologically complex, non-English language, provides a multidimensional taxonomy rather than a simple binary, includes LGBTQ+ representation, and offers transparent documentation of the annotation process and limitations. The paper also provides a broad empirical comparison including Polish encoder models, Polish LLMs, and multilingual LLMs, with detailed error analyses. The authors are unusually candid about annotation difficulties and label instability, which is a strength. However, the central empirical claim rests on ground-truth labels whose stability is questionable, and the reported aggregate agreement is low; this tempers the significance of the benchmark comparisons until the label reliability issue is addressed.

major comments (3)
  1. [Section 4 and Section 4.2, Table 4] The final labels are determined by majority vote over three annotators, but the statistical evidence in Section 4.2 and Table 4 shows that one annotator (Fem1) had pairwise Cohen's kappa values of only 0.14–0.18 against all other annotators, with near-zero agreement on the ambiguous class (0.015–0.053). Because each sentence is labeled by exactly three annotators, Fem1's unreliable vote is decisive whenever the other two annotators disagree, and her documented over-use of the ambiguous label can create an ambiguous majority or force a three-way tie. The overall Krippendorff's alpha of 0.387 is low, and it rises to 0.716 only when Fem1 is excluded. The paper does not quantify how many final labels are determined by Fem1's vote, nor does it report the distribution of label changes if Fem1's annotations were removed from the aggregation. This is a load-bearing issue for the ground-truth labels used in all subsequent experiments, and it is acknowledged in the Limitations section as 'potential instability in ground truth labels' without being mitigated. I request an analysis of the stability of the final labels under alternative aggregation rules (e.g., excluding Fem1, or using soft labels), and if the instability is substantial, the experiments should be re-run on a cleaner label set.
  2. [Section 6, Tables 5 and 6] The central empirical claim—that specialized Polish language models achieve superior performance compared to multilingual alternatives—is supported by macro-F1 scores computed against the majority-vote labels just described. If those labels are not a stable ground truth, the scores in Tables 5 and 6, and the ranking drawn from them, may be measuring noise rather than detection ability. This concern is especially acute for the ambiguous class, which has the lowest annotator agreement (Table 4) and which is exactly the class whose addition causes the sharp performance drop from the Basic to the Core configuration. The paper argues that the degradation reflects the difficulty of finer-grained distinctions, but it could equally be an artifact of an ill-defined label. To support the central claim, the authors should report at least one sensitivity analysis: for example, macro-F1 on the subset of sentences with full annotator agreement, or re-run the main comparisons using labels aggregated without Fem1. Without such evidence, the superiority claim is not yet established.
  3. [Section 9 and Section 3] The released dataset (Release 1.0) contains only 3,704 erotic and ambiguous sentences, which is 15% of the full dataset; neutral sentences and the rare violence/unacceptable classes are excluded. However, the experiments in Tables 5 and 6 use the full 24,768-sentence dataset, including the rare classes and the neutral majority. As a result, the benchmark results cannot be reproduced or independently verified from the public release. This is a significant limitation for a resource paper whose main contribution is a dataset: the authors should either release the full label distribution (even without the text, to address copyright and ethical concerns) or provide a clear protocol for reconstructing the full dataset, and should state explicitly that the public release is only a subset and that the reported results pertain to the full, non-public dataset.
minor comments (6)
  1. [Appendix D, Table 11] The row for PLLuM-Mistral-12B (SFT) in the 1-shot condition contains only three numeric entries instead of four, which appears to be a formatting error: the table should be checked and corrected.
  2. [Section 7, Figure 1] The term 'Type I error percentage' is defined as FP/(FP+FN), which is not a conventional Type I error rate but rather the false positive proportion among all errors; this definition should be stated more prominently in the text and the figure caption should be adjusted accordingly.
  3. [Appendix C, Figure 2] The prompt template does not include the label definitions, and the Limitations section correctly notes this omission; however, the main text describing the LLM evaluation (Section 5.3 and 5.4) should mention this design choice and its potential impact on the reported LLM scores, rather than relegating it only to the Limitations section.
  4. [Table 6] There are inconsistent decimal formats in the table, such as '0.58' alongside '0.580' and '0.42' alongside '0.420'; these should be unified for readability.
  5. [Section 6] The sentence 'for datasets labeled as Extended and Core, which are marked by pronounced class imbalance' contains a typo ('asExtended') and is also somewhat imprecise, since Core has a modest class imbalance compared to Extended and Full; consider rewording.
  6. [References] Several references use nonstandard author formatting, such as 'cjadams' and 'inversion' in the Jigsaw corpus entry; these should be converted to the journal's citation style.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the dataset construction and model benchmarks are empirical measurements, not derivations from their own outputs.

full rationale

The paper's central claims—that forePLay is a new manually annotated Polish dataset and that specialized Polish models outperform multilingual alternatives on it—are supported by direct measurement rather than by a derivation loop. The final labels are produced by majority vote over annotator judgments, with a superannotator resolving three-way ties; these labels are then used as a fixed benchmark, and model scores are computed against that benchmark. No model prediction is fed back into the label definitions, and no parameter is fitted to a subset of data and then reported as a prediction of the same quantity. The annotation-quality discussion is transparent about instability (Krippendorff's alpha of 0.387 overall, 0.716 after excluding Fem1, and the acknowledged 'potential instability in ground truth labels' in the Limitations section), but this is a reliability limitation, not a circularity. Similarly, the absence of any explicit description of a train/test split in the experimental section is a serious correctness and reproducibility concern, since the reported macro-F1 scores could conceivably reflect training performance, but the rules require exhibiting a specific reduction or fitted-parameter-renamed-as-prediction step, and no such step is identifiable from the paper's text. There is no load-bearing self-citation chain, no imported uniqueness theorem, and no ansatz smuggled in via citation; the comparisons to HerBERT, Polish RoBERTa, PLLuM, Bielik, GPT-4o, Mixtral, and Llama are all external model evaluations. The paper therefore warrants a circularity score of 0, with the noted experimental-design gaps treated as correctness risks rather than circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper's central claims rest on annotation and sampling assumptions rather than fitted constants. There are no free parameters in the reported benchmark, and no new postulated entities. The main epistemic costs are the ground-truth assumption under documented disagreement and the sampling assumption that repository tags mark erotic content.

assumptions (4)
  • domain assumption Majority vote with superannotator tie-breaking yields valid ground-truth labels.
    Section 4: final labels by majority vote and superannotator for 830 ties; Section 4.2 reports alpha 0.387 overall and 0.716 only after removing outlier Fem1.
  • domain assumption Sentence-level labels are meaningful without surrounding discourse.
    Appendix B: 'Each sentence was presented in isolation, without additional context'; the ambiguous category absorbs context dependence but does not eliminate it.
  • domain assumption Online repository erotic tags reliably identify erotic texts for sampling.
    Section 3.1: 'To filter erotic content, we relied on tags and category labels related to erotica provided by the online repositories.'
  • domain assumption Five exclusive labels with precedence rules cover the space of erotic discourse.
    Section 4.1 defines the labels and gives u priority over v; the Limitations section concedes the binary treatment of violence and social unacceptability may oversimplify severity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Behind Closed Words: Creating and Investigating the forePLay Annotated Dataset for Polish Erotic Discourse." pith.science (2026). https://pith.science/paper/JIJFJTMD

@misc{pith2026241217533,
  author       = {Pith},
  title        = {Pith review of: Behind Closed Words: Creating and Investigating the forePLay Annotated Dataset for Polish Erotic Discourse},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JIJFJTMD}},
  note         = {Machine review of arXiv:2412.17533}
}
read the original abstract

The surge in online content has created an urgent demand for robust detection systems, especially in non-English contexts where current tools demonstrate significant limitations. We present forePLay, a novel Polish language dataset for erotic content detection, featuring over 24k annotated sentences with a multidimensional taxonomy encompassing ambiguity, violence, and social unacceptability dimensions. Our comprehensive evaluation demonstrates that specialized Polish language models achieve superior performance compared to multilingual alternatives, with transformer-based architectures showing particular strength in handling imbalanced categories. The dataset and accompanying analysis establish essential frameworks for developing linguistically-aware content moderation systems, while highlighting critical considerations for extending such capabilities to morphologically complex languages.

Figures

Figures reproduced from arXiv: 2412.17533 by the authors.

Figure 1
Figure 1. Type I errors percentage across datasets and [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗
Figure 2
Figure 2. Classification prompt template used for evalu [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

33 extracted references · 13 canonical work pages

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Sara Achour. 2016. http://sigtbd.csail.mit.edu/pubs/2016/paper3.pdf A data driven analysis framework and erotica writing assistant . In SIGTBD

  4. [4]

    Nouar AlDahoul, Myles Joshua Toledo Tan, Harishwar Reddy Kasireddy, and Yasir Zaki. 2024. https://arxiv.org/abs/2411.17123 Advancing content moderation: Evaluating large language models for detecting sensitive content across text, images, and videos . Preprint, arXiv:2411.17123

  5. [5]

    Abhishek Anand, Negar Mokhberian, Prathyusha Naresh Kumar, Anweasha Saha, Zihao He, Ashwin Rao, Fred Morstatter, and Kristina Lerman. 2024. Don't blame the data, blame the model: Understanding noise and bias when learning from subjective annotations. arXiv preprint arXiv:2403.04085

  6. [6]

    Gonzalo Molpeceres Barrientos, Rocío Alaiz-Rodríguez, Víctor González-Castro, and Andrew C. Parnell. 2020. https://doi.org/10.2991/ijcis.d.200519.003 Machine learning techniques for the detection of inappropriate erotic content in text . International Journal of Computational Intelligence Systems, 13(1):591--603

  7. [7]

    Steven Bird and Edward Loper. 2004. https://aclanthology.org/P04-3031 NLTK : The natural language toolkit . In Proceedings of the ACL Interactive Poster and Demonstration Sessions , pages 214--217, Barcelona, Spain. Association for Computational Linguistics

  8. [8]

    cjadams, Daniel Borkan, inversion, Jeffrey Sorensen, Lucas Dixon, Lucy Vasserman, and nithum. 2019. Jigsaw unintended bias in toxicity classification. https://kaggle.com/competitions/jigsaw-unintended-bias-in-toxicity-classification. Kaggle

Show all 33 references
  1. [9]

    Thibault Clerice. 2024. https://aclanthology.org/2024.lrec-main.427 Detecting sexual content at the sentence level in first millennium L atin texts . In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC...

  2. [10]

    CohereForAI. 2024. https://huggingface.co/CohereForAI/c4ai-command-r-v01 Command r +

  3. [11]

    Sławomir Dadas. 2023. https://huggingface.co/sdadas/polish-roberta-base-v2 Polish roberta

  4. [12]

    Marie-Catherine De Marneffe, Christopher D Manning, and Christopher Potts. 2012. Did it happen? the pragmatic complexity of veridicality assessment. Computational linguistics, 38(2):301--333

  5. [13]

    Abhishek Kadian

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, and et al. Abhishek Kadian. 2024. https://arxiv.org/abs/2407.21783 The llama 3 herd of models . Preprint, arXiv:2407.21783

  6. [14]

    Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462

  7. [15]

    Weiming Hu, Ou Wu, Zhouyao Chen, Zhouyu Fu, and Steve Maybank. 2007. https://doi.org/10.1109/TPAMI.2007.1133 Recognition of pornographic web pages by classifying texts and images . IEEE Transactions on Pattern Analysis and Machine Intelligence, 29(6):1019--1034

  8. [16]

    Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276

  9. [17]

    Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa. 2023. https://arxiv.org/abs/2312.06674 Llama guard: Llm-based input-output safeguard for human-ai conversations ...

  10. [18]

    Jiaming Ji, Mickel Liu, Juntao Dai, Xuehai Pan, Chi Zhang, Ce Bian, Chi Zhang, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. 2023. Beavertails: Towards improved safety alignment of llm via a human-preference dataset. arXiv preprint arXiv:2307.04657

  11. [19]

    Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, L'elio Renard Lavaud, Lucile Saulnier, Marie-Ann...

  12. [20]

    Sean MacAvaney, Hao-Ren Yao, Eugene Yang, Katina Russell, Nazli Goharian, and Ophir Frieder. 2019. Hate speech detection: Challenges and solutions. PloS one, 14(8):e0221152

  13. [21]

    Crawford, Sanjana Gautam, Sorelle A

    Yaaseen Mahomed, Charlie M. Crawford, Sanjana Gautam, Sorelle A. Friedler, and Dana\" e Metaxa. 2024. https://doi.org/10.1145/3630106.3658932 Auditing gpt's content moderation guardrails: Can chatgpt write your favorite tv show? In Proceedings of the 2024 ACM Conference on Fai...

  14. [22]

    Todor Markov, Chong Zhang, Sandhini Agarwal, Tyna Eloundou, Teddy Lee, Steven Adler, Angela Jiang, and Lilian Weng. 2023. https://arxiv.org/abs/2208.03274 A holistic approach to undesired content detection in the real world . Preprint, arXiv:2208.03274

  15. [23]

    Robert Mroczkowski, Piotr Rybak, Alina Wróblewska, and Ireneusz Gawlik. 2021. https://www.aclweb.org/anthology/2021.bsnlp-1.1 H er BERT : Efficiently pretrained transformer-based language model for P olish . In Proceedings of the 8th Workshop on Balto-Slavic Natural Language P...

  16. [24]

    Krzysztof Ociepa, Łukasz Flis, Remigiusz Kinas, Adrian Gwoździej, and Krzysztof Wróbel. 2024. Bielik: A family of large language models for the polish language - development, insights, and evaluation

  17. [25]

    Ellie Pavlick and Tom Kwiatkowski. 2019. Inherent disagreements in human textual inferences. Transactions of the Association for Computational Linguistics, 7:677--694

  18. [26]

    John Pavlopoulos, Jeffrey Sorensen, Lucas Dixon, Nithum Thain, and Ion Androutsopoulos. 2020. Toxicity detection: Does context really matter? arXiv preprint arXiv:2006.00998

  19. [27]

    Barbara Plank. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.731 The `` problem '' of human label variation: On ground truth in data, modeling and evaluation . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10671--10682, Ab...

  20. [28]

    Huachuan Qiu, Shuai Zhang, Hongliang He, Anqi Li, and Zhenzhong Lan. 2024. https://doi.org/10.48550/arXiv.2403.13250 Facilitating pornographic text detection for open-domain dialogue systems via knowledge distillation of large language models . CoRR, abs/2403.13250

  21. [29]

    Piotr Rybak, Robert Mroczkowski, Janusz Tracz, and Ireneusz Gawlik. 2020. Klej: Comprehensive benchmark for polish language understanding. arXiv preprint arXiv:2005.00630

  22. [30]

    Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530

  23. [31]

    Alexandra Uma, Tommaso Fornaciari, Anca Dumitrache, Tristan Miller, Jon Chamberlain, Barbara Plank, Edwin Simpson, and Massimo Poesio. 2021. https://doi.org/10.18653/v1/2021.semeval-1.41 S em E val-2021 task 12: Learning with disagreements . In Proceedings of the 15th Internat...

  24. [32]

    Jialin Wu, Jiangyi Deng, Shengyuan Pang, Yanjiao Chen, Jiayang Xu, Xinfeng Li, and Wenyuan Xu. 2024. https://doi.org/10.1145/3658644.3690322 Legilimens: Practical and unified content moderation for large language model services . In Proceedings of the 2024 on ACM SIGSAC Confer...

  25. [33]

    Ou Wu and Weiming Hu. 2005. https://doi.org/10.1109/NLPKE.2005.1598819 Web sensitive text filtering by combining semantics and statistics . In 2005 International Conference on Natural Language Processing and Knowledge Engineering, pages 663--667

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.