Pith. sign in

REVIEW 1 cited by

PET: An Annotated Dataset for Process Extraction from Natural Language Text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.04860 v2 pith:ATOCKFQM submitted 2022-03-09 cs.CL

classification cs.CL
keywords extractionprocessannotatedbusinessinformationtextapproachesdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Process extraction from text is an important task of process discovery, for which various approaches have been developed in recent years. However, in contrast to other information extraction tasks, there is a lack of gold-standard corpora of business process descriptions that are carefully annotated with all the entities and relationships of interest. Due to this, it is currently hard to compare the results obtained by extraction approaches in an objective manner, whereas the lack of annotated texts also prevents the application of data-driven information extraction methodologies, typical of the natural language processing field. Therefore, to bridge this gap, we present the PET dataset, a first corpus of business process descriptions annotated with activities, gateways, actors, and flow information. We present our new resource, including a variety of baselines to benchmark the difficulty and challenges of business process extraction from text. PET can be accessed via huggingface.co/datasets/patriziobellan/PET

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What is the Best Process Model Representation? A Comparative Analysis for Process Modeling with Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new dataset and head-to-head comparison of nine process model representations with LLMs finds Mermaid best for general use and BPMN text best for generation.

Pith tools