Pith. sign in

REVIEW 4 cited by

RIRAG: Regulatory Information Retrieval and Answer Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.05677 v2 pith:4LLIPRWO submitted 2024-09-09 cs.CL cs.AIcs.CEcs.ETcs.IR

classification cs.CLcs.AIcs.CEcs.ETcs.IR
keywords regulatorydocumentsanswercompliancegenerationinformationobligationsorganizations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Regulatory documents, issued by governmental regulatory bodies, establish rules, guidelines, and standards that organizations must adhere to for legal compliance. These documents, characterized by their length, complexity and frequent updates, are challenging to interpret, requiring significant allocation of time and expertise on the part of organizations to ensure ongoing compliance. Regulatory Natural Language Processing (RegNLP) is a multidisciplinary field aimed at simplifying access to and interpretation of regulatory rules and obligations. We introduce a task of generating question-passages pairs, where questions are automatically created and paired with relevant regulatory passages, facilitating the development of regulatory question-answering systems. We create the ObliQA dataset, containing 27,869 questions derived from the collection of Abu Dhabi Global Markets (ADGM) financial regulation documents, design a baseline Regulatory Information Retrieval and Answer Generation (RIRAG) system and evaluate it with RePASs, a novel evaluation metric that tests whether generated answers accurately capture all relevant obligations while avoiding contradictions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AUEB-Archimedes at RIRAG-2025: Is obligation concatenation really all you need?

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Concatenating the exact 'obligation' sentences that RePASs extracts yields a near-perfect score (0.947), exposing the metric's vulnerability; a verify-and-refine system with the same oracle scores a more credible 0.639.

  2. MST-R: Multi-Stage Tuning for Retrieval Systems and Metric Evaluation

    cs.IR 2024-12 conditional novelty 5.0 of 10

    A multi-stage retrieval system (MST-R) improves Recall@10 from 0.78 to 0.87 on the ObliQA regulatory dataset, and a passage-concatenation baseline inflates the RePASs answer metric to 0.95.

  3. CTRAG: An In-Context Retrieval-based Framework for Automated Compliance Checking using LLMs

    cs.CL 2026-08 reject novelty 3.0 of 10

    A RAG pipeline with tuned chunking, retrieval depth, and in-context examples reports 78% F1 for automated compliance checking, but the evaluation has no held-out validation and a post-hoc No-Evidence-to-Non-Compliant ...

  4. 1-800-SHARED-TASKS at RegNLP: Lexical Reranking of Semantic Retrieval (LeSeR) for Regulatory Question Answering

    cs.CL 2024-12 conditional novelty 3.0 of 10

    LeSeR, a hybrid dense-plus-BM25 reranking pipeline, achieves competitive recall and mAP on the RegNLP regulatory QA retrieval task.

Pith tools