{"id":"459e07bf-4a5c-427a-ada2-63549bde0bfd","arxiv_id":"2412.11063","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"An agentic workflow using reusable legal tools beats a raw GPT-3.5 baseline on contract retrieval, but the evaluation is compromised because the ground truth was generated with the same tools.","lead":"LAW is an AI system that pairs a large language model with specialized tools to answer questions about mutual fund custody contracts, such as finding termination dates or comparing fee clauses. The authors report large accuracy gains over a simple chatbot, but the evaluation design is circular because the 'correct answers' are generated by the same tools the system uses.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ground-truth labels are generated by LAW's own tools, so the reported 92.9-point lead measures self-consistency, not legal accuracy.","rationale":"The reader's weakest_assumption precisely identifies the load-bearing flaw: the evaluation is circular because ground truth is generated by the same tools LAW uses. I agree with this diagnosis. The paper's Section 7 explicitly states that the hand-coded scripts 'leverage the same tools and text agents that the proposed system has access to.' This means LAW's high scores measure how well its code-generation agent re-invokes those tools, not whether the extracted or computed answers are legally correct. The baseline, by contrast, is a vanilla LLM without tool access, so it is handicapped by design. The absence of independent ground truth makes the 92.9-point termination-date gap uninterpretable as evidence of legal expertise. While the system architecture may be practically useful, the paper's central claim requires valid, independent labels. The concrete test I propose would settle the concern: if LAW still outperforms the baseline and the label scripts match human experts on an independently annotated sample, the claim holds. Otherwise, the reported results are artifacts. Therefore the reader's REJECT verdict remains appropriate; no adjustment is needed.","tokens_in":183,"tokens_out":2434,"duration_ms":30927,"concrete_test":"Select a random sample of 100 contracts (or 100 queries) from the 720-query dataset, covering all task types but emphasizing termination dates. Have two legal experts independently annotate the correct termination dates (and a subset of other fields) using only manual reading of the contracts. Then compare three systems against these human labels: (1) LAW, (2) the gpt-3.5-turbo baseline, and (3) the hand-coded ground-truth scripts. If LAW's accuracy against human labels is not dramatically higher than the baseline's, or if the hand-coded scripts disagree with the human experts on more than, say, 5% of cases, the reported 92.9-point gap is an artifact of label circularity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that LAW significantly outperforms the baseline, with termination-date retrieval at 95.4% vs 2.5% (Table 1). This comparison is valid only if the ground-truth labels are correct and independent of LAW's outputs. Section 7 states the opposite: 'The ground truth answers are generated using hand-coded scripts that leverage the same tools and text agents that the proposed system has access to.' Consequently, LAW is evaluated against labels produced by the very components it orchestrates. If a tool inaccurately extracts or computes a termination date, that error appears identically in both the ground truth and LAW's output, inflating agreement. The baseline, lacking these tools, is scored against tool-generated labels it cannot replicate, so the 92.9-point gap may reflect tool functionality rather than LAW's legal reasoning. The paper provides no independent human validation; Appendix D.1 mentions only a 'small pilot study' without reporting agreement rates or error analysis. Thus the headline result is not evidence of real legal understanding unless the ground-truth scripts themselves are shown to be accurate against independent annotations.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces LAW, a modular agentic workflow that combines a code-generation LLM orchestrator with domain-specific tools (date and party extraction, master-contract retrieval, section-title labeling, lifecycle calculation) and text agents (summarization, comparison), applied to custody and fund services contracts sourced from SEC EDGAR. The authors report experiments on retrieval and analytical queries, claiming that LAW significantly outperforms a gpt-3.5-turbo baseline, with the strongest result being a 92.9-percentage-point advantage for termination-date retrieval (95.4% vs 2.5%, Table 1). The central claim is that orchestration of reusable, guardrailed tools yields accurate and cost-effective automated analysis of long legal contracts.","tokens_in":13111,"tokens_out":5584,"duration_ms":49516,"significance":"If the empirical claims were valid, LAW would be a practical contribution: it demonstrates an agentic orchestration approach to a real domain with a large corpus (17,831 contracts covering 23 years of filings), reusable tools, and a cost argument against fine-tuned legal LLMs. The system design is coherent and the engineering infrastructure is described in reasonable detail. However, the evaluation as presented does not support the headline claims. The ground truth is generated by the same tools and text agents LAW uses, making the reported accuracy measures reflect self-consistency rather than legal correctness. The baseline is weak, not a full retrieval system, and is not tested on all tasks. The absence of independent human validation and statistical testing further limits the conclusions. The paper is therefore best read as a system description with an unsupported comparative evaluation, rather than as a validated empirical study.","major_comments":[{"comment":"The ground truth answers are generated using hand-coded scripts that leverage the same tools and text agents that the proposed system has access to, as explicitly stated in Section 7. This makes the evaluation circular: Table 1 compares LAW's outputs against labels produced by LAW's own components. The headline result, a 92.9-point advantage for termination dates (95.4% vs 2.5%), therefore measures whether the code-generation agent selects the correct tool, not whether the extracted termination dates are legally correct. The paper even states that the procedure \"focuses exclusively on LAW's ability to generate code that correctly orchestrates the tools,\" which contradicts the abstract's claim that LAW \"significantly outperforms the baseline\" in legal tasks. To support the claim, the authors need independent human-verified labels, or a validation study showing that the hand-coded scripts agree with expert annotations; the small pilot study in Appendix D.1, which reports no agreement rates, does not suffice.","section":"Section 7, Dataset Curation"},{"comment":"The baseline is not a comparable system. For retrieval queries, the baseline is given a small set of four relevant and four distractor contracts and asked True/False or extraction questions; it does not perform corpus retrieval, unlike LAW. For \"Compare clause X\" no baseline number is reported, so the general claim of outperformance is not tested for that task. The evaluation also lacks significance tests or confidence intervals: with the stated 20 retrieval queries per combination, reported percentages such as 94.4% and 71.8% are not multiples of 5%, which suggests either a different number of queries or a reporting error. Please report the exact per-cell query counts, variance, and statistical tests before making \"significantly outperforms\" claims.","section":"Section 7, Baseline setup; Table 1"},{"comment":"The human validation is limited to a \"small pilot study\" with no details on the number of participants, contracts, queries, or agreement metrics, so it cannot validate the tool-generated ground truth. Additionally, the section-title classification model used for retrieval has only 46% accuracy (Appendix F, Table 3), yet the end-to-end retrieval hit rates in Table 1 are near or above 90%; the paper should analyze whether retrieval is actually driven by the noisy titles or by other signals, and how title errors propagate to the reported results.","section":"Appendix D.1 and Appendix F"}],"minor_comments":[{"comment":"The text contains a typo: \"force majure\" should be \"force majeure.\"","section":"Section 4.2"},{"comment":"Please define \"Retrieval Hit Rate\" precisely (e.g., exact string match, partial match, or manual inspection); the caption's \"hit rate/recall\" is ambiguous.","section":"Table 1"},{"comment":"The claim that Form 485BPOS filings \"account for a non-trivial (14%) of EDGAR filings\" is ambiguous; please clarify whether this is 14% of all EDGAR filings or 14% of a specific subset.","section":"Section 3"},{"comment":"The statement that tools are \"rigorously guardrailed through unit tests that map their failure modes\" is not supported by any unit-test documentation or failure-mode analysis in the paper; please include at least a summary or a reference.","section":"Section 4"}],"recommendation":"reject","confidential_remarks":"The circular evaluation is the central problem. The paper's empirical contribution is not supported because the ground truth is generated by the same tools and agents that LAW orchestrates, and the baseline is not comparable. I would encourage the authors to re-frame the work as an engineering/system demonstration or to obtain independent human annotations on a sample and re-run the evaluation; if they do, a major revision could be considered. I would also ask the editor to verify the novelty of the \"first legal agentic workflow\" claim relative to the authors' own FlowMind and HiddenTables work, and to check whether the cost-effectiveness claim has any accompanying measurement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is a real system that does something useful: it orchestrates domain tools (date extraction, party matching, section classification) over 17,831 custody contracts from EDGAR. Second, the headline 95.4% vs 2.5% on termination dates does not mean what the abstract implies. The ground truth is generated by hand-coded scripts that use the same tools LAW calls, so the evaluation isolates code-generation/orchestration skill, not legal extraction accuracy. The paper says this plainly in Section 7.\n\nWhat is genuinely new: the dataset is a real contribution, and applying the FlowMind/CodeAct pattern to this domain is a reasonable engineering adaptation. The modular architecture—tools, text agents, caching, retrieval—is sensible and well described. Credit where due: the authors are transparent about their method, and the system design is coherent.\n\nThe soft spots are real and load-bearing. Because the ground truth shares the tools, any tool error appears in both the gold label and LAW's output. The baseline is a raw GPT-3.5 without those tools, so of course it collapses on a task like termination-date calculation that is really a tool call. The 92.9-point gap mostly measures tool capability, not orchestration. There is no independent human validation beyond a small pilot with no reported numbers, no significance tests, and no error analysis of the ground-truth scripts. The T5 section-title accuracy of 46% is weak—the authors acknowledge it, but it underscores that tool accuracy is unvalidated.\n\nThe correct framing is that LAW is a system description, and its evaluation tests whether the code-generation agent selects the right tool calls. That is a legitimate question, but the paper does not answer it with independent labels. The reader's rejection is proportionate to that flaw.\n\nWho this is for: practitioners building legal-document QA pipelines will find the architecture instructive; researchers in evaluation methodology will find a good case study in circularity. It deserves a serious referee, but only if the referee demands major revision: require a human-annotated ground truth (or a clear separation of tool validation from orchestration evaluation), add a baseline that has access to the same tools or an ablated LAW, and report significance tests. With those changes, this could be a solid system contribution.","headline":"LAW is a genuinely engineered system for custody contracts, but its headline claims are inflated by a circular evaluation that scores orchestration against ground truth produced by the same tools.","tokens_in":13658,"tokens_out":2587,"would_cite":false,"duration_ms":24178,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LAW, a legal agentic workflow that orchestrates reusable domain-specific tools and text agents, outperforms direct LLM prompting on custody-contract queries, with a 95.4% hit rate on termination-date retrieval versus 2.5% for the baseline.","keywords":["legal agentic workflow","LLM orchestration","contract analysis","custody services","retrieval-augmented generation","multi-hop reasoning","code generation agent","40Act funds"],"falsifier":"Have independent legal experts, blind to LAW's design, manually label a random sample of the evaluation queries (for example, the actual termination date for 100 contracts); if LAW's accuracy against these human labels falls substantially below its accuracy against the script-generated ground truth, the headline gains are artifacts of evaluation circularity rather than genuine legal competence.","tokens_in":12723,"feed_emoji":"⚖️","tokens_out":6030,"duration_ms":48303,"temperature":0.7,"pith_summary":"This paper introduces LAW, a system that answers questions about custodial and fund-services contracts by having a code-generating language model orchestrate reusable, domain-specific tools and text agents. The authors claim this design significantly outperforms an off-the-shelf LLM prompted directly, with the largest gap on multi-hop tasks such as computing a contract's termination date, where LAW achieves a 95.4% retrieval hit rate against 2.5% for the baseline. The motivation is practical: fine-tuned legal LLMs are expensive, static, and need labeled data, whereas tool-based orchestration is cheaper, updatable, and works with open or closed LLMs. If correct, LAW would automate contract retrieval and comparison work currently done by hand across 23 years of SEC filings.","feed_headline":"Legal agentic AI beats off-the-shelf LLM by 92.9 points","feed_subtitle":"Reusable tools plus a code-writing agent crack multi-hop contract reasoning without fine-tuning.","key_machinery":"The load-bearing mechanism is the code generation agent with its three-tier validation loop: the LLM is prompted, with tool names and examples, to write Python code that calls LAW's APIs; the code is checked for syntax and security, scanned to ensure only existing tools with valid signatures are invoked (hallucination detection), and executed with error feedback routed back to the agent for correction. Around this agent sit the tools: RoBERTa-based span detection plus regex and HTML parsing for dates; fuzzy matching for parties; a date-and-party matcher that links amendments to master contracts; a fine-tuned T5-large classifier that labels paragraphs into 20 section types so clauses can be retrieved by BM25 from an OpenSearch index; and parallel sub-agent summarization and comparison for text that exceeds the context window. This combination lets the system decompose a query like 'find the termination date' into retrieve-effective-date, read-duration, add, and return steps.","core_discovery":"The paper's central claim is that an agentic workflow—replacing direct prompting with a code-generation agent that writes Python calls to a suite of legal-specific tools (date extraction, party matching, master-contract lookup, section labeling) and text agents (summarize, compare)—can handle both simple retrieval and multi-hop analytical queries over thousands of long legal contracts. On a labeled dataset of 720 queries drawn from 17,831 custody contracts, LAW reports near-perfect hit rates for finding master agreements and parties, 95.4% for termination dates, and BERTScore F1 of 89.5 for clause summarization, each well above the gpt-3.5-turbo baseline. The authors also claim the system is cost-effective and extensible because tools are reusable and can be added without retraining the underlying model.","pith_inferences":["The evaluation may overstate real-world accuracy because the ground-truth scripts share components with LAW; a human-labeled test set is the natural next check to separate orchestration skill from tool accuracy.","The simulated noisy-RAG baseline gives gpt-3.5-turbo only four relevant contracts per query, so the reported gap may shrink against a stronger RAG pipeline with better chunking or reranking, especially on the 'explore all contracts' task.","If the toolset were released with its unit-test failure modes documented, other financial institutions could adopt the same legal primitives, converting this single system into a shared library of contract-analysis tools.","The cost argument predicts that maintenance cost stays roughly flat as new contract types are added, whereas fine-tuned LLMs require fresh labels and retraining; this is a testable claim for a follow-up study."],"forward_implications":["A custody bank could answer retrieval and comparison questions across the full 485BPOS corpus, covering contracts that govern trillions in retail and retirement assets, without manual document review.","Because the code agent composes tools at query time, new analytical questions can be fielded without retraining the model, as long as existing tools cover the needed operations.","The system runs on either open-source or closed LLMs, making it a lower-cost alternative to fine-tuned legal LLMs for institutions that already maintain contract databases.","The same orchestration pattern extends beyond custody contracts to other legal document families, such as non-English contracts or similar regulatory filings, by adding new tools."],"supporting_citations":[{"why":"Provides the FlowMind orchestration recipe and prompt pattern that LAW builds on for tool-calling code generation.","marker":"Zeng et al. (2023)"},{"why":"CodeAct supplies the design choice of executable Python code as the agent's unified action space.","marker":"Wang et al. (2024)"},{"why":"RoBERTa is the span-detection model used inside the date extraction tool.","marker":"Liu et al. (2019)"},{"why":"T5 is the architecture fine-tuned for section title classification and generation in the section-labeling tool.","marker":"Raffel et al. (2020)"},{"why":"BERTScore is the F1 metric used to evaluate the summarization and comparison agents against ground truth.","marker":"Zhang* et al. (2020)"},{"why":"HiddenTables is the earlier agentic system for tabular data that LAW extends to legal contracts.","marker":"Watson et al. (2023)"},{"why":"RAG defines the retrieval-augmented generation paradigm that the baseline approximates and LAW's retrieval tools build on.","marker":"Lewis et al. (2020)"}],"fun_headline_variants":["Agentic legal AI beats LLM by 92.9 points on termination dates","LAW: multi-agent workflow outscores gpt-3.5-turbo by 92.9 pts","Code-writing agent plus legal tools crack multi-hop contract reasoning","Reusable legal tools give agentic AI 92.9-point edge over baseline","Legal contracts: agentic system beats direct LLM prompting by 92.9"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ground truth answers for the 720-question evaluation are generated by hand-coded scripts that use the same tools and text agents LAW itself calls, so the reported scores measure how faithfully LAW reproduces its own components' outputs; if those components are inaccurate, the high numbers reflect self-consistency, not legal understanding.","fun_headline_variants_meta":{"raw":{"variants":["Agentic legal AI beats LLM by 92.9 points on termination dates","LAW: multi-agent workflow outscores gpt-3.5-turbo by 92.9 pts","Code-writing agent plus legal tools crack multi-hop contract reasoning","Reusable legal tools give agentic AI 92.9-point edge over baseline","Legal contracts: agentic system beats direct LLM prompting by 92.9"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000799,"raw_usage":{"total_tokens":3474,"prompt_tokens":868,"completion_tokens":2606,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":484,"completion_tokens_details":{"reasoning_tokens":2498}},"tokens_in":484,"tokens_out":2606,"duration_ms":15220,"temperature":1.0,"reasoning_tokens":2498,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:19:26.545582+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have independent legal experts, blind to LAW's design, manually label a random sample of the evaluation queries (for example, the actual termination date for 100 contracts); if LAW's accuracy against these human labels falls substantially below its accuracy against the script-generated ground truth, the headline gains are artifacts of evaluation circularity rather than genuine legal competence.","supporting_citations":[],"review_version":1}