Pith. sign in

REVIEW 2 major objections 5 minor 54 references

The paper argues that classical hadith transmitter grading—complete chains, per-domain narrator grades, weakest-link aggregation, gated corroboration, and content criticism—can be operationalized as claim-level provenance for multi-agent AI

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-07-31 23:01 UTC pith:LVBUOE42

load-bearing objection A genuinely novel and unusually honest transfer of hadith chain-grading methodology to AI provenance; the core mechanisms are validated under synthetic faults, but the load-bearing rule for LLM repair steps is asserted, not tested. the 2 major comments →

arxiv 2607.24117 v1 pith:LVBUOE42 submitted 2026-07-27 cs.AI cs.MA

Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems

classification cs.AI cs.MA
keywords provenancemulti-agent systemsknowledge basestrustepistemologyisnadhadith scienceclaim verification
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper is trying to establish that the classical hadith-science answer to "should I trust transmitted knowledge?"—attach a complete transmission chain to every claim, grade every transmitter per domain, cap chain strength at the weakest verified link, allow independent corroboration to upgrade, and criticize content separately from the chain—can be transferred intact into an operational provenance framework for multi-agent AI knowledge pipelines. Why this matters: current provenance records what agents did, but it does not grade how reliable each transformer is, so users cannot tell whether a specific claim, having passed through a specific chain, deserves trust. In controlled experiments on 20,000 physics-textbook claims, the author reports that weakest-link grading quarantined every claim whose chain contained a rejected narrator, and that independent-chain corroboration fired in every evaluated case across three corpora. The paper is equally explicit about its boundaries: the grading loop recovered only three of four narrator grades and missed the most faulty one, and a matched-coverage comparison was inconclusive because the reference content critic capped serving coverage at 4.8%; the claim, as stated, is that the mechanisms behave as specified where evidence exists, not that end-to-end practical advantage has been shown.

Core claim

On the author's own terms, the discovery is that the classical isnad–rijal protocol—every claim carries its full transmission chain, each transmitter holds a per-domain grade, a chain's grade is its weakest verified link with a bounded repair rule for generative steps, independent chains can corroborate a claim, and content is criticized separately from the chain—can be implemented as a relational schema and decision matrix for multi-agent knowledge systems. In controlled experiments on 20,000 physics-textbook claims, weakest-link grading quarantined every claim whose chain contained a rejected narrator (4,057 claims, each quarantine traceable to the binding grade), and independent-chain cor

What carries the argument

The central object is the isnad—the complete, gap-free, ordered chain of narrators (sources, scrapers, models, humans) attached to each claim—coupled with the rijal registry, which stores each narrator's ordinal grade (reliable/acceptable/weak/rejected) per domain and updates it through a jarh–tadil state-machine loop. The carrying argument is the weakest-link rule refined by transformation type: destructive steps (extraction, chunking) strictly cap the chain, while generative steps can raise the floor only up to their own grade and only when corroboration supports the repair, and can always lower it. Corroboration (mutaba'at) upgrades a claim only through independent, disjoint chains that m

Load-bearing premise

The load-bearing premise is the bounded-repair rule set out in Section 4.1: a generative LLM narrator can repair corrupted upstream information only up to its own grade and only when corroboration supports the repair, and can always lower the chain grade—a rule the evaluation does not test, yet one that governs most real chains that pass through a generative step.

What would settle it

Run a controlled chain in which a known corruption (say, a digit swap or sign flip) is introduced before a high-grade generative narrator, with and without a corroborating chain; the bounded-repair rule predicts the served chain grade never rises above the generative narrator's own grade unless corroboration fires, and that a weak generative narrator can lower it. If a strong generative model passes the corruption while the framework keeps the chain grade high, or if a weak generative step raises the floor without corroboration, the weakest-link refinement is refuted.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A deployed system using weakest-link grading cannot unknowingly serve a claim whose chain includes a narrator graded rejected; each rejection carries the binding grade as its explanation, making trust decisions auditable.
  • Independent-chain corroboration, when it fires, allows a claim carried by multiple disjoint chains to rise from a weak tier without exceeding the sound cap, so legitimate knowledge is not permanently strangled by one weak link.
  • Because grading is per-domain, a narrator rare in one domain may never earn a grade; monitoring per-cell evidence counts becomes a first-class operational risk and not a neutral absence.
  • The unverifiable verdict in the decision matrix is what keeps the framework safe under a weak critic, but it is also the binding constraint on coverage: unless content criticism can render a verdict on real prose, h.asan-tier claims remain in review and coverage stays near the review budget.
  • The grade-recovery failure shows the jarh–tadil loop can miss the most unreliable narrator when calibration evidence is sparse, so calibration coverage per (narrator, domain) cell should be monitored.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the bounded-repair rule for generative narrators is the untested load-bearing premise for real LLM chains; the natural next experiment is to inject an upstream corruption before a high-grade generative step and check whether the grade moves exactly as the rule predicts, with and without corroboration.
  • Editorial inference: corroboration's independence check on narrator identity, model family, and upstream source cannot catch two chains that share a factual error propagated from a common training corpus; constructing such a pair would provide a decisive test of the independence idealization.
  • Editorial inference: if a semantic content critic replaced the word-overlap critic, the matched-coverage comparison would likely become completable, and the paper's own ordering of priorities suggests the risk–coverage advantage of the full framework remains unmeasured until that happens.
  • Editorial inference: the schema's lifecycle columns anticipate supersession but not uncertainty; extending the framework to interval-valued and probabilistic claims would connect it to scientific knowledge bases where most claims are provisional.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper proposes ISNAD, an operational provenance framework that transfers classical Islamic hadith-transmission methodology (isnad, rijal, jarh wa-ta'dil, mutaba'at, matn criticism) to multi-agent knowledge pipelines. Each claim carries a full transmission chain; narrators are graded per domain in a registry; a single chain is graded by its weakest link, with a transformation-type refinement for generative narrators; independent chains can corroborate subject to grade gates and caps; and chain quality is combined with content criticism in a serve/review/quarantine decision matrix. The evaluation injects deterministic faults into 20,000 physics-textbook claims and reports: weakest-link quarantine of all 4,057 claims containing a rejected narrator; partial recovery of narrator grades (3 of 4, missing the highest-fault narrator); and corroboration firing across three corpora with negative controls on the Wikipedia corpus. The paper explicitly reports the grade-recovery miss, the inconclusive matched-coverage comparison, and the uninformative synthetic confidence baseline.

Significance. If the framework's mechanisms hold, this is a useful contribution: it makes claim-level provenance operational by attaching graded transmitter reliability to chains, and provides a concrete relational schema and decision procedure rather than only a formal model. The paper's strengths are its unusually honest reporting, its open-source implementation with 157 passing tests and a leakage firewall against the injection manifest, and its explicit status table separating validated, partial, and inconclusive results. The evaluation is, however, a synthetic mechanism-behavior test rather than a demonstration of practical end-to-end advantage; the paper does not overclaim superiority. The principal correctness risk is the untested bounded-repair rule for generative narrators in §4.1, which is load-bearing for any real chain containing an LLM step.

major comments (2)
  1. [§4.1; §8; Table 8] The bounded generative-repair rule is load-bearing but neither derived nor tested. The rule states that a generative narrator can raise the chain floor only up to its own grade and only when corroboration supports the repair, and can always lower it. This is the aggregation semantic for any chain passing through an LLM, yet §8 exercises only strict-minimum/quarantine (the rejected generative narrator in the §8.2 trace) and corroboration of a weak chain by an independent chain. No condition combines a destructive upstream corruption with a higher-grade generative downstream narrator, with and without corroboration. If the rule is wrong, chain grades and the §4.4 decision matrix are miscalibrated for exactly the chains ISNAD targets. Please either add a targeted experiment — e.g., inject upstream corruption, vary the generative narrator's grade and corroboration status, and measure whether
  2. [§8.5; Table 6] The corroboration 'validation' is weaker than the abstract and Table 8 imply. In v2 and v3, candidate pairs were preselected by the same cosine thresholds that the corroboration check uses (≥0.75/≥0.80), so the 603/603 and 104/104 fire rates are close to a pipeline tautology; the only discriminating evidence is the 8/8 negative controls on v2, and no negative controls were run on v3. The paper honestly notes some of this, but Table 8 still lists corroboration as 'Validated' without the caveat. Please either add a discrimination measure (e.g., manual labels of matched/unmatched pairs and precision/recall) or restate the claim as 'mechanism fires under synthetic conditions,' with the absence of v3 negative controls carried in the abstract.
minor comments (5)
  1. [§8.2] The count 'all 50 (narrator, domain) cells' is unexplained. State how 50 is obtained (e.g., number of domains × seeds) so the reader can reproduce the claim.
  2. [Table 6] The v1 'Match density' cell is blank; fill it in or explicitly state why it is not applicable.
  3. [§8.4; Table 5] No uncertainty is reported for the 4.8% coverage ceiling or for the baseline error rates across the 10 random seeds. Add variance or confidence intervals where the seeds permit.
  4. [§8.6] The reference content critic is named as if it used embeddings, but the implementation is word-overlap plus a negation heuristic. Rename or clarify in the text to avoid misleading readers.
  5. [§8.5; Table 6] The note that v2 candidate pairs were subsampled from 662 to 603 appears only in the table caption. Move this explanation into the main text, since it affects the interpretation of the fire rate.

Circularity Check

2 steps flagged

No significant circularity: the framework is self-contained; the two mechanism 'validations' are partly conformance checks by construction, but the paper labels them as such and the central contribution does not reduce to them.

specific steps
  1. self definitional [§8.2, with decision matrix in §4.4]
    "The weakest-link mechanism performed as specified. Every claim whose chain contained a narrator graded rejected was quarantined—4,057 claims, 29% of the evaluation split—and every quarantine was traceable to the specific narrator grade that caused it."

    This 'validation' is the decision rule itself: §4.4 maps a mawdu'-tier chain to 'Reject and log; quarantine the narrator.' Observing that rejected-narrator chains are quarantined executes the rule rather than testing an independent consequence. The paper's own phrasing ('performed as specified') concedes this is a conformance check. It does not feed back into the framework's derivation, so it is a minor, non-load-bearing definitional element rather than a circular argument.

  2. self definitional [§8.5, Table 6]
    "Corroboration fired 68/136 603/603 104/104"

    Corroboration 'firing' is defined by the framework's own criteria — 'same normalized claim, disjoint narrator sets, independent sources' (§4.3) — applied through the system's semantic matcher and grade gate. A positive fire rate is therefore the rule's output, not an independent discovery. This is mitigated by the v2 negative controls (8/8 no-upgrade) and manual match inspection, and the paper explicitly frames v2 as validating mechanics rather than the independence assumption. The v3 104/104 without negative controls is weaker, but still a mechanism check, not a fitted prediction.

full rationale

The paper's derivation chain is a transfer mapping from hadith methodology to multi-agent systems, not a formal derivation with fitted parameters later relabeled as predictions. The central mechanisms are not fit to outcomes and then re-reported as discoveries. The §8.2 quarantine result and §8.5 corroboration firing are conformance checks of rules already defined in §4.3–§4.4; the paper mostly says so explicitly, calling the weakest-link result 'performed as specified' and the v2 corroboration result a validation of 'mechanics' rather than of the independence assumption. The confidence baseline is openly 'uninformative by construction,' and the unmatched-coverage comparison is reported as inconclusive, not as a win. The unvalidated §4.1 bounded generative-repair rule is a real validation gap for LLM-containing chains, but that is a correctness/scope concern, not circularity: it is asserted, not derived from the conclusion. There is no load-bearing self-citation, no imported uniqueness theorem, and no ansatz smuggled via citation from the authors' own prior work. Overall, the paper's empirical claims are honest and mostly external; the only definitional elements are the mechanism-level conformance checks noted above, which do not undermine the framework's independent contribution.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 0 invented entities

The framework introduces no new physical or metaphysical entities; its components (narrator registry, chain engine, decision matrix) are software abstractions implemented in the reference schema. The free parameters are deliberately open implementation constants; the axioms are domain assumptions about LLM behavior and audit independence that the paper itself flags as idealizations or leaves untested.

free parameters (3)
  • downgrade threshold (reference transition policy) = not fitted; swept 3, 6, 10, 15, 25 in §8.6
    The jarh–ta'dil transition arithmetic is deliberately left open by the framework; the evaluation shows no setting delivers both coverage and grade recovery.
  • corroboration minimum-grade gate and upgrade cap = unspecified
    Left to implementations (§4.3); these determine how much independent corroboration can upgrade a weak chain, and the paper does not calibrate them.
  • designed narrator fault rates = 1%, 2%, 15%, 18%
    Injection-simulation parameters chosen by the authors in §8.1; the grade-recovery result depends on these rates and on per-cell sample sizes.
axioms (5)
  • domain assumption A generative transformer can repair upstream corruption only up to its own grade and only with corroboration
    Introduced in §4.1 without derivation or direct empirical test; load-bearing for weakest-link aggregation on LLM steps.
  • domain assumption Post-hoc audit evidence can label served claims as erroneous independently of the injection manifest
    §8.2 claims grade recovery 'without access to the injection manifest,' but the audit-evidence generation is not specified; if audits use manifest-derived labels, recovery is supervised rather than independent.
  • domain assumption Disjoint narrator identity, model family, and upstream source imply independence for corroboration
    §4.3 and §7 explicitly flag this as an idealization; naive set-disjointness can over-credit correlated chains.
  • domain assumption Claims in scope have determinate truth values
    §7: probabilistic and provisional claims are excluded; contradiction detection between intervals is undefined.
  • ad hoc to paper A simulated perfect reviewer resolves reviewed claims correctly
    §8.1 assumes away human reviewer error, so review-queue precision is not a realistic cost metric.

pith-pipeline@v1.3.0-alltime-deepseek · 19756 in / 15839 out tokens · 139293 ms · 2026-07-31T23:01:06.070621+00:00 · methodology

0 comments
read the original abstract

Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations rather than direct retrieval. Existing provenance work records what happened - execution traces, tool calls, evidence links - and source-reliability estimation is long established (truth discovery, reputation systems). What is missing is an operational framework that attaches graded, per-domain transmitter reliability to claim-level transmission chains, with completeness semantics, transformation-typed aggregation, decoupled content criticism, and serve/review/quarantine routing. Classical Islamic hadith science confronted a structurally similar problem: deciding whether knowledge transmitted through chains of human narrators should be accepted. Over centuries it developed a rigorous methodology - isnad (a complete transmission chain attached to every claim), rijal (systematic grading of each narrator's integrity and precision), weakest-link chain evaluation, corroboration through independent chains, and matn criticism (content evaluated independently of chain quality). This paper transfers that methodology to AI system design. We contribute a formal mapping from hadith-science concepts to multi-agent pipelines, a relational schema implementing claim chains and a graded narrator registry, a decision matrix combining chain grade with content criticism, and an evaluation on 20,000 claims from real physics textbooks. The evaluation validates weakest-link quarantine and independent-chain corroboration; reports a partial failure of the grade-recovery loop, which missed the highest-fault narrator; and reports two analyses as inconclusive, including a matched-coverage comparison the framework could not reach with the reference content critic. The paper is explicit throughout about which claims the evidence does and does not yet support.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

54 extracted references · 4 linked inside Pith

  1. [1]

    Brown, J. A. C. (2017).Hadith: Muhammad’s Legacy in the Medieval and Modern World(2nd ed.). Oneworld

  2. [2]

    Buneman, P., Khanna, S., & Tan, W. C. (2001). Why and where: A characterization of data provenance.Proceedings of ICDT 2001, 316–330

  3. [3]

    Cemri, M., et al. (2025). Why do multi-agent LLM systems fail? (MAST)

  4. [4]

    Chen, Z., Xiang, Z., Xiao, C., Song, D., & Li, B. (2024). AgentPoison: Red-teaming LLM agents via poisoning memory or knowledge bases.NeurIPS 2024

  5. [5]

    Chhikara, P., Khant, D., Aryan, S., Singh, T., & Yadav, D. (2025). Mem0: Building production-ready AI agents with scalable long-term memory

  6. [6]

    Condorcet, N. C. (1785).Essai sur l’application de l’analyse à la probabilité des décisions rendues à la pluralité des voix. Imprimerie Royale

  7. [7]

    lightandmatter.com

    Crowell, B.Light and Matter(series). lightandmatter.com. CC BY-SA

  8. [8]

    Cui, Y., & Widom, J. (2000). Lineage tracing for general data warehouse transformations. Proceedings of VLDB 2000, 71–82

  9. [9]

    P., & Skene, A

    Dawid, A. P., & Skene, A. M. (1979). Maximum likelihood estimation of observer error-rates using the EM algorithm.Applied Statistics, 28(1), 20–28

  10. [10]

    Dempster, A. P. (1967). Upper and lower probabilities induced by a multivalued mapping. Annals of Mathematical Statistics, 38(2), 325–339

  11. [11]

    al-Dhahab¯ ı, Shams al-D¯ ın (14th c.).M¯ ız¯ an al-I‘tid¯ al f¯ ı Naqd al-Rij¯ al

  12. [12]

    L., Gabrilovich, E., Heitz, G., Horn, W., Lao, N., Murphy, K., et al

    Dong, X. L., Gabrilovich, E., Heitz, G., Horn, W., Lao, N., Murphy, K., et al. (2014). Knowl- edge vault: A web-scale approach to probabilistic knowledge fusion.Proceedings of KDD 2014, 601–610

  13. [13]

    L., Gabrilovich, E., Murphy, K., Dang, V., Horn, W., Lugaresi, C., et al

    Dong, X. L., Gabrilovich, E., Murphy, K., Dang, V., Horn, W., Lugaresi, C., et al. (2015). Knowledge-based trust: Estimating the trustworthiness of web sources.Proceedings of the VLDB Endowment, 8(9), 938–949. 23

  14. [14]

    Ebrahimi, S., Dehghankar, M., & Asudeh, A. (2025). An adversary-resistant multi-agent LLM system via credibility scoring.Proceedings of IJCNLP-AACL 2025, 1676–1693

  15. [15]

    arXiv:2606.04990

    From agent traces to trust: Evidence tracing and execution provenance in LLM agents (2026). arXiv:2606.04990

  16. [16]

    Gil, Y., & Artz, D. (2007). Towards content trust of web resources.Journal of Web Semantics, 5(4), 227–239

  17. [17]

    Hartig, O., & Zhao, J. (2009). Using web data provenance for quality assessment.Proceedings of the Workshop on Semantic Web for Provenance Management

  18. [18]

    Mathematical hadith verification with information theory: The HadithRank algorithm

    Hawramani, I. Mathematical hadith verification with information theory: The HadithRank algorithm. https://hawramani.com/mathematical-hadith-verification-with-information- theory-the-hadithrank-algorithm/

  19. [19]

    Hou, Y., et al. (2024). WikiContradict: A benchmark for evaluating LLMs on real-world knowledge conflicts from Wikipedia

  20. [20]

    Of Miracles

    Hume, D. (1748).An Enquiry Concerning Human Understanding, Section X: “Of Miracles.”

  21. [21]

    D., Jennings, N

    Huynh, T. D., Jennings, N. R., & Shadbolt, N. R. (2006). An integrated trust and reputation model for open multi-agent systems.Autonomous Agents and Multi-Agent Systems, 13(2), 119–154

  22. [22]

    al-Shahraz¯ ur¯ ı (13th c.).Muqaddimah (‘Ul¯ um al-H

    Ibn al-S.al¯ ah. al-Shahraz¯ ur¯ ı (13th c.).Muqaddimah (‘Ul¯ um al-H. ad¯ ıth)

  23. [23]

    Jøsang, A. (2001). A logic for uncertain probabilities.International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems, 9(3), 279–311

  24. [24]

    Jøsang, A., & Ismail, R. (2002). The beta reputation system.Proceedings of the 15th Bled Electronic Commerce Conference

  25. [25]

    Juynboll, G. H. A. (1983).Muslim Tradition: Studies in Chronology, Provenance and Author- ship of Early Hadith. Cambridge University Press

  26. [26]

    D., Schlosser, M

    Kamvar, S. D., Schlosser, M. T., & Garcia-Molina, H. (2003). The EigenTrust algorithm for reputation management in P2P networks.Proceedings of WWW 2003, 640–651

  27. [27]

    Karpathy, A. (2026). LLM Wiki. Public gist. https://gist.github.com/karpathy/442a6bf5559 14893e9891c11519de94f (accessed 2026-07-08)

  28. [28]

    Lamport, L., Shostak, R., & Pease, M. (1982). The Byzantine generals problem.ACM TOPLAS, 4(3), 382–401

  29. [29]

    arXiv:2409.13740

    Language agents achieve superhuman synthesis of scientific knowledge (Pa- perQA2/ContraCrow) (2024). arXiv:2409.13740

  30. [30]

    Li, X., et al. (2026). RAPS: Reputation-aware publish-subscribe for LLM multi-agent systems. arXiv:2602.08009

  31. [31]

    Li, X., et al. (2026). TrustTrade: Human-inspired selective consensus for LLM agents. arXiv:2603.22567

  32. [32]

    J., Moebs, W., & Sanny, J.University Physics, Vols

    Ling, S. J., Moebs, W., & Sanny, J.University Physics, Vols. 1–3. OpenStax, Rice University. CC BY 4.0

  33. [33]

    Mghari, M., Bouras, O., & El Hibaoui, A. (2022). Sanadset 650K: Data on hadith narrators. Data in Brief, 44, 108540

  34. [34]

    Min, S., et al. (2023). FActScore: Fine-grained atomic evaluation of factual precision in long- form text generation.EMNLP 2023

  35. [35]

    Moreau, L., Missier, P., et al. (2013). The PROV data model. W3C Recommendation

  36. [36]

    Mosa, M. A. (2025). Synergizing structure and semantics: A knowledge graph-transformer framework for narrator disambiguation in hadith networks.Digital Scholarship in the Human- ities, 40(4), 1085–1100

  37. [37]

    Motzki, H. (2005). The mus.annaf of ‘Abd al-Razz¯ aq al-S.an‘¯ an¯ ı as a source of authentic ah.¯ ad¯ ıth of the first century A.H.Arabica, 52(2), 159–212. 24

  38. [38]

    AI-powered hadith verification: Toward a new model of authenticity in Islamic knowledge transmission (2025).International Journal of Noesantara Islamic Studies, 2(5)

  39. [39]

    Digital takhr¯ ıj hadith as Islamic digital humanities (2025).Digital Muslim Review, 3(1)

  40. [40]

    Pasternack, J., & Roth, D. (2010). Knowing what to believe (when you already know some- thing).Proceedings of COLING 2010, 877–885

  41. [41]

    Pasternack, J., & Roth, D. (2013). Latent credibility analysis.Proceedings of WWW 2013, 1009–1020

  42. [42]

    Prakash, S. (2026). The provenance paradox in multi-agent LLM routing: Delegation contracts and attested identity in LDP. arXiv:2603.18043

  43. [43]

    Ramzy, A., Torki, M., Abdeen, M., Saif, O., ElNainay, M., Alshanqiti, A., & Nabil, E. (2023). Hadiths classification using a novel author-based hadith classification dataset (ABCD).Big Data and Cognitive Computing, 7(3), 141

  44. [44]

    Sabater, J., & Sierra, C. (2001). ReGreT: A reputation model for gregarious societies.Pro- ceedings of AGENTS 2001, 194–195

  45. [45]

    (1950).The Origins of Muhammadan Jurisprudence

    Schacht, J. (1950).The Origins of Muhammadan Jurisprudence. Oxford University Press

  46. [46]

    Schuster, T., Fisch, A., & Barzilay, R. (2021). Get your vitamin C! Robust fact verification with contrastive evidence.NAACL 2021

  47. [47]

    Sequeira, R., Damianakis, S., Iqbal, U., & Psounis, K. (2026). Agent-Sentry: Bounding LLM agents via execution provenance. arXiv:2603.22868

  48. [48]

    (1976).A Mathematical Theory of Evidence

    Shafer, G. (1976).A Mathematical Theory of Evidence. Princeton University Press

  49. [49]

    Souza, R., Gueroudji, A., DeWitt, S., Rosendo, D., Ghosal, T., Ross, R., Balaprakash, P., & Ferreira da Silva, R. (2025). PROV-AGENT: Unified provenance for tracking AI agent interactions in agentic workflows.IEEE e-Science 2025. arXiv:2508.02866

  50. [50]

    Su, H., et al. (2024). ConflictBank: A benchmark for evaluating the influence of knowledge conflicts in LLMs

  51. [51]

    Thorne, J., Vlachos, A., Christodoulopoulos, C., & Mittal, A. (2018). FEVER: A large-scale dataset for fact extraction and verification.NAACL 2018

  52. [52]

    Federal Rules of Evidence, Rule 805: Hearsay within hearsay

    U.S. Federal Rules of Evidence, Rule 805: Hearsay within hearsay

  53. [53]

    Wu, K., et al. (2024). ClashEval: Quantifying the tug-of-war between an LLM’s internal prior and external evidence

  54. [54]

    Yin, X., Han, J., & Yu, P. S. (2008). Truth discovery with multiple conflicting information providers on the web.IEEE TKDE, 20(6), 796–808. 25