Pith. sign in

Paper Citation Record · LEDGER

Statistical Hypothesis Testing for Auditing Robustness in Language Models

As of 20 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2506.07947.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07947 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:28:11.181361Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved15
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29d4592b-fa06-4d7c-803b-04f41a7120f1 · outbound

This paper cites Re-evaluating Evaluation in Text Summarization.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Re-evaluating Evaluation in Text Summarization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.120511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.120511Z digest=sha256:cf280a8b49cc98fef44a79e4db7199be859acc1d125d57e835ac300174abfb7f

Observation 77c6bbc9-e814-4d60-a18c-b83fa3f9270d · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Explaining and Harnessing Adversarial Examples

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.142515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.142515Z digest=sha256:333dd7673d30c98ac5044d33ecb66d6399bbe78a4af49a5b694edf6d8d338b64

Observation 4bbcbe76-0111-4f28-b4b2-36f41321738c · outbound

This paper cites Reducing Gender Bias in Abusive Language Detection.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Reducing Gender Bias in Abusive Language Detection

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.159775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.159775Z digest=sha256:8ca53f4c87f2ca652fccfd6dca738298068f431f4836ad2a4eda697a8d0edff8

Observation 2cd5e4bc-436f-4365-9b47-c76d90088f13 · outbound

This paper cites Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.165754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.165754Z digest=sha256:6deed76d25eaf2e2ee5df8e66c758b5c996aa7de3d04f8665470a79c5f7a9c91

Observation 428fe31f-9dfc-4ecf-9332-74760eb580f7 · outbound

This paper cites RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models.

Statistical Hypothesis Testing for Auditing Robustness in Language Models RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.168723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.168723Z digest=sha256:269cf66f5d31fb547b844f7374cee39b6aaea3136a453754a8c3568ff45411eb

Observation 17e6e69e-9328-4e50-8104-cb8b67cf9315 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Statistical Hypothesis Testing for Auditing Robustness in Language Models ReAct: Synergizing Reasoning and Acting in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.171948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.171948Z digest=sha256:82a9e6f3a7331eb6bd5d3e1a0af513e9f85136b690f139158215d490394e5f23

Observation 1ae77d9e-a123-457e-97d1-4f6f0074a7b8 · outbound

This paper cites MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance.

Statistical Hypothesis Testing for Auditing Robustness in Language Models MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.177892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.177892Z digest=sha256:2646235e8b4b7bc1f7fb642942d91640ea3fa34adec8e753743676167dcdbaea

Observation 4d8d9b65-ca47-4d27-99ad-70cf5a3378af · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Statistical Hypothesis Testing for Auditing Robustness in Language Models BERTScore: Evaluating Text Generation with BERT

Reference 2003

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.174703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.174703Z digest=sha256:4f95e4230b8091801eceec5ddcf4c9b853d04eaa04ba329d8fe3aeabbc9c70ef

Observation b73b1547-da50-4877-909b-94f8ff8fb7fb · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Lost in the Middle: How Language Models Use Long Contexts

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.156193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.156193Z digest=sha256:9f154d2808cf01608ebf55ea9c9e0d9c271c8d2de42ec56ae338c85bec917517

Observation dc38017d-7887-44c6-9ea6-aedfd7cd2b39 · outbound

This paper cites Better Zero-Shot Reasoning with Role-Play Prompting.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Better Zero-Shot Reasoning with Role-Play Prompting

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.152894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.152894Z digest=sha256:1f73c5f2d901619de44d8ce29577e1f64c9fe7757bd4219abd2657a99bd7e69b

Observation 5d19118c-c010-4eab-a08a-2495128f71d3 · outbound

This paper cites The Effect of Sampling Temperature on Problem Solving in Large Language Models.

Statistical Hypothesis Testing for Auditing Robustness in Language Models The Effect of Sampling Temperature on Problem Solving in Large Language Models

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.162864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.162864Z digest=sha256:775ee2836a91e489c8ea6f51238b8d3f0012c9409ce4778d7ebc5fc1be2830b5

Observation 147bf186-03e4-45d1-ab88-b11ee143e6be · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Reasoning with Language Model is Planning with World Model

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.146164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.146164Z digest=sha256:7d402e6d8d1254abaa0fed1f6efb63ff473a354a64b176c0a2243a662cf2efa8

Observation e67ac1fe-e122-4f2a-bcdb-f818cd8f996f · outbound

This paper cites Accountability of AI Under the Law: The Role of Explanation.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Accountability of AI Under the Law: The Role of Explanation

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.135904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.135904Z digest=sha256:55dc0a9cf7c08dfe88a9b4fb3b139bee573611e730f0d9dca28a4f2f4ed79bf6

Observation 85c28076-f0a8-4bda-96a6-365fa331faa0 · outbound

This paper cites The Oscars of AI Theater: A Survey on Role-Playing with Language Models.

Statistical Hypothesis Testing for Auditing Robustness in Language Models The Oscars of AI Theater: A Survey on Role-Playing with Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.128582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.128582Z digest=sha256:bbefa4afe363217ad0999bb70e3bea6951fc6b0a038ceb51efb2fdc22d17c0a7

Observation d4e88715-6dc5-4cf4-8afd-4b96ce44151b · outbound

This paper cites Nuanced metrics for measuring unintended bias with real data for text classification.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Nuanced metrics for measuring unintended bias with real data for text classification

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:28:11.377978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:28:11.124720Z digest=sha256:be54bd66fcb907c4770bc02179e4b0cbb6472e3a7742d4d09ac733eb48f39172

Observation d28078df-5dbd-4f44-8ddd-96eab8434b9c · outbound

This paper cites H., and Beutel, A.

Statistical Hypothesis Testing for Auditing Robustness in Language Models H., and Beutel, A

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:28:11.359472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:28:11.139337Z digest=sha256:38cad388e8009b09a3ddb28891c1d3864250dafec403cc6309d4efe8199e66e7

Observation 0f2037af-455c-43a2-ba48-929b64b276c9 · outbound

This paper cites Inherent Trade-Offs in the Fair Determination of Risk Scores.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Inherent Trade-Offs in the Fair Determination of Risk Scores

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.149472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.149472Z digest=sha256:7b886adadbfe33c84f1251d0081a7b5d5c0e4365cdec6a29377ed6067577b045

Observation cc23a211-5d07-40e3-ab79-55eafc25767c · outbound

This paper cites Measuring and mitigating unintended bias in text classification.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Measuring and mitigating unintended bias in text classification

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:28:11.368946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:28:11.132613Z digest=sha256:0b1250370f4336482dbe1efa1dae33e38f1a8f397d3b834e033ceda40bbc708d

Observation 4664fcd8-01d4-434f-8881-e3dba60bafb1 · outbound

This paper cites Dollar amounts are shown in parentheses for costs exceeding 100¢ (1 USD).

Statistical Hypothesis Testing for Auditing Robustness in Language Models Dollar amounts are shown in parentheses for costs exceeding 100¢ (1 USD)

Reference 2025

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:28:11.349868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:28:11.181361Z digest=sha256:0727c81c9165b5c5f02a738a1fcda3a70d36df971ed602c4560308e334713705

Pith citing papers

No inbound Pith citation observations are available.