Pith. sign in

Paper Citation Record · LEDGER

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing

As of 11 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 2 inbound Pith citation observations for arXiv:2601.18061.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.18061 v3

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T11:46:09.108129Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:16:26.175093Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T04:18:56.704885Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact1
  • verified fuzzy57
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 89705895-d623-41a1-9534-e24224803783 · outbound

This paper cites Clinician-Rated Severity of Nonsuicidal Self-Injury.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Clinician-Rated Severity of Nonsuicidal Self-Injury

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.788739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:37fd23451042f1c468e2e5fcd894f26f7bbe4cfa88bffaf96ae35b251b00c226

Observation 0adde6b6-9e29-4ee8-8ce8-dee16e65a5ba · outbound

This paper cites DSM-5 Clinician-Rated Dimensions of Psychosis Symptom Severity.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing DSM-5 Clinician-Rated Dimensions of Psychosis Symptom Severity

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.802673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:742cc894ab83a2e5c3cf95d2c803ffba917af857b38d620536dd84aa6f10c18c

Observation 1e4bec4c-d8b2-4dc8-90c3-a88efc156446 · outbound

This paper cites DICES Dataset: Diversity in Conversational AI Evaluation for Safety.Advances in Neural Information Processing Systems, 36:53330–53342.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing DICES Dataset: Diversity in Conversational AI Evaluation for Safety.Advances in Neural Information Processing Systems, 36:53330–53342

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.797026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:980374a943c62a46c7a0a7f0bf258fdad3347133a64e3bd365be8d3caaec3f3a

Observation 05f8dcd9-93a8-4a4b-b3a1-d00b890fe36c · outbound

This paper cites Truth Is a Lie: Crowd Truth and the Seven Myths of Human Annotation.AI Magazine, 36(1):15–24.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Truth Is a Lie: Crowd Truth and the Seven Myths of Human Annotation.AI Magazine, 36(1):15–24

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.799885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:2c2f3cb5919dffb56a04fec5b99bd186e21894006acf0853a9d73b673300576c

Observation e9922f6e-93e3-40d3-b196-7528be7f4d82 · outbound

This paper cites Crowd Truth: Harnessing Disagreement in Crowdsourcing a Relation Extraction Gold Standard.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Crowd Truth: Harnessing Disagreement in Crowdsourcing a Relation Extraction Gold Standard

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.777957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:59787f117c3862e46e33e0e85312b6ebf0e9a996cae79fe2f51049e05bace3d0

Observation 4e6e9d3c-fcdb-419a-8271-a99241af5c35 · outbound

This paper cites Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.775559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:1d623352c437c10f379c3c39d91f08f0b90d8dc59e2e5f4f1f44aecfc710a411

Observation f2e61751-f922-45ee-9b08-979a55795f70 · outbound

This paper cites Bech.Clinical Psychometrics.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Bech.Clinical Psychometrics

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.794378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:87d56b676251bde8b83ff3c37242c48e6c598124ef420972153160415dc2ab60

Observation 0d376012-b765-4928-8354-102339370c21 · outbound

This paper cites an unresolved cited work.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:47:49.791417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:349981f41205c7a6b4839cfde3a9de1da8b52890a49775031a8dc29adabbcf29

Observation 55997058-543a-4706-a361-f2c18a69573d · outbound

This paper cites Consensus report of the apa work group on neuroimaging markers of psychiatric disorders.Am Psychiatr Assoc.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Consensus report of the apa work group on neuroimaging markers of psychiatric disorders.Am Psychiatr Assoc

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.771018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:3502f3ef95bf104eb27f871cc4f9b894a60433fcb3a6effc9f5336d4e5245f21

Observation 940f40b0-d336-49c5-9f1a-d098a446facc · outbound

This paper cites Using Thematic Analysis in Psychology.Qualitative Research in Psychology, 3(2):77–101.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Using Thematic Analysis in Psychology.Qualitative Research in Psychology, 3(2):77–101

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.785997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:9e1f2e5526e4e1bed48a4b61ae2b9d1e34393857ff1e46e79c9e0dbd42bfb18b

Observation e77b00b4-e418-4355-9ed3-960f0ca0f5cd · outbound

This paper cites Minton, Abigail Lott, and Jinho D.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Minton, Abigail Lott, and Jinho D

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.780735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:186e9af5536b6fcd22771f202eaf9a5d948c424c17bbe0241c375bb93bafca0d

Observation 5c35b276-3251-4a8b-8e82-3ed483f4f883 · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.692465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:b635d6f728d86dc473d987bd18cbd23f61c45740fa06c5f47023c3dbd435e380

Observation d5270066-a82f-4590-87f7-3912b967b006 · outbound

This paper cites How people use chatgpt.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing How people use chatgpt

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.641532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:dec7c5c9c3537126031e602e035b997260d36913e739b3e9a3585e48a0e19418

Observation e7ba2d63-d3bf-49e4-ad8a-1ce5804a2c25 · outbound

This paper cites Predicting Depression via Social Media.International AAAI Conference on Web and Social Media, 7(1):128–137.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Predicting Depression via Social Media.International AAAI Conference on Web and Social Media, 7(1):128–137

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.653085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:1d7577094ce86e62a9692f969b4d04ae49d89e9a52ee67f8c7f7f1a94c24052c

Observation 28ab72d8-fa4f-4ac1-9997-2d9a277ac1fd · outbound

This paper cites Deep Reinforcement Learning from Human Preferences.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Deep Reinforcement Learning from Human Preferences

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.650228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:e813f3ec41d23c991d09a9a9417ad03278a9688d7952d9c5b652504dde58cf8c

Observation 4e60b9a8-85f4-4021-bc97-fbed1d805f31 · outbound

This paper cites Cicchetti.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Cicchetti

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.661356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:05ed25660f38a1a554ac0c9faad975b09c5768862201faa228b8aae6feb30b91

Observation 5f765e66-41fa-4f04-b2a7-08f9b4020713 · outbound

This paper cites Hashimoto.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Hashimoto

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.658470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:553bfeae1364e79f49ed57e434497e5480f7f2e280f58a43e29ba2191e6edccf

Observation 793adfb1-7bd6-409d-8b4c-9fad441839d2 · outbound

This paper cites Diagnostic and statistical manual of mental disorders.Am Psychiatric Assoc, 21(21):591–643.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Diagnostic and statistical manual of mental disorders.Am Psychiatric Assoc, 21(21):591–643

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.735587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:71dd3524e7aa127bba172dcb202aa46453fbd2a2adfd689f36c1efaea2cba46a

Observation 792b627f-be36-4f30-ae29-b1e0dac0468b · outbound

This paper cites an unresolved cited work.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Unresolved cited work

Reference 19

Resolution
verified exact
doi, observed 2026-05-16T11:47:49.282324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:ed73d6fba843f1afc26a247b6e00b15cd19402c9f6f4908596a3b671443230e2

Observation 212b9a9d-18d0-4aca-8802-7d3bbb48e4a0 · outbound

This paper cites an unresolved cited work.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:47:49.624465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:b6fe6452c349eaa1a35bdb1c69cceff8eb0a3deca4cb45f2c367762d2005df22

Observation 598aa3da-0aeb-4af6-9d2c-1e2b4133470f · outbound

This paper cites Can AI relate: Testing large language model response for mental health support.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Can AI relate: Testing large language model response for mental health support

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.633249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:ef7307f5278dfa911a5cff261221d5beb301c0ef67ade2797adab3d31bceb111

Observation b3d75980-8685-4f1f-b133-7e00f5295efc · outbound

This paper cites Impact of preference noise on the alignment performance of generative language models.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Impact of preference noise on the alignment performance of generative language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.704854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:c93823567be753150dc05ed6d696e6e7a6cec521612f0149f7e8aa483a024b57

Observation 942edd00-ba09-4dff-8d65-bcc7949576ee · outbound

This paper cites Blind spots and biases: Exploring the role of annotator cognitive biases in NLP.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Blind spots and biases: Exploring the role of annotator cognitive biases in NLP

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.619087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:f9da9e5479c13aa4940d316132e66df2acaae366f96f0f5e327d157b2383300d

Observation 26db7dd9-5be6-4967-b51f-cec76322a12d · outbound

This paper cites Goodman, Lawrence H.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Goodman, Lawrence H

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.639057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:d65f5261152289613269367d3fadcea5e34406a965feed0bd0353d5264079f1f

Observation 027577f7-8577-4fcc-9c3a-e8e95fbf536f · outbound

This paper cites Gordon, Michelle S.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Gordon, Michelle S

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.723443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:83fbf20fa121d13cad61263e6d846de8480fa53e225ee9ce630f53879d97b61b

Observation 38e14316-48b9-4210-b0e8-dcd5e5fc5685 · outbound

This paper cites Risks from language models for automated mental healthcare: Ethics and structure for implementation.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Risks from language models for automated mental healthcare: Ethics and structure for implementation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.768121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:69a374305435d5f26b2c4db7a2a35df199cce64fdabacb7a1c62b3b7c789edef

Observation f73d4089-4f83-4d33-ab31-a800c9d1bae7 · outbound

This paper cites Human Feedback is not Gold Standard.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Human Feedback is not Gold Standard

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.689018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:ac5264b12b24674b54b6f2fdec0c7190f78afcb82838bd80a728fcad9b62abd6

Observation 98fac7e5-2c51-4c5d-a5ba-31b4f65ba2b9 · outbound

This paper cites How LLM counselors violate ethical standards in mental health practice: A practitioner-informed framework.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing How LLM counselors violate ethical standards in mental health practice: A practitioner-informed framework

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.783275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:dd824472e85b84c2856a12fd664cdb6cf36e72099293d66ae07276d6da300bda

Observation 33cf71bc-79df-4819-ada5-1bedf766dccc · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.670029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:d5f8b6563e8a5cace3ce71fbb8d80df48718235924760b9a37042d2b90315502

Observation 428a34c9-6494-48b3-bc1d-ece9f3437232 · outbound

This paper cites an unresolved cited work.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:47:49.701487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:53260aa40b742fa0717704335ede23acdd5fdbe9e59eafe9fd175a2ea491e098

Observation f48788c3-4db0-498f-addc-942c10481929 · outbound

This paper cites an unresolved cited work.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:47:49.683843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:985abc7b81c44bd0012612c8e888c16d154a26940e5f3b5b5b54c0d9d65ca0d2

Observation 65cf5188-ab62-47c4-a72b-a9555075c5e4 · outbound

This paper cites Reliability in Content Analysis: Some Common Misconceptions and Recommendations.Human Communication Research, 30(3):411–433.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Reliability in Content Analysis: Some Common Misconceptions and Recommendations.Human Communication Research, 30(3):411–433

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.655706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:6ecaa24439a35a0e0d03f90425d24de687ca23a8a7af2209ba768aaa8a4e9a7a

Observation 4f915837-2483-460e-bffe-798b41f4a387 · outbound

This paper cites Kunstman, Aaron Lulla, Monika Drummond Roots, Manu Sharma, Aryan Shrivastava, Nina Vasan, and Colleen Waickman.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Kunstman, Aaron Lulla, Monika Drummond Roots, Manu Sharma, Aryan Shrivastava, Nina Vasan, and Colleen Waickman

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.698419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:f386b10e976c1dec2eea38ba1106949327c22040e5c8e6f9b5bb8390d5a4cfba

Observation 257c5123-393b-483e-abc0-32495fe4c450 · outbound

This paper cites an unresolved cited work.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:47:49.636158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:4db431432a4bde698d6b2bafcac23e4b07d5b7f769e745d05ef87efb2e32e142

Observation 8d28453c-fef4-4410-a4be-afee8f1fde6f · outbound

This paper cites Hashimoto.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Hashimoto

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.664401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:7d8c6d32846b397c703354f2c51188f775ca0cfce0960eedb11cfe9b562465c7

Observation 4731ff42-453d-448c-b60a-aea94481133b · outbound

This paper cites Bunyi, Adam C.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Bunyi, Adam C

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.627248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:519de5e8052f48e21b4cbb74916996b241f5b6ff7eca04a516fedfbd41f1b897

Observation c6897c9c-874c-422a-8468-6affd4168880 · outbound

This paper cites Sample Size Considerations for Fine-Tuning Large Language Models for Named Entity Recognition Tasks: Methodological Study.Journal of Medical Internet Research AI, 3:e52095.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Sample Size Considerations for Fine-Tuning Large Language Models for Named Entity Recognition Tasks: Methodological Study.Journal of Medical Internet Research AI, 3:e52095

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.681405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:6fd67fd90474516cacf6443ccf9c1ae76ac47c38465f88161c67ed443893565e

Observation bfae3754-10ce-452d-9296-72d1725222a7 · outbound

This paper cites A diagnostic meta-analysis of the patient health questionnaire-9 (phq-9) algorithm scoring method as a screen for depression.General hospital psychiatry, 37(1):67–75.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing A diagnostic meta-analysis of the patient health questionnaire-9 (phq-9) algorithm scoring method as a screen for depression.General hospital psychiatry, 37(1):67–75

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.616428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:b912a91fa9809e4b0f5e28296f2312015e58b59f150ecdee9c07005050a21be5

Observation fb3125eb-1f42-49fc-884e-8dea0af2abc7 · outbound

This paper cites McGraw and S.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing McGraw and S

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.678285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:2084b6886219dcd9f3826f5565acaf5b42d276b6c980dfc64ef2362d4ab05004

Observation fc5bb067-0f13-4b44-9a31-6fe731312cf2 · outbound

This paper cites Ong, and Nick Haber.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Ong, and Nick Haber

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.686452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:32f4ee3b2a4a25720c2bcb8f4e8abe8b6b69440f408b8b8cfe7fb14a3429d8a7

Observation 4d351919-2dd3-4888-99f0-6bac4c34a7b2 · outbound

This paper cites Moyers, Lauren N.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Moyers, Lauren N

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.647305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:35849aabc37f6343405d0a1c612b3936981f77979aaa06304399a9fadff0151c

Observation 59346976-8a19-4392-a242-d53521eaf689 · outbound

This paper cites Department of Veterans Affairs.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Department of Veterans Affairs

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.739563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:a36e4f6747866a3ad794e98a219ffe73af944ed59544cbd29ff6ff8dcd92f49b

Observation 7b07b7ee-fa44-4bde-be7d-1099d2d3f527 · outbound

This paper cites NICHQ Vanderbilt Assessment Scales.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing NICHQ Vanderbilt Assessment Scales

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.720304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:bf16043d94001ca1f78919671b6fdf908ec39b2402b239568b75e087a0cc3dec

Observation 98ca2986-71ec-40ee-8239-f70dbe2cf385 · outbound

This paper cites Depression.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Depression

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.695200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:0967d41aa6275547e6164952ba39f890e2904a50c43a754bcba3ee50a96a88e1

Observation fff6dc70-4901-4d49-8b6a-d778979de163 · outbound

This paper cites Enhancing mental health with artificial intelligence: Current trends and future prospects.Journal of medicine, surgery, and public health, 3:100099.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Enhancing mental health with artificial intelligence: Current trends and future prospects.Journal of medicine, surgery, and public health, 3:100099

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.732269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:d6785aff2eed38dd8b982caa7d1c5a6d6d4a4be6877b8f681786bcb24f6fa829

Observation 00113e04-ee5f-4e6f-ab52-b7563cfa53ba · outbound

This paper cites Christiano, Jan Leike, and Ryan Lowe.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Christiano, Jan Leike, and Ryan Lowe

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.675515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:31ddaab203619cede5ff60a6c87b86ce4fa74f223991b6fc569e828fd4aea645

Observation d4b3fcbe-96fb-417e-9b84-6d5f98ee1736 · outbound

This paper cites Inherent Disagreements in Human Textual Inferences.Transactions of the Association for Computational Linguistics, 7:677–694.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Inherent Disagreements in Human Textual Inferences.Transactions of the Association for Computational Linguistics, 7:677–694

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.644545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:e9597051de70670ec1de7be1953b42ea9610d813bac284a8b95c0d514f6fa642

Observation 70a2be2c-e2e5-49c6-b176-fb9157ec16cd · outbound

This paper cites Red Teaming Language Models with Language Models.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Red Teaming Language Models with Language Models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.672570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:e02c4c47028d6937a5c0ab3f1354c491c30dcd3b2c401b0ba0759afac40d8a27

Observation 7ad7fdb4-d134-4f39-a202-dfbe0a93e440 · outbound

This paper cites Posner, D.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Posner, D

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.630361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:3fb404967490c1231f311ef0d2a2f27b689e7a61a5f3d5f376a4043d9426dce9

Observation ba9f5582-3a8e-4bf6-ba85-08a669a79415 · outbound

This paper cites Prochaska, Erin A.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Prochaska, Erin A

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.729386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:205873662ac79f072b278a8e1e457a89bd4e900e16aeeff0513fbe071f722a3c

Observation e65ef78c-0a0e-4f46-a5cf-4677acdf2c8a · outbound

This paper cites Manning, and Chelsea Finn.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Manning, and Chelsea Finn

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.621831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:890be3b6c2fc80bd40c4b37361588d927a954efd8392ce2aeb8fb0cd745b49f9

Observation b76de69a-9338-4b1e-b092-3f5d05798da4 · outbound

This paper cites Regier, William E.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Regier, William E

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.742505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:7b4820dab2d6c7a839b6c7ad2fb6e7812153906a882eb128ce7a6aabd9d6f813

Observation 04ef73ce-2388-4902-93be-9bdb52f0d312 · outbound

This paper cites Large language models as mental health resources: Patterns of use in the united states.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Large language models as mental health resources: Patterns of use in the united states

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.711024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:9ee68f633f105f5b1b9f5d5d922787bf175a8bc62d1b1eb7b11a6bc229ecd522

Observation 98307c84-6c72-4814-8ecf-597ffa8652cd · outbound

This paper cites an unresolved cited work.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:47:49.745311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:f94454d4d81fd01cd7fb47fb7451f55e8f02777f8bc1d1a0c08a95a19da9554f

Observation 50d4e86b-a636-44a0-a658-f91e3e379ab1 · outbound

This paper cites Lin, Adam S.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Lin, Adam S

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.762597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:63ee501fadc6817ea6259da5ad0e0e2c8c99b88d518bfa041533159110578b12

Observation db133a62-4faa-41a7-82f7-d4b4065e615f · outbound

This paper cites A Computational Approach to Understanding Empathy Expressed in Text-Based Mental Health Support.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing A Computational Approach to Understanding Empathy Expressed in Text-Based Mental Health Support

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.667265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:263a2d2c9310588539fe4690f1cd63463997ddf2237caaffc9ed25e097a7cf5b

Observation 752050e8-e4d8-4045-8ed3-b4cbd6a81428 · outbound

This paper cites an unresolved cited work.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:47:49.707700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:dce9657505ab4c6fa3c4a2b3754eef6ff2a5b59eb283bc4a13948fd451b3426b

Observation 54393a32-3d9c-4bc0-aca0-aa974f33a79c · outbound

This paper cites Clinical Practice Guidelines on using artificial intelligence and gadgets for mental health and well-being.Indian Journal of Psychiatry, 66(Suppl 2):S414–S419.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Clinical Practice Guidelines on using artificial intelligence and gadgets for mental health and well-being.Indian Journal of Psychiatry, 66(Suppl 2):S414–S419

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.726521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:d708f34060b2bf406a9849e7e3beb39de770d4b4e54b682d5fc2ca2d8cdefeda

Observation 6ab45c2b-2947-46ac-9713-ec7a1e4f1c37 · outbound

This paper cites Pfohl, Heather Cole-Lewis, Darlene Neal, Qazi Mamunur Rashid, Mike Schaekermann, Amy Wang, Dev Dash, Jonathan H.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Pfohl, Heather Cole-Lewis, Darlene Neal, Qazi Mamunur Rashid, Mike Schaekermann, Amy Wang, Dev Dash, Jonathan H

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.757209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:013b50b8bcee970949a66490200433777a857fbf51c6284744cf5d64d0c1a5c8

Observation f99fe2f6-d891-48eb-958a-008a368e67a1 · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.717393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:d42723a5836ccb13d91b682275e0074534201e65201ad0043ae398964c13508c

Observation b4f3a450-f6f0-4bea-ae9a-071ef5658955 · outbound

This paper cites A Practical Guide to Fine-Tuning Language Models with Limited Data.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing A Practical Guide to Fine-Tuning Language Models with Limited Data

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.714235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:46f90b7e6b3ce3fc49b075400e512f4234392786f50aa5a0acfc598d847bafd8

Observation d00a5e5b-2d99-4f8a-987b-5806e4dab1c6 · outbound

This paper cites Lukoff, Keith Nuechterlein, R.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Lukoff, Keith Nuechterlein, R

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.754158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:b8f7eb3f26173344ef90de251ed2d5060a42cb536779135031ae0c071eb66d66

Observation 0d9e7534-5b1d-422a-a024-16efb4896549 · outbound

This paper cites Wang, Patricia Berglund, Mark Olfson, Harold A.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Wang, Patricia Berglund, Mark Olfson, Harold A

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.759924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:8090ccf05003b97f7df170d878224fd1bddb890320c8e3722ca88d8d5a73377b

Observation 750f9970-3b04-4c32-a787-ab37a9b73f95 · outbound

This paper cites an unresolved cited work.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:47:49.765353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:1350ba4ca7d925aa93bcec5af9d857f7a30e60c94f2e54848c357d163907d630

Observation e29dff42-8874-474f-8a6e-480b1e3c836b · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Xing, Hao Zhang, Joseph E

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.748254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:b885b5c1adfc3d00feee1a34518019d106b843312d2f3afc4017a80390f4a2c8

Observation 6c6a7069-3c34-4aa4-80cd-17125693c22f · outbound

This paper cites Cold plunges cure psychosis—stop your medication.

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing Cold plunges cure psychosis—stop your medication

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:47:49.751327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:46:09.108129Z digest=sha256:e841eeb08f4208395f097cac171eefb6bdc0cf1af51dfa2271eb0cbdda1c1806

Pith citing papers

Observation cdf42a1e-2a34-4555-8b2b-3949d90c5286 · inbound

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation cites this paper.

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:18:56.830854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T04:18:54.595022Z digest=sha256:94af35e09ae2322e0f224bdff4dd451f6b9a32a39fc82b0d511bf6e4009d85b3

Observation e1bc1de0-ecac-4a49-97c8-886f6aa76f8b · inbound

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation cites this paper.

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T04:16:26.175093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:16:26.175093Z digest=sha256:af22796330deffedaec6e59cff087b1c618733294fc56fd91342920ceddb64c0