Pith. sign in

Paper Citation Record · LEDGER

Interpretable Risk Mitigation in LLM Agent Systems

As of 23 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 3 inbound Pith citation observations for arXiv:2505.10670.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10670 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:11:02.149478Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:51:56.450815Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T00:31:24.634056Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact2
  • verified fuzzy24
  • unresolved49
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c43eafac-f59d-4479-8e8e-f89c0279a822 · outbound

This paper cites Artificial intelligence and the future of work: Evidence from OECD countries.

Interpretable Risk Mitigation in LLM Agent Systems Artificial intelligence and the future of work: Evidence from OECD countries

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.643143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.782535Z digest=sha256:79993e8c21d8c5e9637f2f54e4a51bec726eb27fc4d15f3a28f316814ed9d414

Observation bf3d3a32-c723-4ab7-a21a-243b73e9489b · outbound

This paper cites Dai, Chelsea Finn, Justin Fu, Kanishka Gopalakrishnan, et al.

Interpretable Risk Mitigation in LLM Agent Systems Dai, Chelsea Finn, Justin Fu, Kanishka Gopalakrishnan, et al

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.627827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.787366Z digest=sha256:1fad4d05c22d3f8a2137a5c4aaf35c65ce807a07d8b10ad187002ad550da5d67

Observation 3801ed77-8037-4a96-b87d-1b592e49f25f · outbound

This paper cites Mistral 7b: Open foundation models, 2023.

Interpretable Risk Mitigation in LLM Agent Systems Mistral 7b: Open foundation models, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.612474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.792100Z digest=sha256:7e61a1c2686a9163be4f944f831ef10a2eaf1c0a4001555deeeb1e8f04360060

Observation 54ed5ba8-5bf9-4e48-a28b-7b79fb7685e8 · outbound

This paper cites Playing repeated games with Large Language Models.

Interpretable Risk Mitigation in LLM Agent Systems Playing repeated games with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.796811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.796811Z digest=sha256:d3afb41abea193e6c233643ca65ccf524f447a9bd6e8ee50dead64e1a143b91e

Observation 9b7eff9b-d167-419d-ae41-da6ca8c46fe7 · outbound

This paper cites Concrete Problems in AI Safety.

Interpretable Risk Mitigation in LLM Agent Systems Concrete Problems in AI Safety

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.802560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.802560Z digest=sha256:e3fa78928aa155f3c7d2ed2d961082ed1d210699d78eb27984746a8e23c33479

Observation 16702f8c-2cb7-4d44-a2c2-8387f4b0ae83 · outbound

This paper cites an unresolved cited work.

Interpretable Risk Mitigation in LLM Agent Systems Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:11:03.597680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.807791Z digest=sha256:de24e546708778568d5d54e8027044342ab5aa5b64ae8c5a6d5562fc7db4e6a3

Observation 94b9346b-084f-4af5-a678-a12f9bd12757 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Interpretable Risk Mitigation in LLM Agent Systems Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.813065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.813065Z digest=sha256:865726595d338e7faf825c08db1da58d2aa98603e098f091e62f8b5a15e282ae

Observation 5e6eafe6-a930-4240-a93e-aaa30f5aa5d6 · outbound

This paper cites Emergent tool use from multi-agent autocurricula.

Interpretable Risk Mitigation in LLM Agent Systems Emergent tool use from multi-agent autocurricula

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.580794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.817697Z digest=sha256:a0d23ec8518990ddcca79e2efafcd3507d143a51f492cd75177acdaa95f3993e

Observation 351ad7f1-6ab1-403d-b7ff-ef072b69c117 · outbound

This paper cites Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell.

Interpretable Risk Mitigation in LLM Agent Systems Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.562891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.822090Z digest=sha256:467e556e89330b1c76165510611195219cf498d88fd3e3b94c8ec67492542526

Observation 8a106779-c66d-48c7-8c48-b0f78d8abf43 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Interpretable Risk Mitigation in LLM Agent Systems On the Opportunities and Risks of Foundation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.826932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.826932Z digest=sha256:8578c533134ec3310aaabd602d5652d7eeeac1d59da5094c9fd9276d01fbc42d

Observation 2a7adc57-b623-480d-be0c-b1dd7e69a6ee · outbound

This paper cites an unresolved cited work.

Interpretable Risk Mitigation in LLM Agent Systems Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:11:03.544642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.832003Z digest=sha256:89b4b36f505750537dc380559a5c73360d9e2e8dee99323c2e83b0b3040cce3b

Observation a338b78c-533d-4e4b-9b30-639f80356236 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Interpretable Risk Mitigation in LLM Agent Systems RT-1: Robotics Transformer for Real-World Control at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.836214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.836214Z digest=sha256:7271eb10f1c4ecad6f70d0e33ba6e6d2053f72c397f9f2aa5f09cc6785ffa3e9

Observation 16d22cc2-e43f-4ec2-aff9-61a6d878a1a6 · outbound

This paper cites Playing games with gpt: What can we learn about a large language model from canonical strategic games? SSRN Electronic Journal, 2023.

Interpretable Risk Mitigation in LLM Agent Systems Playing games with gpt: What can we learn about a large language model from canonical strategic games? SSRN Electronic Journal, 2023

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.526152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.840725Z digest=sha256:e1f1c021b901035a7cc332740732fe974a9d57a98a9fa7b85f3a8168a58e8133

Observation fa8b734c-b330-4c95-85f8-426bab50f0df · outbound

This paper cites Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D.

Interpretable Risk Mitigation in LLM Agent Systems Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.508895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.844793Z digest=sha256:3a4d8943b999799733c9c518607e3adcf86817e56259f6daed6d86deca60439a

Observation b60e728a-df0f-4060-9aeb-3d455d436d74 · outbound

This paper cites What can machine learning do? workforce implications.

Interpretable Risk Mitigation in LLM Agent Systems What can machine learning do? workforce implications

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.491497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.848928Z digest=sha256:69b331cdb61985dcdebed6f2af1dcaa627ff50957d6142eeebeaa4a000f82394

Observation e55a86e7-747f-4031-8209-67c0bab512ab · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Interpretable Risk Mitigation in LLM Agent Systems Evaluating Large Language Models Trained on Code

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.853101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.853101Z digest=sha256:44563a09fdfb88c9d8e1e5ecc642bad3748a17bf2f4a7e96d8abfdf628990897

Observation 08b4f22d-3581-40e0-aa27-dd14aaa9041a · outbound

This paper cites Instigating cooperation among llm agents using adaptive information modulation.

Interpretable Risk Mitigation in LLM Agent Systems Instigating cooperation among llm agents using adaptive information modulation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.857296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.857296Z digest=sha256:2c172eae16cfd41558585d7147c0f6115041eef4f8fe146ad3c44c2c89e8b79f

Observation 72433d46-a595-481b-a5ef-2f80a6405e69 · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Interpretable Risk Mitigation in LLM Agent Systems Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.870732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.870732Z digest=sha256:6bba13bfaf3ef1ffd34712d97a897e626cac3729b45f2ce8a024d1b962aea5ca

Observation eadd38c1-fbcd-4b04-a105-2ea6f1e6a5e6 · outbound

This paper cites Reinforcement learning in a prisoner’s dilemma.

Interpretable Risk Mitigation in LLM Agent Systems Reinforcement learning in a prisoner’s dilemma

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.458432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.875550Z digest=sha256:e0d3699950efb846091f5fea687d5f5e8629ca70e757ee66d929e39c48940e4f

Observation 24f6a202-78a8-481d-8db8-1c4b79df229d · outbound

This paper cites Toy Models of Superposition.

Interpretable Risk Mitigation in LLM Agent Systems Toy Models of Superposition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.880133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.880133Z digest=sha256:06624c60b4b516fca0715b1d251458e8c8b6a5b9eff407308462355035a7ff7a

Observation 3da5c266-6d9b-45d6-8d90-16eafac12480 · outbound

This paper cites PoGaIN: Poisson-Gaussian Image Noise Modeling from Paired Samples.

Interpretable Risk Mitigation in LLM Agent Systems PoGaIN: Poisson-Gaussian Image Noise Modeling from Paired Samples

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T21:11:02.834041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.886176Z digest=sha256:b608c7d9f214673e01d21f596742d744fafd42916ae0999a25485aac42091eb4

Observation 9a93bcfe-f4cc-4a52-9e56-a716c5208000 · outbound

This paper cites Not All Language Model Features Are One-Dimensionally Linear.

Interpretable Risk Mitigation in LLM Agent Systems Not All Language Model Features Are One-Dimensionally Linear

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.891128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.891128Z digest=sha256:2d47d6f0b37d363b98c45979e3897450068c4b4325ddb2ca80cef4c80cb38a09

Observation 80ab46e3-96c8-4341-b15b-4ef484f62fc1 · outbound

This paper cites Some experimental games.

Interpretable Risk Mitigation in LLM Agent Systems Some experimental games

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.442507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.896121Z digest=sha256:0e185a2782d5535a4d35a58229ad3c64118b6c039e2d09489bae5994e6aff818

Observation f9da16f4-a983-41b4-b6d2-cccdb6ba0154 · outbound

This paper cites Nicer Than Humans: How do Large Language Models Behave in the Prisoner's Dilemma?.

Interpretable Risk Mitigation in LLM Agent Systems Nicer Than Humans: How do Large Language Models Behave in the Prisoner's Dilemma?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.900891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.900891Z digest=sha256:e3876cc9b7a6b030a24c1693253c6dee8d6736f75e6f14146152227bc6798e88

Observation 88637004-c8d5-4677-890d-d755e3357d38 · outbound

This paper cites Artificial intelligence, values, and alignment.

Interpretable Risk Mitigation in LLM Agent Systems Artificial intelligence, values, and alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.906112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.906112Z digest=sha256:2ec6d821156e684988dfa9802ea6f061515361cc506b385f2c54a760caf4cf62

Observation cb688a1f-ae7d-4316-9fcd-3ddffe2c900b · outbound

This paper cites The Llama 3 Herd of Models.

Interpretable Risk Mitigation in LLM Agent Systems The Llama 3 Herd of Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.910539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.910539Z digest=sha256:b1c06bb7abfceeb4f9cd31334268027730481af1911091c6295ed83d019e970d

Observation 2aea34d7-c236-4ba2-b05c-5266d05cbe3a · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Interpretable Risk Mitigation in LLM Agent Systems Measuring Massive Multitask Language Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.915548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.915548Z digest=sha256:db44e3c4c4758bd5ff4548d4106a78bde799533a9d41cddaf82a82fe846a08c3

Observation daa69867-514e-4500-ac68-9b455cec0cf3 · outbound

This paper cites Measuring mathematical problem solving with the math dataset,.

Interpretable Risk Mitigation in LLM Agent Systems Measuring mathematical problem solving with the math dataset,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.920013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.920013Z digest=sha256:d23741fe8de883dfd9d3b1ce648611561b1e72942d6e0d55247ec967fff16180

Observation ccc257f3-639b-4bd7-b693-d99e2d88ec53 · outbound

This paper cites Cogagent: A visual language model for gui agents.

Interpretable Risk Mitigation in LLM Agent Systems Cogagent: A visual language model for gui agents

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.416984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.930482Z digest=sha256:5d9eb1912aa5c15811c542c7ed0741d804423ac082d3c90b4c209dce26e0be4b

Observation f465e08f-8581-42e6-b51b-25e319bc463c · outbound

This paper cites Non-linear inference time intervention: Improving llm truthfulness.

Interpretable Risk Mitigation in LLM Agent Systems Non-linear inference time intervention: Improving llm truthfulness

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.400826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.935781Z digest=sha256:44d7949477bbc5c3f1954b24d84e2bf2712b9fc76c984b422e44e844dcc04e39

Observation 79c959d9-a3b3-4327-86ed-96acb55e005b · outbound

This paper cites Towards Reasoning in Large Language Models: A Survey.

Interpretable Risk Mitigation in LLM Agent Systems Towards Reasoning in Large Language Models: A Survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.939936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.939936Z digest=sha256:cb56251b99a829a0cee8e561ff681bf805f4c508094740407e520e1066d319fe

Observation f8bf0ff6-e450-4f1c-b07c-950516506ed9 · outbound

This paper cites Large language models for uavs: Current state and pathways to the future.

Interpretable Risk Mitigation in LLM Agent Systems Large language models for uavs: Current state and pathways to the future

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.384589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.944592Z digest=sha256:86aac816c8485736c754ff53320a8f88ab8c941ef948b2fcd694e1d923669143

Observation a7a39f6f-26fb-4c28-b196-bc2a237200ff · outbound

This paper cites Mixtral of Experts.

Interpretable Risk Mitigation in LLM Agent Systems Mixtral of Experts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.949339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.949339Z digest=sha256:6dd27b0f0894380d29083cd31f1f10af848f57cd569511f5a500a69d1e95aa85

Observation 017ef2a9-5520-4eb9-bdc4-f92641554c31 · outbound

This paper cites llama-3-8b-it-res (revision 53425c3), 2024.

Interpretable Risk Mitigation in LLM Agent Systems llama-3-8b-it-res (revision 53425c3), 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.367736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.953994Z digest=sha256:f1785714d4443c972c51c5e671505dab55ee397f9acd92e7664eee21eb06e2e6

Observation 91339d58-f620-44d4-8eb7-b59703b0d07a · outbound

This paper cites an unresolved cited work.

Interpretable Risk Mitigation in LLM Agent Systems Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:11:03.350181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.958191Z digest=sha256:675e21684b51dc9ad019f2afcb429a4211bf3fe50bd39ffd165d96ecc2e2d1a7

Observation 1a302ee4-fd9f-46ae-8fab-c55bc4736872 · outbound

This paper cites Martin, Hans-Theo Normann, and T.

Interpretable Risk Mitigation in LLM Agent Systems Martin, Hans-Theo Normann, and T

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.333745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.962360Z digest=sha256:226f31da3436f6a9408bb2b42b0080a23bc26012dfb41eeb005bb806d9cd5b37

Observation 71a8d99b-1365-4a5b-ac7b-cce9c7fceaf8 · outbound

This paper cites Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al.

Interpretable Risk Mitigation in LLM Agent Systems Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.318008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.966629Z digest=sha256:185006412421ca7e7321b2e924dab76219679a07bc0658dc951323a761340743

Observation 53ae654e-f604-43fa-bec1-7e4bbe48d9bf · outbound

This paper cites Inference-Time Intervention: Eliciting Truthful Answers from a Language Model.

Interpretable Risk Mitigation in LLM Agent Systems Inference-Time Intervention: Eliciting Truthful Answers from a Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.971070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.971070Z digest=sha256:5fff60f8abf3393e8a84a90f1c0b36b9c287e7558800700d3ad7eb78b3b0e19b

Observation accf6756-7ccc-46fa-8ac0-5a95cac9c7e1 · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

Interpretable Risk Mitigation in LLM Agent Systems Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.975880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.975880Z digest=sha256:e26d8a59d2abe7d9479e73939754b62caf36db4ac1b45dbf2f5b97e66c155f2d

Observation e1e40780-c265-4c8d-b588-0f4db9b277d5 · outbound

This paper cites The mythos of model interpretability.

Interpretable Risk Mitigation in LLM Agent Systems The mythos of model interpretability

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.302405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.980577Z digest=sha256:6aed82cb480083803288e57ee17c0003f7efe44c6d16dc260e1fdb2b1a4b71eb

Observation 4fc9c058-9b7f-4030-b591-6b05a8f26e5f · outbound

This paper cites Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing.

Interpretable Risk Mitigation in LLM Agent Systems Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.989753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.989753Z digest=sha256:64b120144d89a692168b6618e726cd86bd378cdbd72c11ffe167b23f67b2ebf0

Observation 3cdffd23-9790-4415-81f6-49a0963613c8 · outbound

This paper cites Large Model Strategic Thinking, Small Model Efficiency: Transferring Theory of Mind in Large Language Models.

Interpretable Risk Mitigation in LLM Agent Systems Large Model Strategic Thinking, Small Model Efficiency: Transferring Theory of Mind in Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.994175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.994175Z digest=sha256:0e243c61f1e3e23cdafff7cfc3578c4d9d3d68fb76472af60c0b47255017c2d0

Observation c86296d2-ed87-49d3-b9f2-78d6f82f7afa · outbound

This paper cites Linguistic regularities in continuous space word representations.

Interpretable Risk Mitigation in LLM Agent Systems Linguistic regularities in continuous space word representations

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.259770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:02.003164Z digest=sha256:cf8f9ca6920797cd92290b7dc1e1ad40934ec4bb447f2c47d67e74eec9f4bd8a

Observation a59a7951-c2e9-43c9-a557-199380bc4da6 · outbound

This paper cites Large Language Models: A Survey.

Interpretable Risk Mitigation in LLM Agent Systems Large Language Models: A Survey

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.008126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.008126Z digest=sha256:37d7137a026c8eca46c3a36ba52fcb1abb0d7d8ea860dc0150a9b80ad4c0d27b

Observation fae7cda6-218c-4394-8746-aa9f8ca079ab · outbound

This paper cites A strategy of win-stay, lose-shift that outperforms tit-for-tat in the prisoner’s dilemma game.

Interpretable Risk Mitigation in LLM Agent Systems A strategy of win-stay, lose-shift that outperforms tit-for-tat in the prisoner’s dilemma game

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.013053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.013053Z digest=sha256:91d3b15568bc02d4be803b8c837e41a4b7fa497c5688fdada2e1522d4e021a27

Observation 4826b429-d15b-417a-b168-ff52c95b5cbc · outbound

This paper cites Training language models to follow instructions with human feedback.

Interpretable Risk Mitigation in LLM Agent Systems Training language models to follow instructions with human feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.018166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.018166Z digest=sha256:b63bfa43074214c2b69cd51995d26e965c1e47781566b87bd0f3e1c9077b07ab

Observation 014ab367-984a-479d-83c0-103ef33837a1 · outbound

This paper cites Cooperation: A systematic review of how to enable agent to circumvent the prisoner’s dilemma.

Interpretable Risk Mitigation in LLM Agent Systems Cooperation: A systematic review of how to enable agent to circumvent the prisoner’s dilemma

Reference 48

Resolution
verified exact
raw_fallback, observed 2026-08-15T21:11:02.613607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:02.023478Z digest=sha256:6cf256f5a317fcba75e7c70e18adb5205949aa64cbcd47b101d1547201270b83

Observation 7a05f4db-7afc-4148-a98b-9ef26cb1000d · outbound

This paper cites Generative Agents: Interactive Simulacra of Human Behavior.

Interpretable Risk Mitigation in LLM Agent Systems Generative Agents: Interactive Simulacra of Human Behavior

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.028670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.028670Z digest=sha256:906119f1bbadc6f5573cf1b79034f51a28b0795aa8a0d02e5095cce36e62b730

Observation de77e093-1471-462a-b0a6-a927f6729dc5 · outbound

This paper cites TinyClick: Single-Turn Agent for Empowering GUI Automation.

Interpretable Risk Mitigation in LLM Agent Systems TinyClick: Single-Turn Agent for Empowering GUI Automation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.033414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.033414Z digest=sha256:16b75ff475d0a997bfab46a863d2da08f3618c86dd3bc96e88f511a4bd1a0acf

Observation 73a4ef3f-38c0-44df-85f6-445319704ac3 · outbound

This paper cites an unresolved cited work.

Interpretable Risk Mitigation in LLM Agent Systems Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:11:03.231251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:02.038063Z digest=sha256:741889572732ffe27e30b98c9071db38b31a482b300551c77d7d40e2e69354ff

Observation 35196dd2-37eb-4576-a6c2-b40e2da03ea4 · outbound

This paper cites Effect of private deliberation: Deception of large language models in game play.

Interpretable Risk Mitigation in LLM Agent Systems Effect of private deliberation: Deception of large language models in game play

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.216654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:02.042680Z digest=sha256:044aaaae570b590ad668e41ac9b4559c57e8dbb3a33b1ca946854a425d2395bd

Observation a70cb5aa-ff92-409f-b95a-3405b51bbf27 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Interpretable Risk Mitigation in LLM Agent Systems GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.047108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.047108Z digest=sha256:668e69bdbe7d9a2727b2761207ee34f363ca94b64305d72019f26b066304a8f1

Observation d5a5b1b9-7b0f-427c-973e-a741a2cd28cd · outbound

This paper cites A primer in BERTology: What we know about how BERT works.

Interpretable Risk Mitigation in LLM Agent Systems A primer in BERTology: What we know about how BERT works

Reference 54

Resolution
malformed identifier
no resolver link, observed 2026-08-15T21:11:02.051703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.051703Z digest=sha256:5caf7b8ae9e50809d195a5a8ec8b3ccfeee7b390c74c4c4d82b6cf87663f57ac

Observation cb2fe07f-6a82-4fb7-bc5a-7b4cefa20efd · outbound

This paper cites Research priorities for robust and beneficial artificial intelligence.

Interpretable Risk Mitigation in LLM Agent Systems Research priorities for robust and beneficial artificial intelligence

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.201675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:02.055985Z digest=sha256:c4a633b97ce2019d3457cb911695dbbd8665743bd53a734b68920781c30cfd4b

Observation 7a2cf6c4-3cb3-42c1-a6c5-e33c3fd83511 · outbound

This paper cites an unresolved cited work.

Interpretable Risk Mitigation in LLM Agent Systems Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:11:03.185336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:02.060734Z digest=sha256:6d41bcd9e83c381754c2ade30c40f3df66b6f767fc8e2bd358d580d1b4e94979

Observation 95101970-bd83-42b3-bfc5-703b311591ce · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

Interpretable Risk Mitigation in LLM Agent Systems Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.064856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.064856Z digest=sha256:16b2ea0b2319509fcbef501ccd0b8a7963ea03d683444acf78ff9614b0986054

Observation 8f6167c9-4a5f-46af-b08b-065532a5c700 · outbound

This paper cites An evolutionary model of personality traits related to cooperative behavior using a large language model.

Interpretable Risk Mitigation in LLM Agent Systems An evolutionary model of personality traits related to cooperative behavior using a large language model

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.170463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:02.069111Z digest=sha256:ab03863ed3bd26d7fc9b0678092480666d6f24de9ae2f2a4e3b84d90cc272803

Observation 0b622b1d-44e4-43db-bc95-870dd9ca5525 · outbound

This paper cites A comparative analysis of the definitions of autonomous weapons systems.

Interpretable Risk Mitigation in LLM Agent Systems A comparative analysis of the definitions of autonomous weapons systems

Reference 59

Resolution
verified exact
doi, observed 2026-08-15T21:11:02.190554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:02.073714Z digest=sha256:fd31bc5e07e2099a466fe5d9eb058ec90e5cbfd08974bb3232d448e266116942

Observation 54cdc25e-dcbe-4a06-a3ca-a0b083ec66e7 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Interpretable Risk Mitigation in LLM Agent Systems Gemma: Open Models Based on Gemini Research and Technology

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.078502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.078502Z digest=sha256:86e88706b1484969c07374ff790b8c6d16fbf14050f7ca827de2239fca14ad88

Observation e70192c0-2fa3-4d26-9e99-25bce4521ab8 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Interpretable Risk Mitigation in LLM Agent Systems Gemma 2: Improving Open Language Models at a Practical Size

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.082737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.082737Z digest=sha256:4b0e3816bdf7f7cd0299d9100c98f4bf0dcf54ff396a9348f2467b0e8d0a1f06

Observation 625f59c6-0eef-406d-9cf0-d8c92c8ef449 · outbound

This paper cites Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet.

Interpretable Risk Mitigation in LLM Agent Systems Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.087316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.087316Z digest=sha256:ea5c6edb2137d516d6bfc794d205c6a4b930d94f831d90b0bb4d0f915f01e992

Observation 7b2ebf77-d4d1-47eb-846b-c88fd4d56a1d · outbound

This paper cites Moral alignment for llm agents.

Interpretable Risk Mitigation in LLM Agent Systems Moral alignment for llm agents

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.143335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:02.092060Z digest=sha256:16a771f01da0fe7fbbc9c462e63e7df6c9927a2e5f2956862a50377636d585b6

Observation f523afad-184f-41ef-9a0e-57e19a11c506 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Interpretable Risk Mitigation in LLM Agent Systems LLaMA: Open and Efficient Foundation Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.098429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.098429Z digest=sha256:a73e10b37d6b130ee2da8cd30e069dab6ec3c3c1dc88da96c3e722df1727b64b

Observation 82ac90b8-2d66-452e-9693-66c10e7d314b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Interpretable Risk Mitigation in LLM Agent Systems Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.104902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.104902Z digest=sha256:488a0e92d73ee904c54b58e96d9ce7d2e3d963edcd6c20102a22f73679042d9b

Observation 20ce1bc9-e4d8-4904-9888-9de2d34c5479 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Interpretable Risk Mitigation in LLM Agent Systems GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.109561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.109561Z digest=sha256:84b8d05572d573c2a304019432bfdfcfd22929918023a6c927362076b26f2ec6

Observation f2942ef8-188f-4e1f-a44b-0747ba8e3471 · outbound

This paper cites Mobile-agent: Autonomous multi-modal mobile device agent with visual perception,.

Interpretable Risk Mitigation in LLM Agent Systems Mobile-agent: Autonomous multi-modal mobile device agent with visual perception,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.127797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:02.114293Z digest=sha256:f6ec9b91c481381947383deb822883dd291eca186605afe990e95c9578d21292

Observation e8c5f846-e91b-401e-80a4-0fd5d73e1194 · outbound

This paper cites Dai, and Quoc V Le.

Interpretable Risk Mitigation in LLM Agent Systems Dai, and Quoc V Le

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:11:03.112312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:02.123585Z digest=sha256:3d6f29712340fbd701b392da665a0487100d85568397e3183d7582e1c8e1f304

Observation 9aac0e58-4508-4dce-b103-c860ac6c28f4 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Interpretable Risk Mitigation in LLM Agent Systems Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.128289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.128289Z digest=sha256:e784b7c5d5b38e5c1bb39b91fedb824e072554b821f5b88bf7b8ceae8b7d13c0

Observation 1cea8ac7-27e1-49cf-b37d-d6744ff11b68 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Interpretable Risk Mitigation in LLM Agent Systems ReAct: Synergizing Reasoning and Acting in Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.133411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.133411Z digest=sha256:12e109c5be4cd28bae23d50a0fc8ecd17221320eaa61ac2037adf5af4713c302

Observation 647ac5ed-565f-4692-8bee-b21bc1834a89 · outbound

This paper cites AppAgent: Multimodal Agents as Smartphone Users.

Interpretable Risk Mitigation in LLM Agent Systems AppAgent: Multimodal Agents as Smartphone Users

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.140648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.140648Z digest=sha256:ffc9db1b1edd25388cc4c38741306ed62b48171bd7b6d5b8014e0e3a6a763ac1

Observation 33d30244-b423-4420-bcc9-fbd665b74b6d · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

Interpretable Risk Mitigation in LLM Agent Systems Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.118923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.118923Z digest=sha256:cb0932424bf65a1d6ab0532aa77530d229365fd2a50a934e3fd16b6fe282cfd0

Observation 10bb76ee-bc2d-40f4-babf-40043502314b · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Interpretable Risk Mitigation in LLM Agent Systems Representation Engineering: A Top-Down Approach to AI Transparency

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.149478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.149478Z digest=sha256:de3cd754c442b347905b0e5c04cce4a47adfa193e7f975aaf60c01be22ac6381

Observation e1633c79-0375-4295-bfa3-8325c644a49a · outbound

This paper cites You Only Look at Screens: Multimodal Chain-of-Action Agents.

Interpretable Risk Mitigation in LLM Agent Systems You Only Look at Screens: Multimodal Chain-of-Action Agents

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:02.145319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:02.145319Z digest=sha256:013fe3a399542f12e0a6157c214bcbd433d7ca25a11f9c77b1104456f200e958

Observation d843c580-4512-4d8f-a3a6-6b41724a5ec6 · outbound

This paper cites an unresolved cited work.

Interpretable Risk Mitigation in LLM Agent Systems Unresolved cited work

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.984893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.984893Z digest=sha256:4e20b00e5435251add6c22a789920ee7cb52bc422b38f2720cc1dc7d32a4c7d8

Observation 671ed4ec-d1ad-4d99-97f8-e0d15322b52b · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Interpretable Risk Mitigation in LLM Agent Systems Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T21:11:01.925637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:11:01.925637Z digest=sha256:2a1027b9de6115620f3840d9b0dbf3b02318c94e1af4425a1480369fa540ce33

Observation 6b1e14ac-f124-489b-84f4-00a88d7ea70d · outbound

This paper cites an unresolved cited work.

Interpretable Risk Mitigation in LLM Agent Systems Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:11:03.474249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.866455Z digest=sha256:8e69029f32e399eb51a36d05ae4c1e5eb85f3183a8d8122161f860c00d3d6a9a

Observation 98fbd528-7341-4dfe-92d5-440307b35277 · outbound

This paper cites an unresolved cited work.

Interpretable Risk Mitigation in LLM Agent Systems Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:11:03.276283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:11:01.998638Z digest=sha256:6086029533cde48ae5ad2669dc83a02a90b94ea6b2a4e6deb0950417257bb98b

Pith citing papers

Observation b72a9458-11dc-47d1-84f0-f77d9122b3d6 · inbound

A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy cites this paper.

A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy Interpretable Risk Mitigation in LLM Agent Systems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:51:56.450815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:51:56.450815Z digest=sha256:ce9e691a59f0c765e1e5e4b4af0c7f11a4f324957fd9078aa4e801d947ebea27

Observation 1b7081ac-e99d-43fc-be2e-190be39eed65 · inbound

LLM Harms: A Taxonomy and Discussion cites this paper.

LLM Harms: A Taxonomy and Discussion Interpretable Risk Mitigation in LLM Agent Systems

Reference 189

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:31:24.636618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T00:29:07.951709Z digest=sha256:eb6cc424930bd73eca14e3b202775dd85faa85436123246c906b31a4a92ddc77

Observation 84d058c8-51a1-4859-970e-8b0c97117aac · inbound

LLM Harms: A Taxonomy and Discussion cites this paper.

LLM Harms: A Taxonomy and Discussion Interpretable Risk Mitigation in LLM Agent Systems

Reference 189

Resolution
unresolved
no resolver link, observed 2026-08-03T18:19:27.337242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:19:27.337242Z digest=sha256:4e62376da4334efae82eb7aaccbb4096187123306b539a8647ffc59c3a41853b