Pith. sign in

Paper Citation Record · LEDGER

What AI Red-Team Evaluations Can and Cannot Prove

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2607.21735.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21735 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:54:48.503377Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d88d1232-5dec-4a04-984b-c4b5a8e0ff33 · outbound

This paper cites Hudson, Ehsan Adeli, et al.

What AI Red-Team Evaluations Can and Cannot Prove Hudson, Ehsan Adeli, et al

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.292789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.292789Z digest=sha256:956d5d689f788930a6d07c32f8f90250fc3f849367bac1539332c0fa4c412ef7

Observation 373d399f-b1e6-43d5-8093-c4405b0438c2 · outbound

This paper cites Model cards for model reporting.

What AI Red-Team Evaluations Can and Cannot Prove Model cards for model reporting

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.298076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.298076Z digest=sha256:57202703ab854563ec46f35d88a4bee465f6c2e10c5fefc3f91764cb99c4b672

Observation 1dc17684-f994-410b-aa15-0a01343edefb · outbound

This paper cites Holistic evaluation of language models, 2023.

What AI Red-Team Evaluations Can and Cannot Prove Holistic evaluation of language models, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.302790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.302790Z digest=sha256:d225594f5e714747ea4cf11fd4f8b1b44d55d858a725e19bb5de9a556950c78c

Observation 77c09251-4538-47c7-85cf-fdae96f89c24 · outbound

This paper cites Gritsenko, et al.

What AI Red-Team Evaluations Can and Cannot Prove Gritsenko, et al

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.307997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.307997Z digest=sha256:26ed5596957376d05bc6fc00d62419851851b1fb6a3eb3973a4231b9226d2fed

Observation dbb6170d-e3d4-476f-b6a1-7d3af53b54bb · outbound

This paper cites Feder Cooper, Solon Barocas, Abhinav Palia, Dan Vann, and Hanna Wallach.

What AI Red-Team Evaluations Can and Cannot Prove Feder Cooper, Solon Barocas, Abhinav Palia, Dan Vann, and Hanna Wallach

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.312735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.312735Z digest=sha256:c2ef0871f7d5c1152e43e436abd0edbe177caf8319c3ca9b39cc102f647cc77d

Observation bfbf40ef-b0e4-49f6-8f4b-3b6555b6659d · outbound

This paper cites The structural safety gener- alization problem, 2025.

What AI Red-Team Evaluations Can and Cannot Prove The structural safety gener- alization problem, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.317460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.317460Z digest=sha256:38a55a62e1ed9526d340a269ee9ef9a9a93065537e28b2e7c9484e87f30a910b

Observation a2c47ca0-2d57-47cc-a63b-e132dcfd7056 · outbound

This paper cites Adding error bars to evals: A statistical approach to language model evaluations, 2024.

What AI Red-Team Evaluations Can and Cannot Prove Adding error bars to evals: A statistical approach to language model evaluations, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.322834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.322834Z digest=sha256:471cf99a47ef3f541addb25f446cb1d71e2d0315dcd7c2685d35af525b2809c1

Observation 3d11f44d-a56b-4035-84e3-33a05813ffa2 · outbound

This paper cites Kim and Anthony R.

What AI Red-Team Evaluations Can and Cannot Prove Kim and Anthony R

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.327008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.327008Z digest=sha256:a13cc2351d4eb3fc76b42f2857f0cbe3bc05b818cc59a32887e40d65e74d92de

Observation 57a5a564-44a7-4132-ae28-c6dccd5a2e43 · outbound

This paper cites Prentice.

What AI Red-Team Evaluations Can and Cannot Prove Prentice

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.331485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.331485Z digest=sha256:cbb15b790b3d618c8e5908dff6df8bc843c9f52d3d56afd714f9f6ef0818fee4

Observation 15d7eebd-aa13-47ce-b215-3e3b5db064f5 · outbound

This paper cites Safety cases: How to justify the safety of advanced AI systems, 2024.

What AI Red-Team Evaluations Can and Cannot Prove Safety cases: How to justify the safety of advanced AI systems, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.336192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.336192Z digest=sha256:b4080eef5400c425d9ad3c0132f2d6cbb57460fde0834590b8abc733372405d8

Observation d1c28ae5-00a6-4d28-bbc5-9371f24a0edf · outbound

This paper cites A sketch of an AI control safety case, 2025.

What AI Red-Team Evaluations Can and Cannot Prove A sketch of an AI control safety case, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.340824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.340824Z digest=sha256:3c334ed327686d961cc5b76ebee5643115181b9cc1bfb86b41ae29b9068636df

Observation bd78ca10-2ea9-4e4b-8876-96ee4677661b · outbound

This paper cites Shadish, Thomas D.

What AI Red-Team Evaluations Can and Cannot Prove Shadish, Thomas D

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.346205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.346205Z digest=sha256:3c7fb860804ccabc27015e9ccb85ad28c58c9bf27a3af47b7275b51557d87684

Observation 8fd1b70a-b96a-4904-8310-6ff63f2a110d · outbound

This paper cites Dulberg, and George A.

What AI Red-Team Evaluations Can and Cannot Prove Dulberg, and George A

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.351912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.351912Z digest=sha256:9a6b8555f2a447c17d5cb8988dcea6db651ebc899d86334019ad2ca9f7be7e30

Observation b985ad45-8b24-404f-bd54-efc7efda56bd · outbound

This paper cites Hanley and Abby Lippman-Hand.

What AI Red-Team Evaluations Can and Cannot Prove Hanley and Abby Lippman-Hand

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.356405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.356405Z digest=sha256:929f8e90a8f1587b99d7afa8fc4b5f557efb11324337293e71c6700dd807ef38

Observation 80f097a5-e73e-4c79-8be4-89db9febbe71 · outbound

This paper cites Brown, and Francis R.

What AI Red-Team Evaluations Can and Cannot Prove Brown, and Francis R

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.360716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.360716Z digest=sha256:2add4500c12b3cbbb3977c6ddd5aa1f27214e2e021a9c3383afe20e9783f586e

Observation 3eef9d01-3c9f-4227-ba1f-356c9a457061 · outbound

This paper cites XSTest: A test suite for identifying exaggerated safety behaviours in large language models, 2024.

What AI Red-Team Evaluations Can and Cannot Prove XSTest: A test suite for identifying exaggerated safety behaviours in large language models, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.365061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.365061Z digest=sha256:1bfb198a11ce0c057215e195b3c7d5a73b73a54c33d4f6de7ef397be8e95e39b

Observation 35546b27-3658-41db-ad14-ecb340c7aaee · outbound

This paper cites SafetyBench: Evaluating the Safety of Large Language Models.

What AI Red-Team Evaluations Can and Cannot Prove SafetyBench: Evaluating the Safety of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.369854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.369854Z digest=sha256:b7564e40575970b4db289a6fb692756f053ba369d01b134a3d4ee9db072cdb20

Observation 9876aa84-411b-4b94-828d-1e58b0411666 · outbound

This paper cites Zico Kolter, and Matt Fredrikson.

What AI Red-Team Evaluations Can and Cannot Prove Zico Kolter, and Matt Fredrikson

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.374594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.374594Z digest=sha256:ea147add96cef7340b752835201fd320bbd132a4e66d64d45b7f190295922fbb

Observation 672f8c41-867c-4c87-b89a-54f65793a14c · outbound

This paper cites HarmBench: A standardized evaluation framework for automated red teaming and robust refusal, 2024.

What AI Red-Team Evaluations Can and Cannot Prove HarmBench: A standardized evaluation framework for automated red teaming and robust refusal, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.378992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.378992Z digest=sha256:317613f366116b70c553a0b1b0b3742ef1aad98744455202ea5f0f9137da9a8b

Observation 5b2fab8a-2f47-4b04-9ebd-f839dc927756 · outbound

This paper cites A StrongREJECT for empty jailbreaks, 2024.

What AI Red-Team Evaluations Can and Cannot Prove A StrongREJECT for empty jailbreaks, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.383131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.383131Z digest=sha256:f697180bfc62f059873c029053c050c9760b94bb8eef50bab5a9ad38b73c9701

Observation 2d6737ef-b31b-4d9d-8da0-0885059f99a3 · outbound

This paper cites an unresolved cited work.

What AI Red-Team Evaluations Can and Cannot Prove Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.387331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.387331Z digest=sha256:82cf9063095204ed9663f33ca4c47c4fe47ae5bffbf92ece227dad7644aec1cc

Observation f37e2ad5-f7d1-4f60-bcc3-33ae49081e2d · outbound

This paper cites John Wiley and Sons, New York, 1965.

What AI Red-Team Evaluations Can and Cannot Prove John Wiley and Sons, New York, 1965

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.391557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.391557Z digest=sha256:c2e3d0df9520e77fa0439f5fa856d342018f16704c7654868a5440d9f49420a9

Observation 6467b702-3a36-4728-acd7-6519488f324e · outbound

This paper cites an unresolved cited work.

What AI Red-Team Evaluations Can and Cannot Prove Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.395950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.395950Z digest=sha256:9903a45677e7ef458448b9f565344a7160f5230ddadc2f1dd970b81487ef0287

Observation ec26b57d-c8a9-40b2-a2eb-5fbf0aad3434 · outbound

This paper cites Model card and evaluations for claude models.

What AI Red-Team Evaluations Can and Cannot Prove Model card and evaluations for claude models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.400904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.400904Z digest=sha256:fe1942f3ca5285558738e0ff50e9b0a1094de9d39f343b99a8d07ac22437636f

Observation 2743c202-13bb-48ad-9ec3-9d6c206cc841 · outbound

This paper cites Sentence-BERT: Sentence embeddings using siamese BERT-networks.

What AI Red-Team Evaluations Can and Cannot Prove Sentence-BERT: Sentence embeddings using siamese BERT-networks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.406066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.406066Z digest=sha256:54da77787ef9f00c04f3f12e9fe860af4927ce6c1ac17e29ac086de850e75da6

Observation 9382bd33-680a-4971-a2c4-76cec48fe29c · outbound

This paper cites LMSYS-Chat-1M: A large-scale real-world LLM conversation dataset, 2024.

What AI Red-Team Evaluations Can and Cannot Prove LMSYS-Chat-1M: A large-scale real-world LLM conversation dataset, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.410572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.410572Z digest=sha256:56d662dc83285ac122758accbff91de0cd67c681c746e3799eac4272b82b9b74

Observation 211ea424-422f-4e2d-b5a6-6eaec1d95e12 · outbound

This paper cites Borgwardt, Malte J.

What AI Red-Team Evaluations Can and Cannot Prove Borgwardt, Malte J

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.415116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.415116Z digest=sha256:e862de04bff29a1036185d5316080dd19ffaaae691f3bb0bf08ab015a587a4b3

Observation d26b55e6-5b01-46cc-9a5f-dc8f24d4fc7a · outbound

This paper cites UMAP: Uniform manifold ap- proximation and projection for dimension reduction, 2018.

What AI Red-Team Evaluations Can and Cannot Prove UMAP: Uniform manifold ap- proximation and projection for dimension reduction, 2018

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.420308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.420308Z digest=sha256:62d15eade1f54db988c216cc9e164e0385d1fdcbc14ed041bc3ca061b372a451

Observation 942aafff-26e4-42eb-bb8c-7f19f5c9ef46 · outbound

This paper cites Choquette-Choo, et al.

What AI Red-Team Evaluations Can and Cannot Prove Choquette-Choo, et al

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.426517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.426517Z digest=sha256:d584ca71ed63b716c9d7dba3423795a9ad1289e281946ed6972216b305483106

Observation 1a80966c-fb58-438f-9e1b-4b4c63f0615c · outbound

This paper cites AutoDAN: Generating stealthy jailbreak prompts on aligned large language models, 2024.

What AI Red-Team Evaluations Can and Cannot Prove AutoDAN: Generating stealthy jailbreak prompts on aligned large language models, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.434251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.434251Z digest=sha256:b381ffffdb7a75292acfa0f314e25657d990c84cab0c3421bd612f54615fdf48

Observation 7130a205-44c7-4b3b-a148-c24277ff8fe8 · outbound

This paper cites Does refusal training in LLMs generalize to the past tense?, 2025.

What AI Red-Team Evaluations Can and Cannot Prove Does refusal training in LLMs generalize to the past tense?, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.440217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.440217Z digest=sha256:5a489c424153c073205f247c2991e0cfbdb43debdbd766c13df9ec64a7b3a44c

Observation 40a1a1e4-0715-4b4e-9a2d-87dba58c35be · outbound

This paper cites an unresolved cited work.

What AI Red-Team Evaluations Can and Cannot Prove Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.446386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.446386Z digest=sha256:fab526245b3fb048f8ec96411bc1c9de9dbf6d776da6c3f83c653bdce44006aa

Observation 0eb41bd9-2072-4694-a6f9-e520ee816c51 · outbound

This paper cites Tree of attacks: Jail- breaking black-box LLMs automatically, 2024.

What AI Red-Team Evaluations Can and Cannot Prove Tree of attacks: Jail- breaking black-box LLMs automatically, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.451336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.451336Z digest=sha256:07bc042c93e9bb2054685b00dedc8b28d68be09ccd1732375da865284ae0fe73

Observation 652b69c2-d087-4678-bab5-18eb0e197c1d · outbound

This paper cites Ignore previous prompt: Attack techniques for lan- guage models, 2022.

What AI Red-Team Evaluations Can and Cannot Prove Ignore previous prompt: Attack techniques for lan- guage models, 2022

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.455476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.455476Z digest=sha256:d1d72a122139a4874eb7854dfc942676a629dee46db0224ceaa425c29c4b9fb1

Observation 66ea6870-7599-4e90-8976-8d8e4930c06f · outbound

This paper cites GPT-4 technical report, 2023.

What AI Red-Team Evaluations Can and Cannot Prove GPT-4 technical report, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.460464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.460464Z digest=sha256:4b5ed9d98934231017e7987ca378e96bf1b2052c249d296a1137193316b827f6

Observation d73ba15e-e881-40ee-a011-21c5730c57e0 · outbound

This paper cites GPT-4o system card, 2024.

What AI Red-Team Evaluations Can and Cannot Prove GPT-4o system card, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.464794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.464794Z digest=sha256:b3e32e98bdbfa13339fff3025d6dbc2c85ba15585eecfed7a4629864035a3ea9

Observation 6aeced3a-b254-4c43-9259-387270fb29bd · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

What AI Red-Team Evaluations Can and Cannot Prove The claude 3 model family: Opus, sonnet, haiku

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.469031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.469031Z digest=sha256:4e8d6d8a42f67f3600e26131ce8bee120479e5d413eb4e5c442ac7a4d759ac7f

Observation b8d0024a-c2f1-4657-910c-c039b8fdf02f · outbound

This paper cites System card: Claude opus 4 and claude sonnet 4.

What AI Red-Team Evaluations Can and Cannot Prove System card: Claude opus 4 and claude sonnet 4

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.473574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.473574Z digest=sha256:20b1190c30b5a36368391d523e7625fbf8e1e003f1d22efc09bc2033806c1952

Observation 9a976dac-022b-4d95-a2b6-79f8b8d1df37 · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2023.

What AI Red-Team Evaluations Can and Cannot Prove Gemini: A family of highly capable multimodal models, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.478226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.478226Z digest=sha256:8f8897b6c40ff5020b18dbcbca193cb8e0f0483633c676b99bea574477b75ca4

Observation 43d610c4-3c39-4a56-b54c-3403433df20b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024.

What AI Red-Team Evaluations Can and Cannot Prove Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.482631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.482631Z digest=sha256:2a51f0d8032cd2cb948e99347d99c64eae029887fd111ae3cb5ec2952b608cdd

Observation 4954885d-4486-4200-a812-e363906d50eb · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models, 2023.

What AI Red-Team Evaluations Can and Cannot Prove Llama 2: Open foundation and fine-tuned chat models, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.486674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.486674Z digest=sha256:224b2c13e09e949464ed8638f636752a70b9527d49a97e6579a5d19d8b7ae411

Observation edfa1227-57cf-4042-8c31-f4869976d8f4 · outbound

This paper cites The llama 3 herd of models, 2024.

What AI Red-Team Evaluations Can and Cannot Prove The llama 3 herd of models, 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.490984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.490984Z digest=sha256:8d6432a847b10b941717db7aa80eac4aac3da6b1454b4218f9e2476fe5db1479

Observation 56e9109f-1aee-4983-9afd-fa36d97d3de9 · outbound

This paper cites Responsible scaling policy.

What AI Red-Team Evaluations Can and Cannot Prove Responsible scaling policy

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.495168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.495168Z digest=sha256:3cd0c0fc20afe762276350e54304d85a3d86f3b3ef5f6c1c53979ca3e233fe43

Observation 2bacd9a9-bb05-4f97-a16d-74edf5a71c39 · outbound

This paper cites Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022.

What AI Red-Team Evaluations Can and Cannot Prove Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.499159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.499159Z digest=sha256:26bc378f760497dbfe1bdfcfc2fb20659419362b3d0564877bf94c76de30e2f8

Observation 8811b220-1865-4f65-acbe-2ba4172ec09b · outbound

This paper cites Red teaming language models with language models, 2022.

What AI Red-Team Evaluations Can and Cannot Prove Red teaming language models with language models, 2022

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.503377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.503377Z digest=sha256:b2d21e45e8434a5d6c84f12429fa806d53c88ac1345ad5a3aa08b23c69e8069d

Pith citing papers

No inbound Pith citation observations are available.