Pith. sign in

Paper Citation Record · LEDGER

What AI Red-Team Evaluations Can and Cannot Prove

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2607.21735.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21735 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:54:48.503377Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d88d1232-5dec-4a04-984b-c4b5a8e0ff33 · outbound

This paper cites Hudson, Ehsan Adeli, et al.

What AI Red-Team Evaluations Can and Cannot Prove Hudson, Ehsan Adeli, et al

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.292789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.292789Z digest=sha256:99c363d6320d9e5c40763578adefccecafeef1c4089b0489b7529c6adae29d0b

Observation 373d399f-b1e6-43d5-8093-c4405b0438c2 · outbound

This paper cites Model cards for model reporting.

What AI Red-Team Evaluations Can and Cannot Prove Model cards for model reporting

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.298076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.298076Z digest=sha256:4a2688f9163822340e09b3aa1457f96d3b6a88ae63eb1bc986ae8d30bc2dda6b

Observation 1dc17684-f994-410b-aa15-0a01343edefb · outbound

This paper cites Holistic evaluation of language models, 2023.

What AI Red-Team Evaluations Can and Cannot Prove Holistic evaluation of language models, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.302790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.302790Z digest=sha256:92308e4042347d5fdac6973030d5afa934c10a1aa93910536553740c814d6d00

Observation 77c09251-4538-47c7-85cf-fdae96f89c24 · outbound

This paper cites Gritsenko, et al.

What AI Red-Team Evaluations Can and Cannot Prove Gritsenko, et al

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.307997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.307997Z digest=sha256:801141b07ef6770af30cd3efbc8a451723781fee5e147528c14873785f54bde0

Observation dbb6170d-e3d4-476f-b6a1-7d3af53b54bb · outbound

This paper cites Feder Cooper, Solon Barocas, Abhinav Palia, Dan Vann, and Hanna Wallach.

What AI Red-Team Evaluations Can and Cannot Prove Feder Cooper, Solon Barocas, Abhinav Palia, Dan Vann, and Hanna Wallach

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.312735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.312735Z digest=sha256:e1d6c92524d855b71e06539b8f5a86a0814e4ba2423ab23e2caf878602e197c6

Observation bfbf40ef-b0e4-49f6-8f4b-3b6555b6659d · outbound

This paper cites The structural safety gener- alization problem, 2025.

What AI Red-Team Evaluations Can and Cannot Prove The structural safety gener- alization problem, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.317460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.317460Z digest=sha256:ce3c310da618bb842b2adf901b845ce4c735201141c4019bdb5b171e38d1962e

Observation a2c47ca0-2d57-47cc-a63b-e132dcfd7056 · outbound

This paper cites Adding error bars to evals: A statistical approach to language model evaluations, 2024.

What AI Red-Team Evaluations Can and Cannot Prove Adding error bars to evals: A statistical approach to language model evaluations, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.322834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.322834Z digest=sha256:e1f0d9a0dd3f6f739bd2176f0c8b44296edffc6ad8c8b7169d62419d45ca94e0

Observation 3d11f44d-a56b-4035-84e3-33a05813ffa2 · outbound

This paper cites Kim and Anthony R.

What AI Red-Team Evaluations Can and Cannot Prove Kim and Anthony R

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.327008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.327008Z digest=sha256:9c4d9fac1312ce542f0a71d55306f96639cf21628c4fda5f3ee0e3df70e5afe0

Observation 57a5a564-44a7-4132-ae28-c6dccd5a2e43 · outbound

This paper cites Prentice.

What AI Red-Team Evaluations Can and Cannot Prove Prentice

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.331485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.331485Z digest=sha256:add91f71d35ab47efe11208ae27619624447973cf4df3ce5d7b7ec701cdac0cd

Observation 15d7eebd-aa13-47ce-b215-3e3b5db064f5 · outbound

This paper cites Safety cases: How to justify the safety of advanced AI systems, 2024.

What AI Red-Team Evaluations Can and Cannot Prove Safety cases: How to justify the safety of advanced AI systems, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.336192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.336192Z digest=sha256:6d6d29a72ef31c86f25bede15748f77136ffe6d34361822e78bd9c2f503db90a

Observation d1c28ae5-00a6-4d28-bbc5-9371f24a0edf · outbound

This paper cites A sketch of an AI control safety case, 2025.

What AI Red-Team Evaluations Can and Cannot Prove A sketch of an AI control safety case, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.340824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.340824Z digest=sha256:4c1f8e0f84feaa9cb9f03c3f47afa877b8e9bb6b340d9292e24c0ba4ff126501

Observation bd78ca10-2ea9-4e4b-8876-96ee4677661b · outbound

This paper cites Shadish, Thomas D.

What AI Red-Team Evaluations Can and Cannot Prove Shadish, Thomas D

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.346205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.346205Z digest=sha256:74349b3d26f5c9d57d95901a047ad009efbcc2c400708ff29da8308cc5c12c26

Observation 8fd1b70a-b96a-4904-8310-6ff63f2a110d · outbound

This paper cites Dulberg, and George A.

What AI Red-Team Evaluations Can and Cannot Prove Dulberg, and George A

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.351912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.351912Z digest=sha256:60ec55d54bd67268f8e1fca942c6c732cc999033d643186de70bfa5aa9c55ab9

Observation b985ad45-8b24-404f-bd54-efc7efda56bd · outbound

This paper cites Hanley and Abby Lippman-Hand.

What AI Red-Team Evaluations Can and Cannot Prove Hanley and Abby Lippman-Hand

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.356405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.356405Z digest=sha256:231ef54e84ffbf7040f229f3104e42791052bcf8e0d31072b5d0f307160dbb5d

Observation 80f097a5-e73e-4c79-8be4-89db9febbe71 · outbound

This paper cites Brown, and Francis R.

What AI Red-Team Evaluations Can and Cannot Prove Brown, and Francis R

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.360716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.360716Z digest=sha256:f1b7c393771ce2614e03d80b6423b3ba574fdedce57a93050389862a9f12162a

Observation 3eef9d01-3c9f-4227-ba1f-356c9a457061 · outbound

This paper cites XSTest: A test suite for identifying exaggerated safety behaviours in large language models, 2024.

What AI Red-Team Evaluations Can and Cannot Prove XSTest: A test suite for identifying exaggerated safety behaviours in large language models, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.365061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.365061Z digest=sha256:234078f098c8e680b58718e930f2b9696c272515239c280269eaf51cf686c7bb

Observation 35546b27-3658-41db-ad14-ecb340c7aaee · outbound

This paper cites SafetyBench: Evaluating the Safety of Large Language Models.

What AI Red-Team Evaluations Can and Cannot Prove SafetyBench: Evaluating the Safety of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.369854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.369854Z digest=sha256:1260af41948219a795c7557978523920b6310bd42b834c3c6328f39112e40a63

Observation 9876aa84-411b-4b94-828d-1e58b0411666 · outbound

This paper cites Zico Kolter, and Matt Fredrikson.

What AI Red-Team Evaluations Can and Cannot Prove Zico Kolter, and Matt Fredrikson

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.374594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.374594Z digest=sha256:91467bec5d628c93ffae89de18b9c9c973e88fd40cafa83519e6a2fd456ae59f

Observation 672f8c41-867c-4c87-b89a-54f65793a14c · outbound

This paper cites HarmBench: A standardized evaluation framework for automated red teaming and robust refusal, 2024.

What AI Red-Team Evaluations Can and Cannot Prove HarmBench: A standardized evaluation framework for automated red teaming and robust refusal, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.378992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.378992Z digest=sha256:481c9095f4232b74c5ae57dcb19e6dcacec2d0611ecaaadf5619c725883a3b2c

Observation 5b2fab8a-2f47-4b04-9ebd-f839dc927756 · outbound

This paper cites A StrongREJECT for empty jailbreaks, 2024.

What AI Red-Team Evaluations Can and Cannot Prove A StrongREJECT for empty jailbreaks, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.383131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.383131Z digest=sha256:a62771427912245e882908ce4d0f0b66c4cd262fdd61e5bb7bf2148f3a02dd4c

Observation 2d6737ef-b31b-4d9d-8da0-0885059f99a3 · outbound

This paper cites an unresolved cited work.

What AI Red-Team Evaluations Can and Cannot Prove Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.387331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.387331Z digest=sha256:1ecbf3423e517acebd2c7f0da88dd7260852214b41e780c1164e796ca3a5485c

Observation f37e2ad5-f7d1-4f60-bcc3-33ae49081e2d · outbound

This paper cites John Wiley and Sons, New York, 1965.

What AI Red-Team Evaluations Can and Cannot Prove John Wiley and Sons, New York, 1965

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.391557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.391557Z digest=sha256:141bd0c250966e66415e84aa6cd61db6316492958fc114a063f2a1c4d403d030

Observation 6467b702-3a36-4728-acd7-6519488f324e · outbound

This paper cites an unresolved cited work.

What AI Red-Team Evaluations Can and Cannot Prove Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.395950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.395950Z digest=sha256:16303a6df54efaf5ed1ba13b3cd63f7d54caa8cba35b76b736e0636597fac985

Observation ec26b57d-c8a9-40b2-a2eb-5fbf0aad3434 · outbound

This paper cites Model card and evaluations for claude models.

What AI Red-Team Evaluations Can and Cannot Prove Model card and evaluations for claude models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.400904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.400904Z digest=sha256:e1dcd412e9ba88a1363aeb4e386d721029fd62f204e3a9bd94ad21378a7f0e67

Observation 2743c202-13bb-48ad-9ec3-9d6c206cc841 · outbound

This paper cites Sentence-BERT: Sentence embeddings using siamese BERT-networks.

What AI Red-Team Evaluations Can and Cannot Prove Sentence-BERT: Sentence embeddings using siamese BERT-networks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.406066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.406066Z digest=sha256:1e9c11effdd7c219f195e9f54c378e88f768227ff82e6794dfbe1561d2a4643e

Observation 9382bd33-680a-4971-a2c4-76cec48fe29c · outbound

This paper cites LMSYS-Chat-1M: A large-scale real-world LLM conversation dataset, 2024.

What AI Red-Team Evaluations Can and Cannot Prove LMSYS-Chat-1M: A large-scale real-world LLM conversation dataset, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.410572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.410572Z digest=sha256:aa40d43b58e9f494a5403c3b5e9b20034fdcf1409637e82bce218d75487b48bf

Observation 211ea424-422f-4e2d-b5a6-6eaec1d95e12 · outbound

This paper cites Borgwardt, Malte J.

What AI Red-Team Evaluations Can and Cannot Prove Borgwardt, Malte J

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.415116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.415116Z digest=sha256:a9b1577eb206ada888bf9e4bf8827bff52d5e969f36f09e7ebbbe9a7c994965d

Observation d26b55e6-5b01-46cc-9a5f-dc8f24d4fc7a · outbound

This paper cites UMAP: Uniform manifold ap- proximation and projection for dimension reduction, 2018.

What AI Red-Team Evaluations Can and Cannot Prove UMAP: Uniform manifold ap- proximation and projection for dimension reduction, 2018

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.420308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.420308Z digest=sha256:1c1c5252329baf7769074885d1a84491a326e79b943e833645b8a4c7bb91bb1b

Observation 942aafff-26e4-42eb-bb8c-7f19f5c9ef46 · outbound

This paper cites Choquette-Choo, et al.

What AI Red-Team Evaluations Can and Cannot Prove Choquette-Choo, et al

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.426517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.426517Z digest=sha256:0371e571e23f7b7ca393c7515b8edde5a8b4358c4a93d188e7d5756706675607

Observation 1a80966c-fb58-438f-9e1b-4b4c63f0615c · outbound

This paper cites AutoDAN: Generating stealthy jailbreak prompts on aligned large language models, 2024.

What AI Red-Team Evaluations Can and Cannot Prove AutoDAN: Generating stealthy jailbreak prompts on aligned large language models, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.434251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.434251Z digest=sha256:de4ebcd6753372efcaa91492eac9d58a9a151278fb623a99d4169e7c3f4f8769

Observation 7130a205-44c7-4b3b-a148-c24277ff8fe8 · outbound

This paper cites Does refusal training in LLMs generalize to the past tense?, 2025.

What AI Red-Team Evaluations Can and Cannot Prove Does refusal training in LLMs generalize to the past tense?, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.440217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.440217Z digest=sha256:4c3fd066f5e2ddaa0b952eab14a7c6e74571e16f6c1e9d335e888617f044e1b1

Observation 40a1a1e4-0715-4b4e-9a2d-87dba58c35be · outbound

This paper cites an unresolved cited work.

What AI Red-Team Evaluations Can and Cannot Prove Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.446386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.446386Z digest=sha256:be4e1507508d6ac6ee49c7cd64e3f9e0a5a9f6ad52d4aeb3b44335e98671c229

Observation 0eb41bd9-2072-4694-a6f9-e520ee816c51 · outbound

This paper cites Tree of attacks: Jail- breaking black-box LLMs automatically, 2024.

What AI Red-Team Evaluations Can and Cannot Prove Tree of attacks: Jail- breaking black-box LLMs automatically, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.451336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.451336Z digest=sha256:5ad7bb50b548581ce725c8ef2da6639b91de32709dc67785f94db9ab5e1c81e3

Observation 652b69c2-d087-4678-bab5-18eb0e197c1d · outbound

This paper cites Ignore previous prompt: Attack techniques for lan- guage models, 2022.

What AI Red-Team Evaluations Can and Cannot Prove Ignore previous prompt: Attack techniques for lan- guage models, 2022

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.455476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.455476Z digest=sha256:fe5a4796db30598d7a1e75be380e9af3442006149c6d5f079daf57a4ee34374f

Observation 66ea6870-7599-4e90-8976-8d8e4930c06f · outbound

This paper cites GPT-4 technical report, 2023.

What AI Red-Team Evaluations Can and Cannot Prove GPT-4 technical report, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.460464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.460464Z digest=sha256:c434dac3dd0d024ae4bde0c269278c38731618e572388127e0ef6130a411a8a2

Observation d73ba15e-e881-40ee-a011-21c5730c57e0 · outbound

This paper cites GPT-4o system card, 2024.

What AI Red-Team Evaluations Can and Cannot Prove GPT-4o system card, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.464794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.464794Z digest=sha256:e12b7258ebf14af4a77b17dc4af5aaffb342715cbcc846425916318217d64271

Observation 6aeced3a-b254-4c43-9259-387270fb29bd · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

What AI Red-Team Evaluations Can and Cannot Prove The claude 3 model family: Opus, sonnet, haiku

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.469031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.469031Z digest=sha256:8cf4a9e574b8e48d825842964ed13ccf324290804db47cfaf76f409068e59738

Observation b8d0024a-c2f1-4657-910c-c039b8fdf02f · outbound

This paper cites System card: Claude opus 4 and claude sonnet 4.

What AI Red-Team Evaluations Can and Cannot Prove System card: Claude opus 4 and claude sonnet 4

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.473574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.473574Z digest=sha256:a2f33a885b44c9ee6c955bce23da822f2f037a65e5effb219cb2ac69c91f2285

Observation 9a976dac-022b-4d95-a2b6-79f8b8d1df37 · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2023.

What AI Red-Team Evaluations Can and Cannot Prove Gemini: A family of highly capable multimodal models, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.478226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.478226Z digest=sha256:3223972d5eff59979473dcd7346a4af76c9e9ec633e75980332213b70992a6b9

Observation 43d610c4-3c39-4a56-b54c-3403433df20b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024.

What AI Red-Team Evaluations Can and Cannot Prove Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.482631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.482631Z digest=sha256:12f2422e11ed6dd0067641d2f95dd3e912303a65b39e1b6eabe8078f172d09a1

Observation 4954885d-4486-4200-a812-e363906d50eb · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models, 2023.

What AI Red-Team Evaluations Can and Cannot Prove Llama 2: Open foundation and fine-tuned chat models, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.486674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.486674Z digest=sha256:a2933747d598041557c6c40d2fe5c988b7f097249f9c5e6e5147c5c39a0c216f

Observation edfa1227-57cf-4042-8c31-f4869976d8f4 · outbound

This paper cites The llama 3 herd of models, 2024.

What AI Red-Team Evaluations Can and Cannot Prove The llama 3 herd of models, 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.490984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.490984Z digest=sha256:53d57b235131ea6ee029ebdadb0ddd1724478bad98780a5c6166297b4951bd6e

Observation 56e9109f-1aee-4983-9afd-fa36d97d3de9 · outbound

This paper cites Responsible scaling policy.

What AI Red-Team Evaluations Can and Cannot Prove Responsible scaling policy

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.495168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.495168Z digest=sha256:0e16c498143ecbea781a662329ce5d82c366537a40ab752a3d7fb177fae34eaf

Observation 2bacd9a9-bb05-4f97-a16d-74edf5a71c39 · outbound

This paper cites Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022.

What AI Red-Team Evaluations Can and Cannot Prove Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.499159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.499159Z digest=sha256:94be8bf94adaaae8240b12cbf7c45a54254787c875421ca943eb20bce9fd6b76

Observation 8811b220-1865-4f65-acbe-2ba4172ec09b · outbound

This paper cites Red teaming language models with language models, 2022.

What AI Red-Team Evaluations Can and Cannot Prove Red teaming language models with language models, 2022

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.503377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.503377Z digest=sha256:7efacf4c33027c9b1e6cb9803d340b2a2aff9b805609dc88909a48a77c7e4fbf

Pith citing papers

No inbound Pith citation observations are available.