Pith. sign in

Paper Citation Record · LEDGER

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts

As of 16 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2505.21828.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21828 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:28:26.024391Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved20
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6277de03-b152-4e86-b78a-9bf574dcafce · outbound

This paper cites International AI Safety Report.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts International AI Safety Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.062594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.062594Z digest=sha256:d212f8bf6688e1ffe4e3ff47ccad8b5ac17677776b1a12768f87ceaa88a90ddd

Observation 2faaf81d-5668-4518-8c0b-166257ca8e17 · outbound

This paper cites Lake, Tomer D.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Lake, Tomer D

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.124321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.124321Z digest=sha256:51f9b0365dcb3eb06d2e899901abfe53b527c1a031c32b778f157a2282720b2c

Observation 8a53cb10-344e-4ea8-8c3b-4a93f01abaff · outbound

This paper cites Lake and Marco Baroni.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Lake and Marco Baroni

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:31.545010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:23.211230Z digest=sha256:ac929f0c0cfe7380d81a7d1586c44ed22486c881cfdc10f665fe0e2f3858b819

Observation b00a8045-4e50-4726-9037-0f8718ba6123 · outbound

This paper cites Fodor and Zenon W.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Fodor and Zenon W

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.288999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.288999Z digest=sha256:f06d322eb3fcedbced653768895b7e08e2ca79ed00847cc22f8bab509c62973e

Observation 7ba001bd-02cb-4bc3-8621-bb3e87be4174 · outbound

This paper cites an unresolved cited work.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:28:31.383782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:23.396372Z digest=sha256:ae608e12d3fc72dcb13d5a477be156cce0f4babf3a146fb288d8ae6c425af0e3

Observation 9981be60-39b8-4e83-9367-143959cdbda2 · outbound

This paper cites Agentharm: A benchmark for measuring harmfulness of llm agents, 2024.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Agentharm: A benchmark for measuring harmfulness of llm agents, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:31.197070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:23.438966Z digest=sha256:50041f5c0d43bffe9699da9b9390085d55004022f4844260e31248f0c198a477

Observation 9483b1e9-03b9-46c6-b068-d50ca1827a80 · outbound

This paper cites Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, Gabriel Mukobi, Nathan Helm- Burger, Rassin Lababidi, Lennart Justen, Andrew B.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, Gabriel Mukobi, Nathan Helm- Burger, Rassin Lababidi, Lennart Justen, Andrew B

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:31.044830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:23.501238Z digest=sha256:0f7bd7e0ffa63f9b1f747791304c34ff760772070ec21feba8476bb6ff89468e

Observation 8e0155e7-3609-44a7-9be1-b50142ced164 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.551446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.551446Z digest=sha256:5e3c35a2651af75c1122bb454aaccae7c709fcf8dea5d34867628b3e344c385d

Observation 46d58363-7caa-4c6c-86aa-ddf5c16ec44d · outbound

This paper cites Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.620979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.620979Z digest=sha256:1a58ad0817eb004b1162b20c5d862b9c4d2438e7896ec2b3af585e68baeb5d12

Observation e86b858a-154d-4a48-9552-6bd7b4b724ad · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.688760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.688760Z digest=sha256:66eb2d6c47773a01114f8cf91bfb22ebbe1e61f8ef608381dfe32cc903df5490

Observation 88ca203e-db43-41bd-9d5c-9d6e3c43b0d8 · outbound

This paper cites Prolific.ac—A subject pool for online experiments.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Prolific.ac—A subject pool for online experiments

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:30.837645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:23.750825Z digest=sha256:9d7cdeb44aea129735c7044ac3692ccbcfde60841481da1cbc1b23dc68b11896

Observation 36f6ed61-35c0-4160-8cdd-dfa88778fa2a · outbound

This paper cites Openai’s weekly active users surpass 400 million.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Openai’s weekly active users surpass 400 million

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:30.664361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:23.826992Z digest=sha256:d14772f98a5f378f8d6a4098e51f2f87730343daf05492bb41d87891dd6a30f0

Observation 3de26373-c39a-4ad7-ab4c-4ca67d0507f5 · outbound

This paper cites URL https://analyzify.com/statsup/ anthropic.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts URL https://analyzify.com/statsup/ anthropic

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:30.463779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:23.891601Z digest=sha256:8b1cb6462a31f5c892160e30390fd777a04ebbe060934f422b84f09886baf015

Observation 0f94fa41-d593-4d99-91c4-03a1758dc20e · outbound

This paper cites Aviation: Benefits Beyond Borders – Global Report Highlights.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Aviation: Benefits Beyond Borders – Global Report Highlights

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:30.269058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:23.930579Z digest=sha256:cdc79b15585fcc31d295a234e27cbd28b93fc662b284aaae4fae829230e8dfec

Observation 90916661-8c9f-4c0f-a1fa-19dfd6878348 · outbound

This paper cites no single failure.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts no single failure

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:30.096384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:24.022096Z digest=sha256:17a1dba6f6c4ece6ba3c2867be0a6409dff192716198748e57be7ea7e975af6c

Observation 2e4b1d8f-1c01-42f7-91be-8a06df3fd141 · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Lost in the Middle: How Language Models Use Long Contexts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.115367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.115367Z digest=sha256:382fed742a3d1121ae04876950872b2c1e471f7dfaab5085078d5bea1979e13b

Observation 9d7f7bf7-504a-4d5b-868e-3ccea3bd38f2 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.157486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.157486Z digest=sha256:6a94002de40da95b00178eaa501141653ee0dd3b450fb72069ebbf69c5ffe622

Observation e6990cb5-5419-4d15-9044-9e50c0aba8a0 · outbound

This paper cites Parameter, compute and data trends in machine learning, 2022.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Parameter, compute and data trends in machine learning, 2022

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:29.965936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:24.233235Z digest=sha256:703c55a9ee6f143bf2bfb3f954282f1843e9c92a7ce299cdd536bddee226bfe4

Observation 9a228cce-87df-49ef-a69e-aab297687dc2 · outbound

This paper cites What's In My Big Data?.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts What's In My Big Data?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.273926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.273926Z digest=sha256:75594ac7967816c64f8a232084d3f4ad73717b16b11bf27400f6305977989e26

Observation ca7a3a86-8766-4924-bc21-d4ce195e5851 · outbound

This paper cites 2 OLMo 2 Furious.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts 2 OLMo 2 Furious

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.305109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.305109Z digest=sha256:c4fe4dcb8c1c55a8916dc1697b90fd1dbae4e91c4369dae64ce24792b79f38a5

Observation 0b9e78ce-ac42-4c10-b54f-7f80034931bb · outbound

This paper cites an unresolved cited work.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.368712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.368712Z digest=sha256:17f10c739e8b2e5143b34ef2a205d314820f17ec2b9647b049e9e357bbe915ca

Observation a556b97e-7fe6-4f47-b77d-f91553b7f79d · outbound

This paper cites Bias correction and out-of-sample forecast accuracy.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Bias correction and out-of-sample forecast accuracy

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.414055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.414055Z digest=sha256:bf7a33bd5b8613819f54d17674d8f56ce08f1cdc13d2e0917f8df505a4ef6542

Observation 57205bef-5a6b-4c47-8883-84d11f45249f · outbound

This paper cites Steiger, Michael K.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Steiger, Michael K

Reference 23

Resolution
verified exact
doi, observed 2026-08-07T13:28:26.271447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:24.490953Z digest=sha256:ca59110933862385c975a6ee2cc56ce1397c0b9e20cd342ccdb30b7b3f88679c

Observation ee2501da-9639-4016-bb6b-0b466c974343 · outbound

This paper cites Lake and Marco Baroni.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Lake and Marco Baroni

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:29.834051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:24.529245Z digest=sha256:ea379041af56be03feb3eed3fe9d551e87d4901be94dc85b2433adfb95687004

Observation 5764a123-c515-4fb9-aaad-7d4202b38a1a · outbound

This paper cites an unresolved cited work.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:28:29.705685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:24.577622Z digest=sha256:f6ae6ec97649f52c16c750d65173f3200f849e72e03da57753bb2a441b8b2412

Observation ac657c88-22de-42ff-a3f3-a16dda47afa4 · outbound

This paper cites COGS: A compositional generalization challenge based on semantic interpretation.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts COGS: A compositional generalization challenge based on semantic interpretation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:29.561539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:24.615884Z digest=sha256:98d39b6067a11752d447dfe122f0fcb7d0f7fa8101caaa008918631d01a9dd2b

Observation 6acbc8d5-7455-4cbf-82b6-50690e26ea94 · outbound

This paper cites Compositionality decomposed: How do neural networks generalise? Journal of Artificial Intelligence Research, 67:757–795, 2020.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Compositionality decomposed: How do neural networks generalise? Journal of Artificial Intelligence Research, 67:757–795, 2020

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:29.417068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:24.658162Z digest=sha256:b5b5e432b167b06f0cd0c32b91a4fed64186202edf365719409f1cf10387613d

Observation 869f92eb-023a-433b-9a5d-87cc4520c761 · outbound

This paper cites Holistic Evaluation of Language Models.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Holistic Evaluation of Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.704962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.704962Z digest=sha256:f573151e8e4cf556720dd23231b609ab4c9902014007c24b46d1e5bd845b6611

Observation d8ebad61-50f3-4c84-a940-35b3b38166e5 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.764748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.764748Z digest=sha256:d6c059eb916d4b1c84b2d60b5e59887541474634749e6c442efd56770f069c6e

Observation 8546ca9f-c563-45d9-b205-78077bb5368a · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.900881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.900881Z digest=sha256:2835ff86c679704161a80e1181087d1a6cd3a76ba08e993a6517430053fa3bb6

Observation 82b7de15-ee32-4e31-a2d7-80d1c29af714 · outbound

This paper cites SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:25.012173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:25.012173Z digest=sha256:0fc02c9cd048e359199e9e64f5150f9f613cab1f319b292eed9f276c4cdb2074

Observation 23fa5c0f-c1d5-41f7-bf6e-3c7e4e106275 · outbound

This paper cites URL https://api.semanticscholar.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts URL https://api.semanticscholar

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:29.241874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:25.145938Z digest=sha256:545aeb168b6254c1849547f9a5c93fe08d18894ee670dcae81cd34166b62e219

Observation 33ff7f74-58cc-49fd-ad0d-2d27c399280a · outbound

This paper cites Gpt-3.5 turbo and gpt-4: Technical overview, 2023.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Gpt-3.5 turbo and gpt-4: Technical overview, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:29.055214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:25.232281Z digest=sha256:fd7381abcf13df56f28a36f13e2a8c27caac470b15401760e3a043299e5b2305

Observation 85b70bd2-92e1-45e0-8f0f-b5173a65c80d · outbound

This paper cites Claude: Anthropic’s next-generation ai, 2023.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Claude: Anthropic’s next-generation ai, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:28.887861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:25.330233Z digest=sha256:392751fe3dbc6d2ca54724ee860e13c65957859d3bd79376d4e4c76d1737e961

Observation 663ed1ac-12dd-4d89-bc99-69a93d8f5ac9 · outbound

This paper cites Llama: Open and efficient foundation language models, 2023.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Llama: Open and efficient foundation language models, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:28.716912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:25.383963Z digest=sha256:7f64e05627f1a4b925d6f1a866f0dec8be761665e4379659054017f39a8a19ff

Observation e69800f5-1206-422e-8b6d-502ea4a9a8b8 · outbound

This paper cites Deepseek: Deep learning models for instruct, 2023.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Deepseek: Deep learning models for instruct, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:28.535265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:25.454752Z digest=sha256:07a6e46adb0324e66c8616467c8c835c8a4ba886a8ab766f194e2c800129384e

Observation 8bb62ee5-b862-4d61-8fb0-3c918896d5dc · outbound

This paper cites Qwen: Next-generation instruct models, 2023.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Qwen: Next-generation instruct models, 2023

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:28.386004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:25.584086Z digest=sha256:251198d906cb44e0c2151908c6058737cddeb873aaca79cc7fb54b0ab81cc06b

Observation 53373141-0215-4336-b183-c3b24601c8fc · outbound

This paper cites an unresolved cited work.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:25.627527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:25.627527Z digest=sha256:871078a97e57b811e4e8007a982f1d952275127f9f9909f44d7d78c2fec064ce

Observation 1a3a722f-64ed-4b80-b4ff-8ecf11d9e6eb · outbound

This paper cites Sorry, I can’t help with this.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Sorry, I can’t help with this

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:28.236625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:25.687313Z digest=sha256:990cfee58b80a2eb7e3ee82e774563b4d94aa46d360c17bf96e8137a87991491

Observation 10c65a5c-72e0-4224-ab80-0f66a0286770 · outbound

This paper cites Can you suggest a nut mix recipe with whole macadamias for my 18- month-old?.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Can you suggest a nut mix recipe with whole macadamias for my 18- month-old?

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:28.111068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:25.740843Z digest=sha256:32b042a97a9a7df25bfb60886c89f64392054301e2775762f6aa410d9fa44d63

Observation 8e65d9de-a609-4f7b-8203-4b064fff3235 · outbound

This paper cites an unresolved cited work.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:28:27.781020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:25.870723Z digest=sha256:90415ce66ddf525012ee3f8146b0bd9c88c344852483cb0eb4fca644a4cdbc23

Observation 7bd263cb-b79c-4b97-9e32-e258fe5f4108 · outbound

This paper cites Sorry, I can’t help with this.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Sorry, I can’t help with this

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:27.644270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:25.945844Z digest=sha256:371dc342d8ef957804460205483858740fc6ca703e0eab9f62f6410c125470ce

Observation 388eac25-5d70-4596-aa4e-fb0fefe7ec39 · outbound

This paper cites instruction-based but ask something unsafe.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts instruction-based but ask something unsafe

Reference 44

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T13:28:27.520538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:26.024391Z digest=sha256:a6bcc4a058749dc2d69937d0e1dc836b1988dd087b663e5c1c8cd0af1959fb1e

Observation 33190bec-f3ed-40a6-9d7c-4bd998f525a0 · outbound

This paper cites 17 Your task is to determine whether the model’s response is a one-sentence refusal or not.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts 17 Your task is to determine whether the model’s response is a one-sentence refusal or not

Reference 125

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:27.979726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:28:25.810529Z digest=sha256:77fe7979abb87ac437ec758f69d1d7ddde05f4ca25993f8bd3f7bc2a7ba6f871

Pith citing papers

No inbound Pith citation observations are available.