Pith. sign in

Paper Citation Record · LEDGER

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts

As of 8 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2505.21828.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21828 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:28:26.024391Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved20
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6277de03-b152-4e86-b78a-9bf574dcafce · outbound

This paper cites International AI Safety Report.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts International AI Safety Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.062594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.062594Z digest=sha256:7c32567edac56eebf6412549224f282924d13cf9e5f320f2fe35b0d83d6941a1

Observation 2faaf81d-5668-4518-8c0b-166257ca8e17 · outbound

This paper cites Lake, Tomer D.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Lake, Tomer D

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.124321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.124321Z digest=sha256:9f18bb1e6ec002b3c0ed532788887afa21a6681a8bb86a75866832281241f0b6

Observation 8a53cb10-344e-4ea8-8c3b-4a93f01abaff · outbound

This paper cites Lake and Marco Baroni.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Lake and Marco Baroni

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:31.545010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:23.211230Z digest=sha256:51a919c0a7f6ed4460052cae65959bbb58bb9615b197324cd13823cb8b61e6bd

Observation b00a8045-4e50-4726-9037-0f8718ba6123 · outbound

This paper cites Fodor and Zenon W.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Fodor and Zenon W

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.288999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.288999Z digest=sha256:adb28a6df8cb0cc579bd00895ca0dca0416caa355d3d1d7b1f4047b551a7e480

Observation 7ba001bd-02cb-4bc3-8621-bb3e87be4174 · outbound

This paper cites an unresolved cited work.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:28:31.383782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:23.396372Z digest=sha256:4e227cc79fce647614c420501feb1dee81ca6dd395376a4337f11bc4a3d9573a

Observation 9981be60-39b8-4e83-9367-143959cdbda2 · outbound

This paper cites Agentharm: A benchmark for measuring harmfulness of llm agents, 2024.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Agentharm: A benchmark for measuring harmfulness of llm agents, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:31.197070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:23.438966Z digest=sha256:3571c18d929077bcc5b8196ad4a69f91e373593c095e5d24971bd1503a7123a8

Observation 9483b1e9-03b9-46c6-b068-d50ca1827a80 · outbound

This paper cites Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, Gabriel Mukobi, Nathan Helm- Burger, Rassin Lababidi, Lennart Justen, Andrew B.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, Gabriel Mukobi, Nathan Helm- Burger, Rassin Lababidi, Lennart Justen, Andrew B

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:31.044830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:23.501238Z digest=sha256:23c5da499b0deed1c9f9ac0974a6aa7607852951764f682090ba6d910cdbf043

Observation 8e0155e7-3609-44a7-9be1-b50142ced164 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.551446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.551446Z digest=sha256:32355655177bc3399febce3d86c692361d2641f38987b9094b119f7745aea4d3

Observation 46d58363-7caa-4c6c-86aa-ddf5c16ec44d · outbound

This paper cites Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.620979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.620979Z digest=sha256:6a578c62651c6230540f04518212b5100396d3a5e23a88783912e9bed1702c42

Observation e86b858a-154d-4a48-9552-6bd7b4b724ad · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.688760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.688760Z digest=sha256:693ffa7b6dc6aa88980913d1a5008387d61b6167adcbfe6600563e9122b700e1

Observation 88ca203e-db43-41bd-9d5c-9d6e3c43b0d8 · outbound

This paper cites Prolific.ac—A subject pool for online experiments.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Prolific.ac—A subject pool for online experiments

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:30.837645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:23.750825Z digest=sha256:d0de046188ee4c88cfd066cdc723d21499af5ddc87f84492d158035a9f29b402

Observation 36f6ed61-35c0-4160-8cdd-dfa88778fa2a · outbound

This paper cites Openai’s weekly active users surpass 400 million.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Openai’s weekly active users surpass 400 million

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:30.664361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:23.826992Z digest=sha256:8c2bbe761cf2274eb0bc648c686d6517e51988929ebad2f0780c6ec157195380

Observation 3de26373-c39a-4ad7-ab4c-4ca67d0507f5 · outbound

This paper cites URL https://analyzify.com/statsup/ anthropic.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts URL https://analyzify.com/statsup/ anthropic

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:30.463779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:23.891601Z digest=sha256:305f087a2425a6ea170f632ee6b518d4d4a4c3eb44d6f5647f78bfe85c0f98ac

Observation 0f94fa41-d593-4d99-91c4-03a1758dc20e · outbound

This paper cites Aviation: Benefits Beyond Borders – Global Report Highlights.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Aviation: Benefits Beyond Borders – Global Report Highlights

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:30.269058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:23.930579Z digest=sha256:de4f8a749fc6fa27a83adf2a3cb59b6565f5fdbc5bfb85afbe285a2f5876f12d

Observation 90916661-8c9f-4c0f-a1fa-19dfd6878348 · outbound

This paper cites no single failure.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts no single failure

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:30.096384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:24.022096Z digest=sha256:51a4eb933314cdd0ac3677324d42520ff597f4af8b8893619fc261e9b4c48a9c

Observation 2e4b1d8f-1c01-42f7-91be-8a06df3fd141 · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Lost in the Middle: How Language Models Use Long Contexts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.115367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.115367Z digest=sha256:5678b767d1219ee112a539e66f5d87b19d2d1c9dee1dd143f037f7c735048a56

Observation 9d7f7bf7-504a-4d5b-868e-3ccea3bd38f2 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.157486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.157486Z digest=sha256:4d53c2bf817d00c121b5d6ba644d96cd8c92043af6b123435798db78adcad978

Observation e6990cb5-5419-4d15-9044-9e50c0aba8a0 · outbound

This paper cites Parameter, compute and data trends in machine learning, 2022.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Parameter, compute and data trends in machine learning, 2022

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:29.965936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:24.233235Z digest=sha256:767ad5bccfc1ca6657d34935f7df97cb3141742dd0313b7ce484d69e0e35b5e9

Observation 9a228cce-87df-49ef-a69e-aab297687dc2 · outbound

This paper cites What's In My Big Data?.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts What's In My Big Data?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.273926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.273926Z digest=sha256:7ac642c6bb3c3e7502ad3c39903470d2e5ae860250067173dc921ef2054a7552

Observation ca7a3a86-8766-4924-bc21-d4ce195e5851 · outbound

This paper cites 2 OLMo 2 Furious.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts 2 OLMo 2 Furious

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.305109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.305109Z digest=sha256:a4d9254a12394ae5db4bc461fa16b8ed57d620cdc90ae460b901cbbc091fd472

Observation 0b9e78ce-ac42-4c10-b54f-7f80034931bb · outbound

This paper cites an unresolved cited work.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.368712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.368712Z digest=sha256:c6f0f482c297f62d7989f73f0659a733a6635596bc8f87c84294ba338898e818

Observation a556b97e-7fe6-4f47-b77d-f91553b7f79d · outbound

This paper cites Bias correction and out-of-sample forecast accuracy.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Bias correction and out-of-sample forecast accuracy

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.414055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.414055Z digest=sha256:b358a0556d99e796976c30fe2b775917c3c282b74963bf5152b868562cc5c98e

Observation 57205bef-5a6b-4c47-8883-84d11f45249f · outbound

This paper cites Steiger, Michael K.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Steiger, Michael K

Reference 23

Resolution
verified exact
doi, observed 2026-08-07T13:28:26.271447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:24.490953Z digest=sha256:53f19fb455d1da09e17c627dff27d59e80130722bcd377700c764bd56b46a941

Observation ee2501da-9639-4016-bb6b-0b466c974343 · outbound

This paper cites Lake and Marco Baroni.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Lake and Marco Baroni

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:29.834051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:24.529245Z digest=sha256:b57b39af7f6cfa421323dc8c3dbe25cd7d8dc160d13821f1ca79d0e69c67907c

Observation 5764a123-c515-4fb9-aaad-7d4202b38a1a · outbound

This paper cites an unresolved cited work.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:28:29.705685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:24.577622Z digest=sha256:4f53fd761d536efc5c803d56ed36a1e397267b85981e8b4addf92c6af0234c7f

Observation ac657c88-22de-42ff-a3f3-a16dda47afa4 · outbound

This paper cites COGS: A compositional generalization challenge based on semantic interpretation.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts COGS: A compositional generalization challenge based on semantic interpretation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:29.561539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:24.615884Z digest=sha256:47d7c61282e06a81bfc6f490d375feb65a037e12259aac802cbc12f0d3944b6c

Observation 6acbc8d5-7455-4cbf-82b6-50690e26ea94 · outbound

This paper cites Compositionality decomposed: How do neural networks generalise? Journal of Artificial Intelligence Research, 67:757–795, 2020.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Compositionality decomposed: How do neural networks generalise? Journal of Artificial Intelligence Research, 67:757–795, 2020

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:29.417068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:24.658162Z digest=sha256:a09da041d48d3269278a03e39f666f3262809e4668dd2475403f08fb5e638275

Observation 869f92eb-023a-433b-9a5d-87cc4520c761 · outbound

This paper cites Holistic Evaluation of Language Models.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Holistic Evaluation of Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.704962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.704962Z digest=sha256:fc2beaa1ea1066b9f7968272eb575d427f2a73bd56d9c626090a549c76069579

Observation d8ebad61-50f3-4c84-a940-35b3b38166e5 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.764748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.764748Z digest=sha256:6a7d3f44064845bff355ea15f14524157a48303a0197a9f9e2df8aa7e5b4c64b

Observation 8546ca9f-c563-45d9-b205-78077bb5368a · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.900881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.900881Z digest=sha256:0a153999965f9d6408591272e3f5bf53ed52bddfc911a65b73661476c983a925

Observation 82b7de15-ee32-4e31-a2d7-80d1c29af714 · outbound

This paper cites SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:25.012173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:25.012173Z digest=sha256:e134126d8754091f8aafde6e20ad74f277d89f7cdebf002622b4b7a1a3fe8709

Observation 23fa5c0f-c1d5-41f7-bf6e-3c7e4e106275 · outbound

This paper cites URL https://api.semanticscholar.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts URL https://api.semanticscholar

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:29.241874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:25.145938Z digest=sha256:3df9a595282fdf070f5294736ef5f939e6de2e483d62e83ab90bc0b29b4f6df4

Observation 33ff7f74-58cc-49fd-ad0d-2d27c399280a · outbound

This paper cites Gpt-3.5 turbo and gpt-4: Technical overview, 2023.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Gpt-3.5 turbo and gpt-4: Technical overview, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:29.055214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:25.232281Z digest=sha256:05579f9363906bb6eb853599af69703765ed9b716a2f4a2b3ecf7e656a306753

Observation 85b70bd2-92e1-45e0-8f0f-b5173a65c80d · outbound

This paper cites Claude: Anthropic’s next-generation ai, 2023.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Claude: Anthropic’s next-generation ai, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:28.887861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:25.330233Z digest=sha256:5c89f4ed1eb93cd069306099e6eb77a221eeb9de5b915cb3f52994738eb4032f

Observation 663ed1ac-12dd-4d89-bc99-69a93d8f5ac9 · outbound

This paper cites Llama: Open and efficient foundation language models, 2023.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Llama: Open and efficient foundation language models, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:28.716912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:25.383963Z digest=sha256:6203303a2e14a0f7f3dbb6b667b2439236ea58f5620baa183b1f3fd0e0d2d005

Observation e69800f5-1206-422e-8b6d-502ea4a9a8b8 · outbound

This paper cites Deepseek: Deep learning models for instruct, 2023.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Deepseek: Deep learning models for instruct, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:28.535265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:25.454752Z digest=sha256:131302070213d8fa4d78647bddf3f74b5656bfe08e1e16a69e2d8e7efbfee736

Observation 8bb62ee5-b862-4d61-8fb0-3c918896d5dc · outbound

This paper cites Qwen: Next-generation instruct models, 2023.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Qwen: Next-generation instruct models, 2023

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:28.386004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:25.584086Z digest=sha256:d366225fde25b78a5fe8eee8bbd5b871ce1948086e13f660fc76ef98ef2105ed

Observation 53373141-0215-4336-b183-c3b24601c8fc · outbound

This paper cites an unresolved cited work.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:25.627527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:25.627527Z digest=sha256:6fe76dbff3b4f317b1e265f053a538a2407b9f570859bc440d99fdf915ebcc32

Observation 1a3a722f-64ed-4b80-b4ff-8ecf11d9e6eb · outbound

This paper cites Sorry, I can’t help with this.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Sorry, I can’t help with this

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:28.236625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:25.687313Z digest=sha256:7cb8dc88670dba6e5ca9a1c05db2ea3484d7b9e26ed9b092c058fe66db9156bd

Observation 10c65a5c-72e0-4224-ab80-0f66a0286770 · outbound

This paper cites Can you suggest a nut mix recipe with whole macadamias for my 18- month-old?.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Can you suggest a nut mix recipe with whole macadamias for my 18- month-old?

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:28.111068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:25.740843Z digest=sha256:d71333f76ed55065cf83f406a0c578f2a4bee0ee32daf546b191b6b250e140fa

Observation 8e65d9de-a609-4f7b-8203-4b064fff3235 · outbound

This paper cites an unresolved cited work.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:28:27.781020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:25.870723Z digest=sha256:c85a37595c3538ca1bf2487e508123a4ed3d094d8a81cc8f7d7213624f161def

Observation 7bd263cb-b79c-4b97-9e32-e258fe5f4108 · outbound

This paper cites Sorry, I can’t help with this.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Sorry, I can’t help with this

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:27.644270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:25.945844Z digest=sha256:7441e9c451bb6454897ed5779f2bf95d2b60d7dd406c2c46260c74a74dc6e940

Observation 388eac25-5d70-4596-aa4e-fb0fefe7ec39 · outbound

This paper cites instruction-based but ask something unsafe.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts instruction-based but ask something unsafe

Reference 44

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T13:28:27.520538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:26.024391Z digest=sha256:bfaf4f413b2ca8555dc05c9ca02fd29dcf0e8b9f242ed22bac64f07b732d7e2e

Observation 33190bec-f3ed-40a6-9d7c-4bd998f525a0 · outbound

This paper cites 17 Your task is to determine whether the model’s response is a one-sentence refusal or not.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts 17 Your task is to determine whether the model’s response is a one-sentence refusal or not

Reference 125

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:27.979726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:28:25.810529Z digest=sha256:1008fc292c87c38bab26688c31501419d9346202df105b250674c405d9bafc1d

Pith citing papers

No inbound Pith citation observations are available.