Pith. sign in

Paper Citation Record · LEDGER

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond

As of 6 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 2 inbound Pith citation observations for arXiv:2603.01589.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.01589 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T16:21:25.563051Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact16
  • verified fuzzy2
  • unresolved7
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69bccb58-f08e-485b-ba64-87e140781be5 · outbound

This paper cites an unresolved cited work.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-15T18:01:26.341189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:fc2b986365f55ac44d4c0725c732da29eff736e1e43871da32aa7e53ba850d1c

Observation a0463868-eace-4862-a74e-b46ceda5927a · outbound

This paper cites We split all 125 tasks into two sets,Gated Public Access setand HighRisk Restricted Access set.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond We split all 125 tasks into two sets,Gated Public Access setand HighRisk Restricted Access set

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:01:26.345078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:aa9a08dad16d932b4abe7eb7e54a681c1aeb58798866613dd03c81c4eb538c35

Observation 50943350-583d-436c-9543-29f2b5d3c95e · outbound

This paper cites an unresolved cited work.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Unresolved cited work

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:01:25.541805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:b98698dab729a1b53a1779723a214cd622979ade6ead8811c8b03f4b5f047b89

Observation b2589ea2-7fa2-4c2d-92ff-72d21cafe9ad · outbound

This paper cites Intern-s1: A scientific multimodal foundation model.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Intern-s1: A scientific multimodal foundation model

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:01:25.535498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:fa586b3fd361fd2411bfbe73c99cc24f0a3a529aeb50a95c678097f1a9f88565

Observation 472464fd-c3a2-4519-919b-01ee2029bf4f · outbound

This paper cites Why Should Adversarial Perturbations be Imperceptible? Rethink the Research Paradigm in Adversarial NLP.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Why Should Adversarial Perturbations be Imperceptible? Rethink the Research Paradigm in Adversarial NLP

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:01:25.528515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:641ac8efde40bc7608b66de52d391513941c33039a5ed3ffbc776eb7a8bbb533

Observation 0fe1c598-9fcc-4c8e-916e-a273ed8c76d0 · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T00:47:04.366365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:f40df424f89d7327744aa134b8b0261a0c55d6ac842619e55e2177dcf2a18308

Observation 7a1ead7c-9d98-424e-8806-31927ff8f8a4 · outbound

This paper cites The Llama 3 Herd of Models.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond The Llama 3 Herd of Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:01:25.554485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:fc26d28bf6d382bb19834eace7c163690ec171695c562bc5edf7ef69fd41d839

Observation a12b6004-4cd2-4988-88ac-fc6ef244fb63 · outbound

This paper cites 2024 , url =.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond 2024 , url =

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:01:25.521960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:0172deec313862eaf7a33aadd312c08676a51fb25c21b1523eed6e3777c9da2f

Observation 36667f83-b39d-4a1a-bc61-5cc16d782a26 · outbound

This paper cites an unresolved cited work.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-15T18:01:26.336964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:6edd65d97bd4704609d39f33b226d7d0c1d0c868af40cd31aea04afb632e21ba

Observation 678099fb-4732-426b-b9a1-8f0e79af9e86 · outbound

This paper cites Control Risk for Potential Misuse of Artificial Intelligence in Science.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Control Risk for Potential Misuse of Artificial Intelligence in Science

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:01:25.547934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:cbb586650658366614eefb7094e9b1735ae4de14af41d57658c28c5aab26a5db

Observation 905da372-f9b8-4398-9af3-2bda31f53ec4 · outbound

This paper cites an unresolved cited work.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-15T18:01:26.332750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:7404ea480a90f8b34bb5d433552e764ae3cf3526a5ac7d719c8322cc616b3a72

Observation 6716e6ff-3368-4e39-8e49-7a5bc2dbdcbe · outbound

This paper cites SoSBench: Benchmarking Safety Alignment on Six Scientific Domains.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond SoSBench: Benchmarking Safety Alignment on Six Scientific Domains

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:01:25.616943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:e04be4e2a205142be2be79de06ef991edef2eb4bf1e48388be0e37312894c456

Observation 3d135be1-68bf-4279-b346-7668da1352cb · outbound

This paper cites an unresolved cited work.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-15T18:01:26.319892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:5b46ad5245c87b9634a9d8b8374bfdb1b5e3fc289efbb65d99e948679c078b57

Observation 07eafd48-7b4a-4fcf-8c25-f10411897d8b · outbound

This paper cites an unresolved cited work.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-15T18:01:26.315522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:be48167feab81920332aa67e9264d85dea583f2ec9cbcaf14aaa5ea74913bff6

Observation 4abfd7fe-5cb2-482c-9ca5-042db629d440 · outbound

This paper cites SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:01:25.610609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:0eb1a9ba05645a23acd75692f64f4c9ea6865705a40967ed4155d995bac4fd37

Observation c8b38e62-a389-450e-970f-bb5c12691c84 · outbound

This paper cites The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:55:50.014661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:7dedd3dd04e0aeaf36458488c5fbfefb5ffaa50d9e71e3ae75adfd5b951fcb1b

Observation 7f150f46-66c7-400c-8f0d-d602b34f1a6e · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:01:25.623098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:f32bfc6e17f204d67dcbc8d1c5b27918226d595be749a12bdcc240435262ed40

Observation 7a9f240b-0e1a-4899-9896-482ff365bc62 · outbound

This paper cites Probing and Steering Evaluation Awareness of Language Models.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Probing and Steering Evaluation Awareness of Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:01:25.603359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:0e054dbfb7ec9c4c1dc6795bb300c7f166d63d16f660d30ba4eea670e3c4ab83

Observation b6022c6a-5560-4088-9723-2352897489c3 · outbound

This paper cites an unresolved cited work.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-15T18:01:26.328132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:203554a2dcb2efdea995b9585a61be8211115067433cbd1bb5bf287117269a82

Observation 614640aa-d0a4-43d6-bf38-ab4e4f7d7006 · outbound

This paper cites an unresolved cited work.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-15T18:01:26.323679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:a710f1a1d6c34d68be3d22bbaa3439b40f5ba407a8c64f270f3658d8d37c1f6c

Observation 6bb84f98-db4b-472d-beeb-68fab9f77420 · outbound

This paper cites Qwen3 Technical Report.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Qwen3 Technical Report

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:01:25.596110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:b395e902709911f4dba1ce2b342a22b225257f26d19cef94b76f94e719001eba

Observation cee15327-83a4-414a-b92b-f201f669c678 · outbound

This paper cites Accessed: 2026-01-29.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Accessed: 2026-01-29

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:01:26.310950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:e7aa8d68517a9bf3f4df679083d4fdb07f4a344817d4eeb9e1eb2f75f7e09913

Observation 93886e3c-a0e1-4be5-8bce-e27eeb57d4ac · outbound

This paper cites Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:01:25.243853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:f0d7a882c15271b0808fadf405b9499cc428ac71d2f749d519ed123d711abe33

Observation dbee5bae-42ba-40d1-a12f-3e827a123fd9 · outbound

This paper cites SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:01:25.575015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:04f05ca7260077ef3e48747240a303eafc6fc0e24506d7f3e0f74ea6fbf64aba

Observation f6f258be-97c4-4423-ad54-4beea4f501d0 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:01:25.588604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:3e40dd59da98a474a68350dcb6beeed88c745de9732deb48c12e77d87d56f3bc

Observation af3b1add-78f6-45e4-a25c-52dd2c13dea4 · outbound

This paper cites Ans.” represent if the questions have corresponding answers. “Rep.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Ans.” represent if the questions have corresponding answers. “Rep

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:01:25.581771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:7e51e086143af348a14bcd20cacff0556e4681ca2431b76fadf820891c3157a3

Observation 0fe6894c-4ced-4bb3-82ff-d16ee20b6920 · outbound

This paper cites an unresolved cited work.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Unresolved cited work

Reference 28

Resolution
malformed identifier
raw_fallback, observed 2026-05-15T18:01:26.349334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:3dad12a26942eaea5c9e2bcacc53f7b3b8ee3e219c16f8ed37b251965cc834ec

Pith citing papers

Observation affd77b5-38e3-41fb-8a5b-5339ac7598ec · inbound

An Early Warning of Emerging Biosecurity Risks in Frontier LLMs cites this paper.

An Early Warning of Emerging Biosecurity Risks in Frontier LLMs SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T16:21:25.563051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:21:25.563051Z digest=sha256:cb80d632499bd041f3d8d869c261b247ddf28815cff91918aa4fc0c8beec5796

Observation dfb99e88-fc71-4050-b015-9d7219c9c2a6 · inbound

SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring cites this paper.

SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T14:45:44.097098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:45:44.097098Z digest=sha256:3fa01f4c228940f8f23b1f562d05404eb9ef03878daad43197061368c5e2eb96