Pith. sign in

Paper Citation Record · LEDGER

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use

As of 18 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2505.17332.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17332 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:52:31.262024Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26b1aae7-672a-45e4-b9c2-5039e9325947 · outbound

This paper cites online" 'onlinestring :=.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:26.322302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:26.322302Z digest=sha256:7d4c075492397c49cd08366ff51b8da5ee1e266e81f972c343af0842790dbae8

Observation 7c087a5a-b0ad-4bcf-b766-8d2e8440c8da · outbound

This paper cites write newline.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:26.464178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:26.464178Z digest=sha256:5dc2c31fac3c6fbebfebb959e2df4c7aa14ad82127586059a7575702c33c356d

Observation 64a2ae5a-527d-47ca-b002-dc740e031058 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:26.615041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:26.615041Z digest=sha256:3d3f6d1aa755169d0cc94165db5472a04a3890c3570943712dc510ca37ccb22d

Observation c2631a25-09ad-4c04-a86b-ff246c00d569 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:52:33.061472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T14:52:26.750767Z digest=sha256:78476ad63331b8ad8a76d95e655a03f249934a5e1a5aef4b8bee43c8e5c0600a

Observation d0965ad8-1162-4095-8b4a-402b8dc9755b · outbound

This paper cites MVTamperBench: Evaluating Robustness of Vision-Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use MVTamperBench: Evaluating Robustness of Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:26.868304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:26.868304Z digest=sha256:7e353d0f6f5e76368f43ed7d73f97f18878b52a5231768e75f2e2fd16dbf7953

Observation e895a1df-56ae-4bd4-a19c-9c5d42c6c30f · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:52:32.928539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T14:52:26.976950Z digest=sha256:9987f7d82bf3a7356c7b368407eab993f67bab38ca846522125727583ac480d4

Observation eb97f4e2-8ef0-4ace-883d-09589e055212 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:52:32.801949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T14:52:27.115403Z digest=sha256:8deddf9f13918fd66fd60b2d66a29eb32c7658651bf51a58549c35fb673bc23a

Observation 550c9ee5-1b49-4834-93b7-21747f8a4949 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.165158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.165158Z digest=sha256:8fe5c3f83144ead739dd1ff190cfce3fa64d7405bdc46c4ee5fcee77d1585be9

Observation 6ee05890-3ee7-49f1-ac40-767856edd6ea · outbound

This paper cites Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.217122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.217122Z digest=sha256:cf85f2dd03b6aa385c041da7a5321be59a312c0fc40e2b77068bc3f4873263cc

Observation 207e7e05-166f-44cc-b4a5-d9f29d35b8d7 · outbound

This paper cites Do, Yan Xu, and Pascale Fung.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Do, Yan Xu, and Pascale Fung

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.279376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.279376Z digest=sha256:8144b7464384fbad71d7333fee2afa269ab11c792e0c46f4eccfe17b5a9314ac

Observation 30a523f3-53cd-48df-9482-3c14107d1c63 · outbound

This paper cites Language Models are Few-Shot Learners.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Language Models are Few-Shot Learners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.438818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.438818Z digest=sha256:8541e637c10759a00e024bcdcab58b6343d3453aff8f0427e331b2400f19bd62

Observation a56ae450-2731-4a80-a54a-990f265db275 · outbound

This paper cites JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.510965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.510965Z digest=sha256:5b0e4de845ecd1b6b4c28ecd5fd01d72d221e12c04bf229c400b0858c5451e23

Observation f7d9ba58-cc46-45fd-a320-59388953a377 · outbound

This paper cites Crosslingual Capabilities and Knowledge Barriers in Multilingual Large Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Crosslingual Capabilities and Knowledge Barriers in Multilingual Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.565303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.565303Z digest=sha256:3423eee3cbfc3b018a3a62e5f784a34aabe197afe4a87bad8ad35e608981dea2

Observation 7f04b227-b7e4-436e-b7d4-918c16af380e · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.646088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.646088Z digest=sha256:c1ee2d7b96704c20038869d7f8601fa8915cf0b3b386ac9449ac8e7c86a8653f

Observation a93f0125-3504-4efb-a012-be1992f6f87d · outbound

This paper cites The Llama 3 Herd of Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.742920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.742920Z digest=sha256:176c1cdab9d4169d272ecb7d242948083e01b003939d551aff6ed3d9b061c5da

Observation 4354f0a9-961f-4be5-b27d-78afba2db335 · outbound

This paper cites Multilingual Large Language Models and Curse of Multilinguality.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Multilingual Large Language Models and Curse of Multilinguality

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.834797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.834797Z digest=sha256:37e6c9e102fc0305257b4d83040e5bb5e135f1fc2f975496e659a7638e267f00

Observation 556dabd3-900e-4d59-a3f7-865468e7a661 · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.948003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.948003Z digest=sha256:88ad1daac7d94cb66a6bbd63993b40e33c9d16b78f2c17803ed454b3b533258a

Observation fe891c49-9c68-4b27-a31d-636d3040cd64 · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.092591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.092591Z digest=sha256:6a161bc7149f0bf65bb7631100fbab779d4ab7c83bd8e0fc6b02bbf188bc3cb4

Observation 289d9c3d-4374-49f1-90c9-abe2b0413ec1 · outbound

This paper cites Mistral 7B.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Mistral 7B

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.202392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.202392Z digest=sha256:06617e2aa9d2bdb5f62bea44083e07b3345b7ea66731783df604e80cd0b5098c

Observation e14fc6fb-c632-433d-b421-a9306e3bc01b · outbound

This paper cites A Survey on Large Language Models for Code Generation.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use A Survey on Large Language Models for Code Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.304469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.304469Z digest=sha256:de84e81a5a9bc32c1e666fb63e633b2330b09870a38a254da8a61a5b6709dd8d

Observation bb0d6a1a-e472-4331-a9bb-dd6f620818a6 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.406960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.406960Z digest=sha256:3120193852866fb8a78202727f354127e767eecc534133405bf94e7ac7226238

Observation d95e5fad-f1cd-4af4-9819-486e74c05bc6 · outbound

This paper cites Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.533469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.533469Z digest=sha256:2ad37f02fd3a6ca784877951e90b72920632cfb0d815d1c7ebf8b25784f91ff5

Observation 7e09b11c-0844-42b0-9106-664da1d3e6ec · outbound

This paper cites SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.676227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.676227Z digest=sha256:04ea0da57e2e11cef55b8c3bea2b69f78f0d0bf39a001d2038221516c20f3fab

Observation 3ffc685f-ab4a-4baa-923c-52ef5ab7142b · outbound

This paper cites Controllable Text Generation for Large Language Models: A Survey.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Controllable Text Generation for Large Language Models: A Survey

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.795914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.795914Z digest=sha256:1d7354971a3b2b42b990ea6d1596623ce410511e05e4302f197f4bb493c7ce9b

Observation e773cb32-b65a-4db0-8695-1cd4132b6e08 · outbound

This paper cites ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.909375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.909375Z digest=sha256:74fc3a7eeaac2e29d698a9b0382438a3b87c6ce346019db481fb8cdd235b66c0

Observation 8ada4d7c-2f56-49aa-8e40-16cf10badd7a · outbound

This paper cites Corporate Communication Companion (CCC): An LLM-empowered Writing Assistant for Workplace Social Media.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Corporate Communication Companion (CCC): An LLM-empowered Writing Assistant for Workplace Social Media

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.995452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.995452Z digest=sha256:6f60c56c22935525c24495ca442fd9a3d9485d76d7659eaba0b986f384edb221

Observation da7635fc-2190-4561-b9c3-f2daa71a078a · outbound

This paper cites A Paradigm Shift: The Future of Machine Translation Lies with Large Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use A Paradigm Shift: The Future of Machine Translation Lies with Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.088283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.088283Z digest=sha256:4a5c20d55c55c5eb5711c2a66a44771bc648cf75a47ad624ba3ed9da68163131

Observation e70cb713-e48b-4f6a-8c0b-d533273190ac · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.187055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.187055Z digest=sha256:47269068c23954a5d4c8f0851c48f1f1500cfa41c52ada08d9c2232e2ac32b71

Observation 96c1bc5c-be2d-424d-b62f-1c8a15e68482 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:52:32.634095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T14:52:29.298139Z digest=sha256:153439a1352206a14128e8926754272699cbaef4b489dd2a7e7989a55247965a

Observation daf4a831-a6f4-4279-b27f-ce3504892117 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.369006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.369006Z digest=sha256:840b3e4f7e47e31ee2dd5e488798f10119a549e6b571b4a30d0217d30b0eef15

Observation af3517cd-6827-4fff-a2d1-0ce3ba1f690d · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.442168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.442168Z digest=sha256:ee3d7aba0e7d4cb96dc8b7805797bef2f54399a5a33c5f8a3f94450001b5e29d

Observation b52a48f3-7d90-42a4-a53f-80ae4217600b · outbound

This paper cites LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.517727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.517727Z digest=sha256:3fa0664d94ad5de51f39d5c36b5a9d5b2a0ae23138c6a6c6971a81707d14dfb4

Observation fa05beb8-01b6-4ffc-96c4-a948895bb029 · outbound

This paper cites Review of reference generation methods in large language models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Review of reference generation methods in large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.573151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.573151Z digest=sha256:ade83f07da417e7316f8a457e5299b567cffd4a940f6e4ccf0806848ab887404

Observation 7090315f-e05e-42f2-8158-447c8d1cf6fd · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.624489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.624489Z digest=sha256:162f8d9ddbeef2e9dc3592627b9d5f989a8fd1460689e487476d2afc0dd08fd3

Observation 5b2fa618-4552-4e31-8336-e1f9e5468b8e · outbound

This paper cites Tokenization Matters: Improving Zero-Shot NER for Indic Languages.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Tokenization Matters: Improving Zero-Shot NER for Indic Languages

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.703994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.703994Z digest=sha256:e4feac683b6547d4684d371f6137ed57f4f9df658493c413b9c92a8540289fa6

Observation 686ffdf8-d601-414b-ac05-0c6020e20bd6 · outbound

This paper cites Clinical QA 2.0: Multi-Task Learning for Answer Extraction and Categorization.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Clinical QA 2.0: Multi-Task Learning for Answer Extraction and Categorization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.883990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.883990Z digest=sha256:3de701e1f1ec1f8c83c4ca9e324f925c600fa32d08843e6420d6697eeb7f9c56

Observation c9cdb51a-fff7-4bf0-8352-94db76df3dcc · outbound

This paper cites Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.947591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.947591Z digest=sha256:87669363ad2b95090d5c26c2a2d550d3697f2a541348dedd554011f8c677ee1e

Observation 69935a2f-d505-4a50-b4ac-003d4e18c356 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.022676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.022676Z digest=sha256:8b9e896ef468cebfbff223b413d361783041079d7415c07e606780b2eca3b5d7

Observation 52b6447c-c493-4b19-89bc-6796518e7df5 · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.095909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.095909Z digest=sha256:18fbaf70a8bdbd65fd2b0ed2f8a9d729fc10480800952141d856e0abd1d716eb

Observation 5058c140-2e80-4716-b643-cb7382fe309d · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.159695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.159695Z digest=sha256:08d0c8376d7c524a8e92c1251b77c1d3ac3ee2cb75a21a4ecf81f3bc99b341ff

Observation 2b1e54c6-b92e-48a4-aefe-4140254a9bb5 · outbound

This paper cites The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.239795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.239795Z digest=sha256:1e9989d7fc41a8e996194a1774b3102e576daaad1a8e4ac82e7988b781d6a403

Observation 4a3b8bd0-cde7-4ee9-a71d-07972626824b · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:52:32.452987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T14:52:30.305517Z digest=sha256:598a976db7fd06aa828ac9a704f9c1ab989f9af2590c03fb78bfeaa1b01ab659

Observation 6875d0ee-2f5a-41e4-a4c4-65e207abcc1f · outbound

This paper cites Text Classification via Large Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Text Classification via Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.375712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.375712Z digest=sha256:f7d57d8c968d88dbfe69ce02c3cb232799ac4dfd89c203d6bc9a7055d3ab75e5

Observation 7e4a57dd-0d8c-41dd-b295-f35e2e2bdbe4 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.469185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.469185Z digest=sha256:b4fb88ef9c5c09e21093c8d7f99675f1c189ce870124058ff90c43a14ef3d5c0

Observation 9ad2ab34-9314-4d9f-966e-fba96dd00264 · outbound

This paper cites ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.530713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.530713Z digest=sha256:7263b1b79010c23436ad8927134a89ca270a530720cc9e942987b9c900fe128a

Observation fff2e369-c818-4d8d-b4a1-3064e957e2bb · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:52:32.284403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T14:52:30.572397Z digest=sha256:b2438a2e05fa933281ac5896dbf392d019788acec14163e59dd696acdff612ed

Observation 8f548a86-a549-49d8-bf54-2da948643e91 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use LLaMA: Open and Efficient Foundation Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.616707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.616707Z digest=sha256:a507aefa1c258ce25a38a2b3d3de2ffa3c9fe2b4cd58fb324bb7f9fa307aab6e

Observation 5661c39f-23e3-44c4-8c81-356c52133435 · outbound

This paper cites All Languages Matter: On the Multilingual Safety of Large Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use All Languages Matter: On the Multilingual Safety of Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.662928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.662928Z digest=sha256:f9f2d81592c8c6a0e5c53d21ba799e7ccc6c8d5c139d5f9fcead7ddeb7ccd775

Observation 8dea832e-a657-4a75-9287-1323bcb16dd5 · outbound

This paper cites Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.727155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.727155Z digest=sha256:cff3ef813dbc372e64febaacbafa798fa476acfc4ebe273693969e5540cfb1a0

Observation a4d6c8c2-935d-472d-a3e9-0023f6223a01 · outbound

This paper cites Adaptable and Reliable Text Classification using Large Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Adaptable and Reliable Text Classification using Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.798154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.798154Z digest=sha256:6e6a3e80198f0a793e4b45be9621ee47d8323f1a17f5e154a0d493ac11032a28

Observation c99df6db-c545-4525-8675-69f32fd72396 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:52:32.143629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T14:52:30.869862Z digest=sha256:6675fd2b6a1397bab645cbe3359d5c1f62a7c1d72d6166965bf620000e83596b

Observation fa468d64-aa82-41e4-93b2-5c3567b007d3 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:52:31.979578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T14:52:30.926313Z digest=sha256:e9c3f49e5be2a98b4e48d20d07fd2ca368a949105e1aed99649bfb7cde0d5473

Observation 6d811496-4a3b-4978-864f-92c658110c44 · outbound

This paper cites SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.965756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.965756Z digest=sha256:9199f0d011907b7c533f11c8eab81756a8ab59257dcd18da7dc46923ee7a6fcb

Observation 4b32d5de-e91b-44ca-bec3-3bf79eb58ad6 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:31.008092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:31.008092Z digest=sha256:993023cfbec62b0634ce45d52ee6f1f9a3010543d6b067361d0011049d8bad1a

Observation 6df9c3e1-659f-4b29-9f90-bbf5a8fe15e0 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:31.085302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:31.085302Z digest=sha256:7079e1858db36badb82ccf7c098ef4a33e2e8343b0df5fcb71a4663466388d0f

Observation a444735e-47da-4991-9cee-16ece478ec49 · outbound

This paper cites SafetyBench: Evaluating the Safety of Large Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use SafetyBench: Evaluating the Safety of Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:31.148227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:31.148227Z digest=sha256:3879257dff7e7009fea2ac7caf78ea91366c7443b3490d360fa0e697a8944e9c

Observation a1c9744b-1e2d-4831-be5f-977c415f43c2 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:31.213008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:31.213008Z digest=sha256:b3565f6727301c4a2295d4cbc42ccfc78f0078456d2d066f39ac8361b5596ef4

Observation 419e77c1-3e29-460a-b529-1c910287c98e · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:31.262024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:31.262024Z digest=sha256:bf874fe56d25f36152c2f78d321d153cf12d7874fc4d381a3d476b5cef592ccf

Pith citing papers

No inbound Pith citation observations are available.