Pith. sign in

Paper Citation Record · LEDGER

S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2405.14191.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.14191 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:12:00.899754Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:10:00.725685Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 246381f3-539f-4c3e-a52a-21eafb6c6b56 · inbound

From Local to Global: A Graph RAG Approach to Query-Focused Summarization cites this paper.

From Local to Global: A Graph RAG Approach to Query-Focused Summarization S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-08-04T01:57:26.797777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-11T05:10:57.816312Z digest=sha256:fcee62f5d419b2f3568a73534ae0c6f1f0a9274931fbee66c8216f21e9888895

Observation 627f8aad-0086-494d-b7bf-38053d7e8e8c · inbound

From Allies to Adversaries: Manipulating LLM Tool-Calling through Adversarial Injection cites this paper.

From Allies to Adversaries: Manipulating LLM Tool-Calling through Adversarial Injection S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:43.479613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:17:43.479613Z digest=sha256:a90a2527f42317ef252769e69e7af029031be1c7b6096a09b415c74d0db59fd7

Observation 3c2d03b5-39a7-4daf-9977-694f7bab023b · inbound

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense cites this paper.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.842188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.842188Z digest=sha256:2a46ae6cdfe22868ecd8be0a4d99c7a44acdc4be7534d7cc5e1382577c432eea

Observation b794624a-4476-4186-9a07-f96a5c272ab1 · inbound

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values cites this paper.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.898841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.898841Z digest=sha256:cd2772f597cabb466c7063b4c9fe7493ded704f83dbb1d1dbc18d0c2ae69f6fb

Observation 2b96bf4d-ed93-4bed-b67b-095f5ad09327 · inbound

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation cites this paper.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:16.948537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:16.948537Z digest=sha256:0c9a3ab55efdd2e47189e7a93e795df6b64fe0c9fcbfd1c865a4d40a4fa90a5d

Observation ed2b1a91-fe68-4fd9-8dba-084e3064c9d9 · inbound

o3-mini vs DeepSeek-R1: Which One is Safer? cites this paper.

o3-mini vs DeepSeek-R1: Which One is Safer? S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T23:32:16.751706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:32:16.751706Z digest=sha256:b63684e54498a56824b6b4e24f3cf222e85eda10df88056301aafed451c47866

Observation 87a1d1c7-2a9d-4572-9231-8c09383047a7 · inbound

Psychometric-Based Evaluation for Theorem Proving with Large Language Models cites this paper.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.653389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.653389Z digest=sha256:2a0703e7a6bff6057a9d78b2c53624c77b802f819e2f63ca826b60728c82bcb7

Observation 545f1f60-e197-4e43-bff9-c3d1015fdb0b · inbound

The Science of Evaluating Foundation Models cites this paper.

The Science of Evaluating Foundation Models S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.916246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.916246Z digest=sha256:86835c2c3747a421b24bee2060d478cd5d87bc9aad2203c2028001a3e39e6606

Observation aa10d285-7172-43b9-88a1-d48d23867685 · inbound

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks cites this paper.

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 179

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:00.899754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:12:00.899754Z digest=sha256:143bf188e24cedebaed4e1618c41f5a4687e7b1b30cc094d0249b038aedff694

Observation 2fe61643-fdf3-4e05-9533-fff34d8e8f5d · inbound

The Aloe Family Recipe for Open and Specialized Healthcare LLMs cites this paper.

The Aloe Family Recipe for Open and Specialized Healthcare LLMs S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:30.053396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:36:30.053396Z digest=sha256:b37c5317fa34d3d916d18c766cd082f99bef521c5811a8ba1addfe82c5276890

Observation d4ac6a66-df10-4627-9818-94933ba03a0e · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.844499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.844499Z digest=sha256:73c947a5fae5b3dc492a282f9f670a7d99639ac2fd1759251a4d5a6fceb4afdd

Observation 303efba4-6986-4882-ab4a-783b4a3ba1d2 · inbound

ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal cites this paper.

ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-08-04T01:57:26.797777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T23:27:52.438709Z digest=sha256:f18696d547b60b864000e910a301e8605027d0133815acf305aa2d15fa21fe1b

Observation d7650cbd-c31f-4368-ad84-0f9410e5f121 · inbound

Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting cites this paper.

Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T12:45:26.905184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:45:26.905184Z digest=sha256:110b84be2ccfb0081751fced184a2e0387d8e91f63376b080efdfc0deb4315b1

Observation 1a5121e0-d0c8-4202-a219-9821cb75e50e · inbound

Multilingual Refusal Alignment for Safer Large Language Models cites this paper.

Multilingual Refusal Alignment for Safer Large Language Models S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-08-04T01:57:26.797777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-04T18:09:17.531727Z digest=sha256:02010c08db8dc20313b21fd80ca06dc217de82014f1815d2ebbb24adb64bc7cf

Observation d42ba51f-292d-4d59-8f56-08e25b9b7ef9 · inbound

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety cites this paper.

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-08-04T01:57:26.797777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T00:46:03.210076Z digest=sha256:45996d1c6953a56113f4f689551bdd847d8164fe5b58b02afdc6dad904e81147

Observation 077f18ff-e8be-4268-a724-7044051ea787 · inbound

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models cites this paper.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:20.542247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:20.542247Z digest=sha256:063f0429d80dce10ee315bba0a9d351f3259dbe8ebe915a615d84289b7bcb519