Pith. sign in

Paper Citation Record · LEDGER

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

As of 15 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 15 inbound Pith citation observations for arXiv:2506.05692.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05692 v3

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:22:13.453350Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T11:46:31.451484Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T04:09:34.777136Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e945359a-f2d0-4f51-8d09-8d28a05c9a3a · outbound

This paper cites GPT-4 Technical Report.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:09.593636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:09.593636Z digest=sha256:16a48fbbdcbb54ca000aa18dfedd6bab86fbd60c9e61803e9dc6a469495fef5d

Observation b50f08a5-c133-4ec9-b83b-a6f13a0c059b · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:09.679322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:09.679322Z digest=sha256:9a7026bc892e6b6b81d3689d0d8a44625959bc86a35c90782c0e974f608311f9

Observation 216ecdb6-fbe3-4a27-a307-06a500fcf67e · outbound

This paper cites Program Synthesis with Large Language Models.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Program Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:09.792008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:09.792008Z digest=sha256:804512b6068384f45f06b55a206e72affe517a41382daea2ad493912e23c2b6a

Observation c1d011af-ea4f-4971-ad9c-a0ce5add58c1 · outbound

This paper cites CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:09.899584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:09.899584Z digest=sha256:8afb3d34396afc49d6d6e007d37f17fe255e2f5250c5752252adb8e76d50a36d

Observation 0e691715-3a62-4b72-bd84-509166a061ed · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:16.061121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:22:09.985448Z digest=sha256:faa795e2692b922d1666e5364a9c263dbf134d9e7320a9323d6f8d7c86ae96c1

Observation 0a4b93a5-80ae-4517-ad1e-80a23e487a0a · outbound

This paper cites Evaluating Large Language Models Trained on Code.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:10.044592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:10.044592Z digest=sha256:85642f478a6c4514c92b4469d7f0c3362a0c9c7b9a65357609a9398faa5283ba

Observation dd6bb4d7-d8d1-4ddd-bc86-80ea1cd806ca · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:10.124029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:10.124029Z digest=sha256:930e550f5fa12372eae24eb26049568abb2e0cc7c9ee2f11ffe0f04b4bd236b3

Observation 75522f84-6e98-4d51-b050-418a766786cd · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:15.924317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:22:10.200771Z digest=sha256:f64ca1bda186485dce6c558d4da073e6295bddd958a10b2fea89cc97af2292bd

Observation 54ac5407-6ff4-4676-afd7-abc8a2c5a588 · outbound

This paper cites Competitive Programming with Large Reasoning Models.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Competitive Programming with Large Reasoning Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:10.290482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:10.290482Z digest=sha256:bae967946bfe79b9d45f5fb0cbb05214183a8037ab2098bedb551bae1d5ca4ca

Observation a74aecec-0ecf-4ba5-b0b6-ed2ff24dcee7 · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:15.779157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:22:10.398449Z digest=sha256:925a4830cee6d9fc0189df0db21a921b6beb694fa6803f51bedd83931681aadd

Observation ca90c4d6-b648-4479-acc6-42d416283997 · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:15.619059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:22:10.472959Z digest=sha256:ef8a833e3daa45dc0d7062a45fe812936b91fd422f60c12bb834d51e1342d18f

Observation 5b2387c1-c766-4429-8dbb-2a32cbc972a9 · outbound

This paper cites A Survey on LLM-as-a-Judge.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code A Survey on LLM-as-a-Judge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:10.547585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:10.547585Z digest=sha256:9d01f392298668c580607145f539e451a15b59889b917172b340677c06d78228

Observation 002a785b-e291-460c-a881-aa33c1898720 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:10.651198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:10.651198Z digest=sha256:e539f8d796ba4f0ea9bb31578087b04e9f1339b7a2c007129f179d4e400d1fbe

Observation 363192a3-d130-4cae-9feb-bcd9bf737d5f · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:10.730473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:10.730473Z digest=sha256:f4afa1cdbcd83c3dea4731420b0f58040ec3798f25f11b1979a6ecb56e4eee70

Observation 36496748-a3f8-43d8-aef6-673899a281eb · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:10.792739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:10.792739Z digest=sha256:64670d5672846e3cfe042fbf98ffc21fba2400ad6531b0473eb62a56840cbfa9

Observation ed26c410-65bb-45b0-9715-9a2930e0d07e · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Measuring Coding Challenge Competence With APPS

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:10.915468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:10.915468Z digest=sha256:f9125277ca45a2c0256a8bf6383b6fabc3c84753bb108bb5906f5fc5a972e15d

Observation 77673158-a6bb-4a09-b582-da48cb1dbac3 · outbound

This paper cites OpenAI o1 System Card.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code OpenAI o1 System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:11.026518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:11.026518Z digest=sha256:58f99cb156350250607a0bd11f78a5afbb390d6f5d7e6a9cf6df9b9bdc5bfb28

Observation 61d8ef6e-bf01-4a7e-b77f-5c445edccef0 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:11.118778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:11.118778Z digest=sha256:4c2de6c38d6e4a1fbe9c046fc8f6fc62d7f4f0c0e0e1bf10e3a44a53b0228b5d

Observation 709a49c8-a7b6-47ae-a80a-b3d99e8bdf4e · outbound

This paper cites Investigating Large Language Models for Code Vulnerability Detection: An Experimental Study.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Investigating Large Language Models for Code Vulnerability Detection: An Experimental Study

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:11.210282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:11.210282Z digest=sha256:21f4c0f1ae08d0a92e4d4a2c59fedee42ff7b27e752877cc2b64189eec6d9e12

Observation a19e61cf-e864-4692-991e-f3e6e21f57a0 · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:11.315557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:11.315557Z digest=sha256:4d9d04eb4a33eba050f0f76d753660ecae07e39bb7c8c164ea3b023d1fd1dc59

Observation b7a52fed-7a47-4295-9330-c45b6301c3f2 · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:15.441597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:22:11.405590Z digest=sha256:e464ff79f5be173bbf91a2a11f12aa32aa6888f4b075e08ca8a6098a0ef323e2

Observation b562392e-e503-40db-ac6c-defa0a008dce · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:15.292832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:22:11.469477Z digest=sha256:b2bbe16582f108c37c1883d041f709f6d5d1547549251a5afae9eace1ce6ea2e

Observation 50a6dc41-d18d-4aef-9878-0a1bbbc5ecaf · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:11.560456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:11.560456Z digest=sha256:5881449e341793daf9ee81c5d255565f52c9bd23c2b01001d40f83e3ac9b5f44

Observation d2c0ba74-3156-4bca-9bce-11714833ca47 · outbound

This paper cites DeepSeek-V3 Technical Report.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code DeepSeek-V3 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:11.638747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:11.638747Z digest=sha256:290bbfe8fe65ff9bbc99504371defbd3749591dfe1da38c55b45d8038c7de921

Observation aca5a2c8-d3f9-4798-8682-4d5db10a7bb6 · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:15.154345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:22:11.743316Z digest=sha256:5f77fe6f574e99aa2a5d28ed06f738ab345a871798ec271b214d46775bb9cdf8

Observation 9fcd9956-69b8-42e5-8fb7-642b93f2fb9e · outbound

This paper cites MarsCode Agent: AI-native Automated Bug Fixing.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code MarsCode Agent: AI-native Automated Bug Fixing

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:11.875131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:11.875131Z digest=sha256:dabe88c105e31125bf885167d8de8f6e7ef2a9af9001bd06709dec005f2c158e

Observation e5f8f7c1-6e73-47dd-a758-bb6ca3d9d522 · outbound

This paper cites StarCoder 2 and The Stack v2: The Next Generation.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code StarCoder 2 and The Stack v2: The Next Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:12.018305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:12.018305Z digest=sha256:49d7b25cbfb9845b1d1d7481d6e383ba771d37101f29a84cbaad2f23c6b1eeac

Observation ff0ee46d-b13f-48a1-9348-cccd1637a82b · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:15.001834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:22:12.109822Z digest=sha256:dcf58bc0db26a4f0adbbd21cf52389046172d23dae3cd7ac21530e7b79ee9578

Observation 498135e9-565c-47c1-8dbd-3381fac4460a · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:14.818106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:22:12.282471Z digest=sha256:45d0d0938f53357b275e569f454c5929eed35b45bf6edfba6665fa64a186224a

Observation 91895404-3140-453a-b3cd-bc9b08e6b070 · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:14.688688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:22:12.423058Z digest=sha256:783aec353b882170fa6618785e49dd6df7a76bf9bd3654de166cb726af799b6f

Observation 93341358-3a43-4082-8ef6-78c9156449c7 · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:14.532416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:22:12.578595Z digest=sha256:e925d7652ed09ad8e5ca586b53e46e9d522ac7d121e7ce5ca1c27bc121381907

Observation 1d4488ce-8ecc-44c4-8328-e777b422a9c9 · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:12.706363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:12.706363Z digest=sha256:660fd67b4ecd3077b2417c805b368401f76803bd5b4341d64b2b5b8a1533c785

Observation 28da7e5f-0abb-41d8-9b7b-0d3af409a614 · outbound

This paper cites CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:12.808455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:12.808455Z digest=sha256:3e8db9a94d43b80c0b87da2f471fd6f3fb471e3f6cab3d97dfb40cc551ff66d1

Observation 333f791c-c1b9-424e-993d-6973fd3be9f2 · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:14.311152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:22:12.902766Z digest=sha256:cd5689739ddfa3d7dd099dd82ef155084ad01900dc62ae591974fd024ad405c3

Observation b9825af3-d8f0-43e2-9445-5be0fcebffdf · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:14.180190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:22:12.994276Z digest=sha256:1351b0f06282094788e828bbf49c2ac7cb204e45651d9d5482df1503705b9bcc

Observation 4ffe3f07-a005-4d2c-ab75-a4fc650549ee · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 36

Resolution
verified exact
doi, observed 2026-08-07T10:22:13.668860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:22:13.105707Z digest=sha256:65b2e236bd5fb81d79622b9d0e5300778f3714bfe281463121cdf6655a738c4c

Observation 0964962b-d080-480e-a05a-7ba431423254 · outbound

This paper cites Code Vulnerability Detection: A Comparative Analysis of Emerging Large Language Models.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Code Vulnerability Detection: A Comparative Analysis of Emerging Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:13.230288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:13.230288Z digest=sha256:fa5c226315e25fe00fa049412d1a62a0f57b8904ad642e4f9c1c02bc26364a1a

Observation 5f7b621f-2d9b-4c06-a54b-0ad68f438aa8 · outbound

This paper cites online" 'onlinestring :=.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code online" 'onlinestring :=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:13.361543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:13.361543Z digest=sha256:74f2ba0e0eeff1aff2bd6f8ca81126749abc090654ad4b8ee3975e71b1cb8496

Observation 4a8c1799-e424-4665-8b53-1d81b39c0d42 · outbound

This paper cites write newline.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code write newline

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:13.453350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:13.453350Z digest=sha256:52e764ed08dce1098c2b6e6e2dc687a2a16361582367f6f010b527462420135b

Pith citing papers

Observation 61f9837f-bff1-4280-b78e-f6453c9785b0 · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.769794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.769794Z digest=sha256:45a06624dc0b76abff26638b0b7bbc832ad8d8fe2a0a9e0686e2a20aa37cb5e1

Observation bd32eeba-1950-480b-aa1a-c8119ef59ad0 · inbound

Extracting Recurring Vulnerabilities from Black-Box LLM-Generated Software cites this paper.

Extracting Recurring Vulnerabilities from Black-Box LLM-Generated Software SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T05:18:43.734775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:18:43.734775Z digest=sha256:43706319ddcb943129e0157f4c822426f7bb625d94fcf773a4110a1cbe8648d8

Observation 80fe48cd-4c88-4d9f-b7c7-86f8fcf6fa8e · inbound

False Security Confidence in Benign LLM Code Generation cites this paper.

False Security Confidence in Benign LLM Code Generation SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:21:26.746518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T06:20:52.774912Z digest=sha256:95224f635e0e7a1a9f349203b9389894b31bae79793bb9a133d9c0d36ba3730d

Observation 90c05cd9-b579-4d74-b9d7-94e4ee5fbf69 · inbound

Taint-Style Vulnerability Detection and Confirmation for Node.js Packages Using LLM Agent Reasoning cites this paper.

Taint-Style Vulnerability Detection and Confirmation for Node.js Packages Using LLM Agent Reasoning SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:59:49.633068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T00:56:42.767929Z digest=sha256:197431f297647e73ca26d670d2a1e0196ee027efb455a80d22b2593f7e0d9bf0

Observation 547f6267-2fff-48c6-a572-25ea7a88cb65 · inbound

On Fixing Insecure AI-Generated Code through Model Fine-Tuning and Prompting Strategies cites this paper.

On Fixing Insecure AI-Generated Code through Model Fine-Tuning and Prompting Strategies SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:09.981257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T09:12:15.486780Z digest=sha256:3893137125dfb2a2309b0c8df843383af821a92c0fac93c7ff49004aa126b800

Observation a3915f51-e380-45ed-95b0-7fae6f255723 · inbound

CoT-Guard: Small Models for Strong Monitoring cites this paper.

CoT-Guard: Small Models for Strong Monitoring SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:39:23.772865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T19:39:16.264272Z digest=sha256:b9e5e34b6e3c7402f3af05038ae4a75210ef3a33112d21447647c6fe16d9693a

Observation 4fd05067-2171-419b-aed8-656c01488385 · inbound

ASSEMBLAGE-DEEPHISTORY: A Cross-Build Binary Dataset with Temporal Coverage cites this paper.

ASSEMBLAGE-DEEPHISTORY: A Cross-Build Binary Dataset with Temporal Coverage SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:41:21.765966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T09:40:21.746161Z digest=sha256:15443423b5a335876b2aab19ad2abb5313888ae9884532e8d4a60938ed54c283

Observation 4da95db5-56eb-40e7-b1fa-eb2d852ed576 · inbound

FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming cites this paper.

FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T04:09:34.779477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T17:11:40.088809Z digest=sha256:e29b58a15648458fce6a040fc07acfd4dafe6a3c9aea942e03437edc5cd5acb4

Observation 9391a7b8-e6cf-48a0-8c26-f446b61fa838 · inbound

An Empirical Study of Security Calibration in Large Language Models for Code cites this paper.

An Empirical Study of Security Calibration in Large Language Models for Code SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:45:43.045640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T05:04:41.223743Z digest=sha256:3ceafea4e6e1026993eb2077fdd43b9ce64d2dacd353475b402a5b14165884ee

Observation 75240d8a-1d5d-40ba-b438-c96b1600794c · inbound

The Illusion of Safety: Multi-Tier Verification of AI vs. Human C++ Code cites this paper.

The Illusion of Safety: Multi-Tier Verification of AI vs. Human C++ Code SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T09:44:42.646014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:44:42.646014Z digest=sha256:8cf88296976109d2ff41f3bef375f3caad464e4b674aa9498f1a9e29c7335ee1

Observation 91d0a92e-fce5-4d73-85f0-61222953f4a8 · inbound

The Illusion of Safety: Multi-Tier Verification of AI vs. Human C++ Code cites this paper.

The Illusion of Safety: Multi-Tier Verification of AI vs. Human C++ Code SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T04:38:06.679657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:38:06.679657Z digest=sha256:04a08471deaa5bc83870a5288ea3e7c61ab3e3532ba4c0d17b89b5354ed64c53

Observation c8762db4-355e-454e-90e2-635ce74067b8 · inbound

The Patchwork Problem in LLM-Generated Code cites this paper.

The Patchwork Problem in LLM-Generated Code SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T01:16:30.101870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:16:30.101870Z digest=sha256:9e789ae91491de8120c11c586d17d6d8c815ac41968a47c4b18f7fbc0e3450e2

Observation b212ebb8-7ab2-4f5b-a7ae-95296fd38590 · inbound

Poster: Rethinking Security in LLM Code Generation through Real-World Risk Scenarios cites this paper.

Poster: Rethinking Security in LLM Code Generation through Real-World Risk Scenarios SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:38.189512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:40:38.189512Z digest=sha256:8a1e6e2d1608997d29e66aa2fb2f669c25474f11dc161858d0d085a390e6a4c3

Observation c6ca3504-9d98-4ad8-8114-0bcf1183b914 · inbound

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation cites this paper.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:57.375647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:57.375647Z digest=sha256:276a362240fd9529ab60794ac560b1ea51edef7b59510128acb022c3f70f5782

Observation 8cffb6aa-3fae-45c6-b8cb-2c79c7085b66 · inbound

Does Fixing Break Security? An Empirical Study of Security Degradation in Iterative LLM-Driven Infrastructure-as-Code Repair cites this paper.

Does Fixing Break Security? An Empirical Study of Security Degradation in Iterative LLM-Driven Infrastructure-as-Code Repair SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T11:46:31.451484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:46:31.451484Z digest=sha256:aa53f74e45def1d11654e66f8c0adccbf76abecc733286f0da3a7500c697cf2a