Pith. sign in

Paper Citation Record · LEDGER

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

As of 8 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 14 inbound Pith citation observations for arXiv:2506.05692.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05692 v3

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:22:13.453350Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:17:26.769794Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T04:09:34.777136Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e945359a-f2d0-4f51-8d09-8d28a05c9a3a · outbound

This paper cites GPT-4 Technical Report.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:09.593636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:09.593636Z digest=sha256:d68c9828845cbdb9dda83014ced371a236a668cc60945ffd1afaf36991b36882

Observation b50f08a5-c133-4ec9-b83b-a6f13a0c059b · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:09.679322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:09.679322Z digest=sha256:c6c3a151cfc653f6e81866924f1705424a2b1b6363acda5100783ad60b094d9b

Observation 216ecdb6-fbe3-4a27-a307-06a500fcf67e · outbound

This paper cites Program Synthesis with Large Language Models.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Program Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:09.792008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:09.792008Z digest=sha256:5acf4551c7819450114dc5e4b5803dffe55981e64aa4495fd778cf415b3f3e83

Observation c1d011af-ea4f-4971-ad9c-a0ce5add58c1 · outbound

This paper cites CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:09.899584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:09.899584Z digest=sha256:5ccf02ac181118f934e159eb6868c7391a50ac1e0a7a61ac91ed6d1b947dc78c

Observation 0e691715-3a62-4b72-bd84-509166a061ed · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:16.061121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:22:09.985448Z digest=sha256:50693f244fd16e1f226b585627aec20405d4f75e28ecfb707be14571335c4666

Observation 0a4b93a5-80ae-4517-ad1e-80a23e487a0a · outbound

This paper cites Evaluating Large Language Models Trained on Code.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:10.044592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:10.044592Z digest=sha256:28c051bda9d31a2e2f5bb2b4216a39a772c86a8459fa82082266dc25227fa917

Observation dd6bb4d7-d8d1-4ddd-bc86-80ea1cd806ca · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:10.124029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:10.124029Z digest=sha256:aab8ceed5d3011ea0b156b4e0f1623c02f84aa7ecf36a08c860c3ec580d89a7a

Observation 75522f84-6e98-4d51-b050-418a766786cd · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:15.924317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:22:10.200771Z digest=sha256:3569893b29d071a64f47cd74a92cbf4bd9656f97ec1b401b8eeb1e8e03b80988

Observation 54ac5407-6ff4-4676-afd7-abc8a2c5a588 · outbound

This paper cites Competitive Programming with Large Reasoning Models.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Competitive Programming with Large Reasoning Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:10.290482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:10.290482Z digest=sha256:6d64e6b8e013c96c03341ce697e5667fa578e99d05b32e89c71fea75951387e0

Observation a74aecec-0ecf-4ba5-b0b6-ed2ff24dcee7 · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:15.779157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:22:10.398449Z digest=sha256:7537f0354b00ee816ba48b993c19b7a1cfb8dca10d655f9136f895b7b875081b

Observation ca90c4d6-b648-4479-acc6-42d416283997 · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:15.619059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:22:10.472959Z digest=sha256:2cce0418bf69b8b64db397acc5c6e44429b78faeae175b28dd0a5972aaa5f960

Observation 5b2387c1-c766-4429-8dbb-2a32cbc972a9 · outbound

This paper cites A Survey on LLM-as-a-Judge.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code A Survey on LLM-as-a-Judge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:10.547585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:10.547585Z digest=sha256:e3326ec2829923daec907ded8dd4d0cbcc96b07332cee9b6c7a6069828c7560b

Observation 002a785b-e291-460c-a881-aa33c1898720 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:10.651198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:10.651198Z digest=sha256:081f34b95fa8abcb2c927e597a719c6d1daee15db9ce95e860375c37044cd9cc

Observation 363192a3-d130-4cae-9feb-bcd9bf737d5f · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:10.730473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:10.730473Z digest=sha256:bb8bfaa7c074120fd94bb789f62dbfe2d2fd821637ede21a52fbf04dc3988c94

Observation 36496748-a3f8-43d8-aef6-673899a281eb · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:10.792739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:10.792739Z digest=sha256:c764c0012e4e35281af2b216b213091c39cc84eca602d5edf9e6b5bdb33afd47

Observation ed26c410-65bb-45b0-9715-9a2930e0d07e · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Measuring Coding Challenge Competence With APPS

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:10.915468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:10.915468Z digest=sha256:3f73ec2a73710abfe966142e158d62dfdbc12d458ec8f7b3eccb94fac0c3695c

Observation 77673158-a6bb-4a09-b582-da48cb1dbac3 · outbound

This paper cites OpenAI o1 System Card.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code OpenAI o1 System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:11.026518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:11.026518Z digest=sha256:f9a1f51ea593722584a9eac23a1a641c9c1313607ed95b37d913e7c412fa5d93

Observation 61d8ef6e-bf01-4a7e-b77f-5c445edccef0 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:11.118778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:11.118778Z digest=sha256:e734905c3c4fe7ad3548e3c834ef98ec2bd188843b5eef7fc992c86a9f45a226

Observation 709a49c8-a7b6-47ae-a80a-b3d99e8bdf4e · outbound

This paper cites Investigating Large Language Models for Code Vulnerability Detection: An Experimental Study.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Investigating Large Language Models for Code Vulnerability Detection: An Experimental Study

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:11.210282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:11.210282Z digest=sha256:207e837db0af41c6940851b304f33b928e870e5a162f679589e21a6a2555bce6

Observation a19e61cf-e864-4692-991e-f3e6e21f57a0 · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:11.315557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:11.315557Z digest=sha256:613012495694a0a69ff97a16f9f375499d2528171a511e842195c4552c1a9be8

Observation b7a52fed-7a47-4295-9330-c45b6301c3f2 · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:15.441597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:22:11.405590Z digest=sha256:af36334b01d8765a60905143c91f81e697b7b25400372143b1deb0fa8356287d

Observation b562392e-e503-40db-ac6c-defa0a008dce · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:15.292832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:22:11.469477Z digest=sha256:049c2293e58b7a76523e1be7f7728cdbd3ba7483b409d4f29f0a25b0cc59e79f

Observation 50a6dc41-d18d-4aef-9878-0a1bbbc5ecaf · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:11.560456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:11.560456Z digest=sha256:78316cac73bbdec111fd136ae20ad5ac7f594ba72101bd778024c49921471c6e

Observation d2c0ba74-3156-4bca-9bce-11714833ca47 · outbound

This paper cites DeepSeek-V3 Technical Report.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code DeepSeek-V3 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:11.638747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:11.638747Z digest=sha256:70cc88a3b61b29bc3cfb52de2b0afb473c03ac7c0a1f7e37b521d1debebc4f44

Observation aca5a2c8-d3f9-4798-8682-4d5db10a7bb6 · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:15.154345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:22:11.743316Z digest=sha256:db16f2496061584d5502ee7be3dd8fff7441db263a5c0005032d0777d260326d

Observation 9fcd9956-69b8-42e5-8fb7-642b93f2fb9e · outbound

This paper cites MarsCode Agent: AI-native Automated Bug Fixing.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code MarsCode Agent: AI-native Automated Bug Fixing

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:11.875131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:11.875131Z digest=sha256:81ed79b61270327a7f503e11506bd6ed73865526570ca4a816cda483783f0855

Observation e5f8f7c1-6e73-47dd-a758-bb6ca3d9d522 · outbound

This paper cites StarCoder 2 and The Stack v2: The Next Generation.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code StarCoder 2 and The Stack v2: The Next Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:12.018305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:12.018305Z digest=sha256:235fcc66fca1a69773ed41cf1f4367e4b205f6441378357c35388aa8b8d5dc81

Observation ff0ee46d-b13f-48a1-9348-cccd1637a82b · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:15.001834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:22:12.109822Z digest=sha256:a62ecf4036080f1bb98b5caa0ccad192e203fc588ea35233d7bc5cb64f0ff2f0

Observation 498135e9-565c-47c1-8dbd-3381fac4460a · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:14.818106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:22:12.282471Z digest=sha256:57031c699e8ba3d86779572bf494f986c4c2cca9ea84b95e9f0b525a155ad9a7

Observation 91895404-3140-453a-b3cd-bc9b08e6b070 · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:14.688688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:22:12.423058Z digest=sha256:cb26335cbcc8d99d3d4bc300ccbfbde52157490c706f0ed6dcd0269c5073127d

Observation 93341358-3a43-4082-8ef6-78c9156449c7 · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:14.532416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:22:12.578595Z digest=sha256:4d8952d327efbe116e7c47e6315373fd722210c7a319729883eebf7cd78c9de6

Observation 1d4488ce-8ecc-44c4-8328-e777b422a9c9 · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:12.706363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:12.706363Z digest=sha256:9de8bc80b6fe7fd0570b87cc05b9a932cca86581be4795b97728947c74ba8c52

Observation 28da7e5f-0abb-41d8-9b7b-0d3af409a614 · outbound

This paper cites CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:12.808455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:12.808455Z digest=sha256:16fd294f2e39331a35994c383cc9b69b1eea23109c3effc25b785f0a21ed480a

Observation 333f791c-c1b9-424e-993d-6973fd3be9f2 · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:14.311152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:22:12.902766Z digest=sha256:49bc210fd6aad491e0211ed24eb5e3afed40d76dee2f6f2a6f9ee4c8962f8965

Observation b9825af3-d8f0-43e2-9445-5be0fcebffdf · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:22:14.180190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:22:12.994276Z digest=sha256:5e0eb874d425249c933b099e94fda185d5865113642defced83fde4e1d076288

Observation 4ffe3f07-a005-4d2c-ab75-a4fc650549ee · outbound

This paper cites an unresolved cited work.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Unresolved cited work

Reference 36

Resolution
verified exact
doi, observed 2026-08-07T10:22:13.668860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:22:13.105707Z digest=sha256:48894c06a1414fd2deb517ffc22818e42852743692f78b088941dc26b00933d1

Observation 0964962b-d080-480e-a05a-7ba431423254 · outbound

This paper cites Code Vulnerability Detection: A Comparative Analysis of Emerging Large Language Models.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code Code Vulnerability Detection: A Comparative Analysis of Emerging Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:13.230288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:13.230288Z digest=sha256:b900b11e475301919cc29df1f0af348f64d1d0aa950ec3c97ebaaf3ee8161af6

Observation 5f7b621f-2d9b-4c06-a54b-0ad68f438aa8 · outbound

This paper cites online" 'onlinestring :=.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code online" 'onlinestring :=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:13.361543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:13.361543Z digest=sha256:6e2f5f89c794a2d2729dbd8d6962610422700cf5901b51a3d3ec05786cc14188

Observation 4a8c1799-e424-4665-8b53-1d81b39c0d42 · outbound

This paper cites write newline.

SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code write newline

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:13.453350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:13.453350Z digest=sha256:8e8975dbdf94668173309683821c816fe9a68cf466303bb5cbc0c5c1df8b30fd

Pith citing papers

Observation 61f9837f-bff1-4280-b78e-f6453c9785b0 · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.769794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.769794Z digest=sha256:8fa26753b88d4a6353c38d16897ddba6276735027e6657b8f205c1e9f8878272

Observation bd32eeba-1950-480b-aa1a-c8119ef59ad0 · inbound

Extracting Recurring Vulnerabilities from Black-Box LLM-Generated Software cites this paper.

Extracting Recurring Vulnerabilities from Black-Box LLM-Generated Software SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T05:18:43.734775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:18:43.734775Z digest=sha256:8cf4b1b48336222ae9255364e2391bb68e6e1e06e7ac06d4961dc66c7c18e547

Observation 80fe48cd-4c88-4d9f-b7c7-86f8fcf6fa8e · inbound

False Security Confidence in Benign LLM Code Generation cites this paper.

False Security Confidence in Benign LLM Code Generation SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:21:26.746518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T06:20:52.774912Z digest=sha256:58f93cddb167cd430fb997309665108fa5ed433e263cf7aaef6475b7f601e847

Observation 90c05cd9-b579-4d74-b9d7-94e4ee5fbf69 · inbound

Taint-Style Vulnerability Detection and Confirmation for Node.js Packages Using LLM Agent Reasoning cites this paper.

Taint-Style Vulnerability Detection and Confirmation for Node.js Packages Using LLM Agent Reasoning SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:59:49.633068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:56:42.767929Z digest=sha256:a5b3fa176c81563203897df3964b9275ca94d84e564c63d904dcff7269ecde63

Observation 547f6267-2fff-48c6-a572-25ea7a88cb65 · inbound

On Fixing Insecure AI-Generated Code through Model Fine-Tuning and Prompting Strategies cites this paper.

On Fixing Insecure AI-Generated Code through Model Fine-Tuning and Prompting Strategies SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:09.981257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T09:12:15.486780Z digest=sha256:d6e0212e12bb3eba9dceb6c063b9a3f325e36032f8b2d027db30419e9a04193f

Observation a3915f51-e380-45ed-95b0-7fae6f255723 · inbound

CoT-Guard: Small Models for Strong Monitoring cites this paper.

CoT-Guard: Small Models for Strong Monitoring SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:39:23.772865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T19:39:16.264272Z digest=sha256:0a5129c782a643cbd02f78b4a07d888dfb2b3b75e821a587fe93ec2dce549ee9

Observation 4fd05067-2171-419b-aed8-656c01488385 · inbound

ASSEMBLAGE-DEEPHISTORY: A Cross-Build Binary Dataset with Temporal Coverage cites this paper.

ASSEMBLAGE-DEEPHISTORY: A Cross-Build Binary Dataset with Temporal Coverage SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:41:21.765966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T09:40:21.746161Z digest=sha256:9ef88720e210b06bf9376aca263df34d049e5389b6b1aff0ca0b1fb1c2cdea3a

Observation 4da95db5-56eb-40e7-b1fa-eb2d852ed576 · inbound

FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming cites this paper.

FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T04:09:34.779477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T17:11:40.088809Z digest=sha256:d69e1b225bf696da0c341e6d822500d5fd1a7ae29a6eeec9f09c4391144c4dd9

Observation 9391a7b8-e6cf-48a0-8c26-f446b61fa838 · inbound

An Empirical Study of Security Calibration in Large Language Models for Code cites this paper.

An Empirical Study of Security Calibration in Large Language Models for Code SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:45:43.045640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T05:04:41.223743Z digest=sha256:a10bf11490ae9fa188c18a5a196c4c3491f634c0d781b2338171ff4b12ef0ccc

Observation 75240d8a-1d5d-40ba-b438-c96b1600794c · inbound

The Illusion of Safety: Multi-Tier Verification of AI vs. Human C++ Code cites this paper.

The Illusion of Safety: Multi-Tier Verification of AI vs. Human C++ Code SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T09:44:42.646014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:44:42.646014Z digest=sha256:b83f18335a11c7fd256f571cb5fa4c6f233f0ef460218bd7fec1d3de0671fda2

Observation 91d0a92e-fce5-4d73-85f0-61222953f4a8 · inbound

The Illusion of Safety: Multi-Tier Verification of AI vs. Human C++ Code cites this paper.

The Illusion of Safety: Multi-Tier Verification of AI vs. Human C++ Code SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T04:38:06.679657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:38:06.679657Z digest=sha256:27c1f5d8fe70522d1642f3937302ea6f1890c07843483162e7410d281fb9a938

Observation c8762db4-355e-454e-90e2-635ce74067b8 · inbound

The Patchwork Problem in LLM-Generated Code cites this paper.

The Patchwork Problem in LLM-Generated Code SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T01:16:30.101870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:16:30.101870Z digest=sha256:8d827662ebb096076e9536109c1316b4e7a00a1e9b81b6b4985bfb6f90ec104b

Observation b212ebb8-7ab2-4f5b-a7ae-95296fd38590 · inbound

Poster: Rethinking Security in LLM Code Generation through Real-World Risk Scenarios cites this paper.

Poster: Rethinking Security in LLM Code Generation through Real-World Risk Scenarios SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:38.189512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:40:38.189512Z digest=sha256:26055800dd0a6ee9d8584f43346a9c3ca12b9a27fb86eb7c5ee648c5e85c704f

Observation c6ca3504-9d98-4ad8-8114-0bcf1183b914 · inbound

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation cites this paper.

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T00:29:57.375647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:29:57.375647Z digest=sha256:1fd076d497017ff02b9b7cc7bc293b5aafbe6811436ee685e689f08295b10373