Pith. sign in

Paper Citation Record · LEDGER

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity

As of 13 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2411.16239.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16239 v3

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:27:06.770620Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d6223f76-b8bf-47fc-9623-9235a82aa926 · outbound

This paper cites The multiple-choice question should provide four answer options, with non-correct options being similar or related to the correct answer.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity The multiple-choice question should provide four answer options, with non-correct options being similar or related to the correct answer

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:07.079989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:27:06.701093Z digest=sha256:408be12d40fbc7db9fcb393603c7c6f0512c441e3c1480cbc052c49e218771c4

Observation 38462fc9-d789-48ce-8a39-a724681605a7 · outbound

This paper cites Holistic Evaluation of Language Models.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity Holistic Evaluation of Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T13:27:06.690101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:27:06.690101Z digest=sha256:08d2f740e7253a6feb1f55a0c566b3c9bc624464ef0255b053d32f1ff6084e31

Observation d0bfada8-5a51-4057-b44e-477677586e40 · outbound

This paper cites Offer four potential answers, making sure that the incorrect options are similar to the correct one.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity Offer four potential answers, making sure that the incorrect options are similar to the correct one

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:07.049896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:27:06.711343Z digest=sha256:0598df5d5e0822d0c9136b829bf57f3c994adbf61fee99f3a83e28643c23b101

Observation 6df72d61-22f0-4855-ab47-4fc5765188aa · outbound

This paper cites Offer yes or no answers.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity Offer yes or no answers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:07.035024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:27:06.716027Z digest=sha256:faaf217b9bb42307e546d500fce494e30b8db7790c8defded62f400de3061cc2

Observation 8c4f9870-5a72-4fed-bbdf-e1c023fc4010 · outbound

This paper cites Provide four possible answers, including the correct one, and make sure the incorrect choices are similar or related to the right answer.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity Provide four possible answers, including the correct one, and make sure the incorrect choices are similar or related to the right answer

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:07.065007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:27:06.706171Z digest=sha256:2303e17c7dfa6d04c8e24caaaf8a6235c978e95ebb175e4d51727866a63c3e7a

Observation d68ade81-76a4-41a5-99d2-0bfc6aa4c32d · outbound

This paper cites For example, if the original question asks about the result of a specific action, the reversed question could ask about the conditions re- quired for that result to occur.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity For example, if the original question asks about the result of a specific action, the reversed question could ask about the conditions re- quired for that result to occur

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:06.942721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:27:06.747016Z digest=sha256:6efa48441fbce4389971c0b4b89fc8e87de7e5ca9a9ff943eb72d603c5a108ca

Observation 98dceb1b-d886-491e-bdd1-fc5ac4448c79 · outbound

This paper cites {Original Question}.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity {Original Question}

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:06.924730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:27:06.751817Z digest=sha256:ebe425760b69de4ab7574cbdbd160f035d9897917fe8357fdb9076334a5fea6a

Observation d10afb71-f4bd-43cc-a5be-7651a4b6a3a4 · outbound

This paper cites Ensure the question still tests the same core concept.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity Ensure the question still tests the same core concept

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:07.020332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:27:06.721287Z digest=sha256:08dc196e769ce670e857713324bd813b75d656a93d57ac493a082ed0db2494e9

Observation 6f5c699c-3d24-4cb3-8fd2-4efc46a1bb4b · outbound

This paper cites {Original Question}.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity {Original Question}

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:07.003477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:27:06.726245Z digest=sha256:9e2d0390f69fe76de59d293b9ef759a5ff32539f25f90d0957f1527981c42e51

Observation 4b3bccfe-1558-42fa-8dc6-b1f553948c51 · outbound

This paper cites Ensure the question remains rel- evant to the original cybersecurity concept.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity Ensure the question remains rel- evant to the original cybersecurity concept

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:06.988704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:27:06.731701Z digest=sha256:ee87e809c9ded74dff07180771ddb50726d11cfa2c84d2847eb46f233363677a

Observation 6f6fb752-538b-4d2b-9ff7-889b47deb4b2 · outbound

This paper cites Ensure the question tests the same concept.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity Ensure the question tests the same concept

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:06.973659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:27:06.737061Z digest=sha256:a858df61d1d554fe7ee383f43e6b9c2878ed1c0227827caab2db6daf76b6ab2c

Observation cfa993c3-36fb-4edb-9287-05d0d6f95d69 · outbound

This paper cites {Original Question}.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity {Original Question}

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:06.958427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:27:06.741964Z digest=sha256:b5a5c7445a083f3491d7aa5925f12589d2adf5cd5f4f8e046b4e7564ac9f3505

Observation 323ec978-3e5a-440d-88a4-d35f6f3171bd · outbound

This paper cites {Original Question}.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity {Original Question}

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:06.909168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:27:06.756418Z digest=sha256:c4906fee7814c1fcb988317e553e80c4d676047e545e29f1056bc94e06c1fbab

Observation 44b35c47-9be7-4d36-9956-5958f5560fa7 · outbound

This paper cites {Original Question} Table 8: Prompts for Question Reformulation Prompts for Dynamic Question Generation, Part 2.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity {Original Question} Table 8: Prompts for Question Reformulation Prompts for Dynamic Question Generation, Part 2

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:06.893060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:27:06.761073Z digest=sha256:491a7e903a85bcb52e8df22e0a117c217089a9d6fe5ffb7da0c8565c21d30484

Observation 916c879d-c879-41e5-b1dc-62767bbb9eee · outbound

This paper cites The questions are: {questions} Reply to me in the following format: ‘‘‘json [’knowledge point 1’, ’knowledge point 2’, ......] ‘‘‘.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity The questions are: {questions} Reply to me in the following format: ‘‘‘json [’knowledge point 1’, ’knowledge point 2’, ......] ‘‘‘

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:06.875943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:27:06.765625Z digest=sha256:fead7e45b70e562d2ec05cd2ce0bf8145ebebb5f81dea97143f72ec546ae9adf

Observation 578995e8-0ca9-487b-8545-69d61c8db5c8 · outbound

This paper cites Then rewrite it as a multiple-choice question by leav- ing one key position blank.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity Then rewrite it as a multiple-choice question by leav- ing one key position blank

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:27:06.859231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:27:06.770620Z digest=sha256:28c3ac91e4554bcb956d6a68ca988bc4c95b0cca82a103e0a898eb80fbf52555

Observation cae6d39c-90b5-48ae-abba-8f30f0f0f551 · outbound

This paper cites A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T13:27:06.683416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:27:06.683416Z digest=sha256:30f1c727d61525b5b78b0b746a55c02de9fe9a32ae73c9a4d52b1531153e9ae4

Observation c1ad2b14-06d5-4e86-a267-0758ad84bb32 · outbound

This paper cites FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models.

CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T13:27:06.695774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:27:06.695774Z digest=sha256:1500bf02a0002c9e6b90985245a24bc347ea61c77e04fe7a993708cb6afadb75

Pith citing papers

No inbound Pith citation observations are available.