Pith. sign in

Paper Citation Record · LEDGER

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

As of 7 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 11 inbound Pith citation observations for arXiv:2507.20439.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20439 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:38:53.742442Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T14:46:30.803463Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact3
  • verified fuzzy1
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 9df98d93-dda8-4a14-b4d0-f058be2027d3 · outbound

This paper cites IEEE Recommended Practice for Software Requirements Specifications.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions IEEE Recommended Practice for Software Requirements Specifications

Reference 1

Resolution
verified exact
raw_fallback, observed 2026-08-06T13:38:55.205137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:47.649609Z digest=sha256:9bde6ad23353e24078654a92e1cbc48677c78134d0c219ef1cd3496723bafcae

Observation a27cb455-ca01-4845-932d-4b37bcdd76f2 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:47.768131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:47.768131Z digest=sha256:adb5a40b752f24c46727e26f9313bb59025be28daf1de8d5e33afbadc52bc7b8

Observation 23de4ca7-f7e8-4054-8996-c4466765f738 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:39:00.084889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:47.886462Z digest=sha256:c09f10b0b7e8c952f9fa6762a3d4b35560507bafbf1a7597e41b52e9441abf88

Observation c6a564b6-62bf-444e-bc17-df086f55d3ee · outbound

This paper cites Program Synthesis with Large Language Models.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Program Synthesis with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:48.144648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:48.144648Z digest=sha256:69b4af376a2a16ee4a46548635bdc4d5d0eca9d75b461c21360c0b53a87d4ad7

Observation 0fac6cd1-51e4-41fa-9b8b-0bb56846416b · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:59.929664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:48.298871Z digest=sha256:7f4be98c5ef802b61689dffcab37251515821ce3e1c529b98f70ccd5cb1114e9

Observation 87fd0afb-5e94-4171-9253-850d8306a7ec · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:59.673488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:48.420423Z digest=sha256:949da1c92fc9dd15357b52a81f0806dcc603ad42856dd09dec161c1c56d4c969

Observation eff87302-791c-471f-ba3b-c2a38e176e51 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Evaluating Large Language Models Trained on Code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:48.655287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:48.655287Z digest=sha256:a849dc4404de7981eb91c9f521f1326e009e4364d02b6c5824308961db6aa709

Observation b13455d5-6388-4554-8959-e2e884814991 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:48.816979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:48.816979Z digest=sha256:1888bc5a390440a0b00d7dd96433e638c5f8a5c928976cd16f5723949e2d3404

Observation b6a20980-cb7d-4729-a40e-20159448581b · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:49.100177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:49.100177Z digest=sha256:6447b9db1a9f52f57663dca2b8b8ce0f96254dba025a17ea7f1990e6f0e1a73b

Observation 365d18e0-4cd8-49f8-9f39-dde997ede50a · outbound

This paper cites Systematic Evaluation of GPT-3 for Zero-Shot Personality Estimation.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Systematic Evaluation of GPT-3 for Zero-Shot Personality Estimation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:49.217874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:49.217874Z digest=sha256:c45dbb2cff07fe783e1f38acb8f0d76cd657547c7008a5678ccf0c2a8b487dc1

Observation 161d3724-ede3-4f83-94cc-e2353ac70e94 · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Measuring Coding Challenge Competence With APPS

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:49.326645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:49.326645Z digest=sha256:a55e4a08e2d1992cf4e7129cc726ce47b04244dabc6e929a8f235ba6fad0d716

Observation 725e18f3-d124-42a1-99bc-91ff4d39a635 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:49.436262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:49.436262Z digest=sha256:e2c06d38421693b01f63ec8c93e22af93a5e28dd7c24b0ad0a6457e0e3564e69

Observation 8c056224-3794-45f4-8d08-3fad5c1eee98 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:49.547770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:49.547770Z digest=sha256:42ec82885170aa6cb42cf33234fe09c6ec9516768247e27dc1ee2cd7c36c0a3d

Observation ca011c1e-cbf9-4062-8c5e-770820f1c436 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:59.481548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:49.766965Z digest=sha256:cb551978b18063750f30f245d9be3e1401175c8529ccb183e0363c194779351d

Observation 2fed75e1-c953-40cd-b241-b8c68755a4c1 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:59.337232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:49.914221Z digest=sha256:0ce6474aa5940380a9c2102c16be730f620412908eb54179d8e5518ffa0218b2

Observation f85e91c9-e344-4da7-833b-a4e2fd9a7e11 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:59.122903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:50.005883Z digest=sha256:ff2e2bcc9c6f8dac244eece60d70c882724b4eeb9b835f7c0c1a9c83dae80017

Observation ee94a9d0-b71b-42cf-8498-dc7de55822a9 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:58.949258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:50.063162Z digest=sha256:afcaa27526a1cb95e164ca9e28f531fbd77c713e581e9b678c857a68ad909204

Observation 5358355d-f85f-46c2-ab56-11a97fdb9406 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:50.268739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:50.268739Z digest=sha256:44b53ae2427d76c644338c2c48d9e2846fd8de66a30341520620ab1f76751edb

Observation 20e2a215-9354-4920-b9b5-6e3907489b99 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 22

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T13:38:54.736511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:50.390088Z digest=sha256:fed876a6009e72b9a98465dabd67b93a08dfa0957969a8804309b3555ddbc0c4

Observation a18c74fa-41c2-4b64-9e60-c3939ed13364 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:50.555980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:50.555980Z digest=sha256:523a5f82db5faf0e3a14fa7c1d6048365bd05175cbc4bb251ed5f6f00e39a4fc

Observation aface219-411c-4fbe-9a10-cfbfea79d036 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:58.616747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:50.824677Z digest=sha256:9fb93f3a4eaebcc23e1b57dc470a028b86993bbb66cbc9cb6afa7fd74789f65d

Observation 982fa63d-cedf-4c65-9275-22cf2d04f7b1 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 25

Resolution
verified exact
doi, observed 2026-08-06T13:38:54.019782Z

Source-reported events for the cited work

correction dated 2022-04-04. Source: crossref record 10.1007/s00766-022-00378-4->10.1007/s00766-021-00367-z:correction, observed 2026-07-11T02:57:50.434999+00:00. This notice travels one citation hop only.

source=pdf_text observed=2026-08-06T13:38:50.698361Z digest=sha256:54dbd03dde57c7a3b2a585fcd75b5e863f01618fa7c587d3232aa042a823f9c6

Observation fa8f5d74-5fa1-4823-8f1e-083426e837c7 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:51.181674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:51.181674Z digest=sha256:6a514e085fb3c8e1d6ea25c990211d83612b0452e3e5decb5820c2473030469d

Observation 39c13da2-363a-4edd-8061-d9b3ceb02ba4 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:58.387181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:51.043563Z digest=sha256:6cc139cbaf19a21c165ccf0c91f6e4814fae5267cf3ff5b5bf94b44396baefa5

Observation 45e8c2d1-c011-4765-8ee6-54bf5031f8ad · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:57.675628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:51.425304Z digest=sha256:db354925eb205a16fc9c2f54d88469750f5f439fc5442c24b32edd65121f197b

Observation 74c4d79d-8833-4016-9175-2a3328b65748 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:58.007097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:51.285484Z digest=sha256:fb9289eaeba2878df5a892a7b7bd47b47c45f2cb7ad079dbb720ea31469067a1

Observation e34977af-ddde-4c41-b37f-5331bc09f3cb · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:57.302392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:51.687667Z digest=sha256:4333b81ac874b1e3ee38152ed2a70abc23ea35d637c9a84674055f8f860216b1

Observation dc21c2dd-ad0a-4abe-bfc5-21ccb202e113 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:57.436680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:51.583331Z digest=sha256:ad6af93fe8e6d4ba1c2fb6697aa2fe6f1770495f08de9ad29b561510e9f5c676

Observation 7f9f76a7-6b47-4e1d-8575-5cde5af1b13c · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:57.053010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:51.934242Z digest=sha256:1aa9ece18934893451a53771543fa79606a9206c22546cfc0bbbd23e83f89421

Observation 7362a76e-f3f8-4a6d-b97a-5d3d60c1cc65 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Code Llama: Open Foundation Models for Code

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:51.815907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:51.815907Z digest=sha256:2b894119039a5727c84faa153bac80a1d1d3eae904621ee73804c3cd744ecea0

Observation 61318ea6-a0f4-41c4-8f1b-7d14faacdcbc · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:56.559707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:52.256132Z digest=sha256:3a5fe36dc48181559161a2dfc4e40bd93f433306d62da590409f726d1d4c6e01

Observation 2542d470-b572-492f-803d-c90b569f308b · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:56.839141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:52.141096Z digest=sha256:4c3470b9a3dd3c177b8fab96fc3db102338f8611ec6e8924c91f85e77f644c84

Observation 353dce04-6766-4e73-870b-f95226b2934c · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:56.312009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:52.526338Z digest=sha256:7d26350b1841508d2e5c9448decd77f6d42ed32528489400b943404e8b23321d

Observation e37dd0df-15a1-48b8-85a3-2fda37ee3973 · outbound

This paper cites Recent Advances in Software Effort Estimation using Machine Learning.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Recent Advances in Software Effort Estimation using Machine Learning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:38:54.305710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:52.374590Z digest=sha256:5678f1536fa4d74ba1458134285e2d4885e7dd7cdac8b7627e7c73d406438b7c

Observation 0cfa7f88-6cea-4a69-8188-34069b47c5d3 · outbound

This paper cites Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:52.828980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:52.828980Z digest=sha256:5d8c4810db0d74d8f81bd2c02b50a424c728355b0104c8ad3bb6e3c72db056f4

Observation 8d00f9a7-7715-4037-9396-90903b31b54f · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:56.058687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:52.695032Z digest=sha256:04a0390ba337993269e9e3f7d02c3480628db9edba7e41d3ff8077aa4b9a14d6

Observation cbc77a68-d9ca-4d9d-a33d-29c36136f676 · outbound

This paper cites DeceptPrompt: Exploiting LLM-driven Code Generation via Adversarial Natural Language Instructions.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions DeceptPrompt: Exploiting LLM-driven Code Generation via Adversarial Natural Language Instructions

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:53.175435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:53.175435Z digest=sha256:824a299a852d779afa765847aa46235660164a163d527def31f6c59c97f5c6d0

Observation 8a3ba20e-0ce5-4251-8cc3-30bc68fb2477 · outbound

This paper cites ReCode: Robustness Evaluation of Code Generation Models.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions ReCode: Robustness Evaluation of Code Generation Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:53.002137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:53.002137Z digest=sha256:981b7211c306042af45eee8c7ee32b145d0de9d0db1263fb64ff6094f7dcf277

Observation 71430e1e-2df9-4de9-8f9f-1a75c545da80 · outbound

This paper cites LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:53.452180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:53.452180Z digest=sha256:22018c0ce0e281cef9e32c3384bd2d530e9a9d711569d6499b623f96e8640e57

Observation f86c9d45-f912-4324-9e17-cf9f78ff88d5 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:55.836809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:53.396905Z digest=sha256:21d71b68941661f7710fa62c9cc82644a8c2d12a700adac837f535d7744cb636

Observation a1570c58-f6a5-43b8-b3bc-45a685e623fc · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:55.486547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:53.742442Z digest=sha256:e68da5832726f1c32e7ae040cf85a8366f99bf846844081eb1bfca68c16479cb

Observation 365c216c-8b0a-48cd-a867-06d1e6cf07da · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:55.691351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:53.611065Z digest=sha256:e564c3cee67e615a029a252b3300eb0ab41838431b557bc5273dff440cd5c75c

Observation f0cc826d-8542-49ee-b570-66ae34b1b1de · outbound

This paper cites Software: Practice and experience 52, 1 (2022), 39–65.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Software: Practice and experience 52, 1 (2022), 39–65

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:58.753242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:38:50.181171Z digest=sha256:8f817c92629b7965408098139eb1780d7ea788dcb71a030117397405fcf0facb

Pith citing papers

Observation 198ab81f-943e-426d-98b9-64b4c293c359 · inbound

Defective Task Descriptions in LLM-Based Code Generation: Detection and Analysis cites this paper.

Defective Task Descriptions in LLM-Based Code Generation: Detection and Analysis When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:16:34.336735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T03:07:06.924437Z digest=sha256:d819e68b38c966fd619c0c5ca0e4172eb95d38887a4cc598da4cf4cc5d6f4559

Observation 70807c01-e5e1-4017-aeca-80197e2fbaef · inbound

When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation cites this paper.

When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:17:08.166144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T03:00:26.137401Z digest=sha256:fdb84ba83fe81d8e8acfb5ec09678a25988c33e331d69c4a4784273daffe34f4

Observation 77d8fa42-aa49-4696-b4b2-df3eb3d40d82 · inbound

ClarifyCodeBench: Evaluating LLMs on Clarifying Ambiguous Requirements for Code Generation cites this paper.

ClarifyCodeBench: Evaluating LLMs on Clarifying Ambiguous Requirements for Code Generation When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:46:48.648730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T08:39:01.987915Z digest=sha256:e38b51c6396b68849b180dfae2d18d837520301bbc7fcba0c1ba5c638ca6fdb7

Observation fe9cb4ff-ad6e-4b67-b6d8-139a902c8018 · inbound

Underspecification does not imply Incoherence: The Risks of Semantic Collapse in Coding Models cites this paper.

Underspecification does not imply Incoherence: The Risks of Semantic Collapse in Coding Models When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:07:47.017340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T09:01:48.501745Z digest=sha256:ce5fa29f1011725882f9e9a43bf89d6a7d54517007b0acfac57b859bb9262ddd

Observation b9f16e6d-8e35-411a-8ed3-6728eba63bfd · inbound

Guiding Human Validation of LLM-Generated Code via Verifiable Literate Programming cites this paper.

Guiding Human Validation of LLM-Generated Code via Verifiable Literate Programming When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:47:49.713356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T08:42:04.804042Z digest=sha256:900b6394613136dac3f1a28d8941b6261b6f56101aa4dd78bb385ae2e1689148

Observation bb9764c9-2676-429e-8063-7ebddb3de03f · inbound

Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse cites this paper.

Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T17:45:46.872944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T17:45:46.872944Z digest=sha256:56178b34ad2bfedd3d6047a702cefc3ef3bf5496bb998ebf29251ab76d3da907

Observation 158d4e41-4d94-43b7-a6fc-b6a5bb766ec5 · inbound

From Failing to Passing: Evolving Natural Language Prompt Optimization Rules for LLM Code Generation cites this paper.

From Failing to Passing: Evolving Natural Language Prompt Optimization Rules for LLM Code Generation When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T08:31:24.850506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:31:24.850506Z digest=sha256:d9f3854e8e953b96e671d3a890ae24d5badcdbb9c90645b03db151a8fee3359e

Observation 1bd8e22d-9906-4da1-90b9-62a267294cc7 · inbound

On the risk of coding before testing: An empirical study on LLM-based test generation workflow cites this paper.

On the risk of coding before testing: An empirical study on LLM-based test generation workflow When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-11T08:12:23.512737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:12:23.512737Z digest=sha256:9af85d6da2938b04d8c2f12355b27b9730d47889330057fe5d80a80ba735c38f

Observation 4d2e8f86-3a11-4e4b-bc0e-b731365ec346 · inbound

Automatically Evolving Prompt Guidelines for Task-Specific Optimization cites this paper.

Automatically Evolving Prompt Guidelines for Task-Specific Optimization When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T14:46:30.803463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:46:30.803463Z digest=sha256:a46d09d293487e87180354cccc9c51528d7c087f4a53b1a335dce399919b27ea

Observation c853259f-c273-4ff6-b6a2-3bac8e2e1e11 · inbound

AssumptionMiner: Extracting, Tracing, and Revising Implicit Assumptions in LLM Code Generation cites this paper.

AssumptionMiner: Extracting, Tracing, and Revising Implicit Assumptions in LLM Code Generation When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T04:15:29.179347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:15:29.179347Z digest=sha256:cbe01ca6c6edbf05d35214824b05293b5475154b2118fbed765dbfdda8e071b6

Observation 2b7984f4-f449-4cc1-9e8d-bd32b8930620 · inbound

VClare: Resolving Imperfect Specifications in LLM-Based Verilog Generation cites this paper.

VClare: Resolving Imperfect Specifications in LLM-Based Verilog Generation When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-30T23:35:19.791382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T23:35:19.791382Z digest=sha256:cd9fc4e6389f85c49a434c025079e6cf6c9da1ee1be248ccfb52c54b88c92995