Pith. sign in

Paper Citation Record · LEDGER

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

As of 9 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 11 inbound Pith citation observations for arXiv:2507.20439.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20439 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:38:53.742442Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T14:46:30.803463Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact3
  • verified fuzzy1
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 9df98d93-dda8-4a14-b4d0-f058be2027d3 · outbound

This paper cites IEEE Recommended Practice for Software Requirements Specifications.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions IEEE Recommended Practice for Software Requirements Specifications

Reference 1

Resolution
verified exact
raw_fallback, observed 2026-08-06T13:38:55.205137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:47.649609Z digest=sha256:99e042383c31c842f8f78c1d9a3bf711ddd45d7984d77f85c6b4192e8b1405c3

Observation a27cb455-ca01-4845-932d-4b37bcdd76f2 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:47.768131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:47.768131Z digest=sha256:adb5a40b752f24c46727e26f9313bb59025be28daf1de8d5e33afbadc52bc7b8

Observation 23de4ca7-f7e8-4054-8996-c4466765f738 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:39:00.084889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:47.886462Z digest=sha256:127b5028ef5f23e76bbc8c5d10cb34815b93cddec668dc8c5ec9e7593601d7b2

Observation c6a564b6-62bf-444e-bc17-df086f55d3ee · outbound

This paper cites Program Synthesis with Large Language Models.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Program Synthesis with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:48.144648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:48.144648Z digest=sha256:69b4af376a2a16ee4a46548635bdc4d5d0eca9d75b461c21360c0b53a87d4ad7

Observation 0fac6cd1-51e4-41fa-9b8b-0bb56846416b · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:59.929664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:48.298871Z digest=sha256:dbf555203d8b0bd885ecbc190eb490a417b8e045d260e49f42acb64ffef80e93

Observation 87fd0afb-5e94-4171-9253-850d8306a7ec · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:59.673488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:48.420423Z digest=sha256:c1c3a3c89bec95c4655a07885f631ec38cb4f808203a267f4b087339d8c8dd3f

Observation eff87302-791c-471f-ba3b-c2a38e176e51 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Evaluating Large Language Models Trained on Code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:48.655287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:48.655287Z digest=sha256:b8dc1ea582ef77b3543ed31a5a0bc8e6f10548ba3ddde6c5e73ccf0e7cfc63de

Observation b13455d5-6388-4554-8959-e2e884814991 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:48.816979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:48.816979Z digest=sha256:1888bc5a390440a0b00d7dd96433e638c5f8a5c928976cd16f5723949e2d3404

Observation b6a20980-cb7d-4729-a40e-20159448581b · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:49.100177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:49.100177Z digest=sha256:6447b9db1a9f52f57663dca2b8b8ce0f96254dba025a17ea7f1990e6f0e1a73b

Observation 365d18e0-4cd8-49f8-9f39-dde997ede50a · outbound

This paper cites Systematic Evaluation of GPT-3 for Zero-Shot Personality Estimation.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Systematic Evaluation of GPT-3 for Zero-Shot Personality Estimation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:49.217874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:49.217874Z digest=sha256:c45dbb2cff07fe783e1f38acb8f0d76cd657547c7008a5678ccf0c2a8b487dc1

Observation 161d3724-ede3-4f83-94cc-e2353ac70e94 · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Measuring Coding Challenge Competence With APPS

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:49.326645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:49.326645Z digest=sha256:a55e4a08e2d1992cf4e7129cc726ce47b04244dabc6e929a8f235ba6fad0d716

Observation 725e18f3-d124-42a1-99bc-91ff4d39a635 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:49.436262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:49.436262Z digest=sha256:e2c06d38421693b01f63ec8c93e22af93a5e28dd7c24b0ad0a6457e0e3564e69

Observation 8c056224-3794-45f4-8d08-3fad5c1eee98 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:49.547770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:49.547770Z digest=sha256:42ec82885170aa6cb42cf33234fe09c6ec9516768247e27dc1ee2cd7c36c0a3d

Observation ca011c1e-cbf9-4062-8c5e-770820f1c436 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:59.481548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:49.766965Z digest=sha256:945df40576a8f3e26561bb916b8d43174d857d2f0d25859479bbaf2caaa03db0

Observation 2fed75e1-c953-40cd-b241-b8c68755a4c1 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:59.337232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:49.914221Z digest=sha256:7a694a99829ac189c190fbde41b6be5e1a3af58892cbf7037eb1845680822c56

Observation f85e91c9-e344-4da7-833b-a4e2fd9a7e11 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:59.122903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:50.005883Z digest=sha256:c6740c1d7d41d60a10eaca0065cc5967e9c8d42cc545e58a5164c3232a830581

Observation ee94a9d0-b71b-42cf-8498-dc7de55822a9 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:58.949258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:50.063162Z digest=sha256:64e08722201dfb3ccd480fe31381e8faf29e9ac422bc76f509daaf5804f164b9

Observation 5358355d-f85f-46c2-ab56-11a97fdb9406 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:50.268739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:50.268739Z digest=sha256:44b53ae2427d76c644338c2c48d9e2846fd8de66a30341520620ab1f76751edb

Observation 20e2a215-9354-4920-b9b5-6e3907489b99 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 22

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T13:38:54.736511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:50.390088Z digest=sha256:13e53669abb018421eb9330aed90044be7a03d2e1dce8bf5dfd85a58e4acce54

Observation a18c74fa-41c2-4b64-9e60-c3939ed13364 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:50.555980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:50.555980Z digest=sha256:523a5f82db5faf0e3a14fa7c1d6048365bd05175cbc4bb251ed5f6f00e39a4fc

Observation aface219-411c-4fbe-9a10-cfbfea79d036 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:58.616747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:50.824677Z digest=sha256:6393da12170acfef6a177bce9ba5fc5ea56b5c0924d5b313754b88a93e50084d

Observation 982fa63d-cedf-4c65-9275-22cf2d04f7b1 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 25

Resolution
verified exact
doi, observed 2026-08-06T13:38:54.019782Z

Source-reported events for the cited work

correction dated 2022-04-04. Source: crossref record 10.1007/s00766-022-00378-4->10.1007/s00766-021-00367-z:correction, observed 2026-07-11T02:57:50.434999+00:00. This notice travels one citation hop only.

source=pdf_text observed=2026-08-06T13:38:50.698361Z digest=sha256:65ad52c87441f7fe07d426eb2f279bb21b7607b19b71500cd121c5b8529fe514

Observation fa8f5d74-5fa1-4823-8f1e-083426e837c7 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:51.181674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:51.181674Z digest=sha256:6a514e085fb3c8e1d6ea25c990211d83612b0452e3e5decb5820c2473030469d

Observation 39c13da2-363a-4edd-8061-d9b3ceb02ba4 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:58.387181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:51.043563Z digest=sha256:33649be92ceaf53af22c3f4dc2b4063e2eba7a0c69670999f9e0bf21e724da2f

Observation 45e8c2d1-c011-4765-8ee6-54bf5031f8ad · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:57.675628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:51.425304Z digest=sha256:7a69b38cb87ed86734d05ecbb385c994fbf00c1d9a4525fbe11671e6e412aeda

Observation 74c4d79d-8833-4016-9175-2a3328b65748 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:58.007097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:51.285484Z digest=sha256:51fbaad22c4dd7fc37c35e95e3d8933f49c85aa5a55900502b53c40c8f61afe2

Observation e34977af-ddde-4c41-b37f-5331bc09f3cb · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:57.302392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:51.687667Z digest=sha256:f2b56cc3f66ca95b4b632fb75f662dcd5e77f4fc9114386a295ae80969a8497c

Observation dc21c2dd-ad0a-4abe-bfc5-21ccb202e113 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:57.436680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:51.583331Z digest=sha256:cccd855b896a1766c26121680ce3454f5e4154638b40e28796ffef755e4e355b

Observation 7f9f76a7-6b47-4e1d-8575-5cde5af1b13c · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:57.053010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:51.934242Z digest=sha256:89284cf629c54b3ca11c7b34ab22edbea83551c0e413a4988e545dcc1073788c

Observation 7362a76e-f3f8-4a6d-b97a-5d3d60c1cc65 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Code Llama: Open Foundation Models for Code

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:51.815907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:51.815907Z digest=sha256:2b894119039a5727c84faa153bac80a1d1d3eae904621ee73804c3cd744ecea0

Observation 61318ea6-a0f4-41c4-8f1b-7d14faacdcbc · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:56.559707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:52.256132Z digest=sha256:279c619f5abfd773e40f97376a2be777a0788eb3fe1979fee6ce380ce439bef8

Observation 2542d470-b572-492f-803d-c90b569f308b · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:56.839141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:52.141096Z digest=sha256:b4bed4f6508d4d7ee6925a5fec7e973444899ca1df80c50a5f1effc1009c6aa2

Observation 353dce04-6766-4e73-870b-f95226b2934c · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:56.312009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:52.526338Z digest=sha256:ae0f936094564c4e5756ba40127aac1fea9ae1e133bcc71e9ba7f3b99339d416

Observation e37dd0df-15a1-48b8-85a3-2fda37ee3973 · outbound

This paper cites Recent Advances in Software Effort Estimation using Machine Learning.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Recent Advances in Software Effort Estimation using Machine Learning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:38:54.305710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:52.374590Z digest=sha256:df070ca464f9cddd043c67a14910403746caaaa0cdf85be989d0dc4d6d9aff8a

Observation 0cfa7f88-6cea-4a69-8188-34069b47c5d3 · outbound

This paper cites Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:52.828980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:52.828980Z digest=sha256:5d8c4810db0d74d8f81bd2c02b50a424c728355b0104c8ad3bb6e3c72db056f4

Observation 8d00f9a7-7715-4037-9396-90903b31b54f · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:56.058687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:52.695032Z digest=sha256:4ea311370ede844ccdf54ff3920983c0e47552d7e58e239c317eb1aaae7f81c8

Observation cbc77a68-d9ca-4d9d-a33d-29c36136f676 · outbound

This paper cites DeceptPrompt: Exploiting LLM-driven Code Generation via Adversarial Natural Language Instructions.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions DeceptPrompt: Exploiting LLM-driven Code Generation via Adversarial Natural Language Instructions

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:53.175435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:53.175435Z digest=sha256:824a299a852d779afa765847aa46235660164a163d527def31f6c59c97f5c6d0

Observation 8a3ba20e-0ce5-4251-8cc3-30bc68fb2477 · outbound

This paper cites ReCode: Robustness Evaluation of Code Generation Models.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions ReCode: Robustness Evaluation of Code Generation Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:53.002137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:53.002137Z digest=sha256:981b7211c306042af45eee8c7ee32b145d0de9d0db1263fb64ff6094f7dcf277

Observation 71430e1e-2df9-4de9-8f9f-1a75c545da80 · outbound

This paper cites LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T13:38:53.452180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:38:53.452180Z digest=sha256:22018c0ce0e281cef9e32c3384bd2d530e9a9d711569d6499b623f96e8640e57

Observation f86c9d45-f912-4324-9e17-cf9f78ff88d5 · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:55.836809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:53.396905Z digest=sha256:8eb2c368ff4f62a6040a120cf6b7f0d28b06d26aae0eea72dd1b249078f3c6cd

Observation a1570c58-f6a5-43b8-b3bc-45a685e623fc · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:55.486547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:53.742442Z digest=sha256:c319f55a08c1203b5963325f28632c4133eb5baf0cc60f2f76253a9eb3ff5564

Observation 365c216c-8b0a-48cd-a867-06d1e6cf07da · outbound

This paper cites an unresolved cited work.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:38:55.691351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:53.611065Z digest=sha256:0c6f6cac4cd16ed651648165f3fa576363bf1925d2179e58bd4dfe9868a38214

Observation f0cc826d-8542-49ee-b570-66ae34b1b1de · outbound

This paper cites Software: Practice and experience 52, 1 (2022), 39–65.

When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions Software: Practice and experience 52, 1 (2022), 39–65

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:38:58.753242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:38:50.181171Z digest=sha256:7893f0de440a879168c40899afeb9562865321f350c9f024d46543609e22cd85

Pith citing papers

Observation 198ab81f-943e-426d-98b9-64b4c293c359 · inbound

Defective Task Descriptions in LLM-Based Code Generation: Detection and Analysis cites this paper.

Defective Task Descriptions in LLM-Based Code Generation: Detection and Analysis When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:16:34.336735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T03:07:06.924437Z digest=sha256:1bb580ff8dbbd9271edcee05a60b11205a93dcd8e690f1c1a105b6ce01a74dfe

Observation 70807c01-e5e1-4017-aeca-80197e2fbaef · inbound

When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation cites this paper.

When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:17:08.166144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T03:00:26.137401Z digest=sha256:ee6a35859a277c3355dc90df42436dd04b45468250b6f2c0c1929451ba1eb0de

Observation 77d8fa42-aa49-4696-b4b2-df3eb3d40d82 · inbound

ClarifyCodeBench: Evaluating LLMs on Clarifying Ambiguous Requirements for Code Generation cites this paper.

ClarifyCodeBench: Evaluating LLMs on Clarifying Ambiguous Requirements for Code Generation When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:46:48.648730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T08:39:01.987915Z digest=sha256:4ff23169c89e6b49e301d65e79d0da712f0cddfd4d70acf7abf77e10898398d5

Observation fe9cb4ff-ad6e-4b67-b6d8-139a902c8018 · inbound

Underspecification does not imply Incoherence: The Risks of Semantic Collapse in Coding Models cites this paper.

Underspecification does not imply Incoherence: The Risks of Semantic Collapse in Coding Models When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:07:47.017340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T09:01:48.501745Z digest=sha256:b5d7f73e6edf249922dafbf433d39e12f571395202cd11eaa14764c2642a38d7

Observation b9f16e6d-8e35-411a-8ed3-6728eba63bfd · inbound

Guiding Human Validation of LLM-Generated Code via Verifiable Literate Programming cites this paper.

Guiding Human Validation of LLM-Generated Code via Verifiable Literate Programming When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:47:49.713356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T08:42:04.804042Z digest=sha256:5351e57f5b4f5ee91df006a2eb9dcb0320528cdd52fa81d6eb8ecedb5b5973dd

Observation bb9764c9-2676-429e-8063-7ebddb3de03f · inbound

Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse cites this paper.

Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T17:45:46.872944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T17:45:46.872944Z digest=sha256:24c395a38f36ed8de9cff27f980bcfc13f576f1f3b5e75adc4c5e8a8e7a7e94c

Observation 158d4e41-4d94-43b7-a6fc-b6a5bb766ec5 · inbound

From Failing to Passing: Evolving Natural Language Prompt Optimization Rules for LLM Code Generation cites this paper.

From Failing to Passing: Evolving Natural Language Prompt Optimization Rules for LLM Code Generation When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T08:31:24.850506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:31:24.850506Z digest=sha256:f609ea29fecbc83dc0a95b86169502bd6516462540f663233d6b8f3724488c90

Observation 1bd8e22d-9906-4da1-90b9-62a267294cc7 · inbound

On the risk of coding before testing: An empirical study on LLM-based test generation workflow cites this paper.

On the risk of coding before testing: An empirical study on LLM-based test generation workflow When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-11T08:12:23.512737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:12:23.512737Z digest=sha256:9af85d6da2938b04d8c2f12355b27b9730d47889330057fe5d80a80ba735c38f

Observation 4d2e8f86-3a11-4e4b-bc0e-b731365ec346 · inbound

Automatically Evolving Prompt Guidelines for Task-Specific Optimization cites this paper.

Automatically Evolving Prompt Guidelines for Task-Specific Optimization When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T14:46:30.803463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:46:30.803463Z digest=sha256:a46d09d293487e87180354cccc9c51528d7c087f4a53b1a335dce399919b27ea

Observation c853259f-c273-4ff6-b6a2-3bac8e2e1e11 · inbound

AssumptionMiner: Extracting, Tracing, and Revising Implicit Assumptions in LLM Code Generation cites this paper.

AssumptionMiner: Extracting, Tracing, and Revising Implicit Assumptions in LLM Code Generation When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T04:15:29.179347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:15:29.179347Z digest=sha256:cbe01ca6c6edbf05d35214824b05293b5475154b2118fbed765dbfdda8e071b6

Observation 2b7984f4-f449-4cc1-9e8d-bd32b8930620 · inbound

VClare: Resolving Imperfect Specifications in LLM-Based Verilog Generation cites this paper.

VClare: Resolving Imperfect Specifications in LLM-Based Verilog Generation When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-30T23:35:19.791382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T23:35:19.791382Z digest=sha256:428d919fd76b5fb72f6e0b957e271a5247e5960f1f9ba9f6545ca7199d9c22c0