Pith. sign in

Paper Citation Record · LEDGER

What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2406.12334.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.12334 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:21:49.563936Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T22:00:41.512733Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0df0e3fd-7ad0-4a67-84ae-ed12264f1c4e · inbound

Adaptive Prompting: Ad-hoc Prompt Composition for Social Bias Detection cites this paper.

Adaptive Prompting: Ad-hoc Prompt Composition for Social Bias Detection What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T15:21:49.563936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:21:49.563936Z digest=sha256:d15d9407f8fb679b8cb8e91b0cf2f2487e561213b6e878d3a45ff2678753565b

Observation b40e3e17-40fd-43cc-8829-2a2244769ae2 · inbound

Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets cites this paper.

Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:21.562642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:21.562642Z digest=sha256:6b38726fb5bc74263648b0a3355271f1a55293ffd38b7fe811b07877462e0f4b

Observation 4f02dd20-f97d-4038-8f17-d0972a1d0d27 · inbound

Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs cites this paper.

Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.027058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.027058Z digest=sha256:87309078c88462be769b373e2dd8fc46c3325a57d80516a1d254d160b5bd2275

Observation 8619f564-aa14-42b3-8b44-b34998980077 · inbound

CEA-LIST at CheckThat! 2025: Evaluating LLMs as Detectors of Bias and Opinion in Text cites this paper.

CEA-LIST at CheckThat! 2025: Evaluating LLMs as Detectors of Bias and Opinion in Text What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:42:30.735057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:42:30.735057Z digest=sha256:d40ee56297f6a2a8411c95c9d707cf3db5457e7c41736deea3439f8b400912f0

Observation 69b5a0ef-14d0-4054-87cb-d1499b3b3125 · inbound

A Conceptual Framework for Requirements Engineering of Pretrained-Model-Enabled Systems cites this paper.

A Conceptual Framework for Requirements Engineering of Pretrained-Model-Enabled Systems What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:09.624223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:09.624223Z digest=sha256:e36343e28567e35d9af55ba638a7495b7a7c79f63324e2eb31fc3bc2ec538e4c

Observation 4f199e73-cd56-421b-91c0-3e0df4fdbbd3 · inbound

Position: Intelligent Coding Systems Should Write Programs with Justifications cites this paper.

Position: Intelligent Coding Systems Should Write Programs with Justifications What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T23:01:59.849579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:01:59.849579Z digest=sha256:ffb54a6a8cc9d9b0836a426af11e22cd621b18b53b2215983661d8bd3844c019

Observation f040e31f-271f-434e-905f-a6f0fe38d6ab · inbound

Reinforcement Learning for Machine Learning Engineering Agents cites this paper.

Reinforcement Learning for Machine Learning Engineering Agents What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T12:24:02.310282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:24:02.310282Z digest=sha256:07a81ce7a1118243e0e75cb82622552dffb106dd18378716420847bc3bd49c09

Observation ce73f51e-41f1-457e-a692-db9a25da401d · inbound

From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models cites this paper.

From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-04T19:27:04.900393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:27:04.900393Z digest=sha256:3256379e88fe9480799cb6c58d35e541a6c2fe687999996c2822cc08b5f19f4e

Observation 658d5516-63eb-40d5-8f06-36bb65bf298a · inbound

When the Code Autopilot Breaks: Why LLMs Falter in Embedded Machine Learning cites this paper.

When the Code Autopilot Breaks: Why LLMs Falter in Embedded Machine Learning What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T17:23:21.474317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:23:21.474317Z digest=sha256:c5718fff74aa7e5fdf1a9ff47b5e98c66988722733a626f4f393b500e096d2fb

Observation 0e4f38a2-3ed7-4f93-a100-f8d063abd32e · inbound

Position: AI Evaluations Should be Grounded on a Theory of Capability cites this paper.

Position: AI Evaluations Should be Grounded on a Theory of Capability What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:00:41.515755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T21:57:55.834632Z digest=sha256:7499b0ce7ae85ed6cafa04c8e826124946fa3215d4ac117e5aab7c1e48bb88f2

Observation 6dbf4a2b-625a-4177-bcbd-c9427ae46443 · inbound

PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data cites this paper.

PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:18:40.050451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T23:15:52.444217Z digest=sha256:f2dd8232beabc0bf2afc3244a050fe1d0b6309c451275242a9b73dc8c7ff947b

Observation 9ef101a1-d5e1-4a5b-911a-080263920c69 · inbound

Information-Consistent Language Model Recommendations through Group Relative Policy Optimization cites this paper.

Information-Consistent Language Model Recommendations through Group Relative Policy Optimization What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:08:36.166115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T22:05:15.926436Z digest=sha256:915ad6d3d23dc584325feba26135d3f16d48e46dd8e4b0943af6639dbc9a7779

Observation b8fb136c-760a-4ab8-af63-520cecba3002 · inbound

Similar Pattern Annotation via Retrieval Knowledge for LLM-Based Test Code Fault Localization cites this paper.

Similar Pattern Annotation via Retrieval Knowledge for LLM-Based Test Code Fault Localization What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:00:56.656182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T02:58:17.585369Z digest=sha256:36bbcae986a989f31e0921ac5de943fbd572b5ce9f57be671027f7ec1ddae940