Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Prompt Sensitivity in Large Language Models

As of 10 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 3 inbound Pith citation observations for arXiv:2502.06065.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06065 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:55:45.557734Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T13:40:04.275438Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T13:43:19.601845Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact5
  • verified fuzzy11
  • unresolved23
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c7831059-049e-483c-b6a7-5ae62f7a32e5 · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:55:46.262863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.426645Z digest=sha256:a80a259aca9761f2953aff6bda7f88655f837aa69b575888bda849574d86d45f

Observation f0db9cb2-d921-4d21-9879-cb757f49591f · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:55:46.254795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.430337Z digest=sha256:0647e50469d10589ddbdb585c345778408dbcb309bad68a5bb38aa4969c6c29c

Observation 5be17900-1104-40a4-9899-7b9681e98a42 · outbound

This paper cites In: Al- Onaizan,Y.,Bansal,M.,Chen,Y.N.(eds.)Proceedingsofthe2024ConferenceonEmpirical Methods in Natural Language Processing.

Benchmarking Prompt Sensitivity in Large Language Models In: Al- Onaizan,Y.,Bansal,M.,Chen,Y.N.(eds.)Proceedingsofthe2024ConferenceonEmpirical Methods in Natural Language Processing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.433342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.433342Z digest=sha256:511446ca99c5ff6fc81b959152eb12dc5ea78f0fdf62782548beee8b6e978148

Observation 2eea4f62-aa24-40e7-8473-e1da6528e81f · outbound

This paper cites In: European Conference on Information Retrieval.

Benchmarking Prompt Sensitivity in Large Language Models In: European Conference on Information Retrieval

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.246742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.436845Z digest=sha256:b2514e111da3f93b20c35ca9da9875c640d27e121282a4975fe74e0d4e8734ca

Observation dfc17dfd-c25a-4437-8f36-400d00802611 · outbound

This paper cites In: Proceedings of the 2024 AnnualInternationalACMSIGIRConferenceonResearchandDevelopmentinInformation Retrieval in the Asia Pacific Region.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 2024 AnnualInternationalACMSIGIRConferenceonResearchandDevelopmentinInformation Retrieval in the Asia Pacific Region

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.238705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.440027Z digest=sha256:65f9a882ca004f3dcbe62306fed37c5e2fdfbe830b8201940f18e7e89b4e4ac2

Observation afea40b4-363b-46dc-9477-8eb6156e6d6b · outbound

This paper cites In: Proceedings of the 31st ACM International Conference on Information & Knowledge Management.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 31st ACM International Conference on Information & Knowledge Management

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.230432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.443134Z digest=sha256:9d7200426dea97216ab7d1e89f77055ff11c6f9ddbc6d4b7963a383cc0690840

Observation 7540a4d6-a0e8-4039-9ebf-3ebec2efc0df · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:55:46.222231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.446388Z digest=sha256:27538edc4074d69cb3a4e818270cd1366d25094f04de32132cadf626ada65e5d

Observation ba1d22d8-203c-426b-a3f1-4a7e7bc106a8 · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 8

Resolution
verified exact
doi, observed 2026-08-08T16:55:45.695016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.449259Z digest=sha256:ba50118339860f6d2eca46fcf7bae8916446b800487275765460fdbef0ced3d2

Observation ca08d4bd-05fa-42ca-88c8-50f68234115a · outbound

This paper cites In: Proceedings of the 28th ACM International Conference on Information and Knowledge Management.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 28th ACM International Conference on Information and Knowledge Management

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.452428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.452428Z digest=sha256:8f37202c8feadeefd9a51f719f2809a41afd02a02fa8a827b617f9513d25016c

Observation 5104d4e0-1104-4669-b20e-731c025fe304 · outbound

This paper cites What's the Magic Word? A Control Theory of LLM Prompting.

Benchmarking Prompt Sensitivity in Large Language Models What's the Magic Word? A Control Theory of LLM Prompting

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.455564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.455564Z digest=sha256:25b44b05cc4af1434f844f95c485a7c8601ef4265d3394d9a96fabf44ed4ea70

Observation abc38505-6812-43a0-a97b-1d0c66d75fdb · outbound

This paper cites In: European Conference on Information Retrieval.

Benchmarking Prompt Sensitivity in Large Language Models In: European Conference on Information Retrieval

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.213922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.459104Z digest=sha256:49196f429d9455dc71dcc8bffaa8a089c93e7a43982120c4f5fc9228bf87d0ae

Observation 522b75aa-f2cc-49d0-92e0-c56cc7f0eea2 · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 12

Resolution
verified exact
doi, observed 2026-08-08T16:55:45.674197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.464402Z digest=sha256:e8cc9015426db0f32f2b92d55035e3e7b0a7d241f745603c775e593ffcc33183

Observation 73b7a0a9-616b-40b3-b352-2b02ab9cff01 · outbound

This paper cites In: European Conference on Information Retrieval.

Benchmarking Prompt Sensitivity in Large Language Models In: European Conference on Information Retrieval

Reference 13

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T16:55:46.205708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.467411Z digest=sha256:c01e784ac93498574e63da0448d27e6428890d6ab56685edf878d25391a0797b

Observation 0737b338-7258-4e24-833d-e987613140a4 · outbound

This paper cites In: Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Informa- tion Retrieval.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Informa- tion Retrieval

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.470420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.470420Z digest=sha256:8d6b172fda88311244a1172b8961a933e0499818af2bcc44e2df27cc9a375b7b

Observation fb873f31-db0a-42d2-b04a-62f139adcc0c · outbound

This paper cites Unveiling and Manipulating Prompt Influence in Large Language Models.

Benchmarking Prompt Sensitivity in Large Language Models Unveiling and Manipulating Prompt Influence in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.473373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.473373Z digest=sha256:d243f1b138b1c1c0669839ed297c971805c66cbe8d2edc403bcaf531a67e9247

Observation 47e15cfe-6d79-4e2d-ba23-a921e1f3f21b · outbound

This paper cites Information13(2), 83 (2022).

Benchmarking Prompt Sensitivity in Large Language Models Information13(2), 83 (2022)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.197515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.476636Z digest=sha256:2a109581b88a2167780210c335929acb4672d1af662ad3a411e97b89b00bf514

Observation fb4d5384-3382-4e8a-8f76-1a495dcddcc6 · outbound

This paper cites IEEE Access11, 76581–76604 (2023).

Benchmarking Prompt Sensitivity in Large Language Models IEEE Access11, 76581–76604 (2023)

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.189086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.479575Z digest=sha256:3c5d0f85559aa3282a97a3b9696882921e63f63710af9beaa3493145fcfe02cf

Observation ba3f28d9-862f-4301-8967-b5991ef90d38 · outbound

This paper cites In: CIKM (2008).

Benchmarking Prompt Sensitivity in Large Language Models In: CIKM (2008)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.180181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.482509Z digest=sha256:42bcffca6e38fb409593296c4ffc320488890134a2fe9cff5fa701ad23b6d81a

Observation e50f2947-2f99-4298-a5e9-46bda323f118 · outbound

This paper cites In: Proceedings of the 33rd ACM International Confer- ence on Information and Knowledge Management.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 33rd ACM International Confer- ence on Information and Knowledge Management

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.171610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.485697Z digest=sha256:1388f3fe7ecd607c80dc067083bba344b002f86f86567b9e2cf09a9cde6bce17

Observation 1108b83d-6dbb-42b0-8f06-b3c492a20636 · outbound

This paper cites Mistral 7B.

Benchmarking Prompt Sensitivity in Large Language Models Mistral 7B

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.488945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.488945Z digest=sha256:dbcfa2722d2206d79f5630d7b678fa1f25a9df905070a501383d587e55f96a93

Observation 26b35b12-ec95-466d-8789-17d125c14868 · outbound

This paper cites In: Barzilay, R., Kan, M.Y.

Benchmarking Prompt Sensitivity in Large Language Models In: Barzilay, R., Kan, M.Y

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.492434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.492434Z digest=sha256:032508a775d92924b1a6229ac6d75b0ba74754a6d5a304bd4b177a24996ea1da

Observation 5d2ef4fe-0748-4656-be65-2175709938e6 · outbound

This paper cites In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.495831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.495831Z digest=sha256:201e413745c717853f898b8d26228f8b9d381125ee476a17b2bf67473a9cc8c9

Observation 3b7ffde7-27fd-4fc6-b0bb-735a7538e3a7 · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 23

Resolution
verified exact
doi, observed 2026-08-08T16:55:45.651096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.499110Z digest=sha256:1d3e4225bd97033ff635df0e41913201ad215142d0ac17d67bb50e73720fde17

Observation b729d661-c7d7-4087-9793-392ae8c4b0ea · outbound

This paper cites Internet Reference Services Quarterly27, 203 – 210 (2023).https://doi.org/10.1080/ 10875301.2023.2227621.

Benchmarking Prompt Sensitivity in Large Language Models Internet Reference Services Quarterly27, 203 – 210 (2023).https://doi.org/10.1080/ 10875301.2023.2227621

Reference 24

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T16:55:45.980823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.502079Z digest=sha256:0f2d51cf81853fb3ff4b0b5f7c85e51e461b9303f1056053cde3888ba96c89f3

Observation 7824525e-6804-474b-89bb-f2d8a9a39010 · outbound

This paper cites https://doi.org/10.18653/v1/2023.findings-emnlp.241, http: //dx.doi.org/10.18653/v1/2023.findings-emnlp.241.

Benchmarking Prompt Sensitivity in Large Language Models https://doi.org/10.18653/v1/2023.findings-emnlp.241, http: //dx.doi.org/10.18653/v1/2023.findings-emnlp.241

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.505163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.505163Z digest=sha256:3807b6d923e91a19714eacf49f2e8ff4a8fdddd4f8b0c017710260acd5f0961b

Observation e257f84d-e8bb-43cd-9f61-7491da07ad2c · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.508370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.508370Z digest=sha256:e8493cba27df492520fd374d350d84fe39e6f0ce1b61cf62dd0c6aee2916295d

Observation d4604f86-0b0a-4cca-890d-2db2442aa080 · outbound

This paper cites Query Performance Prediction using Relevance Judgments Generated by Large Language Models.

Benchmarking Prompt Sensitivity in Large Language Models Query Performance Prediction using Relevance Judgments Generated by Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.511620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.511620Z digest=sha256:c7043febe4e1844a2b26cc6bec8ac0bf659eb57baf190623ff807716fbce9666

Observation 93cfb014-0c4f-48aa-a81e-36368b23297d · outbound

This paper cites The Llama 3 Herd of Models.

Benchmarking Prompt Sensitivity in Large Language Models The Llama 3 Herd of Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.514952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.514952Z digest=sha256:1ccb5e4677bc3864de8fa722dbb1813c3cd235e60c6d6ac9245e07d2ab1db418

Observation 4639ae2e-21b8-4fb9-ab6f-ff2898196f44 · outbound

This paper cites Navigating Prompt Complexity for Zero-Shot Classification: A Study of Large Language Models in Computational Social Science.

Benchmarking Prompt Sensitivity in Large Language Models Navigating Prompt Complexity for Zero-Shot Classification: A Study of Large Language Models in Computational Social Science

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.518162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.518162Z digest=sha256:392defefe93730e1da33aeb20cdd33a4cab3b002b9b39907411181b98c09c9bb

Observation 5f4fe807-9dfd-41fe-8f26-d703dddce000 · outbound

This paper cites Testing LLMs on Code Generation with Varying Levels of Prompt Specificity.

Benchmarking Prompt Sensitivity in Large Language Models Testing LLMs on Code Generation with Varying Levels of Prompt Specificity

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.521371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.521371Z digest=sha256:ff048f59e1bf8f7a489a516510a0383b5837f409b8b97a7ec80992fb77ffee00

Observation 167f177b-1e6d-426a-84e8-7a3332fc9d1f · outbound

This paper cites PQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance Prediction.

Benchmarking Prompt Sensitivity in Large Language Models PQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance Prediction

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-08T16:55:45.792127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.524418Z digest=sha256:fe43a55a0a0b97ea2af52c2b47c1ab52eae132797f7e191e266b9140364ff0cf

Observation 1c0b5fbc-294d-4adf-8817-933f07008425 · outbound

This paper cites Semantic Consistency for Assuring Reliability of Large Language Models.

Benchmarking Prompt Sensitivity in Large Language Models Semantic Consistency for Assuring Reliability of Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.527516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.527516Z digest=sha256:a1b6f973362e831d8eddc77ebf5b00369f65cb55eb3185411064f0e41a69aade

Observation a3f0f6bb-5e0a-4821-b713-52e4edd0a935 · outbound

This paper cites In: Proceedings of the 2024 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 2024 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.531616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.531616Z digest=sha256:b45cb6e2850a8409ed1be051ab2e56e8bf093496c5d2d51289dc890c5b3a3665

Observation 38e2003d-8cca-4612-b730-81796aadcae7 · outbound

This paper cites In: Proceedings of the 32nd ACM Inter- national Conference on Information and Knowledge Management.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 32nd ACM Inter- national Conference on Information and Knowledge Management

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.162750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.534670Z digest=sha256:9a01290e29784613ae59ad97de6a877ef47ae51593eb7b5190c6bb90a14e7331

Observation f4ac3d28-1374-40aa-84f8-2b7d1076fb6c · outbound

This paper cites In: European Conference on Information Re- trieval.pp.30–39.Springer(2024).

Benchmarking Prompt Sensitivity in Large Language Models In: European Conference on Information Re- trieval.pp.30–39.Springer(2024)

Reference 35

Resolution
verified exact
doi, observed 2026-08-08T16:55:45.604745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.537815Z digest=sha256:d23cbea65cd17b961a04048612fc8709efaf6c57ac0a1101200e443728a70671

Observation 11dfe6a2-4267-4f6a-ade1-3359f6563496 · outbound

This paper cites Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting.

Benchmarking Prompt Sensitivity in Large Language Models Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.540970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.540970Z digest=sha256:e94adb5cff844d656059d4e07edd5cf06467f7122a3a102b259e667fa104bdbe

Observation 5502d8ec-7031-42d5-88ac-7a32dcfb7701 · outbound

This paper cites In: Proceedings of the 34th International Conference on Neural Information Processing Systems.

Benchmarking Prompt Sensitivity in Large Language Models In: Proceedings of the 34th International Conference on Neural Information Processing Systems

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.152681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.544449Z digest=sha256:97952653ce8d2f7fb0b04c1adaa36124ef118957ab541019377d089cadae3875

Observation fe230fa3-32ae-4b6d-9ef6-cd5fe05a1870 · outbound

This paper cites eugeneyan.com (Aug 2024),https://eugeneyan.com/writing/llm-evaluators/.

Benchmarking Prompt Sensitivity in Large Language Models eugeneyan.com (Aug 2024),https://eugeneyan.com/writing/llm-evaluators/

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:55:46.142696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T16:55:45.547529Z digest=sha256:43cb22a36835320e7de14b5eb6df65e400591c76ff70106718ccfe790a43be92

Observation d5ffa9f8-6bd1-4e36-88e9-5b82e781e062 · outbound

This paper cites an unresolved cited work.

Benchmarking Prompt Sensitivity in Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.550593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.550593Z digest=sha256:1cfc4b326ef5f52dfe148a494db9c471f9b220dc9beb0d5e6507d6910adb6894

Observation b7267df2-c433-47f8-99f8-1efbc8a7c390 · outbound

This paper cites PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts.

Benchmarking Prompt Sensitivity in Large Language Models PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.553837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.553837Z digest=sha256:2fdc7f24b33ae4c502c9cb0b78c1c5195a93aa21bd28fc962b82a3c181e16ae1

Observation 0396fec9-6d6e-4841-9db6-cc8f4f3a90de · outbound

This paper cites ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs.

Benchmarking Prompt Sensitivity in Large Language Models ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:45.557734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:45.557734Z digest=sha256:269ced6e5949effdc403898f0016fee93988e65020352dc375e5374f2eab02e2

Pith citing papers

Observation 7e8f3c14-1ffe-43e3-98ed-a7d2cb6e2295 · inbound

Understanding the Mechanism of Altruism in Large Language Models cites this paper.

Understanding the Mechanism of Altruism in Large Language Models Benchmarking Prompt Sensitivity in Large Language Models

Reference 192

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:31:02.511069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T01:36:50.329664Z digest=sha256:4a71d9ed303755184f05f9aa4a04097cb5e90f6042f9fcf327707b618b10aae3

Observation ac8129f6-3efa-4cfa-9687-e94cbfea7a27 · inbound

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation cites this paper.

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation Benchmarking Prompt Sensitivity in Large Language Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:46:14.436745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T01:41:42.003483Z digest=sha256:33b914d82048266b2cfcbced1a98a8ba47c5f857ef1a4d84324024821d4d60fd

Observation 9eacf9cf-1ef4-420b-bfd1-49665c342223 · inbound

Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits cites this paper.

Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits Benchmarking Prompt Sensitivity in Large Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:43:19.603517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T13:40:04.275438Z digest=sha256:9c78ef2ecb4ad3e04c47b38dd682277e25b48a87e8c455465bfd9a5cef082362