Pith. sign in

Paper Citation Record · LEDGER

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation

As of 17 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2607.14109.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14109 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T14:42:46.895939Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T03:19:00.468509Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved40
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a8dbedc3-b580-44a4-a9ac-7ad770bbecec · outbound

This paper cites Matharena: Evaluating llms on uncontaminated math competitions, February 2025.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Matharena: Evaluating llms on uncontaminated math competitions, February 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:43.688759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:43.688759Z digest=sha256:c5ff8d1bacd42d36204fdab3965f3f2eeee538d74d515422bd14f8f6aaa9ca18

Observation 8d12022a-c777-4faa-b2d7-cf6bc25524cd · outbound

This paper cites On the Measure of Intelligence.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation On the Measure of Intelligence

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:43.736195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:43.736195Z digest=sha256:937c7c690bf2a80c4439054ae851fd9dcf9afc5f48f852248856d57ab1623756

Observation 417d6c12-b74f-4a4b-a868-83dfde7673f4 · outbound

This paper cites FailureSensorIQ: A Multi-Choice QA Dataset for Understanding Sensor Relationships and Failure Modes.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation FailureSensorIQ: A Multi-Choice QA Dataset for Understanding Sensor Relationships and Failure Modes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:43.792576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:43.792576Z digest=sha256:e0bbcfecd2d23e4e3f6ec480c241a277ea078e92d0af96cb4b2d53184358366d

Observation 62de648b-6cab-427d-b068-09919e94592c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:43.868953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:43.868953Z digest=sha256:72f082fc7404b0ce6e06d19f7437e3a658d359a9c3f8a41924803f1ec379dd72

Observation c81abba5-3dd0-4b90-8299-695c12f88f6d · outbound

This paper cites The language model evaluation harness, July 2024.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation The language model evaluation harness, July 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:43.942648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:43.942648Z digest=sha256:92a0ae21d645c4e7bc6496c695cf12e3ba6de94be1da96a7ab4c3d9e3e4d63bd

Observation 3fa92299-6cb3-4ee8-9aa2-6f6e55ec04e1 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Measuring Massive Multitask Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.008058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.008058Z digest=sha256:86778a2ef19f5924d1881a07126a90c006411cf1e3c32ad57c51b19cd4f2533b

Observation 07c789f4-03fc-4bae-a2b9-ce5bbcb66dd4 · outbound

This paper cites What disease does this patient have? a large-scale open domain question answering dataset from medical exams.Applied Sciences, 11(14):6421, 2021.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation What disease does this patient have? a large-scale open domain question answering dataset from medical exams.Applied Sciences, 11(14):6421, 2021

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.077828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.077828Z digest=sha256:09b58cfb28fdb6a27457c3fdd086662f4254a81a84ec01b4780edf23fd499de3

Observation 6a17f4d5-65a3-4b16-8856-0ccc80f9cbde · outbound

This paper cites Dynabench: Rethinking benchmarking in NLP.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Dynabench: Rethinking benchmarking in NLP

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.125291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.125291Z digest=sha256:b77b6ac5901a5aabe9a65d6b98e7198f1476ed1fc472da900c326981497f5c84

Observation 6da2fbb5-5677-4a57-80d3-639352a73d76 · outbound

This paper cites Large language models are zero-shot reasoners.Advances in Neural Information Processing Systems, 35:22199–22213, 2022.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Large language models are zero-shot reasoners.Advances in Neural Information Processing Systems, 35:22199–22213, 2022

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.201129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.201129Z digest=sha256:65bbc97e25056acfa8257f48f197294be9cc535292bcca9e8b299ccd7dabc1e3

Observation b665a6cf-375d-462c-b5ac-a24d556037b0 · outbound

This paper cites Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.258314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.258314Z digest=sha256:ba8509d60d9bc908198f3d9f6dc0bcf4f80d1a48201fd032ae2439b0b09eb1a6

Observation 0f54f492-b7ab-455a-bc1f-b81c96cb9f13 · outbound

This paper cites Prompt repetition improves non-reasoning llms.arXiv preprint arXiv:2512.14982, 2025.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Prompt repetition improves non-reasoning llms.arXiv preprint arXiv:2512.14982, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.341417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.341417Z digest=sha256:a22f73c5780901e3378173a9be5906ea40e5f23081a13212caa4fb2a75e397db

Observation 41a00ded-0a6b-4172-bbf1-76b0eeb0c10c · outbound

This paper cites Holistic Evaluation of Language Models.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Holistic Evaluation of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.399738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.399738Z digest=sha256:9c4a83ba015e18757c512c20bed89179a8a17692703d4515b432b33dc1dedc52

Observation dd81da62-65e4-44e3-9d57-0260c64f53c7 · outbound

This paper cites Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.484533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.484533Z digest=sha256:3aeab90ea5842b792719df490f73db3ced901607ddd5be2d1a05446f2391c614

Observation bbd846b0-2b40-4227-8ff3-1ac792f89aa5 · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.560800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.560800Z digest=sha256:7da81c3e3596c55d633f410268ec85ca22911e347ca3e1cbb5d908b5735bae70

Observation 0c9c93e9-8fc0-4bda-a6c1-b5b050d53f65 · outbound

This paper cites Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.631644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.631644Z digest=sha256:e23eab4b20cc672f1bd28aca031d80a523ccea0a827842c5e766f651ef9bd3f2

Observation bc0dad1d-790a-46e4-9562-e30006943550 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.690538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.690538Z digest=sha256:b679227fdcb7438f6b04d5f66e7056082f4781ef8ac73f4f8a35a250dd456394

Observation bbe22d41-b35b-4b96-bfe4-2711edc2c830 · outbound

This paper cites State of what art? a call for multi-prompt LLM evaluation.Transactions of the Association for Computational Linguistics, 12:933–949, 2024.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation State of what art? a call for multi-prompt LLM evaluation.Transactions of the Association for Computational Linguistics, 12:933–949, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.771497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.771497Z digest=sha256:80712d734b8e0dc25cd50e78fdf828e9cc3688ed33f68ed3598813e19d8fe2da

Observation 2b98c736-2e7d-4f91-81f7-65991303bab1 · outbound

This paper cites s1: Simple test-time scaling.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation s1: Simple test-time scaling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.851388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.851388Z digest=sha256:476e3fc67969708d9a9d409446058fe176e8baf61c4f0be299ed85461b789b10

Observation 803145ac-f648-4d8e-b804-397b5d9d61b3 · outbound

This paper cites Large language models sensitivity to the order of options in multiple-choice questions.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Large language models sensitivity to the order of options in multiple-choice questions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.933675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.933675Z digest=sha256:3e0feb27d10bf0f7ad69686da7fb6e34e743ff7d86b58d54333f7393854d72a7

Observation 7a2d1432-7497-4207-8d26-67b65e837ee8 · outbound

This paper cites Smith, and Mike Lewis.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Smith, and Mike Lewis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.989333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.989333Z digest=sha256:6228eb3d40d5571677548e404132c3f2586e7172c30c001eab6f7b54f5dff369

Observation 08c60252-8d51-437c-9a79-752328250936 · outbound

This paper cites Qwen3 Technical Report.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Qwen3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.058968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.058968Z digest=sha256:88375378f46a748b399fdd5862de85280161d490a48f4baace6894d15e765ec6

Observation ae5a0261-8753-4e0f-a714-525fe8aca56c · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.126365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.126365Z digest=sha256:a93a4f7d83015cae3a2e4a44ff00a3b6010b476b1c65b8a9392af695b03002f2

Observation 4208621a-99dc-462f-ae7b-af01e99f96ee · outbound

This paper cites Leveraging large language models for multiple choice question answering.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Leveraging large language models for multiple choice question answering

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.188960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.188960Z digest=sha256:8a50608c780fa43f9d2f20bdc277feb251c4feafd5fc63c046d6cd9e73bd0110

Observation d90898b8-5376-45d0-bd23-cfc1466a0610 · outbound

This paper cites Rogers, Inna Goncearenco, Giuseppe Sarli, Igor Galynker, Denis Peskoff, Marine Carpuat, Jules White, Shyamal Anadkat, Alexander Hoyle, and Philip Resnik.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Rogers, Inna Goncearenco, Giuseppe Sarli, Igor Galynker, Denis Peskoff, Marine Carpuat, Jules White, Shyamal Anadkat, Alexander Hoyle, and Philip Resnik

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.271611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.271611Z digest=sha256:8d1c161ed89a0c4e1599d086e902816907eb46581b7aa93cd13ef321ecd07040

Observation 69da3077-c6af-40d9-8c9b-e370aff58097 · outbound

This paper cites Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.351262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.351262Z digest=sha256:7f683024c613f78572a2ccfba078a3b219c9b06edae2febbba25cc7563c9fe57

Observation 3c9b5ff5-58b6-40e5-9167-4b6ec5955431 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.427128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.427128Z digest=sha256:94c8de45cfd2314c23b8bacfee02ff0fd922cce5067813d02b36b06da3d02bb6

Observation 386a2a21-f5f9-4ae0-9201-d03189116cb1 · outbound

This paper cites To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.567640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.567640Z digest=sha256:2c1db9b44dfd9b186cd71d10c3ab263965da2e20793854547e056922c61df1d7

Observation 070c18b0-ffd7-407a-a68b-274a9e4b9031 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.737278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.737278Z digest=sha256:9e83f4dae3001c618460e79fec139e8458238a64549aa881cd00ddd0de392373

Observation c24b6286-216d-48e8-a2b9-698aee799855 · outbound

This paper cites On the self-verification limitations of large language models on reasoning and planning tasks.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation On the self-verification limitations of large language models on reasoning and planning tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.811194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.811194Z digest=sha256:8c132e020b4a2faddf5757a39ca54e880dcf8b4810c08149cdfe5b6e3a546e40

Observation a086f0a8-0857-4400-b802-8181fd1b6235 · outbound

This paper cites Correctbench: A benchmark of self-correction in llms.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Correctbench: A benchmark of self-correction in llms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.891639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.891639Z digest=sha256:bba9aa71d7abaa907fd82ce99a3112c6d94f46992d1e5bb84faac7fda6ea5ad1

Observation 7869f54f-2d87-4e95-bf18-6242d863dbc3 · outbound

This paper cites Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:45.950957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:45.950957Z digest=sha256:1f60fc7b402f3685f520258e14b52c495a68824563d6bbc483206e28bb0130b7

Observation 5d944d42-32fc-4147-8eba-f70080c55be0 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:46.034794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.034794Z digest=sha256:a87361423f56a647b58a54a647d3dd268c69ae5b56a1ea21d60902c90889203b

Observation 7ba8a91c-af99-4941-8964-d5d2feba3625 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:46.087335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.087335Z digest=sha256:b9a88f8bee7d38d276c01f5a5401a6f4e4f328781cb7bab5ae43fc3776d81d90

Observation 435a630f-28ab-479a-9174-543f4c4ada6f · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Chain-of-thought prompting elicits reasoning in large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:46.157373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.157373Z digest=sha256:359603f8b59e5cbdc55742d3c516be31d9179c7fdc102669fce5e9064e48525e

Observation fe702455-6e0e-495a-b1e9-bc7d265db053 · outbound

This paper cites Griffiths, Yuan Cao, and Karthik Narasimhan.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Griffiths, Yuan Cao, and Karthik Narasimhan

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:46.222619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.222619Z digest=sha256:d3e2dafc4663251d6f0867f6048238f956e1e47025690551ff43d825a32cdcee

Observation 1a961236-da12-45fd-bd23-b8616d05477b · outbound

This paper cites Chi, and Denny Zhou.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Chi, and Denny Zhou

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:46.344230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.344230Z digest=sha256:392a8a0ee4db923f61f6771524e90e56c639b326e4c97870ba0d5db8c4b160b6

Observation 4ef20845-7303-4fc7-b64f-855a71505425 · outbound

This paper cites Generate rather than Retrieve: Large Language Models are Strong Context Generators.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Generate rather than Retrieve: Large Language Models are Strong Context Generators

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:46.446916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.446916Z digest=sha256:d8bc6177c7b65eed11d53a1e654b83d77e99b58f0bace46524e875c8717b39b9

Observation dfa6988e-f120-447c-9960-8645380e31a3 · outbound

This paper cites Large language models are not robust multiple choice selectors.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Large language models are not robust multiple choice selectors

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:46.560399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.560399Z digest=sha256:8d2726e501aa18d1f707fb0721a94bec51af04130dd1af1b15060489f8155a98

Observation 14d137af-1c5b-4983-ba1d-0cb7b9eb896b · outbound

This paper cites Scaling physical reasoning with the physics dataset, 2025.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Scaling physical reasoning with the physics dataset, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:46.668559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.668559Z digest=sha256:8a2b876169fdf4283b557ea447ca41d30b9f5649cbf4982d5edc6d8b509cdea1

Observation ab524a00-2979-4d1d-8413-3f01fdfb7fa5 · outbound

This paper cites PromptBench: A Unified Library for Evaluation of Large Language Models.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation PromptBench: A Unified Library for Evaluation of Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:46.785295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.785295Z digest=sha256:2f4c7749ec2d210d9cd44b0357b32c9891b2ffddbc946abb8ab6cc2373e8c94f

Observation 0f23936f-5148-4702-9339-0d8f554e1d0b · outbound

This paper cites Avg. opt.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Avg. opt

Reference 41

Resolution
malformed identifier
no resolver link, observed 2026-08-02T14:42:46.895939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:46.895939Z digest=sha256:0b227bca77bef435c5521a54d0455716c8095c5496e188cb38f49e2abfe6e0e3

Pith citing papers

Observation c09d33ea-403f-4738-85b1-ac4435c38fb5 · inbound

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B cites this paper.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:00.468509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:00.468509Z digest=sha256:2d37ca7f3d7aee9c63f0193509560a51dc8c888f060166d898523d754cceec85