Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T14:42:46.895939Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2607.14109.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T14:42:46.895939Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T03:19:00.468509Z
A source-named dated measurement, never combined with another source.
Source: cited_works
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a8dbedc3-b580-44a4-a9ac-7ad770bbecec · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Matharena: Evaluating llms on uncontaminated math competitions, February 2025
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d12022a-c777-4faa-b2d7-cf6bc25524cd · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation On the Measure of Intelligence
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 417d6c12-b74f-4a4b-a868-83dfde7673f4 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation FailureSensorIQ: A Multi-Choice QA Dataset for Understanding Sensor Relationships and Failure Modes
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62de648b-6cab-427d-b068-09919e94592c · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c81abba5-3dd0-4b90-8299-695c12f88f6d · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation The language model evaluation harness, July 2024
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fa92299-6cb3-4ee8-9aa2-6f6e55ec04e1 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Measuring Massive Multitask Language Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07c789f4-03fc-4bae-a2b9-ce5bbcb66dd4 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation What disease does this patient have? a large-scale open domain question answering dataset from medical exams.Applied Sciences, 11(14):6421, 2021
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a17f4d5-65a3-4b16-8856-0ccc80f9cbde · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Dynabench: Rethinking benchmarking in NLP
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6da2fbb5-5677-4a57-80d3-639352a73d76 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Large language models are zero-shot reasoners.Advances in Neural Information Processing Systems, 35:22199–22213, 2022
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b665a6cf-375d-462c-b5ac-a24d556037b0 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f54f492-b7ab-455a-bc1f-b81c96cb9f13 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Prompt repetition improves non-reasoning llms.arXiv preprint arXiv:2512.14982, 2025
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a00ded-0a6b-4172-bbf1-76b0eeb0c10c · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Holistic Evaluation of Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd81da62-65e4-44e3-9d57-0260c64f53c7 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbd846b0-2b40-4227-8ff3-1ac792f89aa5 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c9c93e9-8fc0-4bda-a6c1-b5b050d53f65 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc0dad1d-790a-46e4-9562-e30006943550 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Can a suit of armor conduct electricity? a new dataset for open book question answering
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbe22d41-b35b-4b96-bfe4-2711edc2c830 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation State of what art? a call for multi-prompt LLM evaluation.Transactions of the Association for Computational Linguistics, 12:933–949, 2024
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b98c736-2e7d-4f91-81f7-65991303bab1 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation s1: Simple test-time scaling
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 803145ac-f648-4d8e-b804-397b5d9d61b3 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Large language models sensitivity to the order of options in multiple-choice questions
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a2d1432-7497-4207-8d26-67b65e837ee8 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Smith, and Mike Lewis
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08c60252-8d51-437c-9a79-752328250936 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Qwen3 Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae5a0261-8753-4e0f-a714-525fe8aca56c · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4208621a-99dc-462f-ae7b-af01e99f96ee · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Leveraging large language models for multiple choice question answering
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d90898b8-5376-45d0-bd23-cfc1466a0610 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Rogers, Inna Goncearenco, Giuseppe Sarli, Igor Galynker, Denis Peskoff, Marine Carpuat, Jules White, Shyamal Anadkat, Alexander Hoyle, and Philip Resnik
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69da3077-c6af-40d9-8c9b-e370aff58097 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c9b5ff5-58b6-40e5-9167-4b6ec5955431 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 386a2a21-f5f9-4ae0-9201-d03189116cb1 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 070c18b0-ffd7-407a-a68b-274a9e4b9031 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c24b6286-216d-48e8-a2b9-698aee799855 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation On the self-verification limitations of large language models on reasoning and planning tasks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a086f0a8-0857-4400-b802-8181fd1b6235 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Correctbench: A benchmark of self-correction in llms
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7869f54f-2d87-4e95-bf18-6242d863dbc3 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d944d42-32fc-4147-8eba-f70080c55be0 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ba8a91c-af99-4941-8964-d5d2feba3625 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 435a630f-28ab-479a-9174-543f4c4ada6f · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Chain-of-thought prompting elicits reasoning in large language models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe702455-6e0e-495a-b1e9-bc7d265db053 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Griffiths, Yuan Cao, and Karthik Narasimhan
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a961236-da12-45fd-bd23-b8616d05477b · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Chi, and Denny Zhou
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ef20845-7303-4fc7-b64f-855a71505425 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Generate rather than Retrieve: Large Language Models are Strong Context Generators
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfa6988e-f120-447c-9960-8645380e31a3 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Large language models are not robust multiple choice selectors
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14d137af-1c5b-4983-ba1d-0cb7b9eb896b · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Scaling physical reasoning with the physics dataset, 2025
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab524a00-2979-4d1d-8413-3f01fdfb7fa5 · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation PromptBench: A Unified Library for Evaluation of Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f23936f-5148-4702-9339-0d8f554e1d0b · outbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Avg. opt
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c09d33ea-403f-4738-85b1-ac4435c38fb5 · inbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.