Pith. sign in

Paper Citation Record · LEDGER

UQ: Assessing Language Models on Unsolved Questions

As of 20 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 1 inbound Pith citation observation for arXiv:2508.17580.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.17580 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:07:19.897165Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T04:09:46.616019Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T12:06:04.092759Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1f01bbf6-b12f-46b1-b9ae-2e3b1a5845af · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

UQ: Assessing Language Models on Unsolved Questions Piqa: Reasoning about physical commonsense in natural language

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.844600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.612229Z digest=sha256:1684cb89c5931e2a2d30579cd6236674ede90fa2a04da10011e64c536f68b20f

Observation 5a54b432-f1ee-46b1-bbef-b1f27ff77a20 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

UQ: Assessing Language Models on Unsolved Questions Evaluating Large Language Models Trained on Code

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.617165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.617165Z digest=sha256:2f26fb718c1acd19c8349269e48778f20b8116dadd73c3c0d3b778d9e2030295

Observation 7c1e60d2-ce2f-4c88-b390-652795b68d92 · outbound

This paper cites Chatbot arena: An open platform for evaluating llms by human preference.

UQ: Assessing Language Models on Unsolved Questions Chatbot arena: An open platform for evaluating llms by human preference

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.832642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.621341Z digest=sha256:6b592326b254287e9faa2dcce41c35c5c71f6976fb7f253ae82cce2349c50990

Observation e4e6170e-de5a-4d34-b437-6df8349d1a78 · outbound

This paper cites ARC Prize 2024: Technical Report.

UQ: Assessing Language Models on Unsolved Questions ARC Prize 2024: Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.625583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.625583Z digest=sha256:a0625f078f4aa36820d23e0e466f9dd6a82d539b4f623bedad96d1eb24facb4c

Observation 2176fa22-ebda-4452-93cd-285dcdc6d287 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

UQ: Assessing Language Models on Unsolved Questions Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.629700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.629700Z digest=sha256:7b94637b51e17dfc7fd73b2b1bd66124e9bbc02be38e1dc13769482713dc910e

Observation 5b463a90-bfec-45d1-b90e-d542c6e823f5 · outbound

This paper cites Chain-of-Verification Reduces Hallucination in Large Language Models.

UQ: Assessing Language Models on Unsolved Questions Chain-of-Verification Reduces Hallucination in Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.634145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.634145Z digest=sha256:8ad5a2397b6c4d3e202d1f536e37ab7fe35ce700e19beb7179a772b2099d4aa7

Observation b0069902-9909-4c98-8a27-57bfafb52b46 · outbound

This paper cites AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback.

UQ: Assessing Language Models on Unsolved Questions AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.638212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.638212Z digest=sha256:43a7b7fc91d00dfb9d0ede7517b4fccc13ceb184947790390dcc97f10533341a

Observation 7f96f91e-c4d8-41ae-a264-112d60ea3305 · outbound

This paper cites Are We Done with MMLU?.

UQ: Assessing Language Models on Unsolved Questions Are We Done with MMLU?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.642257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.642257Z digest=sha256:a088e2ed8283a43e9799e93649694934723f9835115820a922e721e62a5dcb30

Observation f1ab611a-d399-4435-a64c-f27784877ef4 · outbound

This paper cites Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai, 2024.

UQ: Assessing Language Models on Unsolved Questions Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.646336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.646336Z digest=sha256:481c29ce1d17ac73a72c36170b1d3ecc83388a7ad27785d66da9fb23e8edf604

Observation 593bde5f-a009-44e0-89ef-0bb417085f93 · outbound

This paper cites Great Models Think Alike and this Undermines AI Oversight.

UQ: Assessing Language Models on Unsolved Questions Great Models Think Alike and this Undermines AI Oversight

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.653111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.653111Z digest=sha256:44bf5fc92c8187edf0577e79d9601698228311e97e1c93c1dee51b9197aeeb94

Observation cd69e84c-1409-4aa7-8517-78fe5e4d119c · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

UQ: Assessing Language Models on Unsolved Questions Measuring Coding Challenge Competence With APPS

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.657157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.657157Z digest=sha256:d5f1a5f0326a17f6533a9ce14f176ee5d9fd5e3081aa64bbdf56f550731b975f

Observation a09cdff2-45c3-4aa9-bf4b-68c6d2d6426a · outbound

This paper cites Measuring Massive Multitask Language Understanding.

UQ: Assessing Language Models on Unsolved Questions Measuring Massive Multitask Language Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.661079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.661079Z digest=sha256:b1130d2f86eb9d04f2e42bdc0ccef109b0d7b91c79bc44bfb652acdb51077164

Observation fcbcf013-b427-4ecb-95de-df825ee3cb95 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

UQ: Assessing Language Models on Unsolved Questions Measuring Mathematical Problem Solving With the MATH Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.664976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.664976Z digest=sha256:ad61fbaf89eba57525df3a4be3895e90cdf9fff817ebb11c3faed73dd603aa5d

Observation 79c6791a-03b7-4619-aa14-76667eb6461e · outbound

This paper cites MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations.

UQ: Assessing Language Models on Unsolved Questions MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.669042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.669042Z digest=sha256:09bfd48421af79454a64c3d6b38ddc7cd214556274b26c552b6a84f0a84d7f41

Observation 7cfd3e5b-372f-4122-b6cc-97d91e78eacc · outbound

This paper cites Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards.

UQ: Assessing Language Models on Unsolved Questions Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:07:20.365810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.673009Z digest=sha256:5cfb7bcba30824a52ce62730d71fb77b09b1b1afd8ace8026b13c722dc3630ac

Observation 9effe6b5-3a3f-4cf8-b0bf-7c9cf841f69b · outbound

This paper cites GPT-4o System Card.

UQ: Assessing Language Models on Unsolved Questions GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.677086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.677086Z digest=sha256:27606822305a369a774656aa714c2a08d5368dce148423219f41b192807bdd33

Observation 1d031044-a077-4dd3-af42-2e333a3d4d62 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

UQ: Assessing Language Models on Unsolved Questions LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.681246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.681246Z digest=sha256:0df1c245d603676f37c0778b429307fa635d05acdb7036b49c1fbbc9580f10d8

Observation 971742a0-244e-4bb1-ac3d-08e9653d5389 · outbound

This paper cites When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories.

UQ: Assessing Language Models on Unsolved Questions When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.685883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.685883Z digest=sha256:06d0c5f1f288ebb286f1a7d9577f6a3bef1ea7d1c94372b22465bea68fe44231

Observation aa30218c-fcfd-4868-90e8-1ec7d03d2667 · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

UQ: Assessing Language Models on Unsolved Questions Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.690864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.690864Z digest=sha256:8b21722911e10f78c5d9833a94221671867da2e2b4fc520e43e3c2439ab140d2

Observation e32cb678-272d-4129-852e-ff4739ee86ff · outbound

This paper cites Verdict: A library for scaling judge-time compute.

UQ: Assessing Language Models on Unsolved Questions Verdict: A library for scaling judge-time compute

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.695712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.695712Z digest=sha256:46b8f8afad3a8e9ceef33933ae585ba67944222414cfeda22c5469c4bb793261

Observation 190333ed-45ec-463e-9577-a67c0a8fdb35 · outbound

This paper cites Llm-as-an-interviewer: Beyond static testing through dynamic llm evaluation, 2025.

UQ: Assessing Language Models on Unsolved Questions Llm-as-an-interviewer: Beyond static testing through dynamic llm evaluation, 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.805990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.700099Z digest=sha256:71e076961b224c38263cdd5b0d86c8a318e0f0cb7f297ca60f29dd89a186a175

Observation 158af6cb-3842-49c6-a647-0dd64f7d1c0d · outbound

This paper cites Prometheus: Inducing fine-grained evaluation capability in language models, 2024.

UQ: Assessing Language Models on Unsolved Questions Prometheus: Inducing fine-grained evaluation capability in language models, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.703854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.703854Z digest=sha256:4dcd093bf5a48038306c336c2350d166bd1d4fbdd4c43c132018a522139d1b09

Observation 519f40fe-6840-4a3d-98a5-c3c18e81dd31 · outbound

This paper cites Prometheus 2: An open source language model specialized in evaluating other language models, 2024.

UQ: Assessing Language Models on Unsolved Questions Prometheus 2: An open source language model specialized in evaluating other language models, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.785977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.707492Z digest=sha256:d223001d01c3c04e86e3c73ba0186b90a81e68301fe02ccb670371ca2a36b8a8

Observation b93b5b2a-5278-4db2-b564-4f15cb79e6ef · outbound

This paper cites Scaling evaluation-time compute with reasoning models as process evaluators, 2025.

UQ: Assessing Language Models on Unsolved Questions Scaling evaluation-time compute with reasoning models as process evaluators, 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.773770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.711145Z digest=sha256:61b89edb39b017bcdd14a65c04e6b118698b810f92e0b7ebeacaf7e620dfa8da

Observation b77f1e55-ca6b-4450-ac93-948aef2b3bbc · outbound

This paper cites Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov.

UQ: Assessing Language Models on Unsolved Questions Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.761723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.714954Z digest=sha256:0ee78f6195f1c16aeea9d48058e3d1d3237c4b01e113673359ae946d34549118

Observation b75cb40a-367c-4ecf-9060-cdf0a7d783da · outbound

This paper cites The measurement of observer agreement for categorical data.

UQ: Assessing Language Models on Unsolved Questions The measurement of observer agreement for categorical data

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.749131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.718615Z digest=sha256:59c900f5b07c39a36da533c34e61e1dc7dde219943e8967fb56c078e38fd5f20

Observation ddd119d8-450b-4eb6-b495-24bfb2c8b879 · outbound

This paper cites FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets.

UQ: Assessing Language Models on Unsolved Questions FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.722330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.722330Z digest=sha256:afeb1b4eb60600f56a49921ab730eff62e233e9dcb13eac2e6508cc9449d52aa

Observation c5edad0b-6467-463f-8ec4-68fefa0054bc · outbound

This paper cites Holistic Evaluation of Language Models.

UQ: Assessing Language Models on Unsolved Questions Holistic Evaluation of Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.726851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.726851Z digest=sha256:2cc27bac0d5def2b8451e1498f13c6c9bb8d385cc065501dd64c705ab78a6fa6

Observation 374897c1-dee4-4c24-be9c-51e809777187 · outbound

This paper cites Let's Verify Step by Step.

UQ: Assessing Language Models on Unsolved Questions Let's Verify Step by Step

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.730930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.730930Z digest=sha256:21674b9eca387e267267e02c506412a51ff3d1bb1829e4fccec15a7f5e03163c

Observation 8cfc8a7c-e9e9-4a2e-a6f8-b88c3c7ca7da · outbound

This paper cites WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild.

UQ: Assessing Language Models on Unsolved Questions WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.734688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.734688Z digest=sha256:6a83f0ba743d05e12139de2362feae188c1b6ebc381eb28c1d2cdc820ea8be3b

Observation 3f5e8e7a-8fe5-4d40-98c1-2638cc7c8139 · outbound

This paper cites Evaluating Verifiability in Generative Search Engines.

UQ: Assessing Language Models on Unsolved Questions Evaluating Verifiability in Generative Search Engines

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.738927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.738927Z digest=sha256:0c379279d2b47b97f03450f0bd26f9f05bc7bbd792b7d6a2465f079a3622c1eb

Observation 5ef7b271-c72c-4a61-8ea2-4d5fe30ccce6 · outbound

This paper cites The lean 4 theorem prover and programming language.

UQ: Assessing Language Models on Unsolved Questions The lean 4 theorem prover and programming language

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.736175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.742881Z digest=sha256:5a2896986a245aa659d1aecb812b82a051868a1b7e32f2aff05509394961fd4f

Observation b9ea21ba-ee35-4da9-a8cb-eb220e3c4895 · outbound

This paper cites 2024 aime i.

UQ: Assessing Language Models on Unsolved Questions 2024 aime i

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.723466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.747278Z digest=sha256:8c907e8a25b2e07e9070433d3411cd6ef2dbdc8f22780160114bca093451bd05

Observation c95d7ef9-422e-45a1-99ef-461c49e60bd0 · outbound

This paper cites List of open problems in sublinear algorithms.

UQ: Assessing Language Models on Unsolved Questions List of open problems in sublinear algorithms

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.705112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.751098Z digest=sha256:14711c140e99966370a2b77d803c02d5dc6fbfd031f7232ee2b2334888365dcd

Observation 4a628fec-1c3f-4286-8a62-7c19ff390b07 · outbound

This paper cites Introducing deep research.

UQ: Assessing Language Models on Unsolved Questions Introducing deep research

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.687109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.754764Z digest=sha256:d8c7f22f20e77bfa9ab976298da264aad4c5de9eeaac03ea07d51ee0ec969127

Observation 6c25daa6-4998-412a-aae5-bf9e5fbdf080 · outbound

This paper cites Introducing OpenAI o1.

UQ: Assessing Language Models on Unsolved Questions Introducing OpenAI o1

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.671042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.758160Z digest=sha256:52b868ccb9fde4c07d33ceead4a667489275b369681ed7e13d98a949e212e7be

Observation e3eceb3f-2fb3-406b-a739-7e4813bb1314 · outbound

This paper cites Introducing OpenAI o3 and o4‑mini.

UQ: Assessing Language Models on Unsolved Questions Introducing OpenAI o3 and o4‑mini

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.658233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.761596Z digest=sha256:12ed653efafc7f2104e903aba6ec416718141e0144a840d32730794a67c82bd3

Observation deceb3d9-7c72-4a20-9e7f-d229cbc500aa · outbound

This paper cites Llm evaluators recognize and favor their own generations.

UQ: Assessing Language Models on Unsolved Questions Llm evaluators recognize and favor their own generations

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.646999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.765577Z digest=sha256:33cd16afb8d94d81f0b3c9610448da7ee898af3df9e29063207ffc733cd03570

Observation 56d7c659-603b-4be9-bd3b-05300528dfcf · outbound

This paper cites For better or worse, benchmarks shape a field.

UQ: Assessing Language Models on Unsolved Questions For better or worse, benchmarks shape a field

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.635111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.769878Z digest=sha256:31dce7ea63bbb5eca0fd8d3ebb66feb35cee8a45c27b1f7efeeea9ff7b481d61

Observation 5c2371bb-a85c-4f62-af4c-210971e40705 · outbound

This paper cites KILT: a Benchmark for Knowledge Intensive Language Tasks.

UQ: Assessing Language Models on Unsolved Questions KILT: a Benchmark for Knowledge Intensive Language Tasks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.774679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.774679Z digest=sha256:85dd8d92ea0c234cd67d2df3f429cbc9e1e15f6d460ef2053666504747573ef9

Observation 2197aa18-381d-4f4b-8320-10369cff45bd · outbound

This paper cites Humanity's last exam, 2025.

UQ: Assessing Language Models on Unsolved Questions Humanity's last exam, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.622441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.779121Z digest=sha256:4a6923fee32803c69101be295daaf8db66733d0130d10f47ef0e21ec579e1aaa

Observation fd37a82f-a72d-47b4-bf2a-36bb7c1ad53f · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

UQ: Assessing Language Models on Unsolved Questions SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.783306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.783306Z digest=sha256:06444f5cb2f3877b6b891174c0da60907dea040faf878904d967f21989b567fb

Observation 6218fdfb-c2e4-45b0-9ea3-dc64640d69ae · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

UQ: Assessing Language Models on Unsolved Questions Gpqa: A graduate-level google-proof q&a benchmark

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.609957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.787188Z digest=sha256:2ba824998a500c00e2b9e8723665cd3f0076b9d9a0cb1737c722671fcbfa917d

Observation 326f22e5-6d1b-4715-85f1-3cc44efa40da · outbound

This paper cites Measurement to Meaning: A Validity-Centered Framework for AI Evaluation.

UQ: Assessing Language Models on Unsolved Questions Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.792055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.792055Z digest=sha256:c2d7f238d3d5e26ba7e0d4ee196164b4b979226be3be9d71e942a7bfd8ca19a0

Observation fb828bae-2d44-495f-8143-88060708b8ea · outbound

This paper cites The Leaderboard Illusion.

UQ: Assessing Language Models on Unsolved Questions The Leaderboard Illusion

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.796329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.796329Z digest=sha256:4e9150760429b3abec39027fba4edda48d975866ba4fa270e91338a31652764c

Observation 1608eedb-7ee9-4894-8458-1124658ba64f · outbound

This paper cites Stack Exchange.

UQ: Assessing Language Models on Unsolved Questions Stack Exchange

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.598457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.800390Z digest=sha256:66598ff740155af1f838f054dbe43a7810bb57da9ab78857f92db0ff20004020

Observation 0410b2a1-1b7c-4dad-9a8f-d2539af0772e · outbound

This paper cites Terminal‑Bench : A benchmark for ai agents in terminal environments.

UQ: Assessing Language Models on Unsolved Questions Terminal‑Bench : A benchmark for ai agents in terminal environments

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.586555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.804671Z digest=sha256:a3afda3de95da3e9bdbfbd5c0d028e089eb1bdd102d27260d7acb650fbe91539

Observation 0f163747-562a-4de4-97a2-59c4637bf9e5 · outbound

This paper cites Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead.

UQ: Assessing Language Models on Unsolved Questions Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.808716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.808716Z digest=sha256:45fde431d972c9680b286d60ba3f69236526681128404ac041a7fb517d3df3ce

Observation 61312f11-fa5f-47f3-b869-c30b2ab19004 · outbound

This paper cites Supergpqa: Scaling llm evaluation across 285 graduate disciplines, 2025.

UQ: Assessing Language Models on Unsolved Questions Supergpqa: Scaling llm evaluation across 285 graduate disciplines, 2025

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.572183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.812701Z digest=sha256:47824dffbcc8ee56043925c99880bebb0233154c0d00bde02ac43b131795189f

Observation 5716fd16-daec-44b3-8206-cc4a9af77846 · outbound

This paper cites SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems.

UQ: Assessing Language Models on Unsolved Questions SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.816806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.816806Z digest=sha256:79814df4d7a4f9cb0ca6abd75d3648da4f735027d8825f0e90199bfaf02fc35c

Observation d5a02e5a-b954-4fb5-99a2-6e485959b2f8 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

UQ: Assessing Language Models on Unsolved Questions GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.820943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.820943Z digest=sha256:0aa6826e8c272c0d330a261ded74ee0844475a4859323694445910c4fdcd38f3

Observation af631803-152b-46b6-9303-cd34c843b03e · outbound

This paper cites an unresolved cited work.

UQ: Assessing Language Models on Unsolved Questions Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.824905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.824905Z digest=sha256:3019fbb2c1156c92e8e7f2fb06ff2a0e39983dc3ce8549dfe6e9bea2b89a617b

Observation 24f5b672-60a7-49ca-80f8-3ab96fef1669 · outbound

This paper cites PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization.

UQ: Assessing Language Models on Unsolved Questions PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.828751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.828751Z digest=sha256:579bab474381b4ed8ad92d5ad7c9909044ebc09c51641db3a6f0b5a942365399

Observation 77f7bc9b-f90c-4732-aa86-290f9a53859c · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

UQ: Assessing Language Models on Unsolved Questions Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.548388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.832809Z digest=sha256:51fad0eb7f0e929bf411303e8baf336ef1b13b82caaf654af9729a9ab7163338

Observation bbab3f8d-f82d-429a-85a0-c1756cd4f1b6 · outbound

This paper cites Self-Preference Bias in LLM-as-a-Judge.

UQ: Assessing Language Models on Unsolved Questions Self-Preference Bias in LLM-as-a-Judge

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.836673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.836673Z digest=sha256:acb74b48cfc6c95c3287eb5ee485e2d7bd7ccbe5a7c1d28b6e50776378250108

Observation c9f21bba-2b91-4da7-a383-e50b111deb34 · outbound

This paper cites Measuring short-form factuality in large language models.

UQ: Assessing Language Models on Unsolved Questions Measuring short-form factuality in large language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.840580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.840580Z digest=sha256:1d61af165b657af8d32c4462e371e90742f8623267f205bee2d58527c69514cc

Observation 18a23df0-5f8f-4c0a-9cbf-4f3bcb23d3b6 · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

UQ: Assessing Language Models on Unsolved Questions BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.844366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.844366Z digest=sha256:3559144ad4f677445ffab04b918e41fcccc7904d318bb5a7b83c8f6e672f7ad1

Observation b7d7cb80-05cc-4274-92e7-bc59da4121d1 · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

UQ: Assessing Language Models on Unsolved Questions LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.848296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.848296Z digest=sha256:80b65e3eefb9661759d807873c7b659d390e7e5ea5e114ff973e81829e7b1f42

Observation 7ecb45b0-6194-47f7-924f-ae0483dd783e · outbound

This paper cites Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement.

UQ: Assessing Language Models on Unsolved Questions Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.852288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.852288Z digest=sha256:4c5622df982f7462ea2c66944a0a43f12810128e08d5b3fa3753434c22490ffb

Observation 922a6fd3-f114-4d7b-bcbc-f62e11b2bddc · outbound

This paper cites Jimenez, Alex L.

UQ: Assessing Language Models on Unsolved Questions Jimenez, Alex L

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.856397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.856397Z digest=sha256:7c9f8c2eab3d9785e06c25eb0fcbe7c6b38ed30d1378bbae5c182f5c3f6f1f4f

Observation ce5b0cba-cc9a-4e86-bbe2-5c0fc907b01e · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

UQ: Assessing Language Models on Unsolved Questions $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.859994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.859994Z digest=sha256:43d993b2fc854de776261ec152a94e1364236fa887ea80c6a5c5b0d6c55275a7

Observation 4f237b55-1857-4f1e-adad-db3377df5413 · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

UQ: Assessing Language Models on Unsolved Questions Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.863897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.863897Z digest=sha256:b22edc7319ec6f16bf793bbfbaea96f8179d8aca52b2c3e507be4a2c864af100

Observation 66825604-8add-46ae-9958-6f17510d1834 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

UQ: Assessing Language Models on Unsolved Questions HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.868052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.868052Z digest=sha256:3c4a1ded4e092eef81beb43490356d8e0b1509635c1a831dd4e19fd98305b6cd

Observation fd6385f5-84a6-49e2-88ce-eee3f128ed90 · outbound

This paper cites A careful examination of large language model performance on grade school arithmetic.

UQ: Assessing Language Models on Unsolved Questions A careful examination of large language model performance on grade school arithmetic

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.529186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.872281Z digest=sha256:ca744ea475334a65c1dfa9ab9f97b861d78e140b030d6f05d29ab098a4948042

Observation 7adaf2fb-c55c-4003-be7f-b533fa00d918 · outbound

This paper cites Challenges in Trustworthy Human Evaluation of Chatbots.

UQ: Assessing Language Models on Unsolved Questions Challenges in Trustworthy Human Evaluation of Chatbots

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.876480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.876480Z digest=sha256:890cb7ca876df82091c969bc960b6074c27d0086a3a6ee9eb59bbbd539ccd6bf

Observation 2fe1d022-d3b2-45cf-8ff6-2d05aac7e3ba · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

UQ: Assessing Language Models on Unsolved Questions Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.880787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.880787Z digest=sha256:1a8296a2492a8b01980db017278f45b78a2878d265d9bce91ea0adbb7d62de0d

Observation 5dde716a-4b75-4d50-b822-de3b1a0e9cca · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

UQ: Assessing Language Models on Unsolved Questions AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.885288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.885288Z digest=sha256:f8d6e805f3c90966f1cf894c657358ee874523c0e6f3e8aea0dc99fadb010dcd

Observation 73bd8a55-7730-41ab-a739-ca47bfd026d9 · outbound

This paper cites LIMA: Less Is More for Alignment.

UQ: Assessing Language Models on Unsolved Questions LIMA: Less Is More for Alignment

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.889330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.889330Z digest=sha256:896139eacd5d8846b596f414170e52150fb1403be7db9b20c011da579c4483c5

Observation c801d62d-bb02-4a36-ab60-326a1d65e71a · outbound

This paper cites Reinforcing General Reasoning without Verifiers.

UQ: Assessing Language Models on Unsolved Questions Reinforcing General Reasoning without Verifiers

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.893401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.893401Z digest=sha256:fb81abba2333d6d09d347bc6d77786b54f77461c96243d928969cf634233b805

Observation 1adc20b5-9a0a-4593-926a-6c1846e92381 · outbound

This paper cites Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions, 2025.

UQ: Assessing Language Models on Unsolved Questions Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions, 2025

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:20.516012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T17:07:19.897165Z digest=sha256:613ba492a76765de64fa3f662ea93274acd2a57a4afc49d5f965e36f4ec275bd

Pith citing papers

Observation 51e8effd-fb12-4824-8bf6-19061be8cfac · inbound

MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval cites this paper.

MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval UQ: Assessing Language Models on Unsolved Questions

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:06:04.097499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T04:09:46.616019Z digest=sha256:6d260491239688eb9cc7f92d5d8f5e7bbf638fcca5778eb54b5522a4be57abbb