Pith. sign in

Paper Citation Record · LEDGER

HARP: A challenging human-annotated math reasoning benchmark

As of 11 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 8 inbound Pith citation observations for arXiv:2412.08819.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08819 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:34:19.857293Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:01:33.201477Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:46:32.285662Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved33
  • parse uncertain3
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 080e479a-d1aa-401c-a67d-c11e647d55b4 · outbound

This paper cites Chatgpt, 2024.

HARP: A challenging human-annotated math reasoning benchmark Chatgpt, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.821715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.569312Z digest=sha256:4943b14eb05c321ef5822993cbd37b15007d1eb7468a404b99adaf657ca5ccda

Observation 3f96f858-73d7-4a8b-be5b-030891b4d940 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

HARP: A challenging human-annotated math reasoning benchmark Training Verifiers to Solve Math Word Problems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.575547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.575547Z digest=sha256:3ce4c5b78679e5211fcff4e32881a541572253f5c690d852af14bf0dd4e56c7e

Observation 16a420f4-4a86-49af-a7b1-a8ef7c01ce97 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

HARP: A challenging human-annotated math reasoning benchmark Measuring mathematical problem solving with the MATH dataset

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.809078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.582098Z digest=sha256:2b31c740297b24634de67eb83f3461026b02448de069c9be399081af8d62a46e

Observation de6b80d4-fe55-414a-870e-b634fe5182a5 · outbound

This paper cites Learning to reason with llms, 2024.

HARP: A challenging human-annotated math reasoning benchmark Learning to reason with llms, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.587665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.587665Z digest=sha256:011b1f1a14529e6bbc24d53194ab31cb0d54507abfa80990d18c350f5731d0e0

Observation 6e1c5102-e9fe-4f2f-a2cd-1bbb9c9b2027 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

HARP: A challenging human-annotated math reasoning benchmark Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.592819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.592819Z digest=sha256:210d60006ac7ebc9b08a7c513dca2a6500397bfbf3f651b5ccf8b7dbe6f8b047

Observation 8a2e0a48-9b59-45e8-ac1a-f3b8ac19d505 · outbound

This paper cites an unresolved cited work.

HARP: A challenging human-annotated math reasoning benchmark Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:34:20.790293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.598102Z digest=sha256:a7167677c8ec058384c190ea3f00df2fb5db99df0c30154c785465fa4832d0ca

Observation a0bcf2d8-d305-4da1-ab7a-299d8adc1715 · outbound

This paper cites Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves.

HARP: A challenging human-annotated math reasoning benchmark Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.610477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.610477Z digest=sha256:237e5022163c2ba8ef584bbc2057ee3bfe6e0b41cc99f7ff8e872de4954ce35c

Observation f5786bc6-c11f-4e22-817c-333a4cdd35f2 · outbound

This paper cites The Llama 3 Herd of Models.

HARP: A challenging human-annotated math reasoning benchmark The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.615496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.615496Z digest=sha256:2d9a82ddc0594b06c0a9d9f0e76ca15648caacc4175198c1b3b51e6f64406c57

Observation f28e3a18-3242-497d-8408-8726aeec69bd · outbound

This paper cites Data diversity matters for robust instruction tuning,.

HARP: A challenging human-annotated math reasoning benchmark Data diversity matters for robust instruction tuning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.765799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.620043Z digest=sha256:8a23733ed337f87f555a582d47a205afe8820b56616c332e7a5352b553dca6fa

Observation 784b433b-36f8-4e23-9228-587c27d7f4a1 · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

HARP: A challenging human-annotated math reasoning benchmark Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.630654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.630654Z digest=sha256:046e1fc8bdc08ed8adbc7a7be2a48656187df5e883d1c38d4481c032c368fc04

Observation 5c146d61-d15e-4ac8-9d1b-6d92489a81c0 · outbound

This paper cites Data Diversity Matters for Robust Instruction Tuning.

HARP: A challenging human-annotated math reasoning benchmark Data Diversity Matters for Robust Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.625889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.625889Z digest=sha256:e781674cb3a0a817d0b4d6d7f2b6fd7f67cd43d14ddbcada1d215ce4d8db24a7

Observation 1d45a0cf-0e39-4880-ab84-13741379adc0 · outbound

This paper cites Xwin-lm, 9 2023.

HARP: A challenging human-annotated math reasoning benchmark Xwin-lm, 9 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.752907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.639752Z digest=sha256:a0e3c1f01e749345a43cda864a3ae587651b9ee0b868d4b51151df7027eb31d0

Observation 83c817e7-3ec0-462f-89e1-80d6044ca40f · outbound

This paper cites STaR: Bootstrapping Reasoning With Reasoning.

HARP: A challenging human-annotated math reasoning benchmark STaR: Bootstrapping Reasoning With Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.635312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.635312Z digest=sha256:82fb7cb0a58d7eb99b2d95fa3238db808042f89c8890611f6e15aed53d72f04b

Observation 7d9bf9fd-901a-47c9-a242-26fd61d4e88c · outbound

This paper cites Hello, gpt-4o, 2024.

HARP: A challenging human-annotated math reasoning benchmark Hello, gpt-4o, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.727983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.647328Z digest=sha256:cf1c194c19e21621e4c45c5c018b10626a5c1e5d608ef6b69f7ed25e9c5ed424

Observation 229da6c0-15b8-4363-9a31-6067f97830fb · outbound

This paper cites Introducing the next generation of claude, 2024.

HARP: A challenging human-annotated math reasoning benchmark Introducing the next generation of claude, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.739790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.643611Z digest=sha256:34414a76941037d3f97bb115098a4f01df1819ec04f29e8fe6224421393f98ad

Observation 78c4e2cc-60e4-474d-ace9-37c94080dac1 · outbound

This paper cites Solving quantitative reasoning problems with language models.

HARP: A challenging human-annotated math reasoning benchmark Solving quantitative reasoning problems with language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.713844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.654645Z digest=sha256:6ff4c92818f0a4a4900529096d6c40687bf2f8c08aa2c4b99c18f35e9c93e3f8

Observation 81ce59c0-e3f5-4259-bcf3-ab3719a02e2d · outbound

This paper cites Evaluating Large Language Models Trained on Code.

HARP: A challenging human-annotated math reasoning benchmark Evaluating Large Language Models Trained on Code

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.650915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.650915Z digest=sha256:b797bd25e1fe8d71dcd1a4e5ad047e292643432f741e55e351ccfad3740fe85f

Observation 90f8dd49-73ab-4ce4-93e8-3b5545222d3d · outbound

This paper cites Changing Answer Order Can Decrease MMLU Accuracy.

HARP: A challenging human-annotated math reasoning benchmark Changing Answer Order Can Decrease MMLU Accuracy

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.662284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.662284Z digest=sha256:94e1965b80e79330e52e113e900a873a9ebe8a01027cd314155cd564084010de

Observation 5dff412e-a531-43cb-a59d-68ca93ea8d71 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

HARP: A challenging human-annotated math reasoning benchmark Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.658327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.658327Z digest=sha256:17deadc6d926afcd9be7e7388529cd5ce274ea240dc4daed7c9ec39d853f4446

Observation 2c3e9ad9-a431-4c10-b7b2-f6e82ee5b09c · outbound

This paper cites Testing information, 2024.

HARP: A challenging human-annotated math reasoning benchmark Testing information, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.690044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.677792Z digest=sha256:b70966ae67575d9a0ec9497dd648b1210eba41844d697c4b1501768e2737a521

Observation 004c34a2-0270-415e-a792-1cceeb8345ab · outbound

This paper cites Omni-math: A universal olympiad level mathematic benchmark for large language models,.

HARP: A challenging human-annotated math reasoning benchmark Omni-math: A universal olympiad level mathematic benchmark for large language models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.667340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.667340Z digest=sha256:8d31e350b52ad3a42183d05e25bac2a53e812793cb80aaf2a9bf11da60c82b84

Observation 428f568d-3a82-490b-acc5-251c23c4b7d2 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

HARP: A challenging human-annotated math reasoning benchmark Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.672213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.672213Z digest=sha256:0a7b1e6e554c44b1a8b517179f6d22f02f3fec88a38eb667c424f810298480f1

Observation d4eb36ff-f067-4941-b53e-5e2ab3f92be1 · outbound

This paper cites Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku, 2024.

HARP: A challenging human-annotated math reasoning benchmark Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.660892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.690968Z digest=sha256:10373e51db33c2ef253e9b5db86c981dd3aa6cb213e884db3adeddbc09f536fc

Observation 88df54c5-7cd8-4376-8cf3-d44cc4b71f5d · outbound

This paper cites Openwebmath: An open dataset of high-quality mathematical web text, 2023.

HARP: A challenging human-annotated math reasoning benchmark Openwebmath: An open dataset of high-quality mathematical web text, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.675245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.682376Z digest=sha256:5336a46303507fd2cdc04785587360d868ae92a2daf064da472b4c12763a1388

Observation 188d4292-397c-462a-94c2-74e9afd10170 · outbound

This paper cites Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena.

HARP: A challenging human-annotated math reasoning benchmark Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.686581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.686581Z digest=sha256:0a881056aec192f7e25c5659d95abdd70ec6ffb0ae3bf843ebd67c0604f16e1c

Observation 2efcd1d2-87fe-4223-af5d-8a923f59a299 · outbound

This paper cites Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?.

HARP: A challenging human-annotated math reasoning benchmark Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.704545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.704545Z digest=sha256:f89dfcc07ea7b2ad14b83bd2ef544fa7c5e97634470b56c2c7857f3bdc30a8f1

Observation 206d0c5e-20bb-46d0-a8ec-6eb56ffcb2cd · outbound

This paper cites Claude 3.5 sonnet, 2024.

HARP: A challenging human-annotated math reasoning benchmark Claude 3.5 sonnet, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.643416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.695592Z digest=sha256:beea9c758599e1c5b6614d2a1a72f475db591d75e5cf0a7f8188606a32354063

Observation c41a89d4-f75b-4505-aade-5f7fbd999cd4 · outbound

This paper cites Redpajama: an open dataset for training large language models, Oct 2023.

HARP: A challenging human-annotated math reasoning benchmark Redpajama: an open dataset for training large language models, Oct 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.628794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.699736Z digest=sha256:2082fde6b441e543c783561bc9231c6842c34e464929e56ac096ca7b10a6efda

Observation d3d6593b-e74b-4a57-850c-752a03a06229 · outbound

This paper cites Keep Guessing? When Considering Inference Scaling, Mind the Baselines.

HARP: A challenging human-annotated math reasoning benchmark Keep Guessing? When Considering Inference Scaling, Mind the Baselines

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:34:19.960286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.719240Z digest=sha256:5750203e8f21fc0ecb5b97fb685dbfedc4ecee85acc63b3f385a79998dbbc389

Observation 2f63f094-49b2-43eb-ac65-29d216b92852 · outbound

This paper cites Time spent thinking in online chess reflects the value of computation, Oct 2022.

HARP: A challenging human-annotated math reasoning benchmark Time spent thinking in online chess reflects the value of computation, Oct 2022

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.614342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.709025Z digest=sha256:82bac0e8a133976ffc23fa72ab452c661ac83aaa8b732e62d48bf3c930d01d92

Observation b5c65b1d-79b4-44f0-a2c5-7c950230e371 · outbound

This paper cites V-STaR: Training Verifiers for Self-Taught Reasoners.

HARP: A challenging human-annotated math reasoning benchmark V-STaR: Training Verifiers for Self-Taught Reasoners

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.714589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.714589Z digest=sha256:ed738d752bd6060abfc25a88e01e6101e1372644f09ef500866b12f8463550c2

Observation 6cb9b301-1ad3-4483-95a2-a675fc773bc9 · outbound

This paper cites Quantifying Variance in Evaluation Benchmarks.

HARP: A challenging human-annotated math reasoning benchmark Quantifying Variance in Evaluation Benchmarks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.733078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.733078Z digest=sha256:09cba5a1c804107638b4710c18ae12e8c8702c41a0aeeb92f021ddd23105c0d9

Observation 08221698-9596-406e-894b-d2f80709db9e · outbound

This paper cites Testing language models on a held-out high school national finals exam.

HARP: A challenging human-annotated math reasoning benchmark Testing language models on a held-out high school national finals exam

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.723542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.723542Z digest=sha256:44b4a90b7e927c373798a24500dc41c417f6db412b4443849447df225c22d36d

Observation 286fcd5b-e4d4-49c0-82eb-4dbe1ed3d619 · outbound

This paper cites MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data.

HARP: A challenging human-annotated math reasoning benchmark MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.728249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.728249Z digest=sha256:fcd7ff3fdd783f1d1fb26c366c2cbf8bb3f456bc85a24ab4b6bf77be3b5266d4

Observation 6864c502-3dbc-4ecf-b948-b44014a6541f · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

HARP: A challenging human-annotated math reasoning benchmark FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.746349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.746349Z digest=sha256:892a6b26bb96f79fc40bec130cb7acc1bdfbec90a3044a8950cc6ffaffa051b3

Observation 5d1392fc-f062-4c28-ad58-286ae3a831b5 · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

HARP: A challenging human-annotated math reasoning benchmark A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.737720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.737720Z digest=sha256:b2c2eae5ab7495682f90bf7d6400f22aff6be1e53f7958f11a1934fa2c20c1ce

Observation 2fdc50ab-f912-4b5d-90f6-fff5be4a5271 · outbound

This paper cites Scale’s seal research lab launches expert-evaluated llm leaderboards, 2024.

HARP: A challenging human-annotated math reasoning benchmark Scale’s seal research lab launches expert-evaluated llm leaderboards, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.585722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.741799Z digest=sha256:da528b92b9c0cf0fbeed2a1b6b35ac4aa0954370d7263b940403175e7b4fdc6b

Observation 4a6323dd-70cb-4a72-869e-b0d8c4f904b8 · outbound

This paper cites Answer :\ n$ANSWER.

HARP: A challenging human-annotated math reasoning benchmark Answer :\ n$ANSWER

Reference 38

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T17:34:20.549444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.757570Z digest=sha256:e2d418be4623608f7592a2572a7d4452d7c7bffe18ab2575978626bfd9d0a281

Observation 7a0518cd-120a-4150-bcde-84b54b3d16b6 · outbound

This paper cites RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold.

HARP: A challenging human-annotated math reasoning benchmark RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.750428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.750428Z digest=sha256:e7fd7f60b71f16a73a848f509cf1634e88fb0897aad55cd499133c02f17f4db7

Observation cac344a3-8986-427b-9205-5c0e7c7a110d · outbound

This paper cites Lampinen, Stephanie C.

HARP: A challenging human-annotated math reasoning benchmark Lampinen, Stephanie C

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.565437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.754274Z digest=sha256:74be4b33e4e98065dccd1ec57a1c34b632a92cd5ec84ad289e10ad0aea498534

Observation af65e82a-0fe6-4850-bd7c-1df6f7df985a · outbound

This paper cites an unresolved cited work.

HARP: A challenging human-annotated math reasoning benchmark Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:34:20.536236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.762133Z digest=sha256:4f53ac84ce71998e5cd9a46ec02d143917a282afd6b651dd08d3b630e85deeae

Observation bd0dc6b0-997e-4138-bc49-95cd54a1bcc7 · outbound

This paper cites an unresolved cited work.

HARP: A challenging human-annotated math reasoning benchmark Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:34:20.521622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.766398Z digest=sha256:dbafef98f1f2888666eea084d6fbe429c7f13bfa577db721e20c824006db5414

Observation cb5e5df4-8c4e-49f3-a500-3706a1325485 · outbound

This paper cites The area of S is the area of the hexagon with vertices wk, which is 9 √ 3 2.

HARP: A challenging human-annotated math reasoning benchmark The area of S is the area of the hexagon with vertices wk, which is 9 √ 3 2

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.507454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.770514Z digest=sha256:2cba39507b909a1f879c3afa573b31e392c698d01cc60b99f7f59a22d6ecd6c7

Observation 4279aa4d-674f-4a4c-ba98-8d708e9fd9d6 · outbound

This paper cites t5 = 1 t4 = 1.

HARP: A challenging human-annotated math reasoning benchmark t5 = 1 t4 = 1

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.491162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.773776Z digest=sha256:43c52f569829921dc36ce5bcc6d2989493c837c4e69f04fdfccf3618c19d78d5

Observation 719a4869-0873-4673-a2e0-afc6afa9ed32 · outbound

This paper cites an unresolved cited work.

HARP: A challenging human-annotated math reasoning benchmark Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:34:20.463367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.777165Z digest=sha256:026b2ea6778fd5d97d240ea9a4c8be0a6a960ba2dba0a7863c2e0814bac46673

Observation 11f7798d-bfe8-43f0-8f15-7d1daae92721 · outbound

This paper cites t8 = 1 +t8/2 = 1 +t4 = 1 + 3 = 4.

HARP: A challenging human-annotated math reasoning benchmark t8 = 1 +t8/2 = 1 +t4 = 1 + 3 = 4

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.441752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.781287Z digest=sha256:d959b9ee021d9cffc9fe75db24802644669c1989d7dd3fa2b2de7c580dbc2326

Observation aa43fd98-9ec3-4eff-a085-d4852b1b9eed · outbound

This paper cites t11 = 1 t10 = 1 4/3 = 3.

HARP: A challenging human-annotated math reasoning benchmark t11 = 1 t10 = 1 4/3 = 3

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.424343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.785423Z digest=sha256:c67ffccaca245dec2983c14e10c7387b3e8bf56f2497d093351bfc4fe9fe7801

Observation ec2af754-ef09-4dd7-ae48-89f06c0312aa · outbound

This paper cites an unresolved cited work.

HARP: A challenging human-annotated math reasoning benchmark Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:34:20.407534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.789288Z digest=sha256:28f8943b1c9b26c49b15f4322925930c51999d5ca7fec5623378a386780a6a4c

Observation 8bd601f8-f54f-4ef9-8ca3-639c50ebb827 · outbound

This paper cites t14 = 1 +t14/2 = 1 +t7 = 1 +2 3 = 5.

HARP: A challenging human-annotated math reasoning benchmark t14 = 1 +t14/2 = 1 +t7 = 1 +2 3 = 5

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.392257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.793256Z digest=sha256:b2d7a60b63b317097b17628fcaba2828512a2383b36d6f94f21e57fb282f28a1

Observation 994473d6-4b1c-486c-9314-e14f0afd17a0 · outbound

This paper cites an unresolved cited work.

HARP: A challenging human-annotated math reasoning benchmark Unresolved cited work

Reference 51

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T17:34:20.379912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.797215Z digest=sha256:f830acc8ba29766fd252742219a97da5bf923f2aaea37b67f00e64dbb24a2885

Observation ab0c6a93-5e41-4e94-aee8-dec732980ed1 · outbound

This paper cites t17 = 1 t16 = 1.

HARP: A challenging human-annotated math reasoning benchmark t17 = 1 t16 = 1

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.364549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.801034Z digest=sha256:d937410cdf99f22ccbbbd76396c854071566948da82abf505e1c4ef0fc0ef681

Observation 7eb4150f-2fa1-4c84-b4d1-6175dc5aa1b7 · outbound

This paper cites an unresolved cited work.

HARP: A challenging human-annotated math reasoning benchmark Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:34:20.351027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.804881Z digest=sha256:e81596c5c2d8b21b61b00cbac1817a0a7d554e09c0a07d5fbae3c671e1194657

Observation 2f28deb4-8072-48dc-b8f0-1f558e8d1e6f · outbound

This paper cites t20 = 1 +t20/2 = 1 +t10 = 1 +4 3 = 7.

HARP: A challenging human-annotated math reasoning benchmark t20 = 1 +t20/2 = 1 +t10 = 1 +4 3 = 7

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.337684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.808608Z digest=sha256:0c71b9f371395fcabee4d043bf4d55008b3c39798171d62bfad178fe0d3ad26d

Observation 21954628-3140-4ec5-85d4-2f53ca579a08 · outbound

This paper cites an unresolved cited work.

HARP: A challenging human-annotated math reasoning benchmark Unresolved cited work

Reference 55

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T17:34:20.317815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.812767Z digest=sha256:429e67cbd3eb310bc0ff22718661d416dbc151f6e5e7ba277a4ce4f1a3f27ee8

Observation 52713562-15a3-4f72-b1a3-1d57247ae990 · outbound

This paper cites an unresolved cited work.

HARP: A challenging human-annotated math reasoning benchmark Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:34:20.297999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.816515Z digest=sha256:b1e90f0e4035974accfdcbac39164c3226356f2f62919a3ef856d06fcafdf622

Observation 5c1dc497-2e93-4f43-971e-090d5d610f1b · outbound

This paper cites an unresolved cited work.

HARP: A challenging human-annotated math reasoning benchmark Unresolved cited work

Reference 57

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T17:34:20.283851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.820282Z digest=sha256:472c9d91180a6718d669b6d3bbf370d3cbc14667a373241ebab881564c1405a7

Observation 296afc59-afc6-432d-a749-6fa1f685a73c · outbound

This paper cites t25 = 1 t24 = 1 7/2 = 2.

HARP: A challenging human-annotated math reasoning benchmark t25 = 1 t24 = 1 7/2 = 2

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.270810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.824566Z digest=sha256:d52d10ab8c5988531dab5d04c78e1342fce30ebf444efcd7303aa91583f71c41

Observation 100c5cfa-8059-46d5-b9ce-f9cddd92d478 · outbound

This paper cites an unresolved cited work.

HARP: A challenging human-annotated math reasoning benchmark Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:34:20.253505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.828763Z digest=sha256:a0aa349232793190371e27ec6d03061ffecd1dca5c9169de259f3998cdd47158

Observation 4526bc41-104a-47bf-9549-0b241a8874b5 · outbound

This paper cites t86 = 1 +t43 = 1 +1 t42 = 1 + 1 1+t21 = 1 + 1 1+ 3 7 = 1 +1 10 7 = 1 +7 10 = 17.

HARP: A challenging human-annotated math reasoning benchmark t86 = 1 +t43 = 1 +1 t42 = 1 + 1 1+t21 = 1 + 1 1+ 3 7 = 1 +1 10 7 = 1 +7 10 = 17

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.238447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.832931Z digest=sha256:8ef9011de5e4246b312c040093ecd7bc400b50f7fe8d7374752606f9d38754f7

Observation bf11f43d-54ce-40db-96a3-f8908a543937 · outbound

This paper cites t174 = 1+t87 = 1+10 17 = 27.

HARP: A challenging human-annotated math reasoning benchmark t174 = 1+t87 = 1+10 17 = 27

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.223644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.836996Z digest=sha256:5509811868c68ca36032bd07aca0d5a461a7a1f0274631dbe478506b4e1f76cd

Observation 11d80079-384e-4acb-a209-725ca965aa54 · outbound

This paper cites an unresolved cited work.

HARP: A challenging human-annotated math reasoning benchmark Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:34:20.207755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.841092Z digest=sha256:b92d30ef2df65d13c31e446b4ae3fdff6c893ef418c5626a5aaaef30cb1b8385

Observation 52e71fd5-a22b-48c1-914d-0ea903d5f83b · outbound

This paper cites t1392 = 1 +t696 = 1 +61 17 = 78.

HARP: A challenging human-annotated math reasoning benchmark t1392 = 1 +t696 = 1 +61 17 = 78

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.194862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.845251Z digest=sha256:e766ac3226c3deac40927de163619249f1cd6bb478812bed943aa24dc6612ac8

Observation 98d495be-22b7-4d82-b056-41dc07eb986a · outbound

This paper cites an unresolved cited work.

HARP: A challenging human-annotated math reasoning benchmark Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:34:20.181820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.849121Z digest=sha256:1e971b35c8ed3ec58dfeafece78e8b0889386308208ccecceec29e4b0642cc8d

Observation d1ce1236-6d87-4eea-b30a-e348fec83a66 · outbound

This paper cites sinks" (vertices with no outgoing edges) or.

HARP: A challenging human-annotated math reasoning benchmark sinks" (vertices with no outgoing edges) or

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.166612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.852897Z digest=sha256:23c4ddc82d500f18c1337431b1f5fab514c84142b164b4a495116cbb423d435b

Observation 5dd73c09-db56-4040-bbd9-dec686974790 · outbound

This paper cites ## Step 8: Apply the Law of Cosines to triangle O1MO2 Using the Law of Cosines on triangle O1M O2, we can relate the side lengths to the given angle ∠O1M O2 = 60◦.

HARP: A challenging human-annotated math reasoning benchmark ## Step 8: Apply the Law of Cosines to triangle O1MO2 Using the Law of Cosines on triangle O1M O2, we can relate the side lengths to the given angle ∠O1M O2 = 60◦

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:34:20.153138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.857293Z digest=sha256:b37d88f42001593f4c4991aed9fdc8d36d5862a73beb2521071fe9aa668a008f

Observation ec23b7cb-29f8-48eb-ba38-d55ba4d90ce5 · outbound

This paper cites an unresolved cited work.

HARP: A challenging human-annotated math reasoning benchmark Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:34:20.778152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T17:34:19.605190Z digest=sha256:f7303f2db9b0094664126823d7e8446019bbb0b5f621524e46157d5fca1b7923

Pith citing papers

Observation 0a5cc932-9adb-4387-a388-e1f093099b4d · inbound

MathArena: Evaluating LLMs on Uncontaminated Math Competitions cites this paper.

MathArena: Evaluating LLMs on Uncontaminated Math Competitions HARP: A challenging human-annotated math reasoning benchmark

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:10:14.880871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T00:10:14.812539Z digest=sha256:de632b571518b6c73c5828d07cccddca6d8197db896a9d5d8d775a0f3e8b9f35

Observation da373034-26dd-401d-b93d-14e5b682b48f · inbound

A Survey of Deep Learning for Geometry Problem Solving cites this paper.

A Survey of Deep Learning for Geometry Problem Solving HARP: A challenging human-annotated math reasoning benchmark

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T17:01:33.201477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:01:33.201477Z digest=sha256:0b941affa355c06abbb4b478ac6306579c676d242a0fa30d54316250d289572a

Observation 7917bd7d-861b-4f80-9ded-b4b348468bba · inbound

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization cites this paper.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization HARP: A challenging human-annotated math reasoning benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.608109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.608109Z digest=sha256:ea0797e8fbfcdf628f47cdf385876b6f950f360f38da765a6fe92688fa401496

Observation 43746aa5-ec9c-4df8-8e5c-4f0de72a7011 · inbound

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward cites this paper.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward HARP: A challenging human-annotated math reasoning benchmark

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.956023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.956023Z digest=sha256:9c0f1285eb56104c92240f07944b5fabf126a049aee5e4d15943b9f728dcc9da

Observation 43d53863-abe6-482e-8d13-ea891e3004a2 · inbound

EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving cites this paper.

EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving HARP: A challenging human-annotated math reasoning benchmark

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:11:32.573003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T15:08:03.793304Z digest=sha256:c82e6f4f30573f51c27ae66317399d79bec738830e518669233f4ff87a340344

Observation 9f1dda6a-cbfc-4247-b7ba-4190b6b18fff · inbound

Re$^2$Math: Benchmarking Theorem Retrieval in Research-Level Mathematics cites this paper.

Re$^2$Math: Benchmarking Theorem Retrieval in Research-Level Mathematics HARP: A challenging human-annotated math reasoning benchmark

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:11:16.002409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T02:07:39.648269Z digest=sha256:dde9c558745ea9689b6e42fb4016287375b491f48a5bde50bb0fd3a44f4dd713

Observation eff1eb27-23a6-46d0-9ba9-939084390f85 · inbound

RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data cites this paper.

RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data HARP: A challenging human-annotated math reasoning benchmark

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:43:55.081562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-29T19:34:12.081362Z digest=sha256:9e915e2f19e2b467b5f6465099bcbb3eb9882599ed13a3342b00cc5b66fc2c31

Observation 18fdea39-2b56-4a2d-8b13-7633204e5076 · inbound

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions cites this paper.

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions HARP: A challenging human-annotated math reasoning benchmark

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:46:32.287295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T09:45:18.385925Z digest=sha256:09e2d80e4f8db600c6aee100f7854281c142b1608b409baffb2ad7dbbc4e6dd3