Pith. sign in

Paper Citation Record · LEDGER

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation

As of 18 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2506.15455.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15455 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:40:08.960042Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T02:33:04.549851Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T07:31:29.370868Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ddbf1f6e-1568-4b8c-aa5a-c9c0841fd8c5 · outbound

This paper cites R., Bury, G., and de Oliveira, S.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation R., Bury, G., and de Oliveira, S

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:40:09.873549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.689919Z digest=sha256:42544a857dc52d84afb96a7ad92501b8f32750a3c396b0eb5b2ffa28ad04023f

Observation 2381b885-e319-4f6c-8297-873f9ea42ca3 · outbound

This paper cites W., Conway, C.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation W., Conway, C

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.695478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.695478Z digest=sha256:c1c9e0b93d6f243037c2eb19262547764950b54934e3e6e912b7b5c0e75d8af9

Observation d105d418-b29a-4364-9e0c-94f1b0f1805d · outbound

This paper cites Language Models are Few-Shot Learners.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.701122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.701122Z digest=sha256:30dd87bbeb1c5776167304b370ad6ca04753409a5eb3ce3ca7ad95260ec53fb1

Observation 093aa8cc-8502-430e-8dc4-8b91fd908bca · outbound

This paper cites ARC Prize 2024: Technical Report.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation ARC Prize 2024: Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.712175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.712175Z digest=sha256:d2c1ef7338a7533fbde9e103bdd9221d47af1ce76a3959480606c95f568dd7fe

Observation 4dd80ccd-e603-498b-9189-6bff78f6520e · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.718055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.718055Z digest=sha256:7ad19cd128aaba4b4bfbea643221798ca2a91f660d40de99460e15961b606489

Observation c5f468f4-fa2c-475f-9d64-aabb0718792d · outbound

This paper cites Frama-C User Manual.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Frama-C User Manual

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:40:09.858155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.724358Z digest=sha256:18c48612d35122d873ef7473bf8ff4a9d2211fb484d7b03941db5340e1dbac2d

Observation 943f825b-4d28-40e9-b3d8-3642a7307cf8 · outbound

This paper cites and Bj rner, N.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation and Bj rner, N

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:40:09.843181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.729888Z digest=sha256:f34a36b3b737e701d3fe534a0078fce8c9dfa296598b43e37381fcfe2e25b401

Observation 88085062-9a74-4f6e-9941-bfa10bedabfd · outbound

This paper cites Pal: Program-aided language models.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Pal: Program-aided language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.735136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.735136Z digest=sha256:929dfea53f3d12b85e78e1ec3585795a9b6abc741bab7fffe02dcf0b40042bde

Observation 435a24fe-d609-42ce-8590-fbf85262dc76 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.739793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.739793Z digest=sha256:f375e4ca13deae4f5f2d7f53a23dac04123dedf1b225f6d6c79b60b524f9e06c

Observation f64170b5-00a4-4fcb-bc32-35d9dccd9b8d · outbound

This paper cites and Nori, A.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation and Nori, A

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:40:09.816741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.745800Z digest=sha256:be182e60a34069b062d7f6259fef13b37d764b44105c3171599b0adcd7bd275d

Observation f8d38a5b-76cd-4e68-a40e-cdd22d670ef3 · outbound

This paper cites CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.751296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.751296Z digest=sha256:25a6579a523af1a520c40c0917fd0313591ba14f6e89192fd20da4603f84500a

Observation 07508454-5ca7-42f2-b098-135814e255b0 · outbound

This paper cites an unresolved cited work.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:40:09.800701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.756971Z digest=sha256:1cbc9a55694f81202b082205d8987da10e56874771083300ab740e9feffb2f0f

Observation c5ed88e9-ccea-43a8-91e9-faa147a13958 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Measuring Massive Multitask Language Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.761512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.761512Z digest=sha256:bf05ab1e5c5f21cdc108277d9d8441833b6eacb7e704fec889790e6b381eeb7f

Observation 06e2c178-e753-4b85-aafe-13851bc0c9c5 · outbound

This paper cites an unresolved cited work.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:40:09.785238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.766707Z digest=sha256:5576bf43350f04e68bc7963956006ae3fb4f572478b24681644ed9d41de4ca0f

Observation 3d052235-c979-4078-8646-9c3386e45677 · outbound

This paper cites and Chang, K.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation and Chang, K

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:40:09.770621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.771276Z digest=sha256:29951cc6dcfa40983eff136160a7de75be1eba82ee0bd5f03e120de0448b9390

Observation d27b7724-7963-4f98-8403-8e4a37cc09ff · outbound

This paper cites Reasoning Elicitation in Language Models via Counterfactual Feedback.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Reasoning Elicitation in Language Models via Counterfactual Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.775606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.775606Z digest=sha256:2f873d6d19318c2b16d982c624565abcd8952f7750f71a4fc7923c789f1475e8

Observation a7b1354a-4af1-4c76-a11d-3fc8eba53a1e · outbound

This paper cites OpenAI o1 System Card.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation OpenAI o1 System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.780223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.780223Z digest=sha256:fbe8894ca5202021e08a9dbdceb0449d2e752fb90719b646765c0aeb6ce65a06

Observation 642a3b46-ff66-4b87-ba54-aa4de3f64724 · outbound

This paper cites Mixtral of Experts.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Mixtral of Experts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.785242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.785242Z digest=sha256:e1aa08dbc687912e8907639754cb930ea8dd7854f86aa82d74bf389ea6c046a4

Observation 090750a8-9089-4a49-b893-9470231d6c42 · outbound

This paper cites CL adder: A ssessing causal reasoning in language models.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation CL adder: A ssessing causal reasoning in language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:40:09.753221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.790448Z digest=sha256:22c44ec306b6462358b57c8e87c1703145c4bf07c678d6c03c024873c4a355cc

Observation 2e2ef2bd-6003-4140-a8c8-00c2322b0400 · outbound

This paper cites Cladder: Assessing causal reasoning in language models.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Cladder: Assessing causal reasoning in language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:40:09.737804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.795666Z digest=sha256:a59dc40b13e8e06d4edb8920e09c3be16757d94606536db8bc8c79b1683d7fb6

Observation 2beb983d-a1a4-4049-b534-a759dcb5d27a · outbound

This paper cites K., Lal, A., Rastogi, A., Roy, S., and Sharma, R.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation K., Lal, A., Rastogi, A., Roy, S., and Sharma, R

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:40:09.721492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.800802Z digest=sha256:1f75adcfc7b0e045fc41211079fe139e75ee51eab64df1c44b8f3e531120946b

Observation 4ddb2491-b701-47ec-aa87-d8435e90b864 · outbound

This paper cites S., Iwafuchi, T., and Dasaka, S.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation S., Iwafuchi, T., and Dasaka, S

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:40:09.704735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.805630Z digest=sha256:2bd07f68244d98ca64f8f3b7139a8bf63ff205bfcc86fd780b6e86acbf21e4e2

Observation 9a86c7f8-cdb4-41b4-8f8b-aa64fa20bf1f · outbound

This paper cites Evaluating the Robustness of Analogical Reasoning in Large Language Models.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Evaluating the Robustness of Analogical Reasoning in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.810909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.810909Z digest=sha256:ed9ce49bae19fc3c2e65f17f8079a18d2275cdf12805e20585d840ba6cf73b46

Observation d1fc62b6-3321-49b9-9bba-357dee2558d7 · outbound

This paper cites Neuro-Symbolic Data Generation for Math Reasoning.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Neuro-Symbolic Data Generation for Math Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.816339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.816339Z digest=sha256:eb3c38afc44f9dc2778b52a04eca1a84a7f20da8eb9fd7e6250e0e7185ac256c

Observation 59c2d6a5-0701-4cba-8c75-58abf932f19a · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.822965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.822965Z digest=sha256:9500174730f0b4c362ccfc0173575166f4ee440c0799026029f286445335109b

Observation 7768b750-28d6-462d-bcc5-618e68af04a7 · outbound

This paper cites and Krakauer, D.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation and Krakauer, D

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.831764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.831764Z digest=sha256:cb737e032e1c495c996c67053c82e077b0d63ae139d251bf4b1fdc7e427d7a83

Observation 5872c196-5e25-4f84-8f6c-fcd1ae19b23a · outbound

This paper cites an unresolved cited work.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:40:09.689055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.837698Z digest=sha256:a2a18c61280c4f8ad1f66045e5ea4dd944c550f94f149986446f4412a86317e7

Observation 2224da54-95cf-4457-aa8f-41ff87b27d6e · outbound

This paper cites Early access for safety testing, December 2024.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Early access for safety testing, December 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:40:09.673975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.843262Z digest=sha256:2ca4597e8048958adaa662fa9f83aef689d7341f14469009c391e6a7f6dc1aee

Observation 937d3089-5f2d-41ef-be73-f553e88a3b49 · outbound

This paper cites Causality: Models, Reasoning, and Inference.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Causality: Models, Reasoning, and Inference

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:40:09.658383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.850804Z digest=sha256:ea9fb4fbcafbfb62504c1b54429309b7450192ca3e8a0307db49f1e652424d6c

Observation bddd8586-dd46-4de4-9514-4eab8ee8a31d · outbound

This paper cites an unresolved cited work.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:40:09.642596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.855943Z digest=sha256:f41d4ceef8fa06c69f098d7456b61225dc71c268e82e9fedcbfefacfbf35e861

Observation 6487c19c-4d19-4863-8480-e70788761d82 · outbound

This paper cites H., Sch \"a rli, N., and Zhou, D.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation H., Sch \"a rli, N., and Zhou, D

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.861834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.861834Z digest=sha256:fdefd2340d2277fddc78ac14ab42d61f74fc4a0f1c0f5104d3ed2a883d0427fc

Observation 1643a9a5-dc38-4d8a-bdbc-dd73937182e8 · outbound

This paper cites Learning loop invariants for program verification.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Learning loop invariants for program verification

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:40:09.616359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.866989Z digest=sha256:91098742b1a12d18907e34ede59004a6a29d9afd9b5dc4ec950176324f58c3f5

Observation d6d33143-d23d-4653-83cb-12b5ccb956eb · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.873201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.873201Z digest=sha256:1bc10d9149ed8eb8b5b4e10a0ac4086124fbb9dbce660f998a4e4cf37e211298

Observation d0f7b3b7-7f7a-47db-b0b7-a4d04739393c · outbound

This paper cites Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.878802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.878802Z digest=sha256:e0a4cec1b0590c0c5159615951dc54f4aecc34d634adb3c204573e278a490feb

Observation 4ed2bb3d-7c5c-4da1-a6b6-8c9ebe52b141 · outbound

This paper cites an unresolved cited work.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:40:09.600363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.883832Z digest=sha256:ca5afae91746ea968e00a57a246af81e2eddca0143112333f17c0d440fc8ea6e

Observation 0f75384e-28e4-4293-a4f9-85957ff60ff1 · outbound

This paper cites OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.889506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.889506Z digest=sha256:269df89e4b57140d86aa8cf14be703c5177cd7f59481441189d8f785270c9c77

Observation 392b4c00-9a6b-4a73-94e5-3e4382f23feb · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation LLaMA: Open and Efficient Foundation Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.894465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.894465Z digest=sha256:153c5a83a004eaee1151e17d9e07010f1dde0a68fc7a1107e4c1adb124d1ce89

Observation ad0bbeb9-04a3-49cb-92d9-0f6e86c588a1 · outbound

This paper cites Large language models still can't plan (a benchmark for llms on planning and reasoning about change).

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Large language models still can't plan (a benchmark for llms on planning and reasoning about change)

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.899748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.899748Z digest=sha256:9ee8f6ce70bd235dc0a178e7edfc91cc6ef5d459e615fa2bf551d7a492000086

Observation 781fbb60-726c-4f75-9846-b3c7e080dce3 · outbound

This paper cites SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.905494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.905494Z digest=sha256:707d191dde6e3e4891ecf9f27003b161acb6efecacb64126407b59164bb4030f

Observation d32efc91-9715-49a1-8a1c-6d301bf0d49b · outbound

This paper cites S., and Momennejad, I.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation S., and Momennejad, I

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.910833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.910833Z digest=sha256:2693db04e59f5caf8ff55aa7c67c7fd35282b1d32d57bc44d5fcd10617aab3f4

Observation 133e0d5b-1916-4682-8c01-99c2cee94f08 · outbound

This paper cites Ethical and social risks of harm from Language Models.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Ethical and social risks of harm from Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.916015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.916015Z digest=sha256:42b6e88088513a28d88a3aea440a7224b22ec91a373e038a666961d6f640df95

Observation f4654180-7874-405f-bb9a-19f45f7a8b41 · outbound

This paper cites W., and Narodytska, N.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation W., and Narodytska, N

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:40:09.571359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.921815Z digest=sha256:d0060084c3f102abf08a4e3865d630a1383b0afb248059b4ebe43af9685fbcdf

Observation 79947bfb-2d7a-4a18-af7f-dfa847721a23 · outbound

This paper cites Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.926155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.926155Z digest=sha256:d743d32d529e8888c182605ad514df932d6cd80c182e1ab627138452a48a4edb

Observation 893ab7e0-d663-4c55-a758-b0b5384e1de0 · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.930958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.930958Z digest=sha256:0843a19651897cdecc58db44e941f73e567ec4a295944f5a447250b2386e8b37

Observation 4e615cad-e6a8-4121-a107-abada2f49625 · outbound

This paper cites A critical review of causal reasoning benchmarks for large language models.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation A critical review of causal reasoning benchmarks for large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:40:09.549227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.935588Z digest=sha256:222185f8e4a0fccc07f12d2aac65bf081e12a1ac39e16028000cca1eff26da08

Observation 4b8e9cfd-c774-44d1-8dd6-88ea535ca356 · outbound

This paper cites Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:40:09.531174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.940762Z digest=sha256:56bc00df809768cca00d911b4239f78356caacece957db823c47747cdd6b3169

Observation 291ca5db-d6a9-4d6b-a301-bf01ab4ce05f · outbound

This paper cites Darg: Dynamic evaluation of large language models via adaptive reasoning graph, 2024.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Darg: Dynamic evaluation of large language models via adaptive reasoning graph, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:40:09.513963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.945218Z digest=sha256:215d53d5afe4d09b906718a2a60f883cc414d080c6c9404c71d8a0db7e18cb93

Observation a4c3e969-c9c0-40db-a937-ddacf53dfa8e · outbound

This paper cites Don't Make Your LLM an Evaluation Benchmark Cheater.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.949778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.949778Z digest=sha256:b7a0f648813781ade2dc23da80a4785bf8133537b28c710abfacbfd7b5e7afc3

Observation bf2d0fb5-43c7-4713-9fbc-d07cdf0f3425 · outbound

This paper cites Z., Yang, D., and Xie, X.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation Z., Yang, D., and Xie, X

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:40:09.495937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:40:08.955088Z digest=sha256:95c323da2cfc9667bf3fd93c68a933c05e5feb3c4a2f32fa1e24f7b6f0d89f1d

Observation 8dc153de-a210-45bf-b931-341a7a45cb56 · outbound

This paper cites write newline.

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation write newline

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T19:40:08.960042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:40:08.960042Z digest=sha256:2e5cf6ce9c59906a8dd914c1494b4ca389b912e4827b221fde6fefcbb8bed5aa

Pith citing papers

Observation c4061381-fefe-4817-97eb-c3eb4dfb5b28 · inbound

CauSim: Scaling Causal Reasoning with Increasingly Complex Causal Simulators cites this paper.

CauSim: Scaling Causal Reasoning with Increasingly Complex Causal Simulators RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:31:29.417099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T02:33:04.549851Z digest=sha256:10bd0b31f4e089c4673b5a343263981f445c2924793f74b06d31792954006d72