Pith. sign in

Paper Citation Record · LEDGER

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark

As of 11 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2608.04670.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04670 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:21:45.118956Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa7468e2-a0f4-4914-bcae-ec7edfd4b7dd · outbound

This paper cites Lewkowycz, A.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Lewkowycz, A

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.849252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T19:21:44.956027Z digest=sha256:3c242d254082c0d7a2d84adb17c35ac60613abbfbee8da65da4df10ee43aeb06

Observation d1fe282e-2c47-455a-9284-f52d84fc41a1 · outbound

This paper cites Chang, X.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Chang, X

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.835347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T19:21:44.960843Z digest=sha256:5f9b4fa376078cf64a422a8c0d3228ad762eb428226364c26590094e39b40b4d

Observation e731ee4e-e7ee-40ad-a049-d58207ca6228 · outbound

This paper cites an unresolved cited work.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:21:45.822239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T19:21:44.965029Z digest=sha256:2ded0bf88b89f4f00cae45f36634fed6703898b43a0c081bdaeebfd105136086

Observation 3d39048b-ef78-4893-b00a-e0898b2c173c · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:44.969886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:44.969886Z digest=sha256:14a1619deba3807b1dab2927d030c078718816ab8a710305aa9b4175fb76fc6a

Observation f2726dda-cc6f-41a1-acf9-d8277c2e3b49 · outbound

This paper cites The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:44.974685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:44.974685Z digest=sha256:aa6cbe8a15c275fde77b5f9f7ab569af39d2203c11626aa95945dd44c907d7f7

Observation 32961ddf-e62a-4474-9931-7139f5ae7dc2 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:44.979311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:44.979311Z digest=sha256:c34a5a39a6d398385c55e466ea1efcc1958f544efbdaa283b6d1d950641f9e57

Observation d1664dfb-7a8b-43ae-9d8f-25dcb1d79db5 · outbound

This paper cites an unresolved cited work.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:21:45.808707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T19:21:44.984335Z digest=sha256:4d4499a8e9ccff6c9a3d7b4bff173fb9d4a849383bdaa627fbde46cc5e1e2853

Observation e4a11fcf-645a-4aaa-98c9-6ec210916ee1 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:44.988297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:44.988297Z digest=sha256:f4722274aed0c51ac0b7055968fa1ca3856050231144fb9676fac8c12063cb61

Observation 409df814-db12-4d26-87d8-298b6d9b4a98 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:44.992611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:44.992611Z digest=sha256:7a2d5120116d3e0b54b855a283cae09b332b3d08bbdd01043a31ae5c77a4c9b6

Observation eb55a51e-25bd-48a6-a8a3-2e1a23e3fbd3 · outbound

This paper cites an unresolved cited work.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:21:45.794336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T19:21:44.996759Z digest=sha256:a10dc1fb71ff45cf62fbc8752685e1e22fc24d64fbc3f1c99178867e51a6168f

Observation 44e5fd0c-1344-4e75-a719-9bb3ce192717 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.000585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.000585Z digest=sha256:9ea714ec5c946f3052b5155aafeb018241bb8de1444a097bdbc4d085f42f2b4f

Observation adf15a65-4f2c-4d03-a251-96731052994a · outbound

This paper cites On the Measure of Intelligence.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark On the Measure of Intelligence

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.005089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.005089Z digest=sha256:b5461c072d202dd34390004dc0f81080d43d8640875cbf364ea83e153751000f

Observation 5841a2f2-84ac-4dcb-b08f-127a485e314a · outbound

This paper cites ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.009340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.009340Z digest=sha256:a8a8fcee2cabce04f6f18fca14c257450d6f758041b5b66f9006174d2a1218f5

Observation 5ca2b51a-5480-4db7-9ac3-05a694cf0493 · outbound

This paper cites Attanasio, P.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Attanasio, P

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.779829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T19:21:45.013470Z digest=sha256:e75a89b6d035f6eb4140b4a908e8c364bfe8276933f478137b392d0cc9c7ba83

Observation 50242c54-3c37-4d0a-8311-5e6ccd0e23b5 · outbound

This paper cites Evalita-LLM: Benchmarking Large Language Models on Italian.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Evalita-LLM: Benchmarking Large Language Models on Italian

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.017840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.017840Z digest=sha256:8b6ea24ff8bc5966f48729bf3548c4529185134037d971d69ff837b31dfbabe9

Observation 3becf2ce-aee1-4e6f-97dd-7a5fdfa9ea04 · outbound

This paper cites an unresolved cited work.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work

Reference 16

Resolution
malformed identifier
doi_truncated, observed 2026-08-06T19:21:45.166767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T19:21:45.022090Z digest=sha256:c88dfb80ec2642541bd6e86614dbdc2742c2cd876aed80f0c1f35a11949355ed

Observation 0b1b81e7-71ef-4cc2-8dc1-049d87926a49 · outbound

This paper cites Tedeschi, F.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Tedeschi, F

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.766098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T19:21:45.026171Z digest=sha256:047067a04f800d1834f6011d325cd282653058fae5e7c631a0250a3f5b31eda1

Observation 3fa49c86-3971-4e5c-bc3c-ff461ea6fd2c · outbound

This paper cites Khoshtab, D.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Khoshtab, D

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.751518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T19:21:45.029963Z digest=sha256:de9d9ee66c5d4d607a2d28a446f4f00dda8824df9280d2c68ecce609e67e7ba9

Observation b2676a61-15f0-4206-a9e0-216573980de9 · outbound

This paper cites an unresolved cited work.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.033785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.033785Z digest=sha256:40c2a132730363cab6c96a4d2a88a8b4425c666ca3810e92af9930c228092d2d

Observation 0aae267c-af0e-4576-99bd-5c6f3ded4d1f · outbound

This paper cites an unresolved cited work.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:21:45.736703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T19:21:45.038028Z digest=sha256:8dc719503669279a9ce7ac042a62db9a6bb1b5be29ffaac3366d603ff1df4a72

Observation b6c82470-0b85-460f-bb34-6dd550e20627 · outbound

This paper cites an unresolved cited work.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:21:45.721199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T19:21:45.041895Z digest=sha256:a78b171ef4c7acdfbd26de7d316bff64dd2cb80664040fee98243a83409a6a71

Observation 7dfba636-070c-4208-a330-61e7b51bf8c4 · outbound

This paper cites Donthi, M.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Donthi, M

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.706277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T19:21:45.046086Z digest=sha256:c093cec3df9d21ff97bec6902aff9274227bc3ec7be4964df3d419cb4e1faa68

Observation d7760f35-722f-4c16-9172-439e74990dff · outbound

This paper cites an unresolved cited work.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:21:45.691013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T19:21:45.050259Z digest=sha256:d4de7f858bd36c5c9fc9afaf54d0e67973915b64af891eaf042b153c8a84efdb

Observation 6daa9457-2d54-433c-bd73-1eb68edea5b5 · outbound

This paper cites Proverbs Run in Pairs: Evaluating Proverb Translation Capability of Large Language Model.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Proverbs Run in Pairs: Evaluating Proverb Translation Capability of Large Language Model

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:21:45.389273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T19:21:45.054075Z digest=sha256:0225051ad0cabc20e0d75384455def676db028030447ab5dff6a86d92f9eb09d

Observation 09ce6256-3f37-4da6-9d04-c4342578370a · outbound

This paper cites Caramagna, I 200 proverbi italiani più belli e famosi (con significato), 2025.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Caramagna, I 200 proverbi italiani più belli e famosi (con significato), 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.674684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T19:21:45.058791Z digest=sha256:65eb7cf5dd31b5d9eeb6f1f3c0d0c345f6de200c1469967321d18587572dd5ac

Observation 38eded67-ba3d-48fd-939b-b3ca4da3b4a5 · outbound

This paper cites URL: https://openrouter .ai/, accessed: 2025-06-15.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark URL: https://openrouter .ai/, accessed: 2025-06-15

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.658978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T19:21:45.062942Z digest=sha256:e7f637713a62cb0f16896bdb3d3971f46f86e787f5750ea77da09c6c18b9a3e6

Observation e487e10a-0c4b-457c-b815-695b8298c026 · outbound

This paper cites GPT-4o System Card.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark GPT-4o System Card

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.066685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.066685Z digest=sha256:1e71998547a1c302fd981cbb7e17b4972d30b7cf859176d12b080f760a225701

Observation e269fac7-1e18-4536-b981-0c539fa1a946 · outbound

This paper cites DeepSeek-V3 Technical Report.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark DeepSeek-V3 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.071212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.071212Z digest=sha256:b7b02c5eae4ec2ea2194b4bb0033c8d0b9032acf54fc7af5dfba1e742d06fbb2

Observation 5bfe8cad-5a95-4ff0-b287-f69a764c903f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.075878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.075878Z digest=sha256:57471db5866f734dc2a67f179e36e7750aed2541442cbe9946cd4a2a3aa07e7d

Observation ecb7b7ea-3c87-443d-beff-9246f08d6f9d · outbound

This paper cites Qwen3 Technical Report.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Qwen3 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.080370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.080370Z digest=sha256:b79c3044a6bfe3322da6bcc453a7e170e2e6f0e8da4ded9f476847e8b3dfb2c3

Observation 5a462dfd-05cb-47d9-9ff0-71c05c9a5435 · outbound

This paper cites Gemma 3 Technical Report.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Gemma 3 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.084774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.084774Z digest=sha256:a906bf1b9d4c68d4f45d4b7a9fd83a8ca5e19649f871b6eda593e85358a80bde

Observation af5a729b-cfb3-4a72-876e-fd28f7d1accb · outbound

This paper cites Orlando, L.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Orlando, L

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.644039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T19:21:45.089087Z digest=sha256:f4bc0c6fbd606fb686d6dca7707d928b257bcfe053e8e28523fe5fff2c9ce330

Observation 60171fa0-2dd6-4f60-9313-81f3b1345f32 · outbound

This paper cites URL: https://huggingface .co/ mistralai/Mistral-Small-3.1-24B-Instruct-2503.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark URL: https://huggingface .co/ mistralai/Mistral-Small-3.1-24B-Instruct-2503

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.629134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T19:21:45.092781Z digest=sha256:4895b54190402dab599e26a3785e8433cd256f84e35eaee7fc660d2c533a4576

Observation d3a5a1f8-39b5-4461-92ce-32202342d426 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.096571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.096571Z digest=sha256:61fe7e63f5b7595b08e1f7e578cdc98e5b91f886d6dd8a02751d67ae7fd42e67

Observation 005ae45e-b53a-4848-bbfa-7a72dad54bc4 · outbound

This paper cites Etxaniz, G.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Etxaniz, G

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.100736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.100736Z digest=sha256:e0a75b48840b84479bed29606019d3731d8e99f9ec5f0fb968804562c35fe106

Observation 69c0ce65-5727-4e35-b690-39617bbbd0ea · outbound

This paper cites Ranaldi, G.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Ranaldi, G

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:21:45.614492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T19:21:45.105191Z digest=sha256:a95d4dfd51a2cb145914953fd2bd610271f9aa301defb7a72de48c9fb3c19da2

Observation f768a60d-86bc-491b-a4d5-762fe23e970d · outbound

This paper cites an unresolved cited work.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.109647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.109647Z digest=sha256:2cec6ecf92b02d2e8b1761cc6521c60be0868cc097162100a92d44830fe57f24

Observation 0ebb7780-776f-4606-a588-55645408e891 · outbound

This paper cites Reasoning Models Don't Always Say What They Think.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Reasoning Models Don't Always Say What They Think

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.114206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.114206Z digest=sha256:788fbc114ebb2c82f1e5085121f204ac0dc8eaa015704943610bfe7cca1059ca

Observation 0d10fc8b-a605-42a7-8afb-100a10f1f11c · outbound

This paper cites "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models.

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.118956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.118956Z digest=sha256:fd466ee116c09cc9684c29215176cee22dd1806fb3f9c826cd70ec7522f110b8

Pith citing papers

No inbound Pith citation observations are available.