Pith. sign in

Paper Citation Record · LEDGER

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges

As of 8 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 6 inbound Pith citation observations for arXiv:2502.08859.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08859 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:30:43.662827Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:26:56.639836Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T08:15:33.926301Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 82346c5b-96d1-42e8-81cb-8cd0681ce094 · outbound

This paper cites https://puzzledpint.org/.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://puzzledpint.org/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:44.156381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.527786Z digest=sha256:618ffea45bc1b7cccc85b9955f35707d6f35bf3cf22b333d294e0aaf6698e98a

Observation 843cd5f7-0e66-41bd-a6cc-633ae07e1a93 · outbound

This paper cites an unresolved cited work.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:30:44.145614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.531958Z digest=sha256:2c94de7e3033e3558218b701b35011e63fa00e2bad7074d8ad805c46aaf54f20

Observation 5c814eac-f57c-4476-a4be-ec1b1dcaaf93 · outbound

This paper cites Puzzle Potluck.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Puzzle Potluck

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:44.133987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.535582Z digest=sha256:3349fcd3f6642ddfadfe7f84d78f147c2d90ab126b508f8981e929d21ee68a25

Observation 07555f79-d9b5-4272-9dae-060dc3c4e83d · outbound

This paper cites SOME PUZZLES by Mark Halpin.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges SOME PUZZLES by Mark Halpin

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:44.123189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.539340Z digest=sha256:3afcb19050ab1f138ec985551b7660a09e906e9b72e275f59f8f6db4c68e7c40

Observation 5126262f-5fa0-42a8-875d-6c8ad29cd47c · outbound

This paper cites an unresolved cited work.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:30:44.112500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.542911Z digest=sha256:563ce477286d762c5bcf1e977efb813b8e01059a2030d893d5806ceeb0bc442b

Observation 8ae17add-ffd2-4fa9-b4b2-a6c2515e9957 · outbound

This paper cites https://puzzles.mit.edu/.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://puzzles.mit.edu/

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:44.100847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.546828Z digest=sha256:d7d6ab29abe70409caf41c80ec5b072775e33821470832fff487292c064fa6fb

Observation b2dfcc65-8744-4ff9-81e1-10de40e61f08 · outbound

This paper cites Grandmaster Puzzles.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Grandmaster Puzzles

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:44.090026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.550584Z digest=sha256:0fec223fc1babef3a65e04db8c6eeb5552f409c74b61dbf3d085054f7612a6dc

Observation 7cd582dc-3fd8-469c-a10f-7f7fbcb531af · outbound

This paper cites Measuring mathematical problem solving with the math dataset, 2021.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Measuring mathematical problem solving with the math dataset, 2021

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.554415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.554415Z digest=sha256:8b7f6e0f388231495fbaf8986830b306aab76722a6348e1acb7aff6ec247c5be

Observation 391adf69-e187-44ba-acb9-c27e1c67651b · outbound

This paper cites an unresolved cited work.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.557938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.557938Z digest=sha256:e841fc95d511692e0ea03384ce8afd63025dfb90546fd8b737bdbcecb82f2905

Observation 9d881114-a65c-439d-850e-23a6e4529837 · outbound

This paper cites Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai, 2024.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:44.065829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.561371Z digest=sha256:66b59f214dd2ada7f75e1e9e6825a7445342c3d847f02b350abad523951adaf2

Observation 0cd4cd2d-5df6-4280-bfbb-f1383fef72a4 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.564791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.564791Z digest=sha256:3a6978b94c4446101a82997cb69bde78af5ef619be4ea7eb91d08b6203900c1e

Observation 7d64a4a5-0732-4d6c-bbf7-a631a43aa849 · outbound

This paper cites Humanity’s Last Exam, 2025.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Humanity’s Last Exam, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:44.046677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.568159Z digest=sha256:69457466b94b070c99a9363d9cad22daf3bb932d7ddf58e6a95345b731a4ef8a

Observation b177cc5e-d96f-4155-9f66-ae21f995b0b1 · outbound

This paper cites Measuring massive multitask language understanding, 2021.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Measuring massive multitask language understanding, 2021

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.571998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.571998Z digest=sha256:f71c9ccef1d5fd590ced1de3a72ca04d14bd753f8357effe4201ee58b6bffd62

Observation 8f5d7d3e-4517-4098-8eb6-4c80358f5981 · outbound

This paper cites Mmmu: A massive multi- discipline multimodal understanding and reasoning benchmark for expert agi.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Mmmu: A massive multi- discipline multimodal understanding and reasoning benchmark for expert agi

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:44.030082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.575607Z digest=sha256:ad7fb528f8fb6abc9c8b4010c344155535cfdd5588c6d3c1f7603c09719456ad

Observation 65f7a8ab-ff2c-48cc-a61b-b9bc36381e13 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.578906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.578906Z digest=sha256:8563f38f14869a53de0266de186398f9eebc106809a28b421413fd474b51ad96

Observation 460bbc65-3f94-4551-9fc9-f93c875e9445 · outbound

This paper cites Vista: A rubric- based visual task assessment.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Vista: A rubric- based visual task assessment

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:44.013743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.582228Z digest=sha256:3595cce22ce708b72fe86420d83376bdc49b8f42e5d95c816a0024390cd98d59

Observation cfc927e1-817f-462c-a866-d9bbabc01ebe · outbound

This paper cites On the measure of intelligence, 2019.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges On the measure of intelligence, 2019

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.585618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.585618Z digest=sha256:3bf0a6a58a5dbb5e9ac77b4804193fa19b43a0c78d0c3b6cf9a58943a38dfd89

Observation 8c0601a4-d84c-47d8-905a-eef165df06a6 · outbound

This paper cites Lanzendörfer, Yannick Niedermayr, and Roger Wattenhofer.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Lanzendörfer, Yannick Niedermayr, and Roger Wattenhofer

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.997755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.588907Z digest=sha256:c3a21337db187e4b2a0868bed197296d406e4e39c789df03e3fdfc9eb9bc48d5

Observation 5ac52958-4155-4a5f-98e0-eb7b092d71c4 · outbound

This paper cites PuzzlePlex: A Benchmark to Evaluate the Reasoning and Planning of Large Language Models on Puzzles, 2025.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges PuzzlePlex: A Benchmark to Evaluate the Reasoning and Planning of Large Language Models on Puzzles, 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.987521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.592101Z digest=sha256:9f2b56d2a4cc4f4061c8e407f841cf8283cf5a4cb508e52c9c39f866a04ad031

Observation d18500c4-f372-48a4-a7e0-2d5e38939151 · outbound

This paper cites FCoReBench: Can Large Language Models Solve Challenging First-Order Combinatorial Reasoning Problems?.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges FCoReBench: Can Large Language Models Solve Challenging First-Order Combinatorial Reasoning Problems?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.595386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.595386Z digest=sha256:7043c1e5b871c1c6490b0c8059b45e77f08b7dd9114926ae9545362baf8de5d0

Observation 46e6cb43-c7dd-4a51-8eb6-9c7d8dccd28e · outbound

This paper cites Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.599486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.599486Z digest=sha256:e821059381265729bf9711aefae513fe53900a73557a932a581911fde054a320

Observation d7a2bf39-b9b4-4d99-a75d-af9f8f300053 · outbound

This paper cites RiddleSense: Reasoning about Riddle Questions Featuring Linguistic Creativity and Commonsense Knowledge.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges RiddleSense: Reasoning about Riddle Questions Featuring Linguistic Creativity and Commonsense Knowledge

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.603305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.603305Z digest=sha256:84541fa704a4616eeb5221b6b5eb7dd6839544fb661002586b944157cc56e00b

Observation e6bf3606-0828-4f48-9380-73d8c2b2cc1d · outbound

This paper cites https://www.melbunimathsstats.org/puzzlehunt.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://www.melbunimathsstats.org/puzzlehunt

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.977237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.607095Z digest=sha256:b85a9bfdc3f60608e11cd94243d91b11743435eacde82cf5fc0e25ed258d4d56

Observation b1cee88c-c0d9-446e-b0f4-45668caf367e · outbound

This paper cites https://web.archive.org/web/20210725192741/https://www.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://web.archive.org/web/20210725192741/https://www

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T23:30:43.610435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:30:43.610435Z digest=sha256:16feac93b4ecdb2f8919b2ffe43a08a8f63535ee834778a6f83981043f0e4948

Observation 342ba2ad-cb5f-49c1-ae7e-e7f7828dfbdd · outbound

This paper cites https://harvardpuzzles.github.io/.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://harvardpuzzles.github.io/

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.967490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.613709Z digest=sha256:9f4d04196c042ee6e4c507191d18b23455a80a05ac7a6bf1bee6c9eae82e6fe7

Observation c7fdeb96-070a-4dc1-aaeb-f1ba1088d752 · outbound

This paper cites https://www.mezzacotta.net/puzzle/cisra/.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://www.mezzacotta.net/puzzle/cisra/

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.957817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.616875Z digest=sha256:48ddba281aebd771cc64ebf3f7823534910fec71ea355f2b08a45c35bda4e202

Observation 503937d6-4c14-453e-ac0c-00bc546d01df · outbound

This paper cites https://www.janestreet.com/puzzles/archive/index.html.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://www.janestreet.com/puzzles/archive/index.html

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.947801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.620473Z digest=sha256:3a080e282e664c67b41847d83e2a1487d3c6cf2e2a85fa6af8fab92f5321c533

Observation 3a1b91e4-717d-4d1e-af44-e612c8e06b68 · outbound

This paper cites https://gooooogol.theburninators.org/puzzles/.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://gooooogol.theburninators.org/puzzles/

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.938099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.624023Z digest=sha256:6d3f050777cea599dab4dcad09b8156f16689bad3b23a32a9fc39c1cd53a5bee

Observation a1c18164-aca6-46c1-821a-3a837d60b995 · outbound

This paper cites https://playdash.org/.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://playdash.org/

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.928572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.627592Z digest=sha256:688fcb3c978753a0066a7ea73fc28ca37a1b49af86bbdd6f5b9369a2c6d7ca8c

Observation 1776b629-139b-4d48-a687-d23cda9ee5a3 · outbound

This paper cites https://www.baphl.org/.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges https://www.baphl.org/

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.918925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.631361Z digest=sha256:49e44707dbe5e0bbfc44544b36f5d0ce8c96dc6eb4499b03ce748e8630da5e67

Observation f509d110-5f11-4786-99f7-72c04cdb207d · outbound

This paper cites Forbidden actions.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Forbidden actions

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.909101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.634785Z digest=sha256:5ba0ec028955e8045baf0e7b668a069eea83ca1ec38eb3ceb8e521c08c6058ac

Observation 49d8ad0e-f128-40dc-9fbe-5178ccecdb90 · outbound

This paper cites an unresolved cited work.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:30:43.899181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.638190Z digest=sha256:0a2c880869b176952dbabb514f9a01d9ca1a4cf22da9a6c311a5e280183938af

Observation de4a874d-eb72-42cd-87eb-a1f8537c0f43 · outbound

This paper cites an unresolved cited work.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:30:43.889272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.641496Z digest=sha256:c66d613a0c5423e328a1d986734222a0ebecf6d787a20db18ca36da9443d3095

Observation 79fa53c8-27b7-4cb2-8e7c-6eb38e6962f6 · outbound

This paper cites an unresolved cited work.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:30:43.878286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.645200Z digest=sha256:83121ff6a0c0a360fa739aaf0abc43d2b5e0b5761d07a45e3815075daf5a83aa

Observation 708a5c9a-11ba-41b1-9db5-99094dc2a39c · outbound

This paper cites an unresolved cited work.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:30:43.867120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.648403Z digest=sha256:11554f4fd098ea796ed279307615f0ebbeb727d5c1c37147541381425cc7b2de

Observation 25a673f7-2449-4694-8b86-389c4c22039d · outbound

This paper cites Problem Web Wage.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Problem Web Wage

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.856528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.651991Z digest=sha256:66a09c8fe3a67625068b01f948512d5e5e9536d27aaf3298bb97df3a6f19830d

Observation ea411947-d5fb-4d7c-8130-9d9b09a33023 · outbound

This paper cites an unresolved cited work.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:30:43.846067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.655828Z digest=sha256:2b8c1caec0f3563aaa4737b6e84229cd0824c8927032eb767584d5c44580e80a

Observation cd0e5dc9-42fa-4b02-8123-e1e335477d5a · outbound

This paper cites an unresolved cited work.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:30:43.835897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.659550Z digest=sha256:66b9d866a6342bcb78c2a51f18a2f37ff2542003b46712876c1dd3fef926b57a

Observation d052f01f-dada-4f97-b0b3-42046345fa91 · outbound

This paper cites This structured approach to answer formats allows us to extract answers consistently and reduces ambiguity when comparing model outputs to ground-truth solutions.

EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges This structured approach to answer formats allows us to extract answers consistently and reduces ambiguity when comparing model outputs to ground-truth solutions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:30:43.824654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:30:43.662827Z digest=sha256:33516c6458a40e632ff6d318cd28cd301e42c957d212ce684e3c1cb24e4eb406

Pith citing papers

Observation 94a3833b-a525-47d0-b967-5d248b0f689d · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:57:38.309833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:d815f500efa7b65d7ffc697a60694e888e0494a3f5281699097476eef955cfc3

Observation 57530e12-cda5-41fe-bf7f-6574601f213e · inbound

lmgame-Bench: How Good are LLMs at Playing Games? cites this paper.

lmgame-Bench: How Good are LLMs at Playing Games? EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.639836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.639836Z digest=sha256:4ab189391cd68da460f7750d3ec0e38c37c95531fc37f69cd91e8afdaef42cff

Observation e80c223a-4010-4c5b-ac75-717ba7a021e7 · inbound

Sudoku-Bench: Evaluating creative reasoning with Sudoku variants cites this paper.

Sudoku-Bench: Evaluating creative reasoning with Sudoku variants EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:08:30.334466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:08:30.334466Z digest=sha256:86cdc85dcff67be41bb07002bda24280932a2da1498305c5b64ddf1e746778d5

Observation 3b4b129e-89df-44b5-ae33-826b8ebefe91 · inbound

PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts cites this paper.

PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:47:15.164738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T10:43:02.601014Z digest=sha256:3a8fc06706076c81e6ea73b75bdf3ad7cea046ad6ff5952e6aed3ea8bfd561fb

Observation 048dd1b5-73d0-4c8c-aeca-b028b91d41fe · inbound

FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation cites this paper.

FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:15:33.930689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T08:13:05.328746Z digest=sha256:8a277e4ad333794dcc2e26e7d643d36eb633915361e1ba1de42755febddecbd8

Observation 8ee27dfd-b016-448f-9762-eac45709a9c7 · inbound

Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning cites this paper.

Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:39:47.813355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T05:35:32.806871Z digest=sha256:cc27027a38488bfc177c538b2ac04f7b1ee569fa8971c57307e3919d655160e8