Pith. sign in

Paper Citation Record · LEDGER

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains

As of 12 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2607.18438.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.18438 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T15:29:29.577032Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 199fe0a8-c4c4-4def-b0fc-d153adc8b48b · outbound

This paper cites ARC Prize Leaderboard.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains ARC Prize Leaderboard

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:23.070728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:23.070728Z digest=sha256:21136ce402dfef95840b2bac98e61b4bbf8883ed027be4867c3d4ee0eac9a8c0

Observation 98b12e11-c705-42d5-8d7e-4dd29783267c · outbound

This paper cites τ 2-Bench Telecom Benchmark Leaderboard.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains τ 2-Bench Telecom Benchmark Leaderboard

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:23.142006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:23.142006Z digest=sha256:c1c82631adf3f437f50c1806f5378a57d93b208d162778ed32c586e6678d2c2a

Observation 65cf5065-a338-470a-ab92-4fda93111e5c · outbound

This paper cites Humanity’s Last Exam Benchmark Leader- board.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Humanity’s Last Exam Benchmark Leader- board

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:23.285997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:23.285997Z digest=sha256:9cf3db16a4088768cae4a0a197e1902746ed51b9038e4c363891b1444b5ad7b8

Observation c5941784-cf22-451b-b93c-d91dc20f6a0d · outbound

This paper cites Introducing Claude Opus 4.7.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Introducing Claude Opus 4.7

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:23.431791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:23.431791Z digest=sha256:adebf1566aac2f6afc5399673e2c04f63de15d8f91baceb6fcbdc4d4ef7dd96b

Observation a92a0fe3-69d0-4dbc-b7df-5392c0d50925 · outbound

This paper cites Claude Mythos Preview System Card.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Claude Mythos Preview System Card

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:23.571603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:23.571603Z digest=sha256:dfafe475f537fa9dcb7a326a563173bc0d3db841c849b4379ca36c4329f2916b

Observation 46c37784-cc22-4005-b93c-d405ec4e3649 · outbound

This paper cites an unresolved cited work.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:23.682318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:23.682318Z digest=sha256:37004ba6d8cf0374ea13495399a93f33b8434dc6dcb50911380a0fe346ad39f2

Observation 25c5f1eb-6ec9-490e-8c4c-ecdc93a38316 · outbound

This paper cites Introducing GPT-5.5.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Introducing GPT-5.5

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:23.797304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:23.797304Z digest=sha256:cbea0c67e8d5e4c1b03c7e56692f0063208efbec50dc66e0f674f2fa5d9c4037

Observation 86e1502d-9276-42a8-bfd5-614bd3e637c9 · outbound

This paper cites Grok 4.3 Beta.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Grok 4.3 Beta

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:23.914038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:23.914038Z digest=sha256:1e10738a480a871135ae1fb68f8ea4934ab194cc9fc5fc4a691ebe96e23cf603

Observation eb367314-54ce-45e8-a229-cb20f3c374f6 · outbound

This paper cites Qwen3.7: The Agent Frontier.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Qwen3.7: The Agent Frontier

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:24.079864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:24.079864Z digest=sha256:41ece8e1bbe09b063a7a7eaeaea4ee8c29c25bf52c8bd96cccdbe0d77859e5f7

Observation aa2f134a-7b0c-48a9-8009-f1286765e82f · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:24.224184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:24.224184Z digest=sha256:a4f3e41689ae0b12195e41989937ab79c4a94d4a7e44d005649c6ad5bac77b58

Observation 79d822aa-4ddb-4df3-bd1e-ec4c29589558 · outbound

This paper cites WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:24.341223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:24.341223Z digest=sha256:0117365d965a30cce914b2e1dd288a11eabcf88835cfdd28546b444a1595b12d

Observation cf52d656-99a6-493f-ad55-56598de2dd88 · outbound

This paper cites AgentFrontier: Expanding the Capabil- ity Frontier of LLM Agents with ZPD-Guided Data Synthesis.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains AgentFrontier: Expanding the Capabil- ity Frontier of LLM Agents with ZPD-Guided Data Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:24.460230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:24.460230Z digest=sha256:f57a4c19c261fc8c44d23d31952112faaccea4ab81538effafa8267589d21505

Observation be277fbd-77a6-48c9-8b0c-12d907e8234a · outbound

This paper cites Measuring Agents in Production.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Measuring Agents in Production

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:24.584013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:24.584013Z digest=sha256:1b6aedb06da3876b91959a02fdd70851cada7a49b02d059fb015f81642ac8996

Observation 34aec3da-e1c8-4667-92ca-5fb377f91ae9 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:24.722570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:24.722570Z digest=sha256:31f7decdde20209c69960a69e7eaa666e1955893a6f470e8d3514add1bf0c381

Observation ca1b3420-6791-41f2-bfa6-4db8fbad9a91 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:24.846792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:24.846792Z digest=sha256:552249742b64f931cd8b1715a9e3d17dd2345ebab5f580a16e983e6f79e14ad2

Observation 164f11a7-8477-4b31-afce-87a851f1deff · outbound

This paper cites SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:24.950535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:24.950535Z digest=sha256:82fca866e8b9f4619a621afe6651a66054b8d361d97fc290ccd4544a2788802a

Observation 1afa2c31-87b0-4c15-858f-f118faad26df · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.040725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.040725Z digest=sha256:b4c56ea79db49827eb376685b63eaac42c1999dc0cdda8c9fc37c99920470fbd

Observation e6019c4f-e0c5-4d44-8dc9-dfb79436910d · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.119434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.119434Z digest=sha256:99f7d695e670cb389ca2299ddde3c0bd6b480c3421acdd2848d54fa3eaef289c

Observation ba79076e-70f6-42aa-b0c2-6211e74cc6e6 · outbound

This paper cites GPQA Diamond Benchmark Leaderboard.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains GPQA Diamond Benchmark Leaderboard

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.207610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.207610Z digest=sha256:09d15461907360b52e24b6f24b78f08e4db3a2e6ea02d1cbde5ef80729157926

Observation a682bf96-755a-461d-9ae9-0677552ab193 · outbound

This paper cites MATH-500 Benchmark Leaderboard.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains MATH-500 Benchmark Leaderboard

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.254961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.254961Z digest=sha256:fe316d09d63c2665ca6c7a9c15a9f4c5c091e7d69464570c1a40e027e8a0bb3a

Observation 10e94e7b-1cd9-4635-942e-8306e0a16a0c · outbound

This paper cites https: //artificialanalysis.ai/evaluations/aime- 2025.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains https: //artificialanalysis.ai/evaluations/aime- 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.310495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.310495Z digest=sha256:6916d3deb6702f9a64c33953e2adba01c182708f66acdba690dfbb9c6217a5e4

Observation 61df446c-4935-4572-926a-f55994545daa · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.488347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.488347Z digest=sha256:591b3187b6c1b709ff96fdfdb654703ba26215b06f13e26d65cfb27a8fe47814

Observation 0450d210-6410-4c0d-b99f-b0f43082cc93 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.602195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.602195Z digest=sha256:5f6592a063b19f32cec67c7e2a49f48a403d0e050c85c68ceaa4d7ee4df2a5f2

Observation 9a02249a-32d0-481c-aaa8-e66eae75cdca · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.746583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.746583Z digest=sha256:15ffd8f9a4b0ada68c6610ecdc163664e556be2934446d821e5f6f7fc785eb8b

Observation 3e48715e-bf87-4070-8042-f448aa32867b · outbound

This paper cites API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.830028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.830028Z digest=sha256:d1a6d44bb1e181c9453543a4f2d202584f19a435b9bbf958fa47c2526a58ff30

Observation 75bdcedc-0a18-4811-958c-30d6e7c4cf80 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.978104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.978104Z digest=sha256:badc6c669bd80d3e0d2d09b9ac8f5544598da20aaeb4e1d60c908f9de2fd0116

Observation e0e9443d-03ac-42f0-b36f-7d9d8a8b9786 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:26.141497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:26.141497Z digest=sha256:2d9734cf5d624fe4563bfc203db6817146eac7be5a75fc869145188ba6cd1c7b

Observation dbc74aea-868e-4afc-9ae7-3903b33d138c · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains GAIA: a benchmark for General AI Assistants

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:26.306719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:26.306719Z digest=sha256:49f3ef72c387063cb42b352c33f84d92a6c010173ed97f223e4efc059b9ebede

Observation 6721ea98-8a24-4d57-944e-251f92ceba4c · outbound

This paper cites APEX-Agents.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains APEX-Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:26.450130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:26.450130Z digest=sha256:704d2d127a4c7684157e4537138d3f7fd27c20bd937c82e8888bce8a56bf3d5f

Observation df722cae-32da-4264-9865-75bf17af0167 · outbound

This paper cites Are Your LLMs Capable of Stable Reasoning?.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Are Your LLMs Capable of Stable Reasoning?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:26.561138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:26.561138Z digest=sha256:aa464e2d8329c951512cc02eb99be8b73232ce3577356d9a1a8285c60a8c84e0

Observation e43b92a4-133a-4a15-8c65-8b8bb364b280 · outbound

This paper cites Project Glasswing.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Project Glasswing

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:26.693813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:26.693813Z digest=sha256:a109aab2075c9c09075f74961d74a024d75fb0ffa99b0b7fe47fa48b1ae4eeb1

Observation a0ca0af4-d1a1-45b7-a902-76095d03c743 · outbound

This paper cites Claude Mythos Preview: Anthropic’s Frontier Model Explained.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Claude Mythos Preview: Anthropic’s Frontier Model Explained

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:26.838864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:26.838864Z digest=sha256:71fdde6026cf46004d62449027a5fbaee609d2fdb1ae3498cd0bfc1cc961438d

Observation 54ef6eb1-c340-4cc9-8b34-a0f704039f99 · outbound

This paper cites Holistic Agent Leader- board: GAIA.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Holistic Agent Leader- board: GAIA

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:26.968437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:26.968437Z digest=sha256:196806a9915796ca9420b6527b37310665af6673060363983b5d81589d0e5d8e

Observation c8223312-68e0-4dc1-b75a-1ce7a7689831 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.108781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.108781Z digest=sha256:5d50f21e0aa7e2eab6a908bbde5e0696b6c8630472bf2a5ebc0c268d9c2d0483

Observation dd0ef2fa-538b-4ceb-ab22-42df4ed8a441 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.236946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.236946Z digest=sha256:50d6a25824324089ba51ba2bd2dcb1f801ef38f5d84326fb71527ab4322f8492

Observation 832cdfff-61ee-400e-9c61-dcf0b491a644 · outbound

This paper cites Investigating Data Contamination in Modern Benchmarks for Large Language Models.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.374252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.374252Z digest=sha256:57b549704183ffe955b91a6908d680b121003ba23639b9711be79d021161d579

Observation f2cbc7d4-c6f3-4dc7-9736-64b0026e5fb3 · outbound

This paper cites Anthropic API documentation.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Anthropic API documentation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.505166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.505166Z digest=sha256:46c8ebd70d9ec62bb99f63b9538cbb799a0a07ca1bdf6803628db78607c0a78b

Observation bb4d9cfe-0542-4d59-81fc-e38bc39c98dd · outbound

This paper cites Gemini API documentation.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Gemini API documentation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.579197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.579197Z digest=sha256:7debeec1d9fee38954b32b617779c9fe2e4858a29dcc7f9bd4b3044484dcfa69

Observation 2f09ac12-5d15-46c7-9323-d81a42746256 · outbound

This paper cites OpenAI API reference.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains OpenAI API reference

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.635227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.635227Z digest=sha256:57f499a50e0110806e7f4d52f6ddf4b15969844b296a736ab0a5b1120db79f0f

Observation 9015a179-9740-4a4e-9b9e-e292861079af · outbound

This paper cites Adaptive thinking.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Adaptive thinking

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.700320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.700320Z digest=sha256:e0ef5948ff9b119d6bd8ea60faf7d9a7e81e4ef70230736baf377014f274b888

Observation ff719fe2-494c-4888-88d9-75f23ea9945f · outbound

This paper cites Gemini thinking.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Gemini thinking

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.703737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.703737Z digest=sha256:3997c329c296bd5057ff8c66be1575601b88c184d913a1558a65e4a94ac693b7

Observation b9c86265-215c-42b4-98af-cade56917493 · outbound

This paper cites Reasoning models.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Reasoning models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.729612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.729612Z digest=sha256:6860b6959b6db696dd5d128df4aa75cddb9b1dfa90c399e7c13d161ba747f118

Observation da1afee8-2254-41c8-a250-dc17ace86eb9 · outbound

This paper cites AA-Omniscience: Evaluating Cross- Domain Knowledge Reliability in Large Language Models.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains AA-Omniscience: Evaluating Cross- Domain Knowledge Reliability in Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.872812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.872812Z digest=sha256:71c79efef56361a97ddc96c5c3a724d76765b8213b7b17593a72f85abbdaf085

Observation 686673b7-628e-455a-817a-91a3608e4013 · outbound

This paper cites Context Length Alone Hurts LLM Performance Despite Perfect Retrieval.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Context Length Alone Hurts LLM Performance Despite Perfect Retrieval

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:28.016995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:28.016995Z digest=sha256:d39e1776cc0c01eec6c9954a418fffe49cf3501324d8befcc43f3de284672d6d

Observation 924eb265-a4e8-4033-8bd0-665f4c8775e2 · outbound

This paper cites AA-Omniscience: Knowledge and Halluci- nation Benchmark.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains AA-Omniscience: Knowledge and Halluci- nation Benchmark

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:28.123820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:28.123820Z digest=sha256:17dd7c6f5a3b0c111d95cc914e80b8d64d27ca5810554917c3b9e3c74e329758

Observation d31f90aa-ea5f-4814-b6eb-2506aa063eb5 · outbound

This paper cites Why Language Models Hallucinate.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Why Language Models Hallucinate

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:28.235129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:28.235129Z digest=sha256:1efcaea71b530f11ee15e1c8c74a1d96770a3d4a2b99310d0fd534659f718361

Observation fcd7a8fc-2e9e-4786-b250-2b065eb18798 · outbound

This paper cites Claude Fable 5 and Claude Mythos 5.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Claude Fable 5 and Claude Mythos 5

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:28.335808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:28.335808Z digest=sha256:e0f86a044f73b3c3511fc8ce9bc5550ecdef7ea3d762b5397dbfe222bea74389

Observation 80a38cb3-1ead-4f0f-9e68-19d31dca1107 · outbound

This paper cites Artificial Analysis Intelligence Index.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Artificial Analysis Intelligence Index

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:28.505901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:28.505901Z digest=sha256:6b34b69f22cec7e1104c49a2541594831332999b88610057584121860a2fc760

Observation bf7420a9-0ec0-4c21-92ab-0a5a112220d2 · outbound

This paper cites an unresolved cited work.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:29.037989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:29.037989Z digest=sha256:558be8822660efd6c2f2b648669c646dbeaac4fb35622a1c3f88ecdc0060e79b

Observation 744271cb-af13-4ff9-bc2c-d1eb992b9f24 · outbound

This paper cites Artificial Analysis Intelligence Index: Methodology.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Artificial Analysis Intelligence Index: Methodology

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:29.254414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:29.254414Z digest=sha256:d8521d215586a00f64b0478b44f3ba69ae36d31f874353a6fe08eda416dce79d

Observation b76a908b-d87a-4bce-a245-d025cda6e789 · outbound

This paper cites AssetOpsBench: Stirrup Agent.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains AssetOpsBench: Stirrup Agent

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:29.460428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:29.460428Z digest=sha256:f48bd9e6a67a0e69d47451d7fc7ff41070d9455701d8f2bd3a434392d8253f31

Observation a169a86c-7b9a-4bc1-8c97-bfb852aaa9e5 · outbound

This paper cites Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:29.577032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:29.577032Z digest=sha256:f108d08fbe262572c2cefacbc4ac16a41456365508850c94ac4c05b5b1667df5

Observation 1d788fbd-d216-4eab-9658-8c918d5a007b · outbound

This paper cites Humanity's Last Exam.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Humanity's Last Exam

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:28.838974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:28.838974Z digest=sha256:9055c7b02340b315d9a190884ab1afa135ca1907cb70959ce83779f36f70f41a

Pith citing papers

No inbound Pith citation observations are available.