Pith. sign in

Paper Citation Record · LEDGER

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

As of 9 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 8 inbound Pith citation observations for arXiv:2506.10764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10764 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:23:50.683069Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T14:04:57.314444Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:20:07.360446Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5b6495d8-7106-4e4e-9e68-c534f22f34d1 · outbound

This paper cites GPT-4 Technical Report.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.543505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.543505Z digest=sha256:417b0d8700b65a9fe2ba3c8d184088985f1138ed8efd62b82bdfb01918e5f50d

Observation 692304e6-190e-4b51-affa-5494d3e16856 · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.547360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.547360Z digest=sha256:e68e763971c8005f44e732883dbc6fcc96c52ce1a65e2817c8f6cf78af857c57

Observation 2e969ac0-cb4b-40eb-82a3-f510cb7984d6 · outbound

This paper cites Program Synthesis with Large Language Models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Program Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.551498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.551498Z digest=sha256:7fa53d34d0624da74868704693a2a33981422926e530b28b678d344da77e4ca8

Observation 39138021-4c0d-44eb-af22-74039c8e02bb · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.554887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.554887Z digest=sha256:2d8e1e6fb6c919dcf8085d3eec4f19a26961ec593a99e3abf2b3d9c1639a4550

Observation 895e5a4d-eca7-4c88-ba1b-70ad41bb08db · outbound

This paper cites Evaluating Large Language Models Trained on Code.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.561890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.561890Z digest=sha256:575b56b1568960aee2f13c9b83f43aa9329122bd75f097023fada551fb701db2

Observation d36b7c62-4ba5-4106-b181-07d11b280735 · outbound

This paper cites Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1– 113, 2023.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1– 113, 2023

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.565288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.565288Z digest=sha256:a4b8e49a094f3511fb4300cce3f2d5d7809b6f8ae108eb4f77315ed5d06e239d

Observation 4bbc9d8e-a657-4e40-8e9a-35c1b8918355 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.568987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.568987Z digest=sha256:8258f7348a77a432e9bcae25b8cacd8e0d093f956fa39774c31b68464b4a70fd

Observation ac041da5-8e76-4147-b5dc-be0ffe382590 · outbound

This paper cites Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.571999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.571999Z digest=sha256:6f41391fb543dec6450344ab4899a94cd965fd21f3aafeb5572d1931615a3399

Observation 54c5a251-52c7-496e-97db-ec3ed7ebaf2d · outbound

This paper cites NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.574907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.574907Z digest=sha256:3a6275646eb264f76ec481ae8735cc391c8323a449ae68d4819b2e0e2ec346cf

Observation a424da53-d40a-46f5-af3b-ef1a0875be4c · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.577993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.577993Z digest=sha256:550fd7c7ea85dc2f176270ad6f08fcbcd1ac0dd15f8aec0f5723253e31865f89

Observation bc03159f-8cdf-4858-901d-19719e22bfa7 · outbound

This paper cites The Llama 3 Herd of Models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.580934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.580934Z digest=sha256:358ed4324c93acd1260b1d59f1f4d01cc1f175e0e9c4ba64ca7f795d20584242

Observation 7cbf9a37-12b5-4b76-aecd-cc4a7a198749 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.584251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.584251Z digest=sha256:072e5cbfd7d72e167f3093d6bdc642a283cf8448322261e7df11c414114fd17f

Observation 4380e4f1-ea7e-40b0-b55e-c85b8b215dc0 · outbound

This paper cites R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.587173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.587173Z digest=sha256:a2434b07deec3ceeab4af04a3bad7029ed2dc8ff3efdf669f39df36367af58e2

Observation 94684895-c743-40d3-87f1-6e6be6a6bbe9 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Measuring Massive Multitask Language Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.590207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.590207Z digest=sha256:a56e645dbdbde31e408443ffcceeadaad05667e98fbbe3e9b4e7a3db988b574a

Observation 5b1a572f-7661-4758-a832-1cfbdf5d1d36 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.593077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.593077Z digest=sha256:38b97ec8259ef254e5b3a917c261001e544dd86abf6bfe6ddd81b5a09c2529a7

Observation 59c6256c-9287-43d1-96eb-183abc89676b · outbound

This paper cites Mlagentbench: Evaluating language agents on machine learning experimentation.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Mlagentbench: Evaluating language agents on machine learning experimentation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.596213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.596213Z digest=sha256:afadeb8640d04d5c69a9a7d4122e87741920ad78b51b1b2bddc6a367e1b74835

Observation 141b4162-76ee-4f17-8f17-b187f4a13aa0 · outbound

This paper cites Aide: Ai-driven exploration in the space of code, 2025.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Aide: Ai-driven exploration in the space of code, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.599339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.599339Z digest=sha256:a8a87218c54f42265f40bff6eb85865c71f0994ed083495de1bfae2ce80ecffe

Observation 316f20a2-a8c8-4f4d-923d-612cab304516 · outbound

This paper cites Let’s verify step by step.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Let’s verify step by step

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.602556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.602556Z digest=sha256:934937917dcf1652c5f057ec6176c86287f0e0e5d6ad6c2e91d8365a2ca44435

Observation 4ae518c8-e72e-4df0-8705-bd7480778af1 · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Truthfulqa: Measuring how models mimic human falsehoods

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.237836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:23:50.605362Z digest=sha256:d4c4472ca6c7f035275b759935e9ef46a5b95e4549ba11736893bbdc59e243a7

Observation c22a6d6e-7175-4207-896c-7be959e993f2 · outbound

This paper cites Criticbench: Benchmarking llms for critique-correct reasoning, 2024.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Criticbench: Benchmarking llms for critique-correct reasoning, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.227353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:23:50.608019Z digest=sha256:7903ad585f601059178c81e6a15a195df009161984cb9742b4e3c04ea348d846

Observation cbe914bb-a89b-488e-9a4e-23270b03daf5 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems AgentBench: Evaluating LLMs as Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.610680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.610680Z digest=sha256:e892b3fec392b6ad32ce0c49ea0c051a07b664019af7bab42f5179543e4484f4

Observation 6fbbee3c-4310-4576-b5dd-86a596378fa4 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Self-Refine: Iterative Refinement with Self-Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.613897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.613897Z digest=sha256:68a00dc013c8adfbd52bbea07abfdf3dd99a8bc0eb6458599d45b6589ec5e381

Observation b37b2185-8b21-4127-a86d-11fa9a8c23d5 · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.616857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.616857Z digest=sha256:ef2d4a67e00c52bdb62642930b4f3579aa2c8a2ca51545cbe6963467216da09f

Observation 3ad8f589-dcb4-4b9e-bcbd-43b850d2aae7 · outbound

This paper cites Openai o1 system card.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Openai o1 system card

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.217387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:23:50.619574Z digest=sha256:571177bf7d7dfa1cec09c4481cfe0759423639875b7bcd66cd1858c3e7c16357

Observation 70add6e7-969c-42c6-85f3-69a81e02bb88 · outbound

This paper cites Toolbench: An open platform for training, serving, and evaluating large language models as tool agents.https://github.com/OpenBMB/ToolBench, 2023.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Toolbench: An open platform for training, serving, and evaluating large language models as tool agents.https://github.com/OpenBMB/ToolBench, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.207456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:23:50.622398Z digest=sha256:167d650c24e1de6983a54589e7a5ae9789ed05aefe68ebec43c5823e990e81c3

Observation 88b56bbd-2842-4e35-9c4a-b5e20125ad00 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.625582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.625582Z digest=sha256:de208cd2af7b12b1c7419966a441d9196dddc5c132c5747d88a9fcbf4d21d385

Observation aa9aceab-8e31-46f2-a87d-48e6db534039 · outbound

This paper cites ART: Automatic multi-step reasoning and tool-use for large language models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems ART: Automatic multi-step reasoning and tool-use for large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.628403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.628403Z digest=sha256:f4ea58f8ce6fe9dfd32abbc1692c387003ecdbcdd038dea1f7583c9140675433

Observation 04db7ec6-6883-4523-be2d-152247168773 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.631361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.631361Z digest=sha256:9994cd980d6f9b45572334d829e23b73ca4e9bc7294944f5c15959594b96d4a4

Observation b5a94c9a-ccbe-4db5-ba5d-22ef39cabf80 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.634456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.634456Z digest=sha256:db6e4478c04ec8a2622b885d7c887561bde84b28129d072e0576d4f6bb883bfe

Observation cde1fcaf-a9ed-4c4d-b2fc-0222ca212c1b · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.637541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.637541Z digest=sha256:cb3913d4081b5a337cd57e54d280367f0526e90ea0b479e389b8c98b07490bdb

Observation 5c483474-2b68-44a9-af0b-9fb159ef37a7 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.640693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.640693Z digest=sha256:70b695aa5ac4127f066ef5614d0d7ea4dc9a4dd42f63f3ecc305e307fb8939c1

Observation 1cedfafc-06fb-410e-80c2-4218bf49f1f6 · outbound

This paper cites Commonsenseqa: A question answering challenge targeting commonsense knowledge.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Commonsenseqa: A question answering challenge targeting commonsense knowledge

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.192086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:23:50.643779Z digest=sha256:ef79a607ed4c055c0ccea5c355d08c21baa446ebe68f1bacd459b802ba92e8bd

Observation e0742482-4557-4ffc-802a-7fe44fd554b8 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.646976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.646976Z digest=sha256:13d7493000dd938e71ca1c60dc7382ec7f30b2f660ac6f6a77fcc5ae8a142c2a

Observation 0ee32718-1a8f-4d23-80ac-5bf23969389c · outbound

This paper cites an unresolved cited work.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.650424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.650424Z digest=sha256:585a2b2794113ac27b757d85edc3a30c9962ad0f0dc3742e4091f82de5246192

Observation 77c7a4e8-b25b-4364-bbba-1c434b0551d9 · outbound

This paper cites Superglue: A stickier benchmark for general-purpose language understanding systems.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Superglue: A stickier benchmark for general-purpose language understanding systems

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.175816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:23:50.653556Z digest=sha256:f5b731eaaf49f0c3ce9813c77592e745258cd6702791713b052dc27980ee82c0

Observation 30ab32ec-0e81-4648-bd65-b0f1709bc8ab · outbound

This paper cites Glue: A multi-task benchmark and analysis platform for natural language understanding.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Glue: A multi-task benchmark and analysis platform for natural language understanding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.165231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:23:50.656382Z digest=sha256:dce0d8111452d1b05852cd21a5952c733d4f6c047b4e7b43c016cadb3b37c8bf

Observation a3adac98-9a74-4d92-a4ca-8cc85b9b86ed · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Chain-of-thought prompting elicits reasoning in large language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.659110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.659110Z digest=sha256:14cbbb65f6cb29189625d26cfceec84370c90a53099fbc75749fb45ff7f2f8d4

Observation d5d9f157-174c-4510-85da-0e779b3a2316 · outbound

This paper cites Are large language models really good logical reasoners? a comprehensive evaluation and beyond.IEEE Transactions on Knowledge and Data Engineering, 2025.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Are large language models really good logical reasoners? a comprehensive evaluation and beyond.IEEE Transactions on Knowledge and Data Engineering, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.147672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:23:50.661863Z digest=sha256:290a9a1e9fdcc38ded95b66ff74b8292be88a0616ad4ae15ea64fbf4ffcdbcbf

Observation fe4b2975-5b95-4c5c-9a92-7fd4b40436d2 · outbound

This paper cites InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.664728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.664728Z digest=sha256:901e753a45c46fd73a1b094c16620531e39902bbe8b600e820aa1ec10cecdb8d

Observation 651be0e9-89ba-4c4c-92e1-47304ff082b0 · outbound

This paper cites WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.667812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.667812Z digest=sha256:a183039a843c7348ae40f0f4b647faf9a116112813f794d535236bf0fd8e3fdd

Observation fea9e1be-8724-4124-8197-19d1b07a18d8 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.671193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.671193Z digest=sha256:9b3e2f99b686a06f06e9f0011c93bb49998ed743716bbd01647d8ed8e07cd8b8

Observation 08e606c5-141b-4ebd-bbb3-d0b2510a6b24 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems ReAct: Synergizing Reasoning and Acting in Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.674463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.674463Z digest=sha256:b1704316089a5f326695532ae5bbfdbf9bd8d524d50f7adf14c7ee62da0bce2a

Observation 33d97532-bf7c-44fe-8194-c5efb9b6b978 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, 2019.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, 2019

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.677471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.677471Z digest=sha256:f7f6653df0f650dfa8ba11756eecc9907ad7c721451e282f764bb914ac929276

Observation 88981e13-795f-4a16-931d-086f52242ac5 · outbound

This paper cites Iolbench: Benchmarking llms on linguistic reasoning.arXiv preprint arXiv:2501.04249, 2024.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Iolbench: Benchmarking llms on linguistic reasoning.arXiv preprint arXiv:2501.04249, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.680320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.680320Z digest=sha256:6096cdc55b04f35d1651217701ef08cdd82754867fc470776dbdef8f5187162a

Observation 50a6d4b1-5579-4ce8-ae5a-68fe93b66e54 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.683069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.683069Z digest=sha256:2888d1583ae108d2c43b4e1a6a93ba5f5a9ca932a9a042245df248b82fc4212e

Pith citing papers

Observation d1d85105-0d9e-478a-b365-51b8874b8d35 · inbound

What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search cites this paper.

What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:06:05.611346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:19:18.121220Z digest=sha256:dc65bcdff181c581abfc131957ef9bda81a586bd5b1b36ba42682af7c972eac9

Observation adec540f-a8ac-440d-a6f2-7f273a7efedc · inbound

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs cites this paper.

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:46:18.873851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T02:44:33.143247Z digest=sha256:76e65321a00a35050063fc63ba9674f41c02e7eb6f3c7f5d6a0b565f5b377629

Observation d9b968af-82cf-4ae5-90d6-7438f9d985eb · inbound

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents cites this paper.

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:43:05.828295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:41:23.712146Z digest=sha256:36e5ac032774b180b9d92ca753702278d1767a631e5ffffc4e30e1b54c0954c1

Observation 3b67f314-ccd9-4705-bca2-1bc82a5318f9 · inbound

Large Language Models for Operations Research: A Comprehensive Survey cites this paper.

Large Language Models for Operations Research: A Comprehensive Survey OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-05-21T03:59:32.604852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T03:56:29.983335Z digest=sha256:73a76b0066ec185b0d505bab4988aa5f5edd6dd282f0f36abe30eda7f06fecec

Observation 1e605860-bb3a-4446-bdb3-38aadd4f87eb · inbound

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games cites this paper.

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:19:13.672489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:17:02.332687Z digest=sha256:e89ed48c766f4991d4773a120d419c37f97862cc8ac3b9895789869653296f23

Observation a99942cd-2771-402e-bc91-67b96f67ca3e · inbound

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources cites this paper.

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:20:07.361970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T20:19:30.720291Z digest=sha256:2545fd99d4a2c8a673126332ab34170136329d7c22052e65571f779bd6b67fa4

Observation 0d3bf3f0-e0b5-4632-93cb-a063009b55b1 · inbound

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources cites this paper.

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:19:50.751708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T05:18:55.074710Z digest=sha256:c3e2a9e328ce4ff02150f1815ee381410b31cc9dc20e0ef29a1c4c78302ecf26

Observation a9ffee4b-2b47-4112-8e2e-8d49e0cf63f5 · inbound

PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language cites this paper.

PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T14:04:57.314444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:04:57.314444Z digest=sha256:da6b8075c377e566278a32b3924505a93ba84a8d2802327612fd1808b8519abc