Pith. sign in

Paper Citation Record · LEDGER

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 8 inbound Pith citation observations for arXiv:2506.10764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10764 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:23:50.683069Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T14:04:57.314444Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:20:07.360446Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5b6495d8-7106-4e4e-9e68-c534f22f34d1 · outbound

This paper cites GPT-4 Technical Report.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.543505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.543505Z digest=sha256:bb273c8d5e584483e96f15620b3e5a0d29ea663fb6d20e39daddac7eb366788c

Observation 692304e6-190e-4b51-affa-5494d3e16856 · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.547360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.547360Z digest=sha256:a07b44e27d6b704474dd6899e4e2ee896dea929c8f4751c9f9a0eda5d0daf63d

Observation 2e969ac0-cb4b-40eb-82a3-f510cb7984d6 · outbound

This paper cites Program Synthesis with Large Language Models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Program Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.551498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.551498Z digest=sha256:c889e9e8eb72ad4698fc4b3d5edf63c6493cf15162fe63989d098057280dfce2

Observation 39138021-4c0d-44eb-af22-74039c8e02bb · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.554887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.554887Z digest=sha256:6cbbbd7e2f9f5fcaa78fe560b01419ee6c1cc2931e288b08e45433689471b698

Observation 895e5a4d-eca7-4c88-ba1b-70ad41bb08db · outbound

This paper cites Evaluating Large Language Models Trained on Code.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.561890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.561890Z digest=sha256:270011170486619271393b2280a6ff7c75d858de390031f9df7634156521b9cc

Observation d36b7c62-4ba5-4106-b181-07d11b280735 · outbound

This paper cites Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1– 113, 2023.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1– 113, 2023

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.565288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.565288Z digest=sha256:8106f9eaa1bdda9027122f07c268368fd30d41ddabcbe379fafd5580f56a6c85

Observation 4bbc9d8e-a657-4e40-8e9a-35c1b8918355 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.568987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.568987Z digest=sha256:67807d4fe4054d8ad920d96f004e68396a4935aab92b4f7df6158361ec57c805

Observation ac041da5-8e76-4147-b5dc-be0ffe382590 · outbound

This paper cites Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.571999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.571999Z digest=sha256:ca109598199098f701be7b71adf7301816a3c9ab59175acdeb6ae50f7d314358

Observation 54c5a251-52c7-496e-97db-ec3ed7ebaf2d · outbound

This paper cites NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.574907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.574907Z digest=sha256:636f389e847df5c824ec8ede40f9cb7307e214dd0d42aef607f70f03f4908eb3

Observation a424da53-d40a-46f5-af3b-ef1a0875be4c · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.577993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.577993Z digest=sha256:999f852ae55ac70c2ed72b0ebea3eb2519ac9a2e9d4c7d3ee9d7d3e3ef15c20f

Observation bc03159f-8cdf-4858-901d-19719e22bfa7 · outbound

This paper cites The Llama 3 Herd of Models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.580934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.580934Z digest=sha256:c6fac0ecdea506a97e5de48bf90a1b6002665075ccaaca97ceebdff8418b4528

Observation 7cbf9a37-12b5-4b76-aecd-cc4a7a198749 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.584251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.584251Z digest=sha256:a384d39796115693a3483cba171162525027993e4b5808c4eabb95bb7917bd5c

Observation 4380e4f1-ea7e-40b0-b55e-c85b8b215dc0 · outbound

This paper cites R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.587173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.587173Z digest=sha256:c962878fec5855a63df0efb216c9373f6c8d2552d9840ff9cc9422f335199f48

Observation 94684895-c743-40d3-87f1-6e6be6a6bbe9 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Measuring Massive Multitask Language Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.590207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.590207Z digest=sha256:26ad3c83fbdc8f2a98100f29a22f4fba983b65efa97306557891355ecec1f0ef

Observation 5b1a572f-7661-4758-a832-1cfbdf5d1d36 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.593077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.593077Z digest=sha256:8f890eeaae0e6371e584ff10819d368cebe5f609e686c1b622a3dc8cf54efc54

Observation 59c6256c-9287-43d1-96eb-183abc89676b · outbound

This paper cites Mlagentbench: Evaluating language agents on machine learning experimentation.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Mlagentbench: Evaluating language agents on machine learning experimentation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.596213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.596213Z digest=sha256:8e500f4c13954edffa1ccf4a6f2064b64f1de639ce86b0c76ac5d788bce5f466

Observation 141b4162-76ee-4f17-8f17-b187f4a13aa0 · outbound

This paper cites Aide: Ai-driven exploration in the space of code, 2025.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Aide: Ai-driven exploration in the space of code, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.599339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.599339Z digest=sha256:4a734e8d36704ac1b6cb34c384c90c2e6864db1a2e33ba06488eca737fbe6b89

Observation 316f20a2-a8c8-4f4d-923d-612cab304516 · outbound

This paper cites Let’s verify step by step.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Let’s verify step by step

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.602556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.602556Z digest=sha256:262590f0dd06b68859af27d6c34581005e71a4b6616d4ab8c40b78a263b9b055

Observation 4ae518c8-e72e-4df0-8705-bd7480778af1 · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Truthfulqa: Measuring how models mimic human falsehoods

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.237836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:23:50.605362Z digest=sha256:1efdfd428ef170d87c7f50122d5e42ffe805500ac5308e5823df6df4e38fde6f

Observation c22a6d6e-7175-4207-896c-7be959e993f2 · outbound

This paper cites Criticbench: Benchmarking llms for critique-correct reasoning, 2024.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Criticbench: Benchmarking llms for critique-correct reasoning, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.227353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:23:50.608019Z digest=sha256:04865afcd0a4e4073aade32b51323cb7dca7010ea95ce2c7574ee876fa1c68e3

Observation cbe914bb-a89b-488e-9a4e-23270b03daf5 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems AgentBench: Evaluating LLMs as Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.610680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.610680Z digest=sha256:63efb9d314d97bb809f68cc9e627fbddcaf635e86f858f3d16559460c0b12c2e

Observation 6fbbee3c-4310-4576-b5dd-86a596378fa4 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Self-Refine: Iterative Refinement with Self-Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.613897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.613897Z digest=sha256:653f30555e4b329e31516d11d84ddb343655744e2e594f9072994ff831ba5764

Observation b37b2185-8b21-4127-a86d-11fa9a8c23d5 · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.616857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.616857Z digest=sha256:4719645dd8e5739c5c3fe1f2183b343b4ab669d7c08bb8fb1bef14c86d6cdb0d

Observation 3ad8f589-dcb4-4b9e-bcbd-43b850d2aae7 · outbound

This paper cites Openai o1 system card.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Openai o1 system card

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.217387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:23:50.619574Z digest=sha256:1fa5b3c265b1af336ecaa4da4234f265e0cf928ea4812f079ff495c423900efa

Observation 70add6e7-969c-42c6-85f3-69a81e02bb88 · outbound

This paper cites Toolbench: An open platform for training, serving, and evaluating large language models as tool agents.https://github.com/OpenBMB/ToolBench, 2023.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Toolbench: An open platform for training, serving, and evaluating large language models as tool agents.https://github.com/OpenBMB/ToolBench, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.207456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:23:50.622398Z digest=sha256:5f57b1d06d82e89e3606b0bdb3e508152840b3fbcae8abd1292889778b0cd5a5

Observation 88b56bbd-2842-4e35-9c4a-b5e20125ad00 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.625582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.625582Z digest=sha256:2b59034c8627b072be98ca2c4a98fd23baa9bdd97773fa93f772428b0092dcc5

Observation aa9aceab-8e31-46f2-a87d-48e6db534039 · outbound

This paper cites ART: Automatic multi-step reasoning and tool-use for large language models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems ART: Automatic multi-step reasoning and tool-use for large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.628403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.628403Z digest=sha256:117d848ea7588f2f633354ed483c8cec3127fe13f4df906ae8dae352f031eba5

Observation 04db7ec6-6883-4523-be2d-152247168773 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.631361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.631361Z digest=sha256:4b0a9e98ba0c71f35eb9583cea63685b860abbc3d76e860cae5849d278f2e16d

Observation b5a94c9a-ccbe-4db5-ba5d-22ef39cabf80 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.634456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.634456Z digest=sha256:5ee34b722942ceee8c4791e96eb25a4b189f38aed099f3e3b62e3b56f3f27d5f

Observation cde1fcaf-a9ed-4c4d-b2fc-0222ca212c1b · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.637541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.637541Z digest=sha256:099a2c6d18f4cdeed32a69b880cb8bd18f5c98562534250f91ba84095d86f4cb

Observation 5c483474-2b68-44a9-af0b-9fb159ef37a7 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.640693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.640693Z digest=sha256:cd6be6d2314f00d5d1462f16da0bd13d1697d8b7be876b6277b699068fc71f76

Observation 1cedfafc-06fb-410e-80c2-4218bf49f1f6 · outbound

This paper cites Commonsenseqa: A question answering challenge targeting commonsense knowledge.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Commonsenseqa: A question answering challenge targeting commonsense knowledge

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.192086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:23:50.643779Z digest=sha256:d3e012d9cdc2c7d6ba77969c65d2169055f0610687c8ce9d405b3bf8e23d028c

Observation e0742482-4557-4ffc-802a-7fe44fd554b8 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.646976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.646976Z digest=sha256:7b3198e546aa1e7689189e459af4f832941946ebed096de331046b7f7ac1a5ec

Observation 0ee32718-1a8f-4d23-80ac-5bf23969389c · outbound

This paper cites an unresolved cited work.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.650424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.650424Z digest=sha256:1c41b40c5af28bb4a813d23e416d8552fe3b4effa5764e9390fbc1aeb13ac012

Observation 77c7a4e8-b25b-4364-bbba-1c434b0551d9 · outbound

This paper cites Superglue: A stickier benchmark for general-purpose language understanding systems.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Superglue: A stickier benchmark for general-purpose language understanding systems

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.175816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:23:50.653556Z digest=sha256:78814d89710b611a550b0fe31d0fe6718ec417b4920182da18fef8d38172852f

Observation 30ab32ec-0e81-4648-bd65-b0f1709bc8ab · outbound

This paper cites Glue: A multi-task benchmark and analysis platform for natural language understanding.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Glue: A multi-task benchmark and analysis platform for natural language understanding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.165231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:23:50.656382Z digest=sha256:427250e3cec4b1b7fac9624ad0e57d3cb5be8c5f8593a29aecc233393a42203a

Observation a3adac98-9a74-4d92-a4ca-8cc85b9b86ed · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Chain-of-thought prompting elicits reasoning in large language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.659110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.659110Z digest=sha256:d10fe16275785fbb82227cfcad2df6d4ebbb6997a0e276d4d7bebdad8f3552ee

Observation d5d9f157-174c-4510-85da-0e779b3a2316 · outbound

This paper cites Are large language models really good logical reasoners? a comprehensive evaluation and beyond.IEEE Transactions on Knowledge and Data Engineering, 2025.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Are large language models really good logical reasoners? a comprehensive evaluation and beyond.IEEE Transactions on Knowledge and Data Engineering, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.147672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:23:50.661863Z digest=sha256:89614d10b004d406fd13c1e549d408ae5c7a19f6d5abb3e6a54ad1c3a29efcab

Observation fe4b2975-5b95-4c5c-9a92-7fd4b40436d2 · outbound

This paper cites InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.664728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.664728Z digest=sha256:e59dbf4d13eefc454bd47a14ac9d23245f253a97c4ecaa499aa18d19d5a1803a

Observation 651be0e9-89ba-4c4c-92e1-47304ff082b0 · outbound

This paper cites WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.667812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.667812Z digest=sha256:be1f8f859b9b60ed703c86f2d42f889b4df3f9ab4245d6e91b641ef48ed810d8

Observation fea9e1be-8724-4124-8197-19d1b07a18d8 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.671193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.671193Z digest=sha256:11aaf54c789ef026d5e7b860cbef6dc417314b2d7886c299aab74291cd87c979

Observation 08e606c5-141b-4ebd-bbb3-d0b2510a6b24 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems ReAct: Synergizing Reasoning and Acting in Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.674463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.674463Z digest=sha256:c589ed1389e447c1a521b16c2ea24e97df4af07a3b166b5c4c82d1f8f93c5a2b

Observation 33d97532-bf7c-44fe-8194-c5efb9b6b978 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, 2019.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, 2019

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.677471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.677471Z digest=sha256:61c888de667c62c830034eab4f7c1daf15677dc813a1bf96d7a73d64dfb0af2c

Observation 88981e13-795f-4a16-931d-086f52242ac5 · outbound

This paper cites Iolbench: Benchmarking llms on linguistic reasoning.arXiv preprint arXiv:2501.04249, 2024.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Iolbench: Benchmarking llms on linguistic reasoning.arXiv preprint arXiv:2501.04249, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.680320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.680320Z digest=sha256:df9b384f5724d71bde65a25cfdf0b6b172d4d21d76784357336ab91037e6dd89

Observation 50a6d4b1-5579-4ce8-ae5a-68fe93b66e54 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.683069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.683069Z digest=sha256:84c435ea7b6d23b8af97a173c49e4a3cfad99c0f1e1815faa662673b2bc2bcbd

Pith citing papers

Observation d1d85105-0d9e-478a-b365-51b8874b8d35 · inbound

What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search cites this paper.

What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:06:05.611346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T02:19:18.121220Z digest=sha256:b7620bc1863fdb7b22600678f89e5ba13a8cabb242f7443dfd47b4c2b55cf5f7

Observation adec540f-a8ac-440d-a6f2-7f273a7efedc · inbound

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs cites this paper.

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:46:18.873851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-12T02:44:33.143247Z digest=sha256:23343a4b46f94eb26624b85df012f730f03423ae83f1060faac52b028aec3d91

Observation d9b968af-82cf-4ae5-90d6-7438f9d985eb · inbound

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents cites this paper.

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:43:05.828295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T05:41:23.712146Z digest=sha256:99d28f74b913a56a562aea9ac131838ff2d5e2704f9b28dc571ac1ba5b06c635

Observation 3b67f314-ccd9-4705-bca2-1bc82a5318f9 · inbound

Large Language Models for Operations Research: A Comprehensive Survey cites this paper.

Large Language Models for Operations Research: A Comprehensive Survey OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-05-21T03:59:32.604852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T03:56:29.983335Z digest=sha256:520b70136e38a22542e05b4d49bc1f3a0c6d721bd48cfd2f7550dedbfe59820f

Observation 1e605860-bb3a-4446-bdb3-38aadd4f87eb · inbound

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games cites this paper.

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:19:13.672489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T21:17:02.332687Z digest=sha256:4e38b2f7ff08278e4a2155b3b5dd6101c4bd6b0bacf769890e416be542e5f538

Observation a99942cd-2771-402e-bc91-67b96f67ca3e · inbound

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources cites this paper.

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:20:07.361970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-25T20:19:30.720291Z digest=sha256:ba2974864df90b81cd0f9f2336edea19b47d8543bd369bbf57d4fe96474ec16c

Observation 0d3bf3f0-e0b5-4632-93cb-a063009b55b1 · inbound

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources cites this paper.

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:19:50.751708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T05:18:55.074710Z digest=sha256:c041a9f53f3fb375d6393e3471feeb9949ed03cd2ec48de104d8acf8ca6ccf69

Observation a9ffee4b-2b47-4112-8e2e-8d49e0cf63f5 · inbound

PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language cites this paper.

PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T14:04:57.314444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:04:57.314444Z digest=sha256:7fc59a21ef4ad7714b7706bbb4b1bcb9d1043ef0c514f1e355e141b761010495