Pith. sign in

Paper Citation Record · LEDGER

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems

As of 20 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 4 inbound Pith citation observations for arXiv:2505.17968.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17968 v1

Coverage vector

measured 99 of 99 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:40:40.704557Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:08:35.185016Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T00:06:37.955246Z

Reference resolution

99 of 99 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved61
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 82402fab-ecce-4773-907f-cd8bb8afa647 · outbound

This paper cites Large-Scale Bandit Problems and KWIK Learning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Large-Scale Bandit Problems and KWIK Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.381858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.381858Z digest=sha256:88fb0222ef29d92ac504dd742892711cdc77e5dab85ea1ef4182b9e96f4d6694

Observation 50c39405-4d30-48ca-9495-442ea9c4d0ee · outbound

This paper cites Queries and Concept Learning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Queries and Concept Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.448282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.448282Z digest=sha256:90feb4b22b3751414ae56840c1372b58b3071620a6c82775aff072c5e885ed5a

Observation 57f7fe63-fce0-45b3-bd70-3443d251a9f7 · outbound

This paper cites Inductive Inference: Theory and Methods.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Inductive Inference: Theory and Methods

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.491594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.491594Z digest=sha256:0d502489bc6a63f7539b5c6e2c1554814b4f505faf41260557e7d4ba99d1d8ff

Observation d30d95ee-ceca-4221-a946-e5cbf8dc69cb · outbound

This paper cites Claude 3.5 sonnet.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Claude 3.5 sonnet

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.547987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.547987Z digest=sha256:ed986c5b77f29c0681fc56932e367bc92f953c6d44ae09fcd0c4b1dac8c0d952

Observation 9693341c-04d0-41a0-9cb9-4c3f94395333 · outbound

This paper cites Toward Efficient Exploration by Large Language Model Agents.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Toward Efficient Exploration by Large Language Model Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.601633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.601633Z digest=sha256:d5f28aa57f1349bbc166347e55c7432e8b434465299ec5efda96510895390ed0

Observation fcefdc3a-515d-4d59-93b9-9c562e09f279 · outbound

This paper cites A Markovian Decision Process.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems A Markovian Decision Process

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.640312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.640312Z digest=sha256:732ed13a1f29a692f2dda1c3c57cd8230773889ff2a2411601180562cc711927

Observation fc58a8ce-a78f-466e-a746-4cebf820daae · outbound

This paper cites Using cognitive psychology to understand GPT-3.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Using cognitive psychology to understand GPT-3

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.682542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.682542Z digest=sha256:f948dc640c3a6b7a7d43053782b7350694b40ce7a6b5a215185aebc3681efdbe

Observation d291f60f-f390-4fdc-8cb9-c8ac0b12fa8c · outbound

This paper cites Variational inference: A review for statisticians.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Variational inference: A review for statisticians

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.763095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.763095Z digest=sha256:84591379df0e07f3d45514f13155ac02e19fb3f2ef2579fbb7457b373afc037d

Observation d22695ee-a3ea-41f9-8dd8-bf2b9b39962b · outbound

This paper cites R-MAX – A General Polynomial Time Algorithm for Near-Optimal Reinforcement Learning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems R-MAX – A General Polynomial Time Algorithm for Near-Optimal Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.855269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.855269Z digest=sha256:d624648128c70a43666151ba36a75f30d3f233c21d8dfbdcae25ee9a020aa963

Observation 6bc69ef7-5ca6-44d8-8eaf-0c4f7ce9e93c · outbound

This paper cites Language models are few-shot learners.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Language models are few-shot learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.911448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.911448Z digest=sha256:40a93a60fc9d0a8caf1b5f1936c6f5c6871c47c7cbac139de3faf7decddbc159

Observation 43339939-93b2-46e5-a19c-abb7de83b6e5 · outbound

This paper cites Why Do Multi-Agent LLM Systems Fail?.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Why Do Multi-Agent LLM Systems Fail?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.993539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.993539Z digest=sha256:504235a5954ac4b18f4b22d8850811aa079102da13d3000515b7105d8575d704

Observation 6ce3290c-0b2a-44f6-8b06-4dd45832900d · outbound

This paper cites Bayesian Experimental Design: A Review.Statistical Science, pp.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Bayesian Experimental Design: A Review.Statistical Science, pp

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.070856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.070856Z digest=sha256:e2979c323725259960acc67d40f0030fb96ef59c5c709e4a85762ecb10088ec5

Observation 3f6f2e9b-8ebe-4a0e-80b4-0d4cc2750171 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.147832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.147832Z digest=sha256:581a060d0d4c405d367aec7a7bc76dfc2dfc2c9601a41a1603a2e6e5b8f4d840

Observation 51fed5c1-5bc9-4750-9138-b24393c37def · outbound

This paper cites The first crank of the cultural ratchet: Learning and transmitting concepts through language.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems The first crank of the cultural ratchet: Learning and transmitting concepts through language

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.213779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.213779Z digest=sha256:321adbdee13393594ebc52e2a54b5bc8ce3cd48eea1595b1c0cf224aef319abf

Observation 573ec285-8443-4393-8478-1bca2ad57f2b · outbound

This paper cites CogBench: a large language model walks into a psychology lab.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems CogBench: a large language model walks into a psychology lab

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.302731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.302731Z digest=sha256:194e993c25c560a84a95bf7e6b98370cf74d9017e8eda988b4517c70bf265d62

Observation a02f0fd1-c699-42df-a7ca-9513f1fda469 · outbound

This paper cites The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.403417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.403417Z digest=sha256:3b4c3a2d73287db951250bee536dc1dcdaac198f727a04ec0462e6d666c3c795

Observation c841fc75-f7d6-4854-b798-1586206b5ac9 · outbound

This paper cites Uncertainty, Information, and Sequential Experiments.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Uncertainty, Information, and Sequential Experiments

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.481437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.481437Z digest=sha256:e6063729b3f997639629376b705f3e5452e1f25c192079ec21d7dc8354fd8d70

Observation 65f83589-7dfb-43e6-8e21-a237ddcfcf42 · outbound

This paper cites PILCO: A Model-Based and Data-Efficient Approach to Policy Search.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems PILCO: A Model-Based and Data-Efficient Approach to Policy Search

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.569042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.569042Z digest=sha256:83101f0ddb6121dae0bf781eab958726f814fd798f1fed2960cfc842c0501606

Observation 3e2b302c-9af2-4c54-9c27-023bc5314e69 · outbound

This paper cites Aleatory or Epistemic? Does it Matter? Structural Safety, 31(2):105–112, 2009.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Aleatory or Epistemic? Does it Matter? Structural Safety, 31(2):105–112, 2009

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.675725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.675725Z digest=sha256:cad9489c5e3a4d40dd065102368bf2f7e0c522b104b42f468f82d4ed0bee7626

Observation f610df2c-264d-4db8-a8fc-ce0f5c8e817a · outbound

This paper cites Is In-Context Learning in Large Language Models Bayesian? A Martingale Perspective.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Is In-Context Learning in Large Language Models Bayesian? A Martingale Perspective

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.779818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.779818Z digest=sha256:b8fcb13b40637ba80e58547feb275b85e6304dfc375f2a4f3dc54d554c83a2a8

Observation 377321d7-611e-4dd4-93f7-db14bbafc4ee · outbound

This paper cites Variational Bayesian optimal experimental design.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Variational Bayesian optimal experimental design

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.861953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.861953Z digest=sha256:f6023d38f82736844042ee4f71e6636f24a457caf9f7d8277233956a022e62fe

Observation 5a31fd6a-78f1-4d59-baae-dfc5895de5c9 · outbound

This paper cites Baby steps in evaluating the capacities of large language models.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Baby steps in evaluating the capacities of large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.954207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.954207Z digest=sha256:12db0e40c5cfd5f1b65588d96ede62fa7f440ee9df1065b9fee4a33f7a1527bb

Observation d67b3eaf-3302-4617-a054-1155f531c76c · outbound

This paper cites BoxingGym: Benchmarking progress in automated experimental design and model discovery.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems BoxingGym: Benchmarking progress in automated experimental design and model discovery

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.042557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.042557Z digest=sha256:20b53b7c2fc582029d5d23f33e7c4022dd2d0eb5869666eaf523a6c81189fa5e

Observation 9a177777-7675-4fca-8622-89383c8e4bff · outbound

This paper cites Amplify scientific discovery with artificial intelligence.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Amplify scientific discovery with artificial intelligence

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.118874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.118874Z digest=sha256:8ace5eb67ed4fb19636e4274bbb057444b4dea22346e71db0e1b83ff3b2f6c7b

Observation 2aaa365c-eb5e-49a1-963f-e3f9dac601e1 · outbound

This paper cites ANOVA: Repeated measures.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems ANOVA: Repeated measures

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.160962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.160962Z digest=sha256:c0bc082fee31fa2ffef21097a18541ca1fd1d2b45d08cfb35e847a08f25ebdaf

Observation b96652ac-dc3f-48b4-b0cb-b936df006f2b · outbound

This paper cites Towards an AI co-scientist.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Towards an AI co-scientist

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.213138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.213138Z digest=sha256:0f3c8a10c439b429b8336d5e7db14941c81270062e7a5a1857bdbcaca8987f2e

Observation 7d509de6-b5d1-4fe6-a570-73c8714ad341 · outbound

This paper cites The Llama 3 Herd of Models.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems The Llama 3 Herd of Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.277604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.277604Z digest=sha256:05c17fe0f782defebea4317bf1b53ce2539a9b2072765e8cabfc155f3993771c

Observation e94454b9-3bf0-4669-96f1-ec0efc5c3a10 · outbound

This paper cites Bayes in the Age of Intelligent Machines.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Bayes in the Age of Intelligent Machines

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.344865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.344865Z digest=sha256:4cbbeceda21d2fc3aa03cf9542f1001f74ec2cb33fca4f656f6d54cf89b049c2

Observation 796414a5-9386-45a3-9c79-e35aefff5510 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.407975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.407975Z digest=sha256:70130dc055109b3cc1b9630142c067738996d61e013005dc5be28050ed68bf88

Observation 997b1a60-ef79-45a8-81af-4733c6156182 · outbound

This paper cites Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.473655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.473655Z digest=sha256:a9721bb5182de80f1e9d58f302ec343919974546cf7c82562805549c1007df57

Observation c1d76051-ac80-4a40-a808-7ee91b4823a7 · outbound

This paper cites GPT-4o System Card.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems GPT-4o System Card

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.532686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.532686Z digest=sha256:13ccea45e2544be76ec364c40dae3d65f539f37d0be9cd441cf83dcdb9fdaecd

Observation 7974c83d-b72e-4146-a36e-208df11b3aa8 · outbound

This paper cites What Do Learning Dynamics Reveal About Generalization in LLM Reasoning?.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems What Do Learning Dynamics Reveal About Generalization in LLM Reasoning?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.600789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.600789Z digest=sha256:321dd08f429877dd1684d98894decbacba468267ef19319178af2e4b9d271302

Observation 81194d79-d3a4-47e4-882b-f0d6797a95f9 · outbound

This paper cites Using the tools of cognitive science to understand large language models at different levels of analysis.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Using the tools of cognitive science to understand large language models at different levels of analysis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.640066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.640066Z digest=sha256:65967de7e3a8a9fcaf1574cc8db6aa4f0a39c5fb823300cfe0b8fb6319a19789

Observation 1a170178-e8a4-4ea6-adee-ed86d9bcc9b9 · outbound

This paper cites A robust class of context- sensitive languages.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems A robust class of context- sensitive languages

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:49.565656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:34.679985Z digest=sha256:658572bc54e40bc11632a265f29d103df603bddb4bdb452191f31417e2d09d96

Observation 1d339c4b-a002-4688-9b16-c4a20db7b5e8 · outbound

This paper cites Passive learning of active causal strategies in agents and language models.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Passive learning of active causal strategies in agents and language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:49.352850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:34.740445Z digest=sha256:b4e1b5d99845ecd16c96dd0091b1e3977f4f0b7e08b0761a5dc4d8398a115b86

Observation 08d6ff58-e230-40a8-8229-997dc742e396 · outbound

This paper cites Structured chain-of-thought prompting for code generation.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Structured chain-of-thought prompting for code generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:49.168710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:34.788602Z digest=sha256:e52d152045b5ad2e9b75983c0593b91ad65c213d9c1b2276c5b32b7c248df63b

Observation 2396efc4-279d-4b28-9905-b61235b5033b · outbound

This paper cites Reducing Reinforcement Learning to KWIK Online Regression.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Reducing Reinforcement Learning to KWIK Online Regression

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:48.955565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:34.860427Z digest=sha256:c57df22417dd06a022ceaec8894a9060fdaeb50b890d21a495ccf3523da896e6

Observation 4d7ba1ae-a463-4956-8e4a-de594c96296b · outbound

This paper cites Knows What It Knows: A Framework for Self-Aware Learning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Knows What It Knows: A Framework for Self-Aware Learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:48.835093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:34.926262Z digest=sha256:54b587461b136a9c40653d4735c4957105fca23b7db8580bde165bf441ef88a0

Observation 35ebf653-8272-48c8-8b50-723822d32714 · outbound

This paper cites On a measure of the information provided by an experiment.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems On a measure of the information provided by an experiment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.999414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.999414Z digest=sha256:e996958961b9dcd8a865d6acbf1f8f5404b0c0c421be9b99c1c0b29fa82be3f2

Observation 7e304cd7-3ff4-4750-91d1-b9ed954c52af · outbound

This paper cites Learning quickly when irrelevant attributes abound: A new linear-threshold algorithm.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Learning quickly when irrelevant attributes abound: A new linear-threshold algorithm

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:48.661957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:35.093317Z digest=sha256:1e2b98c7c3c5c996d6e377cd63e56ec93b6e03ae34877f1202311bba94cfecbc

Observation 1ce2be8f-e194-42a5-87d0-115a6f41818f · outbound

This paper cites Decoupling Exploration and Exploitation for Meta-Reinforcement Learning Without Sacrifices.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Decoupling Exploration and Exploitation for Meta-Reinforcement Learning Without Sacrifices

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:48.477838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:35.176378Z digest=sha256:93a3fc40564beadf83204a3609227f8fc89c008b89d100cba6ce5f45ceef3bd7

Observation e3fe3d14-33ed-45ca-8b75-7ec14e8b2bc8 · outbound

This paper cites Large Language Models Assume People are More Rational than We Really are.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Large Language Models Assume People are More Rational than We Really are

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:35.241752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:35.241752Z digest=sha256:b6fb5298d30765abfaa2408e893a24120f03deee6869bd74ff802c73850e8595

Observation d8884055-f106-44e7-a599-64123dd7f2b3 · outbound

This paper cites Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:35.314587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:35.314587Z digest=sha256:055b59a5749a7c92ab4a0466b9ed110d69b0777c487521f132c7383229dc9f1f

Observation 0913bfcd-5527-4056-921c-763140c1f2b2 · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:35.395340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:35.395340Z digest=sha256:a8c27621c58cb01c2f9b55afc89b68d5ec020b90e43f0f615e1f8a4c321c12db

Observation c49abd94-62aa-43cb-8152-c385fdd9970a · outbound

This paper cites Deconstructing Long Chain-of-Thought: A Structured Reasoning Optimization Framework for Long CoT Distillation.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Deconstructing Long Chain-of-Thought: A Structured Reasoning Optimization Framework for Long CoT Distillation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:35.476624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:35.476624Z digest=sha256:de0caf58b7b2dc46c5cddd63f438770fde834fae9a1d5af6212e48a11341c335

Observation 676edd7b-4582-4f39-86de-04ef2f90fe85 · outbound

This paper cites Category learning through active sampling.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Category learning through active sampling

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:48.302269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:35.562091Z digest=sha256:f428324c73b8886f92be574a8788a5c3f391e8bb08dd61fa911b76bfea926dcb

Observation 6c60ac6e-d798-4f3c-8d5c-a7c133cbac82 · outbound

This paper cites Is it better to select or to receive? learning via active and passive hypothesis testing.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Is it better to select or to receive? learning via active and passive hypothesis testing

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:48.147093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:35.624883Z digest=sha256:9177a6d78617075937319d02a4543cf86f346492de982287a010b6a765ea8460

Observation 1841677b-fbc5-4fad-ac05-f55ef2ea0065 · outbound

This paper cites Modeling rapid language learning by distilling Bayesian priors into artificial neural networks.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Modeling rapid language learning by distilling Bayesian priors into artificial neural networks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:35.712060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:35.712060Z digest=sha256:7a029b7f0a8613007c2928bb0eaa3599603724ae93bb5e598c101534af871d3d

Observation 05f80f70-7c1a-4c6d-a7a7-c357f7027ba9 · outbound

This paper cites Embers of autoregression show how large language models are shaped by the problem they are trained to solve.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Embers of autoregression show how large language models are shaped by the problem they are trained to solve

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:47.995708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:35.818075Z digest=sha256:bcce1f491b1c75a58bdb31140d7cf7d16770f7bae3f7d9c0a309f05cb29b07f5

Observation e3be7c28-fb95-48a7-a078-4fd94a17d26a · outbound

This paper cites MatPilot: an LLM-enabled AI Materials Scientist under the Framework of Human-Machine Collaboration.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems MatPilot: an LLM-enabled AI Materials Scientist under the Framework of Human-Machine Collaboration

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:35.888952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:35.888952Z digest=sha256:42f729ce4ec03eaf791cf037df70484f30d66db359db6bbbdf4a2c84f2a3d4c3

Observation caa38cd3-48dd-4eef-81c8-b9f9bf08921f · outbound

This paper cites Sparks of Science: Hypothesis Generation Using Structured Paper Data.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Sparks of Science: Hypothesis Generation Using Structured Paper Data

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:36.009484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:36.009484Z digest=sha256:f219c9d8d0f7437a13e59a37531cfb7f7641db9ab393475dce3a2c7ffb9ec31e

Observation 3c36647d-194f-4806-981d-f738f5196c12 · outbound

This paper cites (More) Efficient Reinforcement Learning via Posterior Sampling.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems (More) Efficient Reinforcement Learning via Posterior Sampling

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:47.815242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:36.102354Z digest=sha256:99a0a7436469ec9a2ac772bcdef569325ebb16ed977c402d30795b950dd527d5

Observation 038b21dc-64a8-4b1e-978d-80f4a4284dc1 · outbound

This paper cites Puterman.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Puterman

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:47.645625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:36.190057Z digest=sha256:c9fb352804a4e84c24f6b563a6b7879ff85fab8c8a80d4369b7b61c7dba4ef66

Observation 844a8a34-8b35-471e-8b6e-faea4f154c0c · outbound

This paper cites Towards scientific discovery with generative ai: Progress, opportunities, and challenges.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Towards scientific discovery with generative ai: Progress, opportunities, and challenges

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:47.543996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:36.295907Z digest=sha256:cd1b276ae493938bcdc8cc62fe613cfe8b5e4278063e0faa711165fc3e584ede

Observation 61eec6b9-b181-4567-a92f-b729e4834598 · outbound

This paper cites Diversity-Based Inference of Finite Automata.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Diversity-Based Inference of Finite Automata

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:47.335383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:36.395591Z digest=sha256:a688e1f5a33632584113eabcdb8ebdd0beffd3741497e49d5d13aa43e0a1b5a6

Observation e096bed7-de24-4dad-a5fe-528ebfa59b43 · outbound

This paper cites Inference of Finite Automata Using Homing Sequences.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Inference of Finite Automata Using Homing Sequences

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:47.116883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:36.486680Z digest=sha256:12356aa88976301a271bf61064ab7494be0a7948b6eef86b084071758e24afa7

Observation 391f4971-cca2-42ce-b42e-b7e909f1f4f5 · outbound

This paper cites Jagadish, Marvin Mathony, Tobias Ludwig, and Eric Schulz.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Jagadish, Marvin Mathony, Tobias Ludwig, and Eric Schulz

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:36.544531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:36.544531Z digest=sha256:66e4c8f1716ceef669a90943ab64b5d8b824068adae6d6c6477275417272ea67

Observation ac32d513-3360-4f9d-a0a6-e992cb0f6269 · outbound

This paper cites Symbolic metaprogram search improves learning efficiency and explains rule learning in humans.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Symbolic metaprogram search improves learning efficiency and explains rule learning in humans

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:46.746513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:36.636665Z digest=sha256:e54c330acaa5ee7139535736bb92647b259af30d6cad5907d7ce37cbbc94326e

Observation b397f9fd-5cb7-4c33-a49c-73948ca38d19 · outbound

This paper cites Trading off Mistakes and Don’t- Know Predictions.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Trading off Mistakes and Don’t- Know Predictions

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:46.416593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:36.706230Z digest=sha256:224852bb59d9a5bb5ac3af0def63f05ba8ac262ffda14aab2ab116670bfa0b25

Observation c0a039d9-d524-4528-b9c8-acf3dc5d5e92 · outbound

This paper cites Agent Laboratory: Using LLM Agents as Research Assistants.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Agent Laboratory: Using LLM Agents as Research Assistants

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:36.829520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:36.829520Z digest=sha256:a287f274f83cda6fe4f945fc8d62851f3c12e4f8f8fb7e510edb24e6766ff920

Observation 0c74deb9-8ff3-4bd5-8068-21cb2a079a07 · outbound

This paper cites Active Learning Literature Survey.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Active Learning Literature Survey

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:46.070277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:36.913290Z digest=sha256:9ec6185e02010f7a33664ad36424ec73f7f038bf859676a1802989dbc1e111b0

Observation 439cc0a6-169c-4dfb-a18c-432cc02b15ec · outbound

This paper cites Llm-sr: Scientific equation discovery via programming with large language models.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Llm-sr: Scientific equation discovery via programming with large language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:45.807125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:37.006923Z digest=sha256:46391787043511089f2a51d3af221ad0ba4009b09b8acc006466ae3e9011aaf7

Observation 28738de6-bb3c-4bfb-8ab9-454ce709f0e5 · outbound

This paper cites Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:37.118461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:37.118461Z digest=sha256:074d97896f7e92adb6bd652aa9c3772dd7edf5af28187757c1d2d8829d80151a

Observation 7290d399-1b29-42ab-bc4f-f3030c56f46e · outbound

This paper cites To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:37.243258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:37.243258Z digest=sha256:c6bddaafc31db8612f3e2177c163f3b0e5f71954dc1eaff0464757b4fc8c1564

Observation 971622c1-4ba9-46c2-ab12-73b9b13d43ba · outbound

This paper cites PaperBench: Evaluating AI's Ability to Replicate AI Research.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:37.301272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:37.301272Z digest=sha256:d818c827da113846104ac1d0b6b8a7915912b0a5609c948391fd2e3b4558f960

Observation ae21c6d6-7306-4574-a60b-247ea474c416 · outbound

This paper cites An Analysis of Model-based interval estimation for Markov Decision Processes.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems An Analysis of Model-based interval estimation for Markov Decision Processes

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:45.551239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:37.397807Z digest=sha256:460ac22b2f97a546b570b2c6147b8acea90a91a90dfc4800111cf408e57b1b70

Observation 35f9fa0f-d275-4436-88e0-646ebed51a86 · outbound

This paper cites A Bayesian framework for reinforcement learning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems A Bayesian framework for reinforcement learning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:45.290247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:37.465247Z digest=sha256:ff5de795bd63f2373b17329e554c229f11f1e96964b073ce87dd5c28f94718cc

Observation 87f047f3-9fbf-4cf9-b538-640ffa5b6c2a · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:37.537997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:37.537997Z digest=sha256:8e8963073e136fe6d0aef02ce10f62d5e7c34e6d5dd80909c303fe6975b4b11d

Observation 2c3ffde1-9c2f-435d-84b1-6a78dc4c7403 · outbound

This paper cites Integrated architectures for learning, planning, and reacting based on approximating dynamic programming.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Integrated architectures for learning, planning, and reacting based on approximating dynamic programming

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:45.084955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:37.666221Z digest=sha256:abbdacd4cc6883be6fea49bfd371f986add68dda5a554be207f1cbef5decc69d

Observation e35a777f-eb96-4192-b9e3-aacbbb0f762c · outbound

This paper cites Dyna, an integrated architecture for learning, planning, and reacting.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Dyna, an integrated architecture for learning, planning, and reacting

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:37.794791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:37.794791Z digest=sha256:018cbe462f6eea2140cf3b26051c68d1762fef5a4f1e6a8a9cebcff91c60236c

Observation 58d2fbfd-c0df-4dcc-94c2-ab0efdb77019 · outbound

This paper cites Introduction to Reinforcement Learning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Introduction to Reinforcement Learning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:44.912222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:37.907356Z digest=sha256:699427f4c6c1d7ea011093b8697fe3c34ed3715948b75039b65d9facbc0e552f

Observation 028339c6-93e4-4e83-8fe3-50410e086460 · outbound

This paper cites Agnostic KWIK learning and Efficient Approximate Reinforcement Learning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Agnostic KWIK learning and Efficient Approximate Reinforcement Learning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:44.696029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:38.027925Z digest=sha256:53ba9fd5988dcb8768a1b1f3e0502927635f7fbdd2c19a577ef8dc1a5b067548

Observation 9f2ec0af-3f1c-48dd-9e1d-f7c7ec585536 · outbound

This paper cites Active exploration in dynamic environments.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Active exploration in dynamic environments

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:44.516965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:38.110372Z digest=sha256:2d6b95c4da1eb1e2b0d8063d8b48e46f15374b8d22c4fbde574b582693851592

Observation 3af49d99-a977-495d-b799-48b19f383d99 · outbound

This paper cites Exploring Compact Reinforcement-Learning Representations with Linear Regression.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Exploring Compact Reinforcement-Learning Representations with Linear Regression

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:44.353743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:38.219036Z digest=sha256:f08334c8541176402132edc7b284e015a44f9aa6a92f7c959d9a90fa6fac9a1f

Observation 777acb38-9e0d-463b-9206-3e95660ebe3c · outbound

This paper cites Scientific Discovery in the Age of Artificial Intelligence.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Scientific Discovery in the Age of Artificial Intelligence

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:44.215739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:38.326222Z digest=sha256:98489d1f2f75c3611416669bf39dcc85e5fb6c4bab3dbc727457e341f5d4c06e

Observation b6f8b0d4-7ab8-4cd1-8299-3f881701ac04 · outbound

This paper cites Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:38.414916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:38.414916Z digest=sha256:56930f2ce0c33ac659866904015cdd5d27ee1efbd82b99d53fbaef984c2f5ea0

Observation 92d3dbcd-db87-46df-84ab-0cd4783b1d68 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Chain-of-thought prompting elicits reasoning in large language models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:38.504795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:38.504795Z digest=sha256:201a794898bd2f3e77decd11636a007a40503d7bfc1c41298a75a025052b6a8a

Observation eabe149f-13cf-467b-9b42-111b84afa1ed · outbound

This paper cites An explanation of in-context learning as implicit Bayesian inference.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems An explanation of in-context learning as implicit Bayesian inference

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:44.071384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:38.625895Z digest=sha256:1be8828203690fc80341a77e1a9288096d9eaae0914ba0005a86d3eddb9fdc8f

Observation ba0347c7-6571-4e20-9d66-1b705672e0b1 · outbound

This paper cites Piantadosi.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Piantadosi

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:43.980017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:38.740427Z digest=sha256:3440a6cda9e3dea831f164042a1a4dbdcc332218ad643972406750e22718c207

Observation c6fa9a29-8b5d-4f61-90c6-ee414cc17cf0 · outbound

This paper cites On Benchmarking Human-Like Intelligence in Machines.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems On Benchmarking Human-Like Intelligence in Machines

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:38.848439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:38.848439Z digest=sha256:7b137ba419a60cbf54f3d441133a7d3c279cb8cd6239864baa89c90e9eea0ae5

Observation 1d8b21e7-8c45-47e4-8a79-bc16e3fb95a8 · outbound

This paper cites People use fast, goal-directed simulation to reason about novel games.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems People use fast, goal-directed simulation to reason about novel games

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:38.988889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:38.988889Z digest=sha256:28d6e8340729aa0e870169930f2736d1a0854cafb25bd537e14d84983064e355

Observation 5d74f7bd-c2f4-43c7-943a-c49c357804a9 · outbound

This paper cites Eliciting the Priors of Large Language Models using Iterated In-Context Learning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Eliciting the Priors of Large Language Models using Iterated In-Context Learning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:39.080062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:39.080062Z digest=sha256:a0f461786b88087584f9ed395606eb8e069f385f4c6a31baf14ec6f9d0102e13

Observation cd18fa2d-5a8c-4ce0-b329-e413ad30db57 · outbound

This paper cites Incoherent Probability Judgments in Large Language Models.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Incoherent Probability Judgments in Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:39.192857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:39.192857Z digest=sha256:b5f7bf1f052cd47e67cc407f9c535a5d227a18d1e350955249de5a25089fd14f

Observation fdf9af22-6755-463d-9fae-66dd7bb25d63 · outbound

This paper cites Provide a *thorough reasoning* before performing the action.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Provide a *thorough reasoning* before performing the action

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:43.903505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:39.315535Z digest=sha256:d5849c45a3c8ccbf4e0be28c35eb701478a6f8c77a37e88360470d11f9c430a9

Observation 1b26e766-d0ee-4bc4-9852-8428de509d9a · outbound

This paper cites [5 point].

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems [5 point]

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:43.769540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:39.414387Z digest=sha256:87f3c336cfcff86d31ac1333c1cc064899483b5233cadc2c773f8a274ec7ca53

Observation 402086e6-04fc-44ae-b9b7-d37854a3b984 · outbound

This paper cites You will then output a score based on a set of assessment criteria.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems You will then output a score based on a set of assessment criteria

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:43.565586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:39.489262Z digest=sha256:863913ff1955d6861602caf6ef2b21e6a095856558ed72ae2e4f858b3a971b26

Observation 2bd083ca-20b8-4d2e-ab63-70b2b9c2bba2 · outbound

This paper cites [3 points].

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems [3 points]

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:43.409550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:39.578875Z digest=sha256:a60e78d77a15a3ab5aa629496e134dac027e122d77fa69933590cb517dff7441

Observation 02f91f73-9274-43e4-bdab-e63008affb85 · outbound

This paper cites an unresolved cited work.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:40:43.238105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:39.698105Z digest=sha256:36ab0ff062f481d9de8b78be62a609edcbe7828057d4c97eba386d86f4e8bcac

Observation e0ab5211-cae1-4849-b8df-a67c01c50722 · outbound

This paper cites an unresolved cited work.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:40:43.073924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:39.790190Z digest=sha256:f74828ee84544286faec8d30c26016b8be2cd8d50abeb93a8274197c1c524a5f

Observation f8de0d1e-0dd6-4bfa-82fc-2dc4efb0b8c7 · outbound

This paper cites (Note that there will be multiple a_i 's.).

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems (Note that there will be multiple a_i 's.)

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:42.903921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:39.864719Z digest=sha256:a2dc3a4888f804ed0f4038dc106f9ee4e2062db017607a7139bb29313cde1b9b

Observation 7dde8e8c-00b8-4453-9b44-b0a605a0a75b · outbound

This paper cites an unresolved cited work.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:40:42.815299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:39.964515Z digest=sha256:8d2d1a06f808c6d9679f1e485df82aeaf71e5747630c04d9a462033f4da1d24d

Observation 091783d3-1335-4dc9-b82b-376aa5d26901 · outbound

This paper cites an unresolved cited work.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:40:42.686738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:40.059543Z digest=sha256:64791c8f507a2e12a55c9645c8dad8e08dd7cac9416db87d21b4a33b4e74ebab

Observation 1e0f90f9-c980-41b2-b8ff-5de72fd2854d · outbound

This paper cites The score for this bullet should be the accuracy percentage times the total allocated 6 points [6 points].

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems The score for this bullet should be the accuracy percentage times the total allocated 6 points [6 points]

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:42.535878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:40.163623Z digest=sha256:e8b60c29865fccbf513145174f62fe05a576915563042e0b364eb29e1d4905ae

Observation 5c06d16d-e10c-4406-b32f-7e9fbd3d8dd9 · outbound

This paper cites an unresolved cited work.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:40:42.440082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:40.257316Z digest=sha256:23d349774f94be4d27bc46f8e75c242816160f8c45d23e20003e7c6584ca65d6

Observation 1161a586-9ef7-4cc8-9640-82a58d2a0058 · outbound

This paper cites Win by connecting 3 stones in a column.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Win by connecting 3 stones in a column

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:42.304484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:40.338730Z digest=sha256:b1a6d7cd9ca38f28942b84db687cc1f8fc8342da908b2659143fa2194e546f1d

Observation a46dffa2-aae7-4f09-8573-a72c4beaa0a2 · outbound

This paper cites an unresolved cited work.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work

Reference 96

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T14:40:42.165301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:40.418830Z digest=sha256:bad62bf8323d1a00be5c193bf7cad422144f7172a6570e9e0849181c57d5e697

Observation aa59d39a-2c8a-4d6f-8c50-875550e5f662 · outbound

This paper cites an unresolved cited work.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:40:41.981797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:40.525674Z digest=sha256:975f4ca951c5480f2675dd56a5fb1422f4ba03881e4503a05630ab613237620f

Observation 93d5b542-c867-4482-b680-55190bd4d052 · outbound

This paper cites an unresolved cited work.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:40:41.848711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:40.632213Z digest=sha256:6c53f4ae1a0f9615f752707b6e2b92acb1c33cd884815081508ebffbb92e37e4

Observation 99ce0b0d-b0af-4510-9b0a-d475301ec7dc · outbound

This paper cites AAA", "BBB.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems AAA", "BBB

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:41.698993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:40:40.704557Z digest=sha256:e9cfd06e7c5638b77597765b0344c366dc29931ab6378278f1dfe793227b4bc4

Pith citing papers

Observation 3a5a7a26-a078-46cd-b7f9-9315ee619eec · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems

Reference 293

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:23:15.430452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:edda623ef7f7ebb15438e8a77d415d4143adac80bd40e3f316299df2c7f6769d

Observation f0ecc73d-9b18-4cec-9108-f6ee09342aa6 · inbound

CausalGame: Benchmarking Causal Thinking of LLM Agents in Games cites this paper.

CausalGame: Benchmarking Causal Thinking of LLM Agents in Games Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T20:19:27.650696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:19:27.650696Z digest=sha256:01baf4ac23c694ebbbdfe6d984393001c51e8f2c4870fcb79eaee5876842116e

Observation 5bddc95b-cf75-4d8b-bd2c-3374077b922a · inbound

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents cites this paper.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:06:37.956842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:5ab22d73227479a136896da257b4a081558950bf4451e60e4c779ae082aff49a

Observation a46cf61d-ded2-48af-ba09-31842ab7a8c3 · inbound

DiG-bench: Discovery in Games cites this paper.

DiG-bench: Discovery in Games Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:08:35.185016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:08:35.185016Z digest=sha256:2131149d5905062dcbba440562ea0dfcb4d51687cebe8268a6b6f7a665fe0df0