Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:40:40.704557Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 3 inbound Pith citation observations for arXiv:2505.17968.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:40:40.704557Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-11T20:19:27.650696Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T00:06:37.955246Z
99 of 99 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 82402fab-ecce-4773-907f-cd8bb8afa647 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Large-Scale Bandit Problems and KWIK Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50c39405-4d30-48ca-9495-442ea9c4d0ee · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Queries and Concept Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57f7fe63-fce0-45b3-bd70-3443d251a9f7 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Inductive Inference: Theory and Methods
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d30d95ee-ceca-4221-a946-e5cbf8dc69cb · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Claude 3.5 sonnet
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9693341c-04d0-41a0-9cb9-4c3f94395333 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Toward Efficient Exploration by Large Language Model Agents
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcefdc3a-515d-4d59-93b9-9c562e09f279 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems A Markovian Decision Process
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc58a8ce-a78f-466e-a746-4cebf820daae · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Using cognitive psychology to understand GPT-3
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d291f60f-f390-4fdc-8cb9-c8ac0b12fa8c · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Variational inference: A review for statisticians
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d22695ee-a3ea-41f9-8dd8-bf2b9b39962b · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems R-MAX – A General Polynomial Time Algorithm for Near-Optimal Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bc69ef7-5ca6-44d8-8eaf-0c4f7ce9e93c · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Language models are few-shot learners
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43339939-93b2-46e5-a19c-abb7de83b6e5 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Why Do Multi-Agent LLM Systems Fail?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ce3290c-0b2a-44f6-8b06-4dd45832900d · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Bayesian Experimental Design: A Review.Statistical Science, pp
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f6f2e9b-8ebe-4a0e-80b4-0d4cc2750171 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51fed5c1-5bc9-4750-9138-b24393c37def · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems The first crank of the cultural ratchet: Learning and transmitting concepts through language
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 573ec285-8443-4393-8478-1bca2ad57f2b · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems CogBench: a large language model walks into a psychology lab
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a02f0fd1-c699-42df-a7ca-9513f1fda469 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c841fc75-f7d6-4854-b798-1586206b5ac9 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Uncertainty, Information, and Sequential Experiments
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65f83589-7dfb-43e6-8e21-a237ddcfcf42 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems PILCO: A Model-Based and Data-Efficient Approach to Policy Search
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e2b302c-9af2-4c54-9c27-023bc5314e69 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Aleatory or Epistemic? Does it Matter? Structural Safety, 31(2):105–112, 2009
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f610df2c-264d-4db8-a8fc-ce0f5c8e817a · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Is In-Context Learning in Large Language Models Bayesian? A Martingale Perspective
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 377321d7-611e-4dd4-93f7-db14bbafc4ee · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Variational Bayesian optimal experimental design
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a31fd6a-78f1-4d59-baae-dfc5895de5c9 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Baby steps in evaluating the capacities of large language models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d67b3eaf-3302-4617-a054-1155f531c76c · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems BoxingGym: Benchmarking progress in automated experimental design and model discovery
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a177777-7675-4fca-8622-89383c8e4bff · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Amplify scientific discovery with artificial intelligence
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aaa365c-eb5e-49a1-963f-e3f9dac601e1 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems ANOVA: Repeated measures
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b96652ac-dc3f-48b4-b0cb-b936df006f2b · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Towards an AI co-scientist
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d509de6-b5d1-4fe6-a570-73c8714ad341 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems The Llama 3 Herd of Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e94454b9-3bf0-4669-96f1-ec0efc5c3a10 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Bayes in the Age of Intelligent Machines
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 796414a5-9386-45a3-9c79-e35aefff5510 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 997b1a60-ef79-45a8-81af-4733c6156182 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1d76051-ac80-4a40-a808-7ee91b4823a7 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems GPT-4o System Card
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7974c83d-b72e-4146-a36e-208df11b3aa8 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems What Do Learning Dynamics Reveal About Generalization in LLM Reasoning?
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81194d79-d3a4-47e4-882b-f0d6797a95f9 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Using the tools of cognitive science to understand large language models at different levels of analysis
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a170178-e8a4-4ea6-adee-ed86d9bcc9b9 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems A robust class of context- sensitive languages
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1d339c4b-a002-4688-9b16-c4a20db7b5e8 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Passive learning of active causal strategies in agents and language models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 08d6ff58-e230-40a8-8229-997dc742e396 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Structured chain-of-thought prompting for code generation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2396efc4-279d-4b28-9905-b61235b5033b · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Reducing Reinforcement Learning to KWIK Online Regression
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4d7ba1ae-a463-4956-8e4a-de594c96296b · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Knows What It Knows: A Framework for Self-Aware Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 35ebf653-8272-48c8-8b50-723822d32714 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems On a measure of the information provided by an experiment
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e304cd7-3ff4-4750-91d1-b9ed954c52af · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Learning quickly when irrelevant attributes abound: A new linear-threshold algorithm
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1ce2be8f-e194-42a5-87d0-115a6f41818f · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Decoupling Exploration and Exploitation for Meta-Reinforcement Learning Without Sacrifices
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e3fe3d14-33ed-45ca-8b75-7ec14e8b2bc8 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Large Language Models Assume People are More Rational than We Really are
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8884055-f106-44e7-a599-64123dd7f2b3 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0913bfcd-5527-4056-921c-763140c1f2b2 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c49abd94-62aa-43cb-8152-c385fdd9970a · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Deconstructing Long Chain-of-Thought: A Structured Reasoning Optimization Framework for Long CoT Distillation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 676edd7b-4582-4f39-86de-04ef2f90fe85 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Category learning through active sampling
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6c60ac6e-d798-4f3c-8d5c-a7c133cbac82 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Is it better to select or to receive? learning via active and passive hypothesis testing
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1841677b-fbc5-4fad-ac05-f55ef2ea0065 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Modeling rapid language learning by distilling Bayesian priors into artificial neural networks
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05f80f70-7c1a-4c6d-a7a7-c357f7027ba9 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Embers of autoregression show how large language models are shaped by the problem they are trained to solve
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e3be7c28-fb95-48a7-a078-4fd94a17d26a · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems MatPilot: an LLM-enabled AI Materials Scientist under the Framework of Human-Machine Collaboration
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caa38cd3-48dd-4eef-81c8-b9f9bf08921f · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Sparks of Science: Hypothesis Generation Using Structured Paper Data
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c36647d-194f-4806-981d-f738f5196c12 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems (More) Efficient Reinforcement Learning via Posterior Sampling
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 038b21dc-64a8-4b1e-978d-80f4a4284dc1 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Puterman
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 844a8a34-8b35-471e-8b6e-faea4f154c0c · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Towards scientific discovery with generative ai: Progress, opportunities, and challenges
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 61eec6b9-b181-4567-a92f-b729e4834598 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Diversity-Based Inference of Finite Automata
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e096bed7-de24-4dad-a5fe-528ebfa59b43 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Inference of Finite Automata Using Homing Sequences
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 391f4971-cca2-42ce-b42e-b7e909f1f4f5 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Jagadish, Marvin Mathony, Tobias Ludwig, and Eric Schulz
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac32d513-3360-4f9d-a0a6-e992cb0f6269 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Symbolic metaprogram search improves learning efficiency and explains rule learning in humans
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b397f9fd-5cb7-4c33-a49c-73948ca38d19 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Trading off Mistakes and Don’t- Know Predictions
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c0a039d9-d524-4528-b9c8-acf3dc5d5e92 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Agent Laboratory: Using LLM Agents as Research Assistants
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c74deb9-8ff3-4bd5-8068-21cb2a079a07 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Active Learning Literature Survey
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 439cc0a6-169c-4dfb-a18c-432cc02b15ec · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Llm-sr: Scientific equation discovery via programming with large language models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 28738de6-bb3c-4bfb-8ab9-454ce709f0e5 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7290d399-1b29-42ab-bc4f-f3030c56f46e · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 971622c1-4ba9-46c2-ab12-73b9b13d43ba · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems PaperBench: Evaluating AI's Ability to Replicate AI Research
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae21c6d6-7306-4574-a60b-247ea474c416 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems An Analysis of Model-based interval estimation for Markov Decision Processes
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 35f9fa0f-d275-4436-88e0-646ebed51a86 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems A Bayesian framework for reinforcement learning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 87f047f3-9fbf-4cf9-b538-640ffa5b6c2a · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c3ffde1-9c2f-435d-84b1-6a78dc4c7403 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e35a777f-eb96-4192-b9e3-aacbbb0f762c · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Dyna, an integrated architecture for learning, planning, and reacting
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58d2fbfd-c0df-4dcc-94c2-ab0efdb77019 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Introduction to Reinforcement Learning
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 028339c6-93e4-4e83-8fe3-50410e086460 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Agnostic KWIK learning and Efficient Approximate Reinforcement Learning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9f2ec0af-3f1c-48dd-9e1d-f7c7ec585536 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Active exploration in dynamic environments
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3af49d99-a977-495d-b799-48b19f383d99 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Exploring Compact Reinforcement-Learning Representations with Linear Regression
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 777acb38-9e0d-463b-9206-3e95660ebe3c · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Scientific Discovery in the Age of Artificial Intelligence
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b6f8b0d4-7ab8-4cd1-8299-3f881701ac04 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92d3dbcd-db87-46df-84ab-0cd4783b1d68 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Chain-of-thought prompting elicits reasoning in large language models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eabe149f-13cf-467b-9b42-111b84afa1ed · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems An explanation of in-context learning as implicit Bayesian inference
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ba0347c7-6571-4e20-9d66-1b705672e0b1 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Piantadosi
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c6fa9a29-8b5d-4f61-90c6-ee414cc17cf0 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems On Benchmarking Human-Like Intelligence in Machines
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d8b21e7-8c45-47e4-8a79-bc16e3fb95a8 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems People use fast, goal-directed simulation to reason about novel games
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d74f7bd-c2f4-43c7-943a-c49c357804a9 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Eliciting the Priors of Large Language Models using Iterated In-Context Learning
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd18fa2d-5a8c-4ce0-b329-e413ad30db57 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Incoherent Probability Judgments in Large Language Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdf9af22-6755-463d-9fae-66dd7bb25d63 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Provide a *thorough reasoning* before performing the action
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1b26e766-d0ee-4bc4-9852-8428de509d9a · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems [5 point]
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 402086e6-04fc-44ae-b9b7-d37854a3b984 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems You will then output a score based on a set of assessment criteria
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2bd083ca-20b8-4d2e-ab63-70b2b9c2bba2 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems [3 points]
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 02f91f73-9274-43e4-bdab-e63008affb85 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e0ab5211-cae1-4849-b8df-a67c01c50722 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f8de0d1e-0dd6-4bfa-82fc-2dc4efb0b8c7 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems (Note that there will be multiple a_i 's.)
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7dde8e8c-00b8-4453-9b44-b0a605a0a75b · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 091783d3-1335-4dc9-b82b-376aa5d26901 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1e0f90f9-c980-41b2-b8ff-5de72fd2854d · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems The score for this bullet should be the accuracy percentage times the total allocated 6 points [6 points]
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5c06d16d-e10c-4406-b32f-7e9fbd3d8dd9 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1161a586-9ef7-4cc8-9640-82a58d2a0058 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Win by connecting 3 stones in a column
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a46dffa2-aae7-4f09-8573-a72c4beaa0a2 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation aa59d39a-2c8a-4d6f-8c50-875550e5f662 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 93d5b542-c867-4482-b680-55190bd4d052 · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 99ce0b0d-b0af-4510-9b0a-d475301ec7dc · outbound
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems AAA", "BBB
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3a5a7a26-a078-46cd-b7f9-9315ee619eec · inbound
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems
Reference 293
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f0ecc73d-9b18-4cec-9108-f6ee09342aa6 · inbound
CausalGame: Benchmarking Causal Thinking of LLM Agents in Games Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bddc95b-cf75-4d8b-bd2c-3374077b922a · inbound
Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.