Pith. sign in

Paper Citation Record · LEDGER

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review

As of 10 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2607.28631.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28631 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:54:35.911023Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 54804b4f-9eb0-45f0-b60f-e6bfadbd4119 · outbound

This paper cites LitLLM: A Toolkit for Scientific Literature Review.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review LitLLM: A Toolkit for Scientific Literature Review

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:33.811783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:33.811783Z digest=sha256:227e028fc91f83898ce06cf4590b5893fd5a10b78021858126a8f8617ddf8339

Observation 3948c35b-512d-4fc0-802e-fad64cf328b4 · outbound

This paper cites Towards an AI co-scientist.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Towards an AI co-scientist

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.085990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.085990Z digest=sha256:52321461e02c233331cb23326758e43a261daf27c0b59968f45776d1ed95c9e5

Observation 7b4f0955-c022-4461-901f-79f3b7df63b2 · outbound

This paper cites Automated Algorithm Selection: Survey and Perspectives.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Automated Algorithm Selection: Survey and Perspectives

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.368864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.368864Z digest=sha256:b0c04205e8f2165ccfdc96f2c982df196df9cd111cb0d17b48bf661e15aa1130

Observation 856ba543-d7a2-4d26-bf19-60ee6784b964 · outbound

This paper cites ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.900318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.900318Z digest=sha256:6afb23a032ef01f7ab8f4de2080daf05b27d98f9b5af173e555aa7f9d963c87e

Observation 0fb7e291-34af-482a-83d1-dc3f9e634416 · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.012634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.012634Z digest=sha256:14cc43540147578e22797e3e862bca38e8159ed4ea59ac6364907cd01cd06f4b

Observation e718b5ec-23ee-477c-b25e-534170c877eb · outbound

This paper cites AlphaEvolve: A coding agent for scientific and algorithmic discovery.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review AlphaEvolve: A coding agent for scientific and algorithmic discovery

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.107501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.107501Z digest=sha256:6879857bd8307ef83e50614db4f1fa0ad2674008b7b87e5b9a136e8fba87f7de

Observation ef9c0efc-42a4-4cdb-952c-0dc5b32b13fc · outbound

This paper cites S., Bartley, N., Urbanowicz, R.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review S., Bartley, N., Urbanowicz, R

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.227993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.227993Z digest=sha256:a73cf6cdbb2db29c696ba5927d13263fa5b17d5f1888ea6f53537d1c0a3138e9

Observation 38d1699a-f476-4fd9-8792-b7a0a14e5d22 · outbound

This paper cites Physics Supernova: AI Agent Matches Elite Gold Medalists at IPhO 2025.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Physics Supernova: AI Agent Matches Elite Gold Medalists at IPhO 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.281409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.281409Z digest=sha256:1b2944cedd7ace2877cd2fa077c0f2094d5048a22f3e2bc1b142871b35cc38aa

Observation a17570f3-d27a-4ef4-894a-7d79718cf68c · outbound

This paper cites Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.499762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.499762Z digest=sha256:f663cc2bc7f3d3bd06a3f9f4c9bf2b5f95972dbc79e3600ac3bc72f45e8cc39d

Observation 7a715175-85ea-4ed5-80a9-87f8a3baa9dc · outbound

This paper cites AI-Researcher: Autonomous Scientific Innovation.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review AI-Researcher: Autonomous Scientific Innovation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.646218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.646218Z digest=sha256:657fa9b4c198a27c3c8650b7cb016269d18daea0c0e8b916d4c8fce3a016544e

Observation 1ba4f510-578c-47c7-95db-22d733d556ac · outbound

This paper cites CycleResearcher: Improving Automated Research via Automated Review.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review CycleResearcher: Improving Automated Research via Automated Review

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.758261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.758261Z digest=sha256:136b73580727700dcc49cf7e04531869b44be97d1ad89e1ec1606007c67733b6

Observation 030a3167-4768-451a-a90b-2dc9bb6a2219 · outbound

This paper cites Neural Architecture Search with Reinforcement Learning.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Neural Architecture Search with Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.911023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.911023Z digest=sha256:020fc853db67eb720111df36ba1369ec05dcffbf25314875593bb2da931c987a

Observation 9b545b39-bfaf-41a9-af63-2d92395c6395 · outbound

This paper cites Agent Laboratory: Using LLM Agents as Research Assistants.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Agent Laboratory: Using LLM Agents as Research Assistants

Reference 1976

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.395915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.395915Z digest=sha256:87c1c49158debd55a440bf6df3495b9fab2b9fb0c54ff959055ce2891ea0b9f0

Observation c6bb8387-0b31-44ef-883d-8451320b7df3 · outbound

This paper cites The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.833406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.833406Z digest=sha256:59e809a6403106f405a2d87e524bdb0d123c1128b0e7a9e1afa6642bc7411928

Observation 6ef9b525-834e-445f-8a9d-f8e3f98ee58f · outbound

This paper cites Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.631818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.631818Z digest=sha256:08957627a3bd8c7ae25a6afb0a427818e0b6fc05ac04ca48b05c13dfbd190700

Observation 1a9add53-6d65-48c7-a0e3-f7120dbbb5ab · outbound

This paper cites Crafting papers on machine learning.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Crafting papers on machine learning

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.512917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.512917Z digest=sha256:26eebaea6ee2497d684cc7c135d988ec8f0f4ffb077cdcf390c9a51f7ff426f5

Observation 9e2a4b00-8ce0-401b-aefc-c6667a5cf512 · outbound

This paper cites DARTS: Differentiable Architecture Search.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review DARTS: Differentiable Architecture Search

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.776651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.776651Z digest=sha256:83070f74c939ad059a19cd1f9afbd8e3a5d76f87daa9dc0887d6e65a12dc39b8

Observation 76e81f91-4eb0-4b5d-b9a3-76b062e91b94 · outbound

This paper cites AgentReview: Exploring Peer Review Dynamics With LLM Agents.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review AgentReview: Exploring Peer Review Dynamics With LLM Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.236060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.236060Z digest=sha256:dcc0695a5156aa7296cafac10174028fee547fdd8cdf11b7186ef4467db3f0f3

Observation f9339de1-69b2-47e7-be8b-28c7717bcce9 · outbound

This paper cites K., Cucerzan, S., and Hwang, S.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review K., Cucerzan, S., and Hwang, S

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:33.952910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:33.952910Z digest=sha256:0a965b9b524c621b6ec7546e0926a45ae771867dfe1ea102a16c064ebce47aba

Pith citing papers

No inbound Pith citation observations are available.