Pith. sign in

Paper Citation Record · LEDGER

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review

As of 7 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2607.28631.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28631 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:54:35.911023Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 54804b4f-9eb0-45f0-b60f-e6bfadbd4119 · outbound

This paper cites LitLLM: A Toolkit for Scientific Literature Review.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review LitLLM: A Toolkit for Scientific Literature Review

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:33.811783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:33.811783Z digest=sha256:7ed7bb236235d53790131d46bd746e80295b58b1c739017522b3315633ed2367

Observation 3948c35b-512d-4fc0-802e-fad64cf328b4 · outbound

This paper cites Towards an AI co-scientist.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Towards an AI co-scientist

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.085990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.085990Z digest=sha256:8cc8abb5ef1568b303f605dd1e525e66dc83c57b3c6bc47816df8d0dc298bcd4

Observation 7b4f0955-c022-4461-901f-79f3b7df63b2 · outbound

This paper cites Automated Algorithm Selection: Survey and Perspectives.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Automated Algorithm Selection: Survey and Perspectives

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.368864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.368864Z digest=sha256:d8c08df948e9439e4a588a57bb708848d2801f1c7d75c34e919f37b7e28c45d4

Observation 856ba543-d7a2-4d26-bf19-60ee6784b964 · outbound

This paper cites ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.900318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.900318Z digest=sha256:ba425efb6759b7c56923d4c0a618dd3268e1830f60b9d213b37b51932302efdc

Observation 0fb7e291-34af-482a-83d1-dc3f9e634416 · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.012634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.012634Z digest=sha256:c0a1358251601bb751fd72dc8726d7e998cc46cd6eed6796d466a307abb3adc1

Observation e718b5ec-23ee-477c-b25e-534170c877eb · outbound

This paper cites AlphaEvolve: A coding agent for scientific and algorithmic discovery.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review AlphaEvolve: A coding agent for scientific and algorithmic discovery

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.107501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.107501Z digest=sha256:4f06bb7eac75336ff5a35d800bda584b015d3f942cedb42266adad3c611e23e5

Observation ef9c0efc-42a4-4cdb-952c-0dc5b32b13fc · outbound

This paper cites S., Bartley, N., Urbanowicz, R.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review S., Bartley, N., Urbanowicz, R

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.227993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.227993Z digest=sha256:84ddf1aef36f03738b827061fae9723024820dd166d20e0339c33b2932c8f146

Observation 38d1699a-f476-4fd9-8792-b7a0a14e5d22 · outbound

This paper cites Physics Supernova: AI Agent Matches Elite Gold Medalists at IPhO 2025.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Physics Supernova: AI Agent Matches Elite Gold Medalists at IPhO 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.281409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.281409Z digest=sha256:591150850022a23ef4788a60ea8ba5c499c80ce8c1a578532d5205fbe89d346a

Observation a17570f3-d27a-4ef4-894a-7d79718cf68c · outbound

This paper cites Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.499762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.499762Z digest=sha256:76af38eeeab90865aed1a5a3f3462d065de56f51871af279b6a074945c84ddca

Observation 7a715175-85ea-4ed5-80a9-87f8a3baa9dc · outbound

This paper cites AI-Researcher: Autonomous Scientific Innovation.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review AI-Researcher: Autonomous Scientific Innovation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.646218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.646218Z digest=sha256:feb6a9f916a8311792ed9f95155e3991d992604055e2c777348240d24dc4d98e

Observation 1ba4f510-578c-47c7-95db-22d733d556ac · outbound

This paper cites CycleResearcher: Improving Automated Research via Automated Review.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review CycleResearcher: Improving Automated Research via Automated Review

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.758261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.758261Z digest=sha256:df00854a4f012f74cf199d151843129baa9d09ede4e9c84a5d70fc55621672b9

Observation 030a3167-4768-451a-a90b-2dc9bb6a2219 · outbound

This paper cites Neural Architecture Search with Reinforcement Learning.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Neural Architecture Search with Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.911023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.911023Z digest=sha256:a697190506aec87237804e82850ca6401ed164727bcda34f5488fab72bd2b6a4

Observation 9b545b39-bfaf-41a9-af63-2d92395c6395 · outbound

This paper cites Agent Laboratory: Using LLM Agents as Research Assistants.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Agent Laboratory: Using LLM Agents as Research Assistants

Reference 1976

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.395915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.395915Z digest=sha256:094593432143eea3f9edf4b8b468986caa147c999bea32eab25593adfa6646f6

Observation c6bb8387-0b31-44ef-883d-8451320b7df3 · outbound

This paper cites The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:35.833406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:35.833406Z digest=sha256:419ba433e0db4c4131d6c5117310f7690d5424f353eb1ca827a1dc15dd33d03a

Observation 6ef9b525-834e-445f-8a9d-f8e3f98ee58f · outbound

This paper cites Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.631818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.631818Z digest=sha256:c5aa1ffa4829dfe55fb0a0fdfea49aa12a4d5eb9dd4a00b5a219234c9988a35c

Observation 1a9add53-6d65-48c7-a0e3-f7120dbbb5ab · outbound

This paper cites Crafting papers on machine learning.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review Crafting papers on machine learning

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.512917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.512917Z digest=sha256:2a75cced917edc2652a4e666c69edbc97ebadc4c5af4b3755c7158ace0de36c8

Observation 9e2a4b00-8ce0-401b-aefc-c6667a5cf512 · outbound

This paper cites DARTS: Differentiable Architecture Search.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review DARTS: Differentiable Architecture Search

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.776651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.776651Z digest=sha256:79ada05404256d740a9a5dfb75b919dce93df4eafb3da5cd147ca51bea0529e6

Observation 76e81f91-4eb0-4b5d-b9a3-76b062e91b94 · outbound

This paper cites AgentReview: Exploring Peer Review Dynamics With LLM Agents.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review AgentReview: Exploring Peer Review Dynamics With LLM Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:34.236060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:34.236060Z digest=sha256:fc38e20c4ec6c6eeb4fc6f4bc51994eab69767287775b9caaf89c7dd97047811

Observation f9339de1-69b2-47e7-be8b-28c7717bcce9 · outbound

This paper cites K., Cucerzan, S., and Hwang, S.

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review K., Cucerzan, S., and Hwang, S

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T00:54:33.952910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:54:33.952910Z digest=sha256:991918b6f00cba0f39fc4111271daef9f076e4ab5c1d217d6a650cf9713521de

Pith citing papers

No inbound Pith citation observations are available.