In short
Lu et al. (Nature, March 2026) present The AI Scientist, an agentic system that executes the entire machine learning research cycle without human intervention: hypothesis generation, code implementation, experiment execution, result analysis, manuscript writing, and evaluation by an automated reviewer. Across three submissions to an ICLR 2025 workshop, one paper received an average score of 6.33 out of 10 (above the acceptance threshold), before being withdrawn according to the pre-established protocol. The other two did not reach the threshold.
In short: imagine a miniature research lab running with no one inside. It poses a question, writes the experiment’s code, runs the computations, reads the results, drafts the article, and even grades itself the way a scientific committee would. All in a single day, for a few dollars of compute. Out of three submitted articles, one cleared the bar — one out of three is not zero, but it is not Nature either.
What the system concretely does
The AI Scientist operates in four sequential phases.
Phase 1 — Ideation
The system maintains a growing archive of research ideas within a machine learning subdomain provided by the operator. For each idea, it generates a title, a central hypothesis, and an experimental plan, with self-evaluated scores for feasibility, interest, and novelty (scale 1–10). Novelty is automatically verified via the Semantic Scholar API: ideas too close to existing work are filtered out before advancing in the pipeline.
Phase 2 — Experimentation
Two modes coexist. In template-based mode, the system receives an existing code template (for example, a transformer training run on Shakespeare) and modifies it to test the hypothesis. In template-free mode, it generates code from scratch and drives a parallelized four-stage tree search: preliminary investigation, hyperparameter tuning, execution of the research agenda, ablation studies. Each node in the tree corresponds to a Python script executed in real time. Execution errors generate debugging nodes; produced visualizations are critiqued by a multimodal model (GPT-4o) to detect mislabeled axes or misleading figures.
Generating a complete paper in template-free mode takes anywhere from several hours to more than 15 hours depending on complexity.
Phase 3 — Writing
The manuscript is produced section by section from a conference LaTeX template. The system queries Semantic Scholar over 20 rounds to build the “related work” section and textually justifies each citation before inserting it. Multiple automated proofreading passes improve consistency. LaTeX compilation errors are corrected automatically (up to five cycles).
Phase 4 — Automated evaluation
The Automated Reviewer simulates the NeurIPS review process: it produces five independent reviews and then a meta-review (area chair role) giving a binary accept/reject decision and numerical scores (rigor, presentation, contribution). Validated on 1,000 ICLR papers (2017–2024): balanced accuracy of 69%, compared to 66% for human reviewers according to the NeurIPS 2021 consistency experiment. On post-cutoff papers (2025, absent from training), accuracy falls to 66%, which remains comparable to humans.
In short: the four phases reproduce the usual rhythm of a research project, but in a 5-to-15-hour cycle instead of several months. Phase 2 (experimentation) is the most expensive — the search tree can have up to 30 nodes, each node is a Python script that runs, fails, debugs, retries. The code is not written “blindly”: a multimodal model critiques the produced graphics to spot mislabeled axes.
The real peer review test
The central experiment in the paper consists of submitting three fully generated manuscripts to the ICBINB workshop (negative results in deep learning) at ICLR 2025 — with prior agreement from the organizers and an ethics committee (IRB, UBC). Reviewers knew some submissions were AI-generated, without knowing which ones. The workshop had an acceptance rate of 70%, compared to 32% for the main ICLR 2025 conference.
One paper received scores of 6, 7, and 6 (average: 6.33), above the workshop’s average acceptance threshold. It covered a negative result — the workshop’s central theme. According to the organizers, it would have been accepted without the pre-established withdrawal protocol. The other two papers did not reach the threshold.
The team of human researchers who reviewed the three papers internally concluded that none met the standard for publication at the main ICLR conference.
In short: “1 out of 3 accepted” sounds impressive, but context tempers it. The ICBINB workshop accepted 70% of human submissions (very lenient), not 32% like the main conference. And the Nature team itself judged that none of the three papers held up at the main ICLR standard. The bar cleared is the bar of a benevolent workshop — not the bar of a serious scientific publication.
Quality correlation with model and compute
The paper documents two trends measured by the Automated Reviewer:
- Quality as a function of the base model: significant correlation (R² = 0.517, p < 0.00001) between the release date of the underlying model and the quality of generated papers. GPT-4 (2023) produces scores around 2; Gemini 2.5 and recent models approach 5–6. Models tested include GPT-4, GPT-4o, o1, o3, Claude Sonnet 3/3.5/3.7/4, Gemini 1.5/2.0/2.5.
- Quality as a function of compute budget: the more nodes allocated to tree search (from 5 to 30 nodes), the more scores improve, with a clear trend up to n = 30 (the tested limit).
What resists automation
The paper explicitly lists observed failure modes:
- Naive or insufficiently developed ideas: the system generates plausible but sometimes shallow hypotheses.
- Incorrect implementation: the code may not precisely realize the proposed idea.
- Lack of methodological rigor: some papers present experiments without sufficient ablations or adequate comparisons.
- Implementation errors: undetected bugs affecting the validity of results.
- Bibliographic hallucinations: inaccurate or nonexistent citations.
- Layout problems: figures duplicated between the body and appendix.
The authors emphasize that the question of “conceptual leaps” — truly novel ideas that change a field — remains open. The system produces variations and extensions of existing work, not paradigm-shifting breakthroughs.
Furthermore, the system is limited to computational experiments. Extension to other disciplines (chemistry, biology) requires automated laboratories that are advancing but remain costly.
Identified ethical concerns
The authors identify five main risks: saturation of the peer review system, artificial inflation of academic credits, unattributed use of others’ ideas, displacement of researcher jobs, and conduct of unethical experiments. The paper stresses the absence of established disclosure norms in the scientific community, and treats the systematic withdrawal of submissions as a necessary precaution in this transitional state.
Key takeaways
- The AI Scientist automates the entire machine learning research cycle: one generated paper passed the first round of peer review at ICLR 2025 (workshop, 70% acceptance rate).
- The quality of produced papers increases measurably with the quality of the underlying model and the allocated compute budget.
- An automated reviewer achieves balanced accuracy comparable to human reviewers (66–69%) on ICLR data.
- Current limitations are structural: shallow ideas, hallucinations, lack of rigor on difficult cases — and none of the three submitted papers met the standard for a main conference.
- Extension to experimental domains outside computer science (chemistry, biology) is ongoing, contingent on the availability of automated laboratories.