To content
Fakultät für Informatik
Master Seminar

AutoAI and AutoResearch: Automated Design of Models, Agents, Algorithms, and Scientific Workflows

Seminar description

Large language models are increasingly embedded in agentic systems that address complex tasks in machine learning and science, including model development, experiment design, algorithm discovery, idea generation, literature analysis, biological design, and laboratory planning. In these systems, an LLM is not used merely as a stand-alone predictor. Instead, it operates within an agentic harness that provides tools, memory, prompts, roles, communication structures, evaluators, and iterative feedback.

Designing these harnesses is currently a substantial human engineering effort. Given an executable representation of a system, suitable benchmarks or evaluators, and sufficient computational resources, however, parts of this design process can themselves be automated. Prompts can be optimized through reflection or evolution; agent roles and communication topologies can be searched automatically; optimization and machine-learning algorithms can be generated as code; and entire scientific workflows can be organized by autonomous research agents.

This seminar studies the emerging intersection of AutoAI, AutoML for and with agents, automated agent-system design, algorithm discovery, and AutoResearch. It covers several levels of automation:

  • optimization of prompts and agent instructions;
  • automatic design of agent workflows and multi-agent topologies;
  • evolution of programs, optimizers, and machine-learning algorithms;
  • self-modifying and recursively improving agents;
  • autonomous hypothesis generation, experimentation, analysis, and scientific writing;
  • applications in mathematics, physics, machine learning, and biological design.

Particular attention will be paid to prominent success stories such as FunSearch, AlphaEvolve, automated AI research, autonomous quantum-physics research, and agentic biological design. At the same time, the seminar will critically examine the empirical foundations of the field. Recent studies show that multi-agent systems can suffer from specification errors, miscommunication, ineffective verification, high variance, benchmark overfitting, and unnecessary complexity. In some settings, carefully constructed single-agent systems, conventional AutoML, majority voting, or simple sampling baselines remain competitive with—or outperform—elaborate agentic systems.

The central questions of the seminar are therefore not only whether AI can automate parts of AI and scientific research, but also:

  • What exactly is being automated: prompts, workflows, code, algorithms, hypotheses, experiments, or evaluation?
  • Which parts of the system and search space remain designed by humans?
  • What feedback signals and evaluators make autonomous improvement possible?
  • Do discovered systems generalize beyond the benchmarks on which they were optimized?
  • How should compute cost, inference cost, variance, and failed runs be reported?
  • How can we distinguish genuine discovery from benchmark overfitting, data contamination, or rediscovery?
  • When are multiple agents beneficial, and when do they merely add cost and failure modes?
  • What forms of human oversight, reproducibility, scientific integrity, and safety are required?

Students will independently study recent research literature, present and critically discuss a selected topic, compare competing methodological perspectives, and prepare a written seminar paper. Where feasible, seminar papers may include a small reproduction study, an ablation, a comparison with a simple baseline, or a critical examination of an evaluation protocol.

Preliminary Reading List

  • Trirat et al. — AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoML, ICML 2025
  • Fernando et al. — PromptBreeder: Self-Referential Self-Improvement via Prompt Evolution, ICML 2024
  • Agrawal et al. — GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning, ICLR 2026
  • Büssing et al. — MO-CAPO: Multi-Objective Cost-Aware Prompt Optimization, arXiv 2026
  • Zhou et al. — Multi-Agent Design: Optimizing Agents with Better Prompts and Topologies, ICLR 2026
  • Hu et al. — EvoMAS: Evolutionary Generation of Multi-Agent Systems, ICML 2026
  • Zhang et al. — Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents, ICLR 2026
  • Zhang et al. — Hyperagents, arXiv 2026
  • Romera-Paredes et al. — Mathematical Discoveries from Program Search with Large Language Models, Nature 2024
  • Novikov et al. — AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery, arXiv 2025
  • Lange et al. — ShinkaEvolve: Towards Open-Ended and Sample-Efficient Program Evolution, ICLR 2026
  • Li et al. — LLaMEA-BO: A Large Language Model Evolutionary Algorithm for Automatically Generating Bayesian Optimization Algorithms, GECCO 2026
  • Goldie et al. — DiscoGen: Procedural Generation of Algorithm Discovery Tasks in Machine Learning, ICML 2026
  • Yamada et al. — The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search, arXiv 2025
  • Lu et al. — Towards End-to-End Automation of AI Research, Nature 2026
  • Arlt et al. — Towards Autonomous Quantum Physics Research Using LLM Agents with Access to Intelligent Tools, arXiv 2025
  • Maus et al. — Purely Agentic Black-Box Optimization for Biological Design, arXiv 2026
  • Cemri et al. — Why Do Multi-Agent LLM Systems Fail?, arXiv 2025
  • Gideoni, Risi, and Gal — Simple Baselines Are Competitive with Code Evolution, arXiv 2026

Registration

Registration via email to Matthias Feurer and Rean Fernandes