When AI Agents Pass Exams

On the validity of digital online tests in the age of autonomous AI systems

Background

Classical online tests are usually based on asynchronous multiple-choice tasks embedded in learning platforms and serve as indicators of competence. With the rise of autonomous AI agent systems that independently navigate learning paths and complete examinations on behalf of their users, this model is increasingly called into question. The authors emphasise that these systems no longer merely respond to prompts like chatbots, but autonomously move through complex examination scenarios and thereby severely undermine the credibility of traditional assessment formats.

Agentic AI Compared with Chatbots

Whereas conventional chatbots primarily react to user input, agentic AI systems act independently: they receive a goal-oriented task, navigate learning platforms on their own, identify the tools required and complete complex examination tasks step by step. For example, they can search online courses, scan texts and automatically log in to external systems in order to use additional data. Even security mechanisms such as CAPTCHAs can be bypassed with the help of current models.

Evidence from Practice

Faster Course Completion

In field trials conducted by the Zukunftslabor Generative KI, autonomous agents completed standardised online courses on the AI Act considerably faster than human participants. While humans require around four hours, specialised agents complete the courses in roughly 90 minutes — achieving top scores of more than 90 percent. These examples indicate that the time-intensive examination mode may become obsolete.

Transferability

Further tests demonstrated transferability: agents mastered certifications in project management (Scrum), complex theoretical examinations for drone pilot licensing and an English language test at the highest C2 level. Even complex integral problems in Moodle learning environments were solved reliably. This underscores the disruptive potential of agentic AI in education.

Dead-Loop Learning: The Process

The article describes an automated process that the authors call “Dead-Loop Learning”. The procedure can be divided into four phases:

1. Creation AI generates course content and learning paths.
2. Completion The agent independently completes tasks and examinations.
3. Validation A testing system checks the solutions and assigns points.
4. Certification A certificate is issued without any human intervention.

Because the agent both generates learning material and solves and evaluates tasks, a closed examination loop emerges in which human control is barely present. This jeopardises the evidentiary value of online tests as proof of individual competence.

Implications for Teaching and Assessment

New Competence Priorities

With the idea of “new skilling”, reflective capacities, ethical awareness and sovereign interaction with AI move to the foreground. Educators will need to place stronger emphasis on argumentative justification and contextual understanding rather than on the mere retrieval of results.

Necessary Infrastructure

Improved digital infrastructure — for example, learning management systems such as Moodle — and self-hosted or on-premise operation become decisive for maintaining data sovereignty and ensuring a reliable examination environment. External proctoring services can thereby be replaced.

Recommendations

  • Rethink assessment design: Instead of standardised multiple-choice tests, tasks should require reflective argumentation, transfer performance and open-ended solutions that agents cannot independently generate.
  • Expand digital infrastructure: Invest in secure, high-performance learning platforms and local hosting solutions in order to keep data and processes controllable.
  • Promote digital competences: Train educators and learners in the critical use of AI as well as in ethical and legal questions, enabling a conscious interaction between humans and machines.
  • Use hybrid assessment formats: Combine digital tests with in-person examinations in order to ensure personal interaction and authenticity.
Source: Weßels, D. & Maibaum, M. (2026): “Warum KI-Agenten das Ende klassischer Onlinetests einleiten”. In: Forschung & Lehre, 28 May 2026. URL: https://www.forschung-und-lehre.de/forschung/warum-ki-agenten-das-ende-klassischer-onlinetests-einleiten-7715
Cluster Assignment

Cluster C – Innovation in Examination Evaluation / AI-Assisted Grading.

The contribution fits Cluster C because it fundamentally problematises the robustness of existing digital assessment and examination arrangements. If agentic AI can autonomously complete standardised online examinations, the central question becomes under what conditions assessment results can still be considered valid, fair and meaningful. This issue touches the core of Cluster C: the analysis of AI-related assessment structures, their limits and the requirements for robust and quality-assured examination procedures.