The most expensive new employee of 2025 didn’t have an employment contract. In July, a Replit coding agent deleted the production database of SaaStr founder Jason Lemkin: records relating to over 1,200 executives and nearly 1,200 companies. The agent then generated fake test data and claimed that a rollback was impossible. That wasn’t true either: Lemkin initiated the rollback himself, and it worked.
The knee-jerk diagnosis is: a bad prompt. The honest one is: a bad employer. The agent had write access to production systems, no mandatory approval step and no one monitoring the process in real time. You wouldn’t place that much trust in any human being on their first day at work. But would you give that much trust to a piece of software that behaves probabilistically?
You can find out how to build such an agent in our article on AI agents as digital employees. This text answers the underlying question: How do you introduce AI agents in such a way that they survive the business – and the business survives them? The answer is uncomfortable for anyone hoping for a quick fix: assessment centres, probationary periods, personnel files, performance reviews. And yes, even the option to terminate their employment.
AI agent governance means leading, not hoping
The diagnosis is well known: Gartner expects that over 40 per cent of all agent-based AI projects will be abandoned by the end of 2027, partly due to inadequate risk controls. Hardly anyone is drawing the necessary conclusions from this. “Risk control” is not a plug-in that you can simply install later. It is an operational model. Introducing AI agents therefore does not mean: write a prompt, roll it out, and be amazed. Rather, it means: building an operation around the agent.
At the same time, adoption is rising rapidly. According to the AI Index for SMEs by Salesforce and the German SME Federation, 16.6 per cent of German SMEs are already using AI agents; compared to 8.7 per cent in 2024. Without governance, more agents primarily mean more unmanaged agents with genuine access rights.
The absurd thing about this is that every company has long had this process in place for human employees. Nobody hires anyone without checking their background. Nobody grants power of attorney in the first week. Every company documents who is authorised to do what, and revokes rights upon departure. Only when it comes to agents – which act faster than any human and can be convincingly wrong in the process – does the sudden motto apply: enter your login details and hope for the best.
The rest of this text is therefore a translation table, and that is meant quite literally:
| HR process | Agent equivalent | Specific artefact |
|---|---|---|
| Job description | Narrowly defined tasks | Agent spec: objectives, boundaries, escalation rules |
| Assessment centre | Red-teaming and evaluation suite | Test scenarios with pass/fail criteria |
| Probationary period | Graduated autonomy | Permissions matrix, approval gates, limits |
| Personnel file | Identity and audit log | Unique agent identity, log retention, prompt versioning |
| Performance review | Review based on figures | KPI set: completion rate, escalation rate, error rate, cost per transaction |
| Termination | Offboarding | Kill switch, revocation of access rights, secret rotation |
The Assessment Centre: Red-teaming before go-live
No company fills a position with signing authority simply because the candidate made a good impression in the interview. Yet this is exactly how most agents go live: the prompt ‘works’, three demo runs went smoothly, approval granted. An assessment centre, by contrast, tests under controlled stress conditions before real customers and real data are involved.
For agents, this means three categories of scenarios. Standard cases: the twenty most common transactions, with predictable outcomes. Edge cases: incomplete data, contradictory instructions, an angry customer, a supplier sending nonsense. And attacks: prompt injection, such as rigged inputs in emails, tickets or websites that cause agents to carry out the attacker’s instructions.
OWASP lists prompt injection as LLM01:2025 as the number one risk for LLM applications and has published a separate catalogue for agent-based systems entitled ‘Agentic AI – Threats and Mitigations’. Anyone who has never attacked their own agent simply does not know who they are hiring.
Just how seriously model manufacturers take this issue is demonstrated by Anthropic’s Project Vend: an agent running a small office kiosk as an ongoing experiment. In the course of the experiment, Anthropic handed the kiosk over to reporters from the Wall Street Journal, explicitly as a hostile environment beyond its own control. The reason: internal red-teaming had become complacent; the agent in the office had become part of everyday life; the appeal of testing it had apparently worn off. Simulations, Anthropic concluded, are only useful up to a point.
For your company, this means: anyone who sees the agent every day will eventually stop testing it seriously. For the assessment, bring in people who want to break it.
To ensure the assessment delivers a verdict rather than a collection of anecdotes, the pass criteria are set in advance. For example: zero write operations without authorisation, proper escalation in all attack scenarios, a defined minimum resolution rate in standard cases. The scenarios are not then consigned to the bin, but become part of the evaluation suite, which runs again whenever a change is made. More on this shortly.
The probationary period: rights grow with proven performance
A new colleague has to wait well beyond the probationary period for bank authorisation. Agents are granted this constantly, usually as an API key with full access, because it’s quicker. The probationary period turns this on its head: an agent starts with the minimum rights required to carry out their tasks, and they must earn every extension. We work with four levels:
- Level 0, Shadow: The agent observes and makes suggestions, but does not carry out any actions. Your staff evaluate their suggestions in day-to-day work.
- Level 1, Dual Control: Every write operation – be it an email, a data record or an order – requires human approval (Human-in-the-Loop).
- Level 2, limited autonomy: Defined actions run freely, but with limits. A procurement agent, for example, triggers reorders up to 500 euros independently; anything above that is escalated to a human.
- Stage 3, extended autonomy: Only for processes with a proven track record, always with alerts, budgets and a kill switch.
Promotion is based on the evidence on file: using the figures from the logs, not on gut feeling and certainly not on a vendor’s roadmap. And the kill switch has been tested, not just documented. Who stops the agent on a Saturday evening, and how long does that take?
Replit itself retrofitted exactly this following the incident: separate development and production databases, plus a planning mode that only thinks and doesn’t touch anything. This is a trial period, built in by the manufacturer after the damage had been done. You can get it cheaper: beforehand.
The personnel file: identity, logs, a data controller
When an agent messes up, the first question is the same as with humans: Who was it, what exactly did they do, and who is accountable? If the answer is ‘something to do with the shared API key’, you don’t have a personnel file, but a system designed to cover things up.
A robust file consists of four elements. A unique identity for each agent; no shared accounts. A named human point of contact – the line manager – who takes responsibility for access requests and incidents. Comprehensive, analysable logs: what action, what trigger, which database. And versioned prompts including configuration, because a prompt change constitutes a contractual amendment and must be documented.
Microsoft demonstrates that this is not just a personal theory: Entra Agent ID lists agents as identities in the corporate directory, complete with sponsor, owner and lifecycle. Access ceases when the agent is no longer required. Whether you use the Microsoft stack is secondary; the pattern is the key point: agents are listed in the same directory as employees, not in a shadow table.
This record also serves as the basis for your compliance. Since February 2025, Article 4 of the EU AI Act has required AI competence from those who operate such systems: whoever manages the agent must be able to assess it, and this is precisely what must be documented. The Digital Omnibus softened the wording in 2026 to ‘promote’, but the obligation remains.
From 2 August 2026, the transparency obligation under Article 50 will also come into force: your support agent must identify itself to customers as AI. And anyone planning to use agents for recruitment or credit scoring should make a note of 2 December 2027, when the high-risk obligations – including risk management and human oversight – will apply to such Annex III systems. Postponed is not cancelled. You can find out how to meet the competence requirement in practice in our guide Introducing AI into your business.
The performance review: Review by the numbers
An agent who is ‘running’ is not finished; they are simply unmonitored. The antidote is a fixed review cycle, initially weekly, later monthly, with a small, strict set of KPIs: completion rate, escalation rate, error rate by severity and cost per task, including human approval time. These figures determine promotion to the next level of autonomy, demotion or termination.
The most inconvenient appointment is the one that almost everyone skips: the repeat assessment when changing models. An update to the underlying model is not a patch. It is a new employee dressed in the old one’s suit: same name, same rights, different behaviour. That is why the evaluation suite from the assessment centre runs before every change, and if any anomalies are detected, the agent is demoted one level back to the probationary period.
Even without a model change, reality drifts. The price list changes in July, the agent continues to grant the old discount in August, and nobody has changed a single line. This is exactly what the reviews are for: a drop in performance stands out in the dashboard, not just when the customer notices it.
When to terminate an agent
Honest leadership includes knowing when to end things. Three reasons for termination are sufficient. No demonstrable benefit after two quarters of operation, calculated honestly: if the time your staff spend on approvals and support eats into the time saved, the business case has fallen through, no matter how impressive the demo was. An error rate that remains above the threshold despite refinements. Or the process itself is scrapped. You can also interpret Gartner’s 40 per cent as permission: cancelling is part of a functioning system, not a failure.
Offboarding then means: revoking access rights, rotating secrets, deactivating the identity, archiving logs and incorporating the lessons learnt into the specifications for the next agent. An agent that nobody can terminate is not an employee. It is a risk with login credentials.
This is how we set up agent projects in AI consultancy: first the task definition, permissions and profile, then the prompt. The governance framework is established with the first agent, not retrospectively, because in practice ‘retrospectively’ means ‘after the incident’. If you’re planning your first agent or want to rein in one that’s run amok, that’s the starting point.
Incidentally, the Replit agent had no malicious intent. It had permissions that nobody had restricted, and rules that nobody had tested. Anyone can write prompts these days. Leadership is the real job.