Skip to content
Logo von nextlevels
Request a project

Agent = Harness + Model: Why the ‘model’ question is the wrong debate

KI & Automation

The first question in almost every AI agent project is: Which model should we use? GPT, Claude or Gemini? It’s the wrong question. It feels important because the model is the visible part: expensive, measured in benchmarks, featured in every newsletter release. Yet it has almost no bearing on the quality of your agent.

Because a productive agent consists of two parts. There’s the model – the language model – which understands and generates text. And there’s the harness, the actual architecture surrounding it: all the programme code that provides context for the model, equips it with tools, checks its results and saves its state.

The model is the engine. The harness is the rest of the car. You can swap out an engine. But a car without steering, brakes or a fuel tank is still just sitting in the garage.

Formula Agent = Harness + Model: interchangeable model chip in the harness frame
Formula Agent = Harness + Model: interchangeable model chip in the harness frame

The model has become a commodity

The price shows just how little the model weighs. Andreessen Horowitz has measured the decline: The computing time required to achieve the quality delivered by GPT-3 in 2021 has fallen from around 60 US dollars per million tokens to about 6 cents. A factor of a thousand in three years, and for consistent performance, the price is falling roughly tenfold per year.

The research institute Epoch AI confirms the trend across several benchmarks, with a median price fall of around 50-fold per year. Computational intelligence isn’t getting more expensive. It’s being given away for free.

In today’s figures: a usable model such as Llama 3.1 8B costs around five US cents per million input tokens via a provider such as Groq. Five cents. For a million words’ worth of mental work, in the broadest sense. At the upper end, you’ll still pay a multiple of that for a top-of-the-range model. But the curve is pointing in one direction, and the lower threshold at which ‘good enough’ begins is slipping lower every quarter.

For the architecture, this means two things. Firstly: the model is the fastest-aging component in the system. Whatever you choose today will be neither the best nor the cheapest in six months’ time.

Secondly: In a well-built agent, changing the model is a matter of a single configuration line, not a migration. You treat the model like a graphics card that you upgrade without having to rebuild the computer. That is precisely the point. If the change is painful, you’ve introduced the dependency in the wrong place.

LLM price decline: 1,000 times cheaper in three years, around 10 times a year
LLM price decline: 1,000 times cheaper in three years, around 10 times a year

The Accounting Agent runs on the low-cost model

An example that brings the theory to life. Take an agent that pre-posts incoming invoices: read the document, identify the supplier, assign items to the correct accounts in the chart of accounts, and escalate for approval in case of uncertainty. A task on which, in many medium-sized businesses, someone spends hours every month.

This work does not require a ‘Grübel’ model. The difficult part is not the thinking, but the reliable extraction of data, the reconciliation with the chart of accounts, and the application of the rules that this particular business has for this specific supplier.

This isn’t achieved by the brilliance of the model; it’s achieved by the Harness: the integration with DATEV or the ERP system, the validation rules, the memory of previous entries for the same supplier, and the approval loop for anything above a threshold value. A low-cost model fills in the fields reliably. One that’s ten times more expensive fills in the same fields, only at a higher cost.

Work out the model costs roughly. Per document, with context, tool calls and structured output, this might amount to three to five thousand tokens. With ten thousand documents a month, that’s several tens of millions of tokens. Even calculated generously, allowing for repetitions and validation loops, the pure model processing time for this agent remains at just a few euros a month. Less than the cost of the lunch the team has to celebrate the go-live.

Same account assignment screen, different price: a comparison of the budget model and the Frontier model
Same account assignment screen, different price: a comparison of the budget model and the Frontier model

So where does the money go, and where is the quality? In everything other than the model itself: in the clean interface with the accounts, in the rules, in the memory, in the controls. Swap the cheap model for the most expensive Frontier model in this agent, and it will hardly improve, because the model was never the bottleneck.

This applies to well-defined, structured tasks, and that is what most agents do that provide real value in small and medium-sized enterprises. For open-ended research or complex reasoning, you need a more sophisticated model; an agent that builds a robust competitive analysis from ten quarterly reports is not a case for the five-cent model. The accounting agent, however, is.

The four elements that really determine quality

If not the model, then what? Four things. And you build all four yourself; you don’t buy them as part of a provider’s subscription.

AI agent architecture: a harness comprising context, memory, tools and evaluations around the pluggable model
AI agent architecture: a harness comprising context, memory, tools and evaluations around the pluggable model

Context is the first – and most frequently underestimated – lever. A model is only ever as good as what is in its window at the moment a decision is made. The accounting agent that sees exactly the right section of the chart of accounts and supplier history will post the entries correctly. The same agent, however, if you dump the entire jumble of data into its prompt, will merely offer advice. The discipline behind this now has a name – Context Engineering – and it is the part of AI agent architecture where the most quality is gained and lost today.

Memory is the difference between a demo and an employee. Without persistence, every run starts from scratch: a very fast intern who forgets every morning what he learnt yesterday. An agent with a memory recalls supplier patterns, previous corrections and the status of an ongoing process. It improves over weeks without anyone needing to touch the model. It is precisely this ability to learn from use that turns software into a colleague.

The tool integration determines whether the agent has ‘hands’. A model that cannot access DATEV, ERP or CRM is merely a chatbot with good pronunciation. Here, the Model Context Protocol (MCP) has established itself as an open standard through which agents and line-of-business systems communicate with one another.

The real reason why this matters strategically: if your integrations are built as MCP servers, the model remains pluggable. The standard sits between the agent and the system. The specific LLM remains interchangeable. This way, you keep your options open for the next model change – which is guaranteed to come.

Finally, there’s self-improvement, and this is the aspect that no competitor can catch up on with a mere model upgrade. An agent doesn’t get better simply by using a more expensive model. It improves by measuring and refining: test cases based on real documents, defined success criteria, and feedback loops in which every human correction feeds back into the system.

The effect adds up. After six months, one accounting agent still escalates every third document for manual approval, whilst the other escalates only one in twenty. Same model version, different evaluation process. After half a year, two companies using the same model are running two completely different agents.

What the Harness Thesis means for your AI agent architecture

The practical consequence is inconvenient, because it pushes the exciting question to the back and the boring one to the front.

Build in a model-agnostic way. No prompt, no business logic should be rigidly tied to a single provider. If switching from one LLM to another involves more than just a configuration change, that’s an architectural flaw, not a law of nature. MCP and a clean abstraction layer are your safeguard against a provider who might double their prices or discontinue a model tomorrow.

Put your energy into context, memory, connectivity and evaluation. That’s where the quality lies, and that’s where the lasting competitive edge lies.

And stick to this order: first the process, then the automation, then the agent. Anything that can be described as a fixed workflow belongs in classic AI automation, before an agent is even needed. We have described exactly what this structure looks like – with task allocation, the orchestrator-worker pattern and guardrails – in AI Agents as Digital Employees.

Model selection thus becomes what it should be: a late, cost-effective, reversible decision. Start with the cheapest model. Only upgrade where your own tests prove that the extra cost translates into better results, not simply because the release is currently making headlines. If you wish to take this approach with external support, it is precisely these architectural decisions that form the core of our AI consultancy.

The debate about the right model is the loudest and the least important one in the room. The companies whose agents work rarely have the best model. They have the best harness. The model is a commodity. The harness is your product.

Ready for the next step?

Put what you've learned into practice — we'll support you.

Related posts