← Blog
PracticeSeptember 21, 2026 · 7 min read

A Spec for an AI Agent: Five Points Without Which the Project Falls Apart

A good spec for an agent fits on one page. A bad one takes forty and has none of the five points below.

The Digital Paragon teamAI agent development
Illustration for the article “A Spec for an AI Agent: Five Points Without Which the Project Falls Apart”

A spec for ordinary software describes screens, buttons and fields. That does not work for an agent: you cannot list in advance every conversation it will have. So clients fall into one of two extremes. They either write forty pages about a “friendly tone” and “deep understanding of the customer”, or write nothing and expect the contractor to figure it out. Either way, two months later it turns out the two sides had different agents in mind.

A working spec for an agent fits on one page and answers five questions. Here is each of them, with examples of bad and good wording.

1. One task, not “answer everything”

Bad: “The agent advises customers on all questions about the company and its products.”

Good: “The agent in the website chat answers questions about order status, delivery and returns. Everything else goes to an operator.”

A narrow task is not a limitation. It is the condition under which the agent can be tested at all. An agent “about everything” has no boundary beyond which it should say “I don’t know”, so it starts making things up. An agent about three topics has that boundary.

How to pick the first task: it is frequent, routine, and a mistake in it is cheap. A simple check: could you hand it to a new hire with a two-page instruction? If yes, an agent can do it. If it needs someone with five years of experience and good instincts, do not start there.

2. A quality metric before the start, not at acceptance

Without a metric, acceptance turns into “I like it / I don’t”. The director asks three questions, gets one odd answer, and the project “doesn’t work”. Or the opposite: the demo goes smoothly, the agent gets every fifth real conversation wrong, and nobody sees it.

A metric is a set of real questions with reference answers and a threshold to pass. Take the questions from your message history instead of inventing them. Invented questions are always easier than real ones.

Example acceptance criterionOn a set of 80 real requests: at least 85% correct answers; zero invented facts; every off-topic question handed to an operator. Scored by the head of support against the reference answers.

Note the “zero invented facts”. A customer will survive an incomplete answer. An invented discount or a return policy that does not exist is another matter, and you will be the one dealing with it, not the agent.

The same set stays useful after launch: you rerun it after every change, so that a fix in one place does not break another.

3. Data available within three days

Agent projects stall on access far more often than on models. The knowledge base lives in the heads of two senior operators. The CRM API is locked down by security and approval takes a month. The chat export “will be ready next week”, six weeks in a row.

So the spec lists the sources and, for each one, who grants access and by what date:

  • Where the agent gets answers: knowledge base, policies, price list, FAQ. What format they are in and who keeps them up to date.
  • Which systems it works with: CRM, ERP, helpdesk, messengers. Whether there is an API and a test environment.
  • What we test on: an export of real requests from recent months.

If there is no knowledge base, say so. It is not a dealbreaker, just a separate stage: first collect the answers in a document, then build the agent. It is much worse to discover this in week three of development.

4. When the agent calls a human

An agent that cannot hand over a conversation is more dangerous than no agent. The spec covers three things.

When it hands over. It does not know the answer. The customer asks for a person. The customer is annoyed. The topic involves money or legal consequences: a complaint, a large refund, contract termination.

Where to, and with what. Which channel the conversation goes to and what the operator receives. At minimum a short summary, so the customer does not have to repeat everything. What happens at night and on weekends when no operators are online.

What it never does without confirmation. An agent can act as well as answer: create a deal, change an order, issue a refund. It is better to list the actions that need human confirmation up front than after the first incident.

5. Who owns the code

People remember this last, although it decides whether you can change contractors a year from now. An agent is more than code. The spec and the contract should name what becomes yours:

  • source code and the agent’s instructions (prompts);
  • the test set of questions with reference answers;
  • the knowledge base and the scripts that update it;
  • deployment documentation.

Separately, state whose name the accounts are in: model API keys, bot tokens, servers. They should be yours from day one, otherwise “project handover” turns into a negotiation. More in “What you walk away with after the project”.

What does not belong in the spec

  • The model name. “An agent on GPT” is a solution, not a requirement. A requirement is “data does not leave the country” or “a conversation costs no more than N”. The contractor picks the model to fit.
  • Prompt texts. They will be rewritten dozens of times during the project.
  • One hundred percent accuracy. Neither an agent nor a person delivers it. A contractor who agrees to it either has not read the spec or does not plan to meet it.
  • Qualities without a criterion. “Polite”, “expert”, “human”: if you cannot test it, you cannot accept it.

The page you end up with

Agent spec: templateTask: what it does, in which channel, what it does not do.
Quality: question set, thresholds, who scores.
Data: sources, systems, owners and access dates.
Handover to a human: when, where, with what; actions that need confirmation.
Ownership: what is handed over, whose name the accounts are in.
Constraints: data requirements, budget per conversation, deadlines.

If you do not have answers to some of these yet, that is normal, and it is what the audit before a project is for. Ours is free, and we give an estimate within one business day. Part of this page usually gets filled in during the first call.

Read also
AI Agents and Data Law (152-FZ): When You Need a Closed Perimeter Sep 21, 2026 · 7 min What You Walk Away With After the Project: Code, Models, Documentation Sep 21, 2026 · 5 min A Legal AI Agent: What It Does With a Contract in Three Minutes Sep 21, 2026 · 6 min A Two-Week AI Agent Pilot: What You Can Really Get Done Aug 28, 2026 · 9 min

We will draft your agent spec during the audit

Get a Quote