Specification as code: How I keep AI under control

Specification as code: How I keep AI under control

Natural language can’t give AI an unambiguous specification. I solve this with specification as code: schema, tests, and rules written as code instead of English. This is how I use AI to build software, and why I do it this way - after 23 years in IT systems design.

Conventions

  • Any prompt referenced in this article is written in English, since English is the dominant natural language in IT projects.
  • The examples use TypeScript, but the same principles apply to any language and work the same way there.
  • TL;DR: browse the images, then jump to the summary. Read the whole article if you want to understand why and how.

The problem

The translation gap: human intent passes through a fallible AI translator on the way to output, and AI has read about many tools but never lived with the consequences of a choice

Compared to a human, AI has unlimited memory and never gets tired of tedious thinking. This makes up for the limits of the human brain. But we and AI are fundamentally different, so we don’t share a natural language. AI acts as its own translator, and that leads to many wrong assumptions and misread intentions. Even a correctly understood intention can still be carried out badly, because AI is fallible. There’s also the matter of correctly understanding the context behind an intention — that context can substantially change the desired outcome.

AI can read about ten databases, but it has never run any of them in production for two years. It cannot say, “We used both, and this one broke under load” — it can only repeat what other people wrote, and those texts are also written for promotion, by people without real, hands-on experience. Real architecture decisions come exactly from the part AI does not have: seeing how a choice behaves years later, after the team changed and the data grew. So AI will not replace the architect. Ask it to find problems in your decision, not to make the decision.

Natural language

Natural language spec drifts away from the code with every iteration; specification as code stays exactly in sync because it is the code

AI converts natural language (for example English), the language of the prompt, into code. This process carries a large margin of error, because it adds an extra layer of abstraction without a clear, one-to-one conversion pattern. Natural language wasn’t built to be precise. In daily life, people use natural language in a loose and general way. Programming languages, by contrast, were refined over the years to make mistakes hard to make and meaning unambiguous. That was never the goal of natural language.

It’s no accident that reading documentation has been a frustrating experience for years — we read it, follow it exactly, and still can’t get the result we want. Even between two humans, just writing and reading causes plenty of misunderstandings. Add differences in experience, culture, and how ambiguous words get interpreted across languages and cultures, plus the fact that only part of the world speaks English natively - writing an unambiguous specification becomes even less realistic.

Writing an unambiguous specification in natural language is impossible, no matter how much time we spend on it. Keeping a large, complete natural-language specification consistent across iterations and changes is such a complex process that it’s practically unachievable. Over time, after enough changes, the documentation ends up contradicting itself in places and leaving gaps unaddressed. Another problem is documentation falling out of sync with the code — something that’s also unrealistic for a person to prevent, especially when the specification is meant to be extensive and to describe the entire application unambiguously as a single source of truth for AI.

That’s exactly why code exists. Code is unambiguous — its design aims to prevent mistakes, and each iteration describes intent in finer detail. That’s something natural language can’t achieve. Of course, code takes more effort than natural language, but that’s precisely because it forces you to be unambiguous.

AI isn’t a magical entity. This is a program that runs in a strictly defined and well-understood way. It’s easy to get swept up in vibe coding, judging it by visual results alone, without checking whether it actually computes correctly in every case for the given logic and holds up against errors. You can build systems with vibe coding, but they won’t hold up at an enterprise or even professional level. Vibe coding fits companies with no budget, or cases where system errors are acceptable and a perfect result isn’t required. It would be irresponsible to build, say, an ERP integration (accounting and inventory) or loan interest calculations with vibe coding. But a small script that sends a newsletter automatically once a week, with no human involved, is acceptable — especially at a small company with no budget.

As a rule of thumb: use specification as code when mistakes are costly or hard to spot by eye. Skip it for throwaway scripts, prototypes, and anything you can fully verify visually in seconds.

That doesn’t mean we should never use prompting without writing code. I do it myself, but for prototyping and MVPs — not for a product headed to production. Even when I write specifications in English myself, at some point, as the specification grows, I lose control over its consistency. Even for small programs, a specification written for AI ends up being extensive, because it has to be complete down to the details. In a team, where multiple people edit the documentation, keeping it consistent becomes completely impossible.

Delegating to AI

Ideally, we’d want to delegate 100% of the work to AI. We can, but then we lose control over AI, and as a result, over what the system does.

To simplify, regardless of what the system does, we can distinguish:

  • input data - human control
  • processing data - delegated to AI
  • rules for processing data - human control
  • output data - human control
  • tests - human control

A single system can have multiple input -> processing -> output pipelines, and that’s fine.

It’s hard to imagine anyone wanting to use, say, an API whose input and output data structure could change unpredictably with every version. On the other hand, nobody cares how the system processes the data — what matters is the expected result. Ideally, we wouldn’t want to deal with how that result is achieved at all.

AI can’t identify its own mistakes — above all, misunderstood human intentions, misunderstood context, and wrong assumptions. Tests written by AI will test its own wrong assumptions and give a false sense of control.

The rules for processing data are a critical part of the system’s specification. We want to keep control over them and not leave them to AI. I’ll explain what I mean in detail in the Project Architecture section.

What’s the goal of tests for AI?

  • verifying that changes work as intended
  • verifying that changes didn’t alter the system in unexpected places
  • specification

This isn’t new, but in AI projects, tests as specification have never been this important. A human can’t read through all the tests and keep them in mind every time while coding, but AI can.

That doesn’t mean we can’t use AI to write code for tests, input data, and output data. It means that in these areas we have to give our full attention — read every line and think it through. Here, we’re the ones writing the code, and AI supports us. But the code that processes the data is a balance between control and speed of making changes — we might not even read it, as long as we control the input data, output data, the rules for processing data, and the tests.

Project Architecture

Control pipeline: input and output data are schema-defined and human-controlled, rules for processing are human-controlled, processing itself is delegated to AI, and tests verify the output and reject failures

I load a list of files into the AI: files that are the source of truth, and files that are 100% written by AI.

Here’s a generic example of AGENTS.md. I built it after asking the AI itself what wording it understands best and which naming it prefers. The name *.spec.* is already reserved for another purpose, so using it would cause complications.

## canon files

Files named `*.canon.*` are the canonical specification: the most precise description of the project in human readable form. They are the source of truth for the app code — the code must conform to them, never the other way around.

List of canonical files:

```bash
find . -type f -name '*.canon.*'
```

Read all canon files always before changing code.

## AI files

Files named `*.ai.*` are 100% AI-written. They must conform to the canon files, never the other way around.

Treat the existing code in these files as disposable output, not as a pattern to follow or preserve. It is an artifact of a previous run, made for a different task under instructions that may have since changed — it may embed wrong or outdated assumptions. Do not let it bias or anchor your implementation. When changing an `*.ai.*` file, re-derive the implementation from the canon files and current instructions, rather than extending or mimicking what is already there.

```bash
find . -type f -name '*.ai.*'
```

### Other files

Other files are coded by humans and AI together.

Instead of find, you can just write out the list of files. The wording doesn’t need to match mine exactly. AI needs to understand which files are the source of truth and which are the output of earlier AI runs — code in the latter doesn’t need to be preserved, extended, kept backward-compatible, or treated as a specification, because it’s an artifact of previous AI runs.

For a more complex system, have canon files in multiple directories.

File architecture: canon files are the source of truth and constrain the AI-written files, never the other way around

Schema

Modern programming languages have Schema, which is the best form of data specification in code. It simply gives you the full set of features:

  • data structure
  • data type
  • values
  • data validation
  • random test data generation

This is an unambiguous specification for AI. For example, it tells the AI that amounts must be stored as BigInt, not Float, which is critical for finance, because of rounding. The interest rate must be a value from 0 inclusive to 100 exclusive, and the loan start date must be in the future. AI knows which data comes in and goes out, and by reading the names it understands the context. We have precisely defined data, and AI writes code that transforms input into output.

AI can quickly and cheaply validate its own mistakes, because Schema returns a clear reason for each failure. Even without a single line of test code, the Schema stage alone already catches problems with types, values, and the dependencies between them.

Schema wasn’t created for AI. Professional projects should be built this way regardless, but now that AI is here, the benefits are even greater. I personally have used Schema in every project for years, and I recommend it to everyone.

Illustrative excerpt for loan interest calculations:

/*
The code below describes the structure of a file to import.
This gives AI validation, random data generation, and context understanding.
It's not just information about the structure, but also names, e.g. that statutoryInterestRate can only be positive. AI can't draw such conclusions from the import file alone.
*/

export const NbpRate = Schema.Struct({
	effectiveFrom: PlainDate,
	referenceRate: posPercent,
	statutoryInterestRate: posPercent,
	maxInterestRate: posPercent,
	statutoryDefaultInterestRate: posPercent,
	maxDefaultInterestRate: posPercent,
});

export const NbpRates = Schema.Struct({
	lastSync: PlainDate,
	rates: Schema.Array(NbpRate),
});
Schema.Array(NbpRate);

...

/*
Specification of the function's input for interest calculations. It's a code snippet, but it shows what Schema is for.
*/

export const LoanCalculatorInputs = Schema.Struct({
	nbpRates: NbpRates,
	disbursements: Schema.Array(LoanTransaction),
	repayments: Schema.Array(LoanTransaction),
	maturityDate: PlainDate,
	asOfDate: PlainDate,
	interestRateMode: InterestRateMode,
	defaultInterestRateMode: DefaultInterestRateMode,
});

/*
periodStart and periodEnd are the same data type, but I use descriptive names instead of PlainDate.
We also validate periodStart against periodEnd. Similarly, we'd add e.g. a 100% cap for referenceRate and other conditions the data must satisfy.

For AI, this is a specification and an understanding of context, such as date inclusive vs exclusive, and especially the names.
*/

export const periodStart = PlainDate;
export const periodEnd = PlainDate;
export const periodRate = posPercent;

const AccrualPeriodsBasicFields = Schema.Struct({
	periodStart: periodStart,
	periodEnd: periodEnd,
	days: Schema.Number.pipe(Schema.int(), Schema.positive()),
	referenceRate: posPercent,
	interestBalance: NonNegPLN,
	periodRate: periodRate,
}).pipe(
	Schema.filter(
		({ periodStart, periodEnd }) => Temporal.PlainDate.compare(periodStart, periodEnd) <= 0 || 'periodStart must be <= periodEnd',
	),
);

/*
Thanks to the type name periodStart instead of plainDate, I can write briefly and unambiguously for AI that the first unknown day for the NBP rate is specifically periodStart. This means I don't have to spell out what AI should do with that date.
*/

export function firstUnknownNbpRateDay(nbpRates: NbpRates): periodStart {
	return nbpRates.lastSync.add({ days: 1 });
}

Tests

It’s worth testing fixed invariants: preconditions on the input, postconditions on the output, and correlations among the output data itself. For example, in loans, late-payment interest can’t be charged before the payment date, and principal interest can’t be charged after it. The sum of interest and transactions must equal the balance. AI can look at a few fixtures (sets of input and output data) and find these invariants on its own, then write the test code.

Tests like these, which check invariants in the data, are useful not just for testing code, but also for catching human error when writing tests for fixtures. A typo, a bad copy-paste, a miscalculation - any of these is enough to cause a mistake.

Once we have invariant tests in place, we can add property-based testing: generating random data from a Schema and checking it against those invariants. This way we can automatically discover, for example, that a remaining principal balance of 0.001 - a valid value - leads to an error in the interest calculation because of rounding, or that the system allows an end date before the loan’s start date. This is implemented so that each test run generates N random data sets and verifies them, since the amount of random data can be treated as effectively infinite.

Property-based tests are mainly meant to catch corner cases, because the data is random and so it tests every scenario, not just the ones a human picked. They also often catch cases where the code throws an unexpected exception on input a human wouldn’t think to test.

An example of random data from a Schema (trimmed here for brevity). It’s generated according to a pattern you control.

import { Arbitrary, FastCheck } from 'effect';
FastCheck.sample(Arbitrary.make(LoanCalculatorInputs), 1)[0];
{
	"nbpRates": {
		"lastSync": "2047-10-06",
		"rates": [
			{
				"effectiveFrom": "2016-01-08",
				"referenceRate": 0.02,
				"statutoryInterestRate": 3.52,
				"maxInterestRate": 7.04,
				"statutoryDefaultInterestRate": 5.52,
				"maxDefaultInterestRate": 11.04
			},
			{
				"effectiveFrom": "2016-01-14",
				"referenceRate": 97.59,
				"statutoryInterestRate": 101.09,
				"maxInterestRate": 202.18,
				"statutoryDefaultInterestRate": 103.09,
				"maxDefaultInterestRate": 206.18
			}
		]
	},
	"disbursements": [
		{ "date": "2017-11-08", "amount": "18831399n" },
		{ "date": "2047-02-26", "amount": "50018634n" }
	],
	"repayments": [
		{ "date": "2025-01-18", "amount": "61535385n" },
		{ "date": "2068-01-24", "amount": "63493469n" }
	],
	"maturityDate": "2027-08-26",
	"asOfDate": "2099-12-23",
	"interestRateMode": "maxInterestRate",
	"defaultInterestRateMode": "maxDefaultInterestRate"
}

Rules

Part of the specification doesn’t fit into the Schema or the tests. Here it’s a matter of choice whether we want to keep it in English or as code. The rules file is code, 100% controlled by a human.

How do I decide what goes into rules and what I delegate to AI?

  • If something is important and I want full control over it.
  • Some guidelines are easier for me to write as code than in English — in English they’d take too many lines and be hard to express.
  • As the system matures and I become an expert in the domain it covers, I add more rules to the file. That’s the natural order of things — I gain certainty about how the system should behave, and I have to document it somewhere.

Here’s an example of rules for how AI should allocate loan-calculator repayments across principal, interest, and default interest. The code below isn’t just a function in the code — it’s a specification as code for the AI. That’s how you should think about it.

// rules.canon.ts

export function allocateRepayment(
	defaultInterestBalance: NonNegPLN,
	interestBalance: NonNegPLN,
	amount: NonNegPLN,
): { defaultInterestAmount: NonNegPLN; interestAmount: NonNegPLN; principalAmount: PLN } {
	const defaultInterestAmount = amount < defaultInterestBalance ? amount : defaultInterestBalance;
	const afterDefaultInterest = amount - defaultInterestAmount;

	const interestAmount = afterDefaultInterest < interestBalance ? afterDefaultInterest : interestBalance;
	const principalAmount = afterDefaultInterest - interestAmount;

	return { defaultInterestAmount, interestAmount, principalAmount };
}

Honestly, this concept originated as an AI recommendation of what would be clearest for itself.

Summary

Summary: AGENTS.md points to canon files (schema, rules) and tests (fixtures, invariants); AI also needs intent, context, current and desired state, and the ability to iterate, all in the same session

The specification as code becomes part of the system, so it is always a consistent and unambiguous source of truth. The method is simple, requires no extra tools, has no vendor lock-in, can be applied in any programming language, optimizes AI token cost, and is effective.

Summary in bullet points, as a reminder:

  • AGENTS.md is loaded at the start of every session, describing which files are the source of truth and which are coded by AI.
  • Canon files
    • Schema
    • Rules
  • Tests
    • Fixtures (input + output)
    • Invariants (recommendation)

For AI to generate the desired code, it needs:

  • to understand the intent
  • to understand the context
  • to understand the current and desired state of the application
  • to be able to iterate while writing code, fix and improve it based on compile-time and runtime errors, and use a virtual web browser or other interface, depending on what kind of system it is

All of the above must be present in the session at the same time. If you neglect, for example, loading context (including the canon files) into the AI session, then even though the AI understands the intent, it will deliver code that does the wrong thing.

Iterations

Iterations: edit the schema, rules, or fixtures, optionally review complex changes with AI, then AI implements from the diff and description, repeating for the next change

I do it like this (each step only when the change touches it):

  • I change the Schema file
  • I change the rules file
  • I change / add fixtures and tests
  • if I changed something more complex where I could have made a mistake, I consult my specification changes with AI before it starts implementing them.
  • I prompt AI to run git diff --cached and update the code, along with a short description in English of what the change involves. This gives AI information about the current and desired state.

To be clear: AI can write any of the above changes for me from a natural-language prompt. I give this my full attention, so the specification code stays readable, simple, and correct — this is the source of truth for the whole system.

Cost

A side effect of specification as code is a significant cost optimization. The AI doesn’t burn tokens reading and interpreting specification in natural language. There is no extra tooling needed to process a task. The specification is genuinely part of the system, so the context the AI holds in memory is optimized to its limits. AI can’t silently ship code that violates the encoded part of the specification: schema, types and invariants fail loudly at compile time or on the first run. Part of the code iteration happens outside the AI’s cloud, since compilation and tests run locally and give fast feedback.

Bonus

I start every new AI change in a separate session. I always keep project memory and cross-session memory turned off. Remembered information goes stale and only creates confusion. If something matters, I write it into AGENTS.md. Every new session is a blank page that I fill in.

AGENTS.md should contain the tools and commands AI should use: a list of MCP servers, skills, libraries, UI guidelines, and so on. This matters.

I use natural language for prototyping, especially with new tools. Once I know what I want to achieve and I get to know the tools better, I switch to specification as code. When I use tools I already know on a new project, I write the specification as code almost from the start, because I already know the architecture and patterns I’ll use.

AI lets me build things that used to be out of reach for me, because they required time, energy, and becoming an expert in every single tool and library. With AI, the technical implementation of advanced functionality in tools I’ve never used before can take just a few minutes. The time it takes to learn, read documentation, search for solutions, and so on has dropped from months to hours, sometimes even minutes.

The developer’s role is shifting, among other things, toward:

  • system architecture
  • being a domain expert
  • writing specifications for AI
  • controlling code quality

This is a simplification, but system architecture experience is starting to matter more than specializing in, say, Spring and Java. Why? Because, as I mentioned in the previous paragraph, years of experience with specific libraries no longer carry the value they used to. AI won’t replace an architect, because on its own it makes poor decisions. That happens because AI can’t have experience — its recommendations are based only on the data used to train the model: marketing, blog posts, random people’s posts, and so on, but never real experience of actually using these tools.