AI readiness: where are you?
People ask when they should start using AI in their marketing. The answer is that they started the day somebody opened their first Google Ads or Meta account.
An AI model is already deciding your bids, your queries, your placements and your creative combinations. You are running at the most automated level there is. It is just that the AI model is not yours, the rules are not yours, and none of them will tell you what it decided on. Not Google, not Meta, not TikTok.
So the question is never whether you are using AI. It is whose. These ten levels are how far you have got with your own.
Levels 1 to 5: worked with you
1 · You know how far each number can be trusted. Google says 1,000 purchases, Meta says 10,000, GA4 says 2,850, and your own webshop says 2,500. One of those four is the place the money actually changed hands. Level one takes the measurement apart — GA4, the Google Tag Manager containers including the server-side one, the conversion signals, consent, the feed, the order data — and says of each number how much its meaning will carry in the area it came from. Some will hold anything. Some will hold a direction and nothing more.
And some of it is measurement error nobody has ever checked. Not through carelessness: a number that arrives every morning stops being questioned. The damage is not in the report, either — Google and Meta optimise on that same number, so when it is wrong the system learns, accurately and tirelessly, in the wrong direction.
This is the one level with something to check against. A tag either matches the specification or it does not. Above here there is no specification, only comparison.
2 · You see what the market does, not what it says. Level one reads what you report about yourself. This one reads what can be observed from outside: demand, who is visible on what, and where you actually stand in organic results.
That step is the whole argument in one move — what a company says about itself is the least reliable register there is, and that includes yours.
What it cannot do, said here rather than in a footnote: these readings are relative. They give you rank and direction, never a market-wide figure. Shopping, Performance Max, Display, YouTube and Demand Gen are outside them, and in Hungary that is most of the money.
3 · You know which way you may go, and why. The channels read: the bidding strategies, what the budgets are doing, what the campaigns are separated by, the search terms, the placements — Google Ads and Meta. What comes back is not a list of improvements. It is a reading of where the road is open and where it is closed, with the numbers each one rests on attached to it.
And every finding is labelled with how much it can bear. Some are specification violations — a country field filled with a county code, so products drop silently out of the feed. That holds up against a hostile reader, because a specification either matches or it does not. Others are professional judgement: a bidding strategy, the pacing of a budget, the quality of a placement. Those are arguable in good faith, and we say so beside them, so you do not read them with the weight of the first kind.
Which is why we ask before calling anything a fault. Most of what looks wrong in an account was chosen by somebody. A campaign optimising for store visits instead of purchases is a fault or a decision depending on what the shop is for. Where there is a specification, we check. Where there is none, we ask.
What this level cannot do alone: it tells you which channels are open. Whether the campaigns are cut up the right way also needs your goals and how the company actually works — and that is level four.
4 · You know where your company and its data disagree. What arrives here is the mess: brand documents, guidelines, directions somebody set two years ago, decisions nobody wrote down. It goes in unordered and comes out ordered. What that buys you is two sentences no system can say without it: this goes against a decision you already made, and this is what is missing for it to be consistent with one.
And held against the three levels beneath it, it produces the one comparison nobody else can make: here is what you say you are, and here is what your data says you are. An agency has the channel data and no corpus. A brand consultancy has the corpus and no data. We ran it on ourselves first, and not as a figure of speech. Our own brand material, our positioning history and our decisions sit in the same kind of registry we build for a client: what we are allowed to claim, a register of which words we use and which ones we have retired, and every revisable document versioned. That last part is in there because we learned it the expensive way — six generations of our own positioning were written over the top of each other, and the earlier ones are gone. This page was written against that registry, and where a sentence had nothing behind it in the record, it did not go out.
5 · You know which of your decisions worked. Every decision goes in with what you expected of it: the premise, the number that was supposed to move, and by when. Months later somebody goes back and asks what actually happened. The answer is that it held, that it did not, or that it cannot be told from what was recorded — and the third is an answer, not a failure. Before-and-after is not proof. Where something else moved at the same time, the honest reading says so instead of producing a figure.
Two different things sit on this level. Building the store is a project, done once. Going back to ask whether a decision held is the recurring half, and it is the only one of the ten levels that recurs. It is also the only one worth more every year. In year one the record is thin and the reading is short. By year three it holds decisions, evidence and outcomes including the wrong ones — and that is the one thing in this whole line that cannot be bought, hired, or trained into a model.
You can start here without levels one to four, under one condition: we read level one first, scoped only to confirm or refuse it. A memory built on numbers nobody checked will answer confidently about nothing. If the read confirms, you start at five. If it refuses, you start at one after all — and the read has paid for itself by saying so before anything was built.
Levels 6 to 10: worked with LLMs
6 · We load your company into AOS Cloud. Everything ordered at levels one to five goes in: the numbers with how far each can be trusted, the market reading, the directions that are open, your own brand and decision material, and the record of what came of each decision. The system and the tooling are ours. The material is yours.
What changes is what a question gives back. Ask an LLM today where a webshop should set its free-shipping threshold and you get the general industry answer. Ask after level six and it comes back about you: what happened the last two times you moved it, what the competitors charge, and which of your own numbers can bear the question at all.
This is capability, not work done. A good answer still needs somebody to ask the question, ask it properly, and turn it into something that can be decided. That is level seven.
7 · We operate our system for you. AOS Cloud connected over MCP to whichever LLM we run it on, and worked by us. We ask the questions, we produce the analyses, we prepare the decisions. What reaches your desk is a decision with the thing it rests on attached — not a login.
The difference from six is who does the work. At six the capability is there and somebody still has to use it. At seven that somebody is us, week after week, so the questions get asked whether or not anyone on your team had time this month. The difference from eight is that at eight your own people are doing it.
And one thing a level-seven engagement has to bring down from higher up the ladder: anything that runs continuously needs its boundaries written down before it runs — what may be settled without asking, and what has to come back to a person. That is level nine, and a good seven already carries part of it.
8 · Your colleagues use our system well. The same system over MCP, worked by your people. Well is doing the work in that sentence, because an LLM always answers. The question is whether it answers correctly, and whether the person at the keyboard can tell.
So four things get taught, in this order. How to run it as it stands, so there is a result inside the first hour. What it can answer, and what it cannot. How it is built, which is what lets somebody carry it into work we have never seen. And how to put your own material into it, which is what makes it yours.
The second one is the centre, and nobody else teaches it. What your team learns there is that a setting is not a fault until you know it was not chosen, that whoever produced a piece of work does not get to judge it, and that how firmly something was said is part of what was said. The package itself is open source. What is sold is the teaching.
9 · Agentic systems with governed decisions. An agent is workforce. A company has values, the values become an HR policy, and the policy settles who gets hired and what they may decide on their own. Nobody runs a team without that. Almost everybody runs agents without it.
Level nine is the same document for them: what an agent may decide alone, what it has to hand up, and what it has to record so that afterwards you can still say who decided, on what basis, and who has to explain it. It is an entrance rather than a rung — a five-person company hiring agents instead of people can start here and never touch the eight levels below.
10 · You use your own LLM model. Open weights, tuned on your own material, running on your own infrastructure. The reason to want it is not the one people expect. The usual worry is whether your data gets out. The larger exposure runs the other way: how much of what you decide on has already been shaped by somebody else. That is far harder to catch, because nothing looks wrong. Every level below this one can check for it. Ten is the only one that removes it.
It is also the only real answer to where the data sits, which matters if you are in Europe and your client data currently travels through an API you do not control.
We have not built one. No fine-tune, no local deployment, no adapter in production — said here rather than discovered later. And the buyer at this level is a different company from the buyer at level one: an enterprise with its own machine-learning people, not a marketing team.
The levels build on each other, and only the first two can swap. Level two needs nothing from your systems — all of it is observable from outside — so it can run first, or while access to everything else is still being arranged. Nothing above those two can be reordered.
Why no agents below nine
Not one. And the reason is not caution. It is that an agent does not hesitate.
A person who gets a number that looks wrong stops and asks somebody. An agent proceeds, at speed, and writes a confident account of what it did. So every defect underneath it stops being a defect and becomes a policy.
Run one on a company where none of the eight is in order and you can name in advance what it will do. It will optimise against numbers nobody ever checked. It will take what the company says about itself as the market. It will not know which directions were closed, and it will find good reasons for a closed one. It will invent the company positions it cannot find written down, and they will be plausible. It will repeat a decision that already failed twice, because nothing recorded that it failed. And it will answer about your industry in general, because nothing of yours was ever loaded.
None of that looks like a malfunction, and that is the whole problem. The output is well argued, it is fast, and it is wrong in a direction nobody is checking.
What breaks if you skip one
This is the more useful of the two lists, because it is what to reach for when somebody wants to start in the middle.
- Skip 1 and every level above inherits numbers nobody checked.
- Skip 2 and you optimise beautifully for a market that has moved.
- Skip 3 and you do good work on the wrong thing.
- Skip 4 and what the company knows stays in people's heads, so every decision starts again.
- Skip 5 and nothing accumulates, because you never find out whether a decision held.
- Skip 6 and every answer is about your industry in general, because nothing of yours is in there.
- Skip 7 and every answer needs somebody to go and produce it, so most questions never get asked.
- Skip 8 and you depend on us forever.
- Skip 9 and the agents run without knowing what good looks like, and afterwards nobody can say what decided.
- Skip 10 and your data trains somebody else's model.
What "your own model" actually means
Open weights first. A model is a very large set of learned numbers, called the weights. Some are published: you download them and run the model on your own machine, and nothing you type leaves your network. The rest are reached over an API, which means you send your text to somebody else and their machine answers. Level ten is only possible with the first kind.
Then the tuning, and this is where the expectation usually goes wrong. A full fine-tune adjusts the whole model and has to be redone from scratch whenever the base model is replaced. A LoRA adapter is a small set of additional weights trained on top of a base that stays frozen — far cheaper to train, and cheap enough to throw away and redo as your material changes. Either way, what the tuning changes is how the model writes and what it reaches for: your vocabulary, the way you frame a decision, the things your company treats as settled.
What tuning does not do is make it know facts reliably. That is retrieval, and retrieval is levels five and six — the record, indexed, and handed to the model when somebody asks. Getting these two the wrong way round is the common and expensive mistake: a company fine-tunes a model on its documents expecting it to remember them, and gets fluent sentences in its own voice with the details invented.
And it will not beat the frontier models at reasoning. That is not what it is for. It is for two things. Your material never leaves your infrastructure. And the thing does not change underneath you — the same question you asked in March comes back different in June because a vendor shipped an update and told nobody, and at level ten that does not happen.
Where to start
You do not buy a level. What gets sold is a shape, and each shape cuts across several levels at once: a read, done once or on a cycle · a build, where something gets installed and then exists · teaching, so your own people can run it · or somebody watching it continuously, month after month.
And there are four honest entrances, not one.
If you cannot tell which of your own numbers to believe, you start at one. Most people start here, and it is the only part of this that a stranger ever searches for.
If your measurement is sound and nobody has read it, you start at two and three. The data is already there. What is missing is the reading.
If we already work together and your material is in order, you start at six, because the loading is simply the next thing that happens.
If you are five people hiring agents instead of employees, you start at nine, and you may never need the eight below it.
What decides is access, not budget. For the entrance at level one that means Viewer rights in your GA4 and read access in your Google Tag Manager — the most routinely granted pair in the industry. No admin rights, no developer, nothing changed while we look, and you can revoke it the day it is finished.
And wherever you start, we look at what is underneath it, because a reading built on an unchecked layer is confident about nothing. Looking is not rebuilding, and it does not commit you to rebuilding.
Tell us what you are trying to fix. That is the whole first step.