Book a call

7 Oct 2026

Tokens make code

For thirty years software was configured to fit, and the business changed to fit the software. That deal is ending because code is becoming cheap. This is what a business does with that: a shared glossary, a process described as it is and as it could be, dashboards, command centers, the limits, and what happens when an agent acts with more access than it should have.

We start with something that has nothing to do with AI, because it explains why the rest matters now.

Thirty years of configuring

Every software product of the last thirty years has been sold on the same deal. You buy the product and configure it as close to your need as it will go. Where it does not fit, you change the way your people work. You do not change the code. The code belongs to the vendor and is expensive to change, so the organisation bends.

The deal made sense when code was expensive. It does not hold any more, because the cost of writing code is heading towards zero. Software is not free, but the step from "this is how we want it to work" to "it works like that" has become small. What stays expensive, and matters more every month, is understanding the business well enough to describe how it should work.

The glossary comes first

Start with the words. In most companies the same word means different things in different departments. Take customer. Sales says it is a company that has signed. The warehouse says it is a delivery address. Support says it is the person who wrote in this morning. Invoicing says it is whoever is named on the invoice.

The same happens with order, case, product and project, the five terms we start with at any client. Nobody notices until data from two systems is put side by side, and then every number is wrong.

The fix is a glossary, sometimes called a conceptual model: the five terms, what each department means by them today, and the one definition the whole company will use. It is a few pages of plain text, and it is the first deliverable of any build.

Many clients cannot change the ERP or the CRM. But with dashboards and command centers you can lift the business out of those systems and let people work with their own words instead of the vendor's. The screen says "customer" and means one thing, and code behind it translates to whatever each system calls it. That is why the glossary has to exist first. It is what the code is built against.

The process is the foundation

The glossary gives the words. The process description gives the flow. Together they are the specification.

When nobody has to think about data storage or screen layout, what remains is the process. How it works today, the as-is. How it could work, the to-be. And a third question that is new: in which steps of the to-be does intelligence go, meaning AI. Most steps do not need it. Moving a record from one status to another is plain code. Reading a messy document and deciding what it is needs judgement, and that is where a model goes.

Process mapping is decades old. What has changed is where the output goes. It used to end up in a slide. Now it is the direct input to a build, and it is the first thing we do in FastTrack.

Tokens make code

Now to the AI. The most valuable thing you can make with tokens is code.

You can also make text and slides with tokens, and those are useful. But a summary is read once and forgotten. Code is written once and keeps running. It can replace a spreadsheet that someone updates every Friday or show two systems side by side. Use AI only for documents and you get faster documents. Use it for code and you get things that keep working while you are asleep.

That is Emberloom's thesis.

What you need to know about the technology

The short version of the AI basics course we run in our workshops: four ideas are enough.

The token. A model does not read words. It reads tokens, which are pieces of words. In English one token is about three quarters of a word, so a 200-word description is roughly 270 tokens. Everything is measured in tokens: what the model can hold, how fast it answers, what it costs.

The context window. Think of it as a desk. Everything on the desk, the model can see. Anything else does not exist for it. A contract that is not in the window cannot be quoted.

Prediction. The model predicts the most likely next piece of text given everything before it. That is all it does. The result is often very good, but there is no internal sense of truth behind it and no signal that says "I do not know". When it lacks the facts, it writes what a plausible answer would look like. Ask for the notice period in a supplier contract it has not seen and you get a specific, confident, invented number. Paste the contract into the window and you get the right one. This is called hallucination. A clever prompt will not remove it, and a lower temperature setting only makes the wrong answers consistent.

Where it fails. Models are poor at arithmetic and counting, so you let code do the sums. They do not know what happened last week unless you give them the data. They sometimes skip step three of a five-step process and lose an instruction deep in a long document. The design responses are code for anything that must be exact, the data in the window, small steps, and the critical instruction near the top.

Code handles the part that has to be right. The model handles the part that needs reading and judgement. That is why code is the valuable output.

Dashboards: the gap between your systems

Every company has an ERP, a CRM, a ticketing tool, a mailbox and a pile of spreadsheets. Each knows part of the truth. Nobody sees the whole.

A dashboard closes that gap. You write code that pulls information from two or more systems into one live view that belongs to this company. Open orders from the ERP next to open complaints from the support mailbox.

Until recently a dashboard that joined two systems meant a project with a developer, so everything small stayed in someone's spreadsheet. Now a person who knows what they need can describe it, a model writes the code, and a working first version can exist the same afternoon. The cost moves to knowing which systems to join and what the view should show. That is where the glossary and the process description pay off.

From our own delivery: we have built dashboards and command centers for product information management, for SEO text production, for CRM handling and for B2B ordering. In these cases the client saved tens of thousands of SEK, in the larger ones hundreds of thousands of SEK, and in some cases could decommission a system they had been paying for. These are our own results, not a market figure.

Most systems carry fifty features the client pays for and three they use. In two cases we replaced a larger system outright, a CRM system in one and the image handling in a DAM in the other. We wrote only the code for the basic functions, where the client needed nothing more. That will not work everywhere. But "do we keep paying for this system" now has a real alternative. Our cases show what this looks like in practice.

Train the business to build

If the only people who can build these things are consultants, you have made a new dependency. It is better than a three-month project, but still a queue.

The more useful step is to hand the ability to the people in the organisation. The person who runs the supplier invoice process knows what they need to see better than any outsider does, and with simple means can create the tools that make their own team more effective.

We think of it the way people once learned Excel. Nobody trained as a programmer to use a spreadsheet. They built something small for their own work, and it spread. Writing a clear order for a model is a skill of the same size. When a person watches a model join data from two systems into one live view in minutes, they realise they could have asked for it years ago, and they start seeing the gaps in their own work. In our experience that moment is the real start, and the point where they are ready for actions.

The command center: the human decides

A dashboard shows. A command center also acts, and it also judges.

The first addition is actions: the tool can change something in a process. It sends the reminder, sets the status to on hold, creates the task, writes the draft reply. The second is judgement: the tool uses AI to assess whether an invoice looks right, how urgent a ticket is or what the next sensible step is for a contract. The person gets a proposal for the next action, with a reason, and approves or changes it. The tool reads, matches and prepares. The person decides.

What is new is that this is cheap enough to build for use cases with a small payback. One command center saves little. Many, each built for a team that needed exactly that, save a lot.

Take supplier invoices, described the way we would describe them for a build.

As-is. Invoices arrive as PDFs in a shared mailbox. One person types the supplier, amount and reference into the ERP, finds the purchase order, checks it against what was ordered and received, and forwards it to the approver. The stuck ones are found when a supplier calls.

To-be. The mailbox feeds one screen. Each invoice is a row with supplier, matched purchase order, approver and days waited, and a proposed action with a one-line reason: approve and pass on, ask the supplier, ask the buyer, hold. One click carries it out.

Where intelligence goes. Two steps: reading the PDF into fields, and judging a mismatch, such as forty units on the invoice and thirty-six on the delivery note. Everything else is plain code, because it has to be exact. The model is never allowed to do the sums.

That is about ten lines and a complete specification. The head of the process can write it in an afternoon.

Agents, briefly

From there agents can run a business process continuously and save people's time. An agent is not smarter than a prompt. It is a prompt in a loop, with tools and rules: something happens, it checks a condition, takes an action, verifies the result and waits for the next event.

That is the next level, and not the whole story. Most of the value is in the earlier steps, which are cheaper and safer. We move a process to an agent only once people trust the command center's proposals and the exceptions are rare.

What has to happen first

Three prerequisites, and the order matters.

First, people learn the tools that generate code. As simple as learning to write an order or speak a need, the way people once learned Excel: how to describe what you want, how to read what comes back, how to ask for a change. Without it, the consultant stays the bottleneck. That is what FastTrack is for.

Second, guardrails. Generated code has to be tested and has to meet the requirements you set: automated checks that run every time something changes, review by a person who can read the result, and a clear rule for what may reach the people who use it. The glossary and the process description are what the checks are written against. This is GuardRails.

Third, a safe zone. A place where data is stored safely and access to it is controlled. When people start building their own tools, the data they point those tools at has to sit somewhere with rules about who may see what. Without it, the first useful dashboard is also the first leak. This is SafeZone.

Limits as well as possibilities

This is where it fails.

The failures from the technology section show up in practice as quiet errors: an invoice read slightly wrong, a summary that leaves out the one clause that mattered, a proposal that sounds sure and rests on stale data. The cure is the same each time, with a human approving anything hard to reverse. Volume is not value: a system that answers every ticket quickly and gets a large share wrong is damage at scale.

Three public cases show what the limits look like. Klarna announced in February 2024 that its AI assistant did the work of 700 agents. In May 2025 the CEO told Bloomberg that "cost unfortunately seems to have been a too predominant evaluation factor" and that the result was lower quality, and Klarna started recruiting human agents again. In July 2025 an AI coding tool from Replit deleted a production database for the SaaStr community, ignored an instruction to freeze the code and invented data. Replit's CEO called it "Unacceptable and should never be possible", and the company began separating development and production databases automatically. In October 2025 Deloitte Australia refunded over A$97,000 of a government contract after its report turned out to contain a fabricated quote from a court judgment and references to research papers that do not exist. What the system says and does is the company's own act.

There are other limits. Where the client needs the fifty features, or where scale or integration is the point of the product, you keep the product. Some processes should be stopped before anyone writes a line. Models change, so every tool needs an owner.

A realistic AI transformation is not a programme with a single launch date. It is a glossary for the five terms, written in ordinary meetings. One process described as-is and to-be. A first dashboard that joins two systems. A command center on the process with the clearest owner and the most repetitive judgement. Then more, built by the people who own those processes, with guardrails and safe zone in place before the number of tools grows. That pace is our judgement from our own delivery, not a benchmark, and we do not promise a timeline.

When the agent leaves the sandbox

An agent does not know where the lines are. It acts with the permissions it is given, no more and no less. If it can delete in the database, it can delete in the database, and the Replit case above is what that looks like on a bad day.

The second problem is prompt injection. A model cannot reliably tell your instructions from text it happens to read. So any document the agent opens, an invoice PDF, a supplier email, a web page, can carry instructions, and the agent may follow them. For a command center that reads invoices from a shared mailbox, that is not theoretical.

The public record is long enough. In July 2025 someone got a command into the Amazon Q extension for Visual Studio Code, written in plain English, telling the agent to clean a system to a near-factory state and delete cloud resources. Amazon says it would not have run, and no customer harm was disclosed. In February 2026 an attacker used prompt injection against an AI triage agent at Cline to reach its release pipeline, and a changed package was live for about eight hours. Cline says no malicious code reached users. In November 2025 Anthropic reported that a group it assesses as state-sponsored used Claude Code to attack about thirty targets, companies and government agencies among them, with the agent doing most of the work. These incidents were fetched on 2026-10-07 and are as reported by the sources.

This is why the three prerequisites are the answer and not a checklist: people who can read what the agent did, guardrails that test and review before anything reaches users, and a safe zone where you decide what the agent may reach. If what an agent can touch sits inside a place with rules, a bad instruction has a small blast radius. If the agent runs with your own login, it has all of your reach.

Here is the rule we give every client. Give the agent the least access that does the job. Start read-only, and add the right to change something one action at a time, only after the proposals have been right for a while. Name one person who can pull the plug without asking anyone, with the access to do it today.

Monday

This is what we would do on Monday.

Pick one spreadsheet, or one process that someone repeats every week. Choose one that a person updates by hand from two other places.

Write what it does in ten lines. Where the information comes from, which words it uses, what the person does with it, and what happens next. If the five words are not agreed yet, write down the disagreement instead of hiding it. Mark each step as code or judgement.

Name the owner. One person who knows the process and is allowed to change it. If you cannot name one, that is your first finding.

Choose dashboard or command center. If people only need to see the situation, build a dashboard. If someone takes the same kind of action every time they look at it, build a command center with a proposed action and a reason.

Then take the ten lines to the owner and build the first version together. If you want a partner for it, book an introduction or see how we work.

Listen and read further

Stefan, who runs Emberloom, tells this story in two podcast episodes.

The concepts used here are explained role by role in the Emberloom Share concept library.

Ready to see this for your business

Book a free introduction. We go through where you are, what you want to solve, and how Emberloom AI fits.