In Practice: Building an AI Company | Lesson 16: When to Hire, and the Six Months Before That

I wrote the job description at one in the morning at the end of a week in which everything had gone slightly wrong at once. It was a good job description. It was specific, it was honest about the stage, and it described a person who would have solved every problem I had experienced in the previous five days.

I saved it as a draft and did not post it.

Three weeks later I read it again, in a normal week, and it described three different jobs held together by exhaustion. Two of those jobs did not exist. The third one did, and it was not the one I had emphasized.

The bad reason, and the good one

The bad reason to hire is that you are tired. Tiredness is real, it deserves to be taken seriously, and it is not a hiring signal, because it is neither durable nor specific. It tells you that last week was hard. It does not tell you which work will still be there in six months.

The good reason is that a particular, repeatable piece of work has existed roughly unchanged for a full quarter, you can describe what somebody would do with it in their first ninety days, and doing so takes a founder off the critical path.

Those are three separate tests and it is worth applying them individually.

Durability. Has this job existed, in recognizable form, for three months? A bottleneck that is two weeks old is frequently a symptom of something else, and hiring against it institutionalizes the symptom.

Specificity. Can you write what they do in week one, month one and month three, and state how you would know by month three whether it was working? If you cannot, you are not hiring for a job. You are hiring for relief, and relief is not a role.

Removal. Does this hire take a founder out of a critical path, or does it add a person who needs a founder?

The third test is the one people fail, and it fails quietly. Every early hire consumes founder time before they return any, typically for two to four months. That is a genuine investment with a genuine payback period, and it is worth modeling as one. Two hires made in the same month can consume more founder capacity than they release for an entire quarter, which is exactly the quarter you were hoping to get back.

The benchmark that replaced headcount

Headcount used to be a credibility signal. It is now closer to the opposite, and the change happened quickly.

The extremes are genuinely striking. Companies in this category have reported annual revenue in the billions with headcounts measured in dozens, producing revenue per employee figures in the tens of millions of dollars, numbers that had no precedent in software before 2024. Nobody should plan to be an outlier. What matters is that the outliers reset the reference point, and investors now routinely ask for revenue per employee alongside growth and burn multiple.

The practical consequence for a founder is a reversal of an old instinct. Announcing that you are forty people used to say we are real. It now invites the question of why it takes forty, and if the honest answer is that twelve could not have done it, that is a fine answer. If the honest answer is that you hired because you had money, that will surface in the burn multiple from Lesson 10.

The order that usually works

The first hire is rarely a generalist, despite the persistent belief that early companies need someone who can do a bit of everything. In practice a generalist at a five-person company becomes a second version of the founders, and you did not need a second version of the founders.

The first hire is usually whoever owns the thing that is currently breaking and that no founder wants to own permanently.

Two roles come up repeatedly at this stage and both are good bets. The first is somebody who sits with customers and makes the product work in their environment: integrations, data mapping, the specific way this customer’s systems are wrong. This removes founders from implementation, which is usually where founder time is disappearing, and it feeds directly into the integration moat from Lesson 14.

The second is somebody who owns evaluation and data quality: the golden set, the scoring, the incident analysis. It is unglamorous, it is the thing founders keep deferring, and it is the difference between a product that improves and one that drifts.

The hires that are usually premature at this stage are a marketing generalist with no distribution thesis to execute, a junior engineer hired to help rather than to own something, and any operations leadership role before there are operations to lead.

Where they sit, and what that costs you

If your founders are in one country, your entity is in another, and your first hire is in a third, you have a set of obligations that nobody warns you about until you have already created them.

Classification is a real risk, not a formality. Engaging somebody as a long-term full-time contractor, doing core work, under your direction, using your systems, with no other clients, is the exact fact pattern that authorities in many jurisdictions treat as employment regardless of what the contract says. The consequences are back-dated social contributions, penalties and sometimes employment rights you did not know you had granted. Getting this wrong is expensive and it is entirely avoidable.

An employer of record buys speed at a price. A third party employs the person locally and invoices you. It solves compliance, payroll and local benefits without you registering an entity in their country. It costs a monthly fee per person and you give up some control over terms. For your first hire in a new country it is almost always the right first move.

Equity does not travel well. This is the one founders get most wrong, out of genuine goodwill. Option treatment varies enormously by country. In some jurisdictions the taxable event occurs at exercise, at the value on that date, with no liquidity available to pay it. A grant that is a gift in one country can be a liability in another, and the person receiving it will not find out until it is too late to restructure.

Ask before you grant. A thirty-minute conversation with somebody who knows the local treatment, per country, before the offer goes out. It is a small cost and it prevents the specific situation where a loyal employee ends up worse off for having been given something.

Hire against a bottleneck that has survived a quarter, never against a week that felt hard.

Tomorrow: the term sheet, and the difference between buying time and selling a say.

In Practice: Building an AI Company | Lesson 15: The Security Questionnaire You Cannot Answer Yet

It arrived as a spreadsheet attachment with a two-week deadline and a cheerful covering note. One hundred and eighty questions across nine tabs. It came from the first customer who was actually going to matter, the one whose name would have changed every subsequent conversation, and it arrived four days after they said they wanted to move forward.

We could answer about sixty of the questions honestly. Another forty we could answer with work. The remainder asked for things that take months of elapsed time and cannot be produced by working harder.

That is the lesson, and it is the whole lesson. Some things you can build under pressure. Trust is not one of them, because trust is largely made of time you have already spent.

Trust has a lead time

Take the most commonly requested attestation. A Type I report assesses whether your controls are appropriately designed at a single point in time. Preparation, remediation and the audit itself typically consume two to three months, and it is achievable from a standing start if you commit.

A Type II report assesses whether those controls actually operated over a period, and the period is the point. It requires an observation window, commonly three to twelve months, during which evidence accumulates. There is no version of this you can accelerate with money or urgency. If a buyer requires Type II and you begin the day they ask, you are a minimum of six months from an answer, and their procurement cycle is not going to wait.

This is why the certification belongs on the roadmap during the quarter when you have no customers asking for it. It is the least satisfying work available and it is pure elapsed time, which makes starting early the only lever that exists.

What the hundred and eighty questions are actually asking

Behind the spreadsheet there are about six real questions, and every tab is a restatement of one of them.

  • Where does our data physically go, and through whose hands?
  • Who inside your company can see it, and how do you know?
  • Is it used to improve anything, including anybody’s model?
  • What happens, specifically, if you are breached?
  • What happens to our data if you cease to exist?
  • Who is personally accountable for these answers?

Answer those six well and most of the spreadsheet resolves. The reason founders find the exercise so painful is not the volume. It is that the questions require you to know your own data flows precisely, and a company that has been shipping quickly for a year frequently does not.

One item deserves separate attention because it is new and it catches people. Your model provider is now a subprocessor. That means they appear on a list your customer’s legal team will read, alongside your hosting provider and your error monitoring service. It means when you change providers, or add a second one for routing as described in Lesson 7, you may have a contractual notification obligation. Founders who treat model selection as a purely technical decision discover this at renewal.

The three documents you need before the deal, not during it

A data processing agreement with a current subprocessor list. Templates exist and are fine as a starting point. What is not fine is a subprocessor list assembled from memory in an afternoon. Keep it accurate as you add services, because you will be asked to attest to it.

A one-page security overview. Written by you, honest about what exists and what does not. Encryption in transit and at rest, access control, logging, where data resides, who has production access, what your review process is. One page, dated. This document alone will short-circuit a surprising proportion of the spreadsheet, because a reviewer who can see you have thought about it clearly asks fewer follow-ups.

An incident response commitment with a number in it. Notification within a stated number of hours, to a stated contact, with a stated escalation path. The commitment matters more than the sophistication of the plan behind it.

Underneath all three sits the actual prerequisite: a diagram of where data enters, where it rests, where it is processed, where it leaves, and how long each copy persists. If you cannot draw it, you cannot answer honestly, and answering dishonestly is a much worse day later.

The questions that did not exist three years ago

Modern questionnaires now carry a section that is specific to this category, and it is worth preparing real answers rather than reassuring ones.

  • Is our data used to train models? The answer needs to be a contractual commitment that flows through to your provider terms, not a setting somebody toggled. Buyers increasingly ask you to evidence it.
  • Where is inference performed? A geographic question with regulatory consequences. Know the answer for every provider and every routing path, including the cheap tier you added last month.
  • What is retained at the provider, and for how long? Including any abuse monitoring retention, which is frequently longer than teams assume.
  • Does a human ever review the content? Under what circumstances and with what controls.
  • How do we learn when the model changes? This connects directly to Lesson 13. A buyer in a regulated function needs to know that the thing they validated is the thing that is running.
  • How is the system classified under applicable AI regulation? European obligations in particular now impose documentation and transparency requirements that vary with the risk category of the use case. Treat this as an architecture question rather than a legal one, because the answer determines what you have to log and that is not something to retrofit.

Saying no honestly is a stronger position than founders expect

The instinct when the spreadsheet arrives is to make every answer as green as possible. It is the wrong instinct, and reviewers are extremely good at detecting it.

The position that works is honest maturity. We do not hold a Type II report. Our observation period began in March and completes in September. Here is our current control set, here is our independent penetration test from last quarter, and here is what we are willing to commit to contractually in the interim.

Buyers accept that far more often than founders anticipate, because every security reviewer has dealt with small vendors before and what they are assessing is partly whether you are the kind of company that tells them things. The answer they will not forgive is the one that turns out to have been optimistic, discovered during an incident, when the conversation is no longer about procurement.

We lost that first deal. Not because of the answers, but because the elapsed time we needed did not fit inside their quarter. The work we did to answer the sixty questions properly meant the next one, five months later, took nine days.

Trust has a lead time. Start it before the deal that needs it.

Tomorrow: the job description you wrote and did not post.

In Practice: Building an AI Company | Lesson 14: Your Moat Is Whatever You Would Have to Rebuild

Print the architecture diagram. Take a pen. Cross out the box in the middle, the one with the model provider’s name on it, and everything that connects directly to it.

Now look at what is left standing on the page.

That is the exercise, and it takes about four minutes. The first time we did it, the remaining diagram was noticeably thin, and there was a silence in the room that told me everybody had reached the same conclusion at the same speed.

Whatever survives that pen is the company. Whatever falls over was rented, and rented on terms set by somebody else.

What survives

Six things, in rough order of how hard they are to reproduce.

The data you accumulated by operating. Not licensed data, which your competitor can license too. The record of what your product suggested, what the user accepted, what they corrected it to, and what turned out to be right three weeks later. This exists only because you ran, it is impossible to buy, and it is the reason the same architecture in your hands and a competitor’s hands produce different products after eighteen months.

The integrations. Reaching into the systems where work already lives, reading real state, writing back correctly, handling the version of that system from four years ago that your customer still runs. It is slow, tedious, unrewarding work. It is a moat for exactly those reasons.

The evaluation harness and the golden set. This one surprises people, and it should not. That spreadsheet from yesterday encodes months of accumulated domain judgment about what a correct answer looks like in your specific field, contributed by practitioners, refined by incidents. A competitor with unlimited engineering cannot generate it. They have to earn it the same way you did, and it takes the same amount of time.

The workflow state. Approvals, audit trails, who saw what and when, what happens when the approver is away. This is the difference between a capability and something a business can actually operate, and it is the substance of what makes a product hard to remove.

The regulatory position. Certifications you hold, agreements you have signed, jurisdictions you can serve. These take months of elapsed time that cannot be compressed with money or effort. Tomorrow’s subject.

Cost engineering. A twenty point gross margin advantage over a competitor selling something similar is a strategic asset, not a finance detail. It lets you price lower, spend more on acquisition, or survive a downturn that removes them. Routing, caching and context discipline are not housekeeping. Compounded, they are a moat.

What does not survive

It is worth being blunt about the things founders describe as differentiation which are not.

Prompts. A prompt is copyable in one screenshot and can be obsoleted by one model update. Prompt engineering is a skill, and skills are not moats. If your defensibility discussion involves the phrase our prompts, there is no defensibility discussion.

Feature sets. Parity now takes weeks rather than quarters. A feature list is a snapshot of who shipped most recently.

Being early. Being early is a head start. Head starts are valuable and they expire. The question is what you converted the head start into while you had it, and for a lot of companies the honest answer is more features.

Your choice of model. Whatever you picked, your competitor can pick it too, this afternoon.

Design for portability, deliberately

The practical consequence is an abstraction layer between your application and whichever provider serves your inference. Not for architectural elegance. For three concrete reasons: you can move if pricing changes, you can route between tiers for margin as described in Lesson 7, and you can keep operating if a provider has an incident during your customer’s business hours.

There is a real cost and it should be stated rather than glossed. An abstraction that reduces every provider to a common denominator throws away the distinctive capabilities that made you choose one in the first place. Providers differentiate, and the differentiated parts are frequently the useful parts.

The rule I have settled on is to abstract the interface but not the capability. Standardize how you call, log, retry and measure, so that swapping is an engineering task rather than an archaeology project. Then, where a specific provider capability genuinely gives your product something, use it deliberately and write down what you would lose and what the fallback is. A dependency you have chosen and documented is a decision. A dependency you drifted into is a liability that surfaces during a pricing negotiation you did not expect to have.

The uncomfortable version of the exercise

Sometimes you cross out the box and very little remains. At month six that is completely normal and not a cause for alarm. You have not had time to accumulate anything yet.

At month thirty it is the most important fact about your company, and it will not be fixed by shipping faster.

The move is to treat the moat as a deliverable rather than a hope. Pick the two items from the surviving list that are most achievable in your situation, put them on the roadmap with dates and owners, and protect that work from the customer request that arrives on Thursday. Data capture is usually the cheapest and highest return, because it costs almost nothing to log and cannot be reconstructed retrospectively. Integrations are next, because they are pure elapsed time and the earlier you start the further ahead you are.

Nobody will ask you for this work. No customer will thank you for it. It is the difference between a company and a very good implementation.

Ask what you would have to rebuild if you switched model providers on Monday. Whatever is left is the company.

Tomorrow: the security questionnaire, and why trust has a lead time.

In Practice: Building an AI Company | Lesson 13: Build the Eval Harness Before You Build the Demo

The golden set is a spreadsheet with about two hundred rows. Each row has an input taken from real usage, the output a competent human produced for it, and a short note on why that output is right. It is the least impressive artifact this company has ever produced. Nobody has ever asked to see it in a meeting.

It is also the only reason we found out, on a Tuesday in the middle of an otherwise ordinary month, that the product had been getting worse for eleven days.

Nobody had deployed anything. Nothing had broken. The model underneath had been updated, quality on one category of task had moved, and every dashboard we had was green because every dashboard we had measured uptime, latency and error rates. All three were excellent. The product was worse and the instrumentation had no opinion about it.

A problem with no equivalent in ordinary software

Conventional software is deterministic. The same input produces the same output, so a regression is a failing test and a failing test is a red build. The whole discipline of continuous integration rests on that property.

An AI product does not have it. The same input produces different outputs. Worse is not a boolean, it is a movement in a distribution, and you cannot detect a movement in a distribution without a reference point.

There is a second difference which is stranger and which founders reliably underestimate. Your product can change without you changing it. The capability you built on is maintained by somebody else, on their schedule, for their reasons. A model update can improve nine things and degrade the tenth, and the tenth might be the one your customers actually bought.

In ordinary software, the ground does not move. Here it does, and you need an instrument that notices.

What a golden set actually is

One hundred to three hundred real cases, drawn from real usage, each with an output that somebody competent has agreed is correct.

Four things make the difference between a useful one and a decorative one.

It is built by the person who does the job, not by an engineer. The engineer knows what the system does. The practitioner knows what right looks like, and those are different pieces of knowledge. If you have design partners from Lesson 6, this is one of the most valuable things to ask them for, and it is a reasonable thing to pay for.

It contains the hard cases deliberately. A set made of representative examples tells you about the average, and the average was never the problem. Include the ambiguous inputs, the malformed ones, the ones where two answers are defensible, and the ones where the correct behavior is to refuse or escalate. That last category is the one most sets omit and the one that matters most for an accountable product.

It is versioned like a product asset. Not a file in a test directory. It has an owner, it grows when incidents reveal a gap, and every entry records who decided and when.

It is small enough to run constantly. Two hundred cases you run on every change beats two thousand you run quarterly. The value is in the frequency.

Three gates, on every change

Once the set exists, it becomes a gate, and it should sit in front of three questions rather than one.

Quality. Score the change against the golden set. Block on any regression beyond an agreed threshold. Not a warning in a log, a block.

Cost. Compute cost per query for the new path and compare it against the current one before merging. A feature that improves quality by two percent and raises cost per query by four hundred percent is a decision somebody should make consciously. Without this gate, that decision gets made by an engineer at eleven at night and discovered by you in an invoice.

Latency. The same discipline. A change that adds two seconds is a product change even if nobody wrote a product ticket.

The critical scope point: these gates run on changes to prompts, retrieval configuration and model selection, not only on changes to application code. A prompt is code. It is the most frequently edited and least reviewed code in most AI companies, and it is routinely modified directly in production environments by people who would never do the same to a function.

Scoring without pretending

Exact matching works for extraction and classification. It does not work for anything generative, and pretending otherwise produces an evaluation that punishes correct answers for being phrased differently.

The practical approach is rubric scoring with a model acting as judge: define what a good answer contains, have a model grade against that rubric, aggregate. It is fast, it is cheap, and it works well enough to be genuinely useful.

It is also biased in ways that are easy to forget. Models grading models tend to reward fluency, length and confidence, which are three properties of exactly the output described in Lesson 12. So two safeguards are not optional. Sample a fixed percentage for human review every cycle and budget the hours properly. And track the rate at which the humans disagree with the judge. That disagreement rate is your instrument for whether your instrument still works, and when it drifts you have an evaluation problem rather than a product problem.

Pin your versions

Two habits close the loop on the eleven day problem.

Pin explicit model versions rather than accepting whatever the latest alias resolves to. Then, when a new version appears, run the golden set against it, look at cost and latency alongside quality, and adopt it as a deliberate change with a date attached.

And log the model version, prompt version and retrieval configuration with every single request. When a customer says it used to do this correctly, that logging is the difference between a two-hour investigation and a fortnight of theories.

Building this before the demo feels wrong. The demo is what gets you the meeting. But the demo is a claim, and the harness is the only thing that lets you keep making the claim after the fourth model update, the second retrieval rewrite and the day you route half your traffic to a cheaper tier to protect your margin.

If you cannot detect a regression, you do not have a product. You have a demo that has not broken yet.

Tomorrow: the architecture diagram, and the question of what you would have to rebuild on Monday.

In Practice: Building an AI Company | Lesson 12: AI Slop Is a Distribution Problem, Not a Taste Problem

The churn email was four lines long and entirely polite. They were not renewing. The team had enjoyed working with us. And then the sentence that took a while to stop thinking about: the output was fine, but the team had started pasting things into a general assistant instead because it was faster.

Not the output was wrong. Not it was too expensive. It was fine, and fine was available elsewhere at no cost.

That is the actual mechanism of slop, and it is not an aesthetic complaint. Slop is a commercial condition.

A definition that is actually useful

Slop is not bad output. Bad output is easy to detect and easy to fix, and customers will tell you about it loudly.

Slop is output that is fluent, plausible, competently structured and completely substitutable. It reads well. It is not wrong. It could have been produced by anybody with access to a general purpose model and ten minutes.

The distinguishing property is not quality. It is substitutability. That is why slop is a distribution problem: it does not damage your reputation, it damages your reason to exist. Customers do not leave angry. They leave the way that email left, politely, having found the same thing on a shorter path.

Three places it lives

In the product. The dangerous property here is that product slop does not hurt acquisition at all. Acquisition is driven by the promise, and the promise demos beautifully. Slop hurts retention, exclusively and on a delay, which means your dashboard looks healthy for two quarters and then does not. If your churn is concentrated at the four to six month mark among engaged users, this is the first thing to check.

In the codebase. Yesterday’s subject. Code that is syntactically clean, passes linting, compiles, and encodes no judgment about the system it lives in. It is the same failure in a different medium: plausible, fluent, undifferentiated, and expensive later.

In the marketing. The most common and the least examined. Content production used to be a genuine moat because it was expensive: you needed writers, editors, subject knowledge and time. All of that collapsed. Which means volume is no longer an advantage, because your competitor has the same collapse available to them.

Publishing a great deal of competent, undifferentiated content in 2026 is paying for noise. Worse, it trains your audience that your name is attached to things not worth reading, which is a hard association to undo.

The substitution test

There is one exercise I recommend to everyone and almost nobody wants to run.

Take ten real tasks from real customer usage. Have somebody outside the product team complete them twice: once with your product, once with a general purpose assistant and a well-written prompt. Then show both sets of results to a customer without telling them which is which.

You are not looking for a win. You are looking for whether they can tell. If they cannot reliably distinguish the two, your differentiation lives entirely in convenience and interface, and you should read Lesson 5 again before spending another quarter on features.

Most teams avoid this test, and the avoidance is itself informative. It is a cheap experiment with a high information yield, which is exactly the profile of an experiment people skip when they suspect the answer.

What the opposite looks like

Four properties, and they are all consequences of decisions rather than of model quality.

  • It is specific to this customer. Their data, their history, their naming conventions, the constraint they mentioned in month two. A general model cannot produce this because it does not have the inputs. This is the data asset from Lesson 5 doing its job.
  • It knows what it does not know. Output that flags uncertainty, declines to guess, or routes to review when the confidence is low. This is genuinely difficult to build, it is what an accountable product does, and it is the property customers cite most often when explaining why they trust one tool over another.
  • It is short. Slop is verbose because verbosity is free and length signals effort. Concision requires a decision about what matters, and a decision about what matters requires knowing the domain. Length is the most reliable surface indicator of undifferentiated output.
  • It encodes an opinion. Somewhere in the product there should be a judgment you made about how this work should be done, which a general tool would not make because it has no stake in the outcome. That opinion will lose you some customers. It is also the reason the others chose you.

Three ways this goes wrong

You respond to slop by adding features. The instinct when retention softens is to ship more. If the underlying problem is substitutability, more features produce a larger substitutable product. The fix is depth in one place, not breadth across five.

You measure satisfaction instead of substitution. Satisfaction scores stay high right up until the churn email, because customers are satisfied. They are just also indifferent, and no standard survey question distinguishes those two states. Ask instead what they would do if you disappeared on Monday. The answers cluster into we would have a serious problem and we would manage, and only one of those is a business.

You let volume stand in for differentiation in your marketing. Publishing five undifferentiated pieces a week costs you time, budget and reputation, and returns nothing that compounds. One piece a month containing something only you know, because you operate the product and see the data, is worth more than all five, and it is the thing that gets cited rather than skimmed.

We won that customer back eventually, though not with the same product. What changed was that the output started carrying information from their own history that a general tool had no way to reach. It was less impressive and considerably harder to leave.

If a model can produce your output, it will produce your competitor’s too. Sell the part it cannot.

Monday: the golden set, and how you find out your product got worse.

In Practice: Building an AI Company | Lesson 11: Vibe Coding Is a Prototyping Tool That Keeps Getting Promoted

The pull request was fourteen hundred lines across nine files. The description said add user flow. It was approved four minutes after it was opened, by somebody who was in a meeting at the time, and it went to production that afternoon.

Nothing bad happened. That is the important part. Nothing bad happened for about eleven weeks.

I want to be careful here, because there is a lot of moralizing about this subject and most of it is written by people who are not shipping anything. The productivity gains are real. Teams report forty to sixty percent improvements on bounded tasks and I believe them, because I have seen it. The problem is not the tool. It is that a tool designed for exploration keeps getting promoted into a role it was never designed for, and the promotion happens silently.

What the data actually says

The most careful measurement available comes from an analysis of two hundred and eleven million lines of code changed between 2020 and 2024 across repositories at large technology companies. Two findings matter.

Copy-pasted code rose from 8.3 percent of changed lines to 12.3 percent, a relative increase of roughly half. Over the same period, refactored lines fell from around twenty-five percent to under ten percent. In 2024, for the first time on record, copy-pasted code exceeded refactored code.

That inversion is the whole story in one statistic. Refactoring is how a codebase stays comprehensible. Duplication is how it stops being comprehensible. The tools made it cheaper to add and no cheaper to consolidate, so the ratio moved, and it moved in the direction that compounds.

A separate study of around eight million pull requests found technical debt rising by thirty to forty percent after teams adopted AI coding tools. And in a detail I find quietly telling, the researcher who coined the term vibe coding in early 2025 had, by early 2026, publicly described it as past its moment and shifted to advocating a model with substantially more human oversight in the loop. The phrase did not survive its first year before its inventor moved on.

Where the line goes

The useful framing is not whether to use these tools. It is which code is allowed to be produced this way.

Generate freely. Interface components, internal dashboards, one-off scripts, data migrations you will inspect anyway, test fixtures, prototypes, anything you intend to throw away, anything with no access to production data. This is the majority of code by volume and the speed gain here is close to free.

Never without dedicated review and automated analysis. Authentication. Authorization, which is a separate and more frequently broken thing. Payment handling. Anything touching personal data. Anything that writes to production. Anything that decides what a user is allowed to see.

The failure modes in that second category are boringly consistent. Across incidents reported through 2025 and into 2026, the same handful of causes recur: databases left open by default configuration, row-level security never enabled, keys hardcoded into code that reaches the client, backend services exposed without authentication. None of these are exotic. All of them are the kind of thing that a generated solution produces because it satisfies the request as stated and nobody stated the rest.

The ninety-day reckoning

There is a recognizable arc and it runs on roughly a quarterly clock.

Weeks one to four, velocity is extraordinary and everybody is delighted. Weeks four to ten, small inconsistencies accumulate: three different ways of handling errors, two competing patterns for the same operation, functions that do four things because it was easier to extend than to separate. Around week ten to twelve, the first incidents arrive, and they are hard to diagnose because nobody wrote the code and the person debugging it is reading it for the first time under pressure.

By day ninety, teams commonly report spending twenty to thirty percent of sprint capacity on defects traceable to generated code. The velocity gain from month one is gone. Then it goes negative, because every subsequent change is harder in a codebase nobody trusts.

This is now a market. Rescue engineering, the business of taking a product that was built quickly and rebuilding it into something maintainable, is a recognized specialty with recognized pricing, commonly quoted between fifty thousand and five hundred thousand dollars depending on how long the product ran before somebody looked underneath it. Of the several thousand startups that shipped production applications using these tools through 2025, a large majority reportedly needed some form of partial rebuild or rescue work within the following year.

The cost is not the rebuild. The cost is the six months you spend doing it instead of building.

The protocol that actually works

Five rules, and none of them require slowing down much.

  • Somebody must be able to explain it. This is the only rule that really matters. If nobody on the team can explain how authentication works without opening a chat window, you do not have authentication. You have an arrangement that has not yet been tested.
  • Cap diff size on anything sensitive. A fourteen hundred line change is not reviewable and everybody knows it, which is why it gets approved rather than read. Small changes get read.
  • Static analysis on every path that handles credentials, permissions or personal data. Automated, in the pipeline, blocking. This costs nothing per run and catches the recurring failures listed above.
  • Write the tests by a different route than the code. Generated code validated by generated tests derived from the same description will agree with itself confidently. Specify the tests from the requirement, not from the implementation.
  • Keep your junior engineers and give them maintenance work. This one is strategic rather than tactical. The people who will be able to maintain generated systems in two years are the people who spend this year reading and modifying code they did not write. That is exactly the work that is being automated away from them, and teams that eliminate it are removing the pipeline for the skill they will most need.

We did eventually go back and read the fourteen hundred line change. It was mostly fine. It also contained a permission check that ran on the client and not on the server, which meant the check was a suggestion. Finding it took an afternoon. Finding it after somebody else had would have taken considerably longer and involved a lawyer.

You can ship code you do not understand. You cannot run a company you do not understand.

Tomorrow: the churn email, and the three kinds of slop.

In Practice: Building an AI Company | Lesson 10: Burn, Runway and the Twenty-Month Reality

The runway model is one tab in one spreadsheet and there is exactly one cell in it that anybody actually looks at. Everything else exists to produce that cell. It contains a month and a year, and once you have seen it you cannot unsee it.

For a long time I treated that cell as the deadline. It is not the deadline. It is roughly six to nine months after the deadline, and confusing the two is how founders end up negotiating from a position where the only honest answer to how much runway do you have is a number that makes the terms worse.

The real deadline is earlier than the cell

Raising money takes time. Building a list, first meetings, partner meetings, diligence, documents, close. Three months is fast. Six months is normal. Longer is common and not a sign of failure.

You also cannot raise well with two months left. Investors ask, you answer, and the conversation changes shape immediately. Not because anyone is cruel, but because a company that must close in six weeks has a different set of options than one that could walk away, and everybody in the room understands that.

So the working deadline is the money-out month minus the fundraising duration minus a buffer. If your cell says November and raising takes five months, your real deadline is around April, with the assumption that you will not be building product in April. Founders consistently discover this in month three of a raise.

The twenty-month problem

There is a specific timing assumption baked into a lot of seed-stage plans and it has quietly stopped being true.

The median gap between closing a seed round and closing a Series A has stretched considerably, on recent measures to around twenty months. That is a substantial change from the twelve to fifteen month planning assumption most founders inherited from earlier cohorts, and it is not a temporary market condition so much as a rise in the bar. Series A now expects repeatability, typically somewhere in the range of one to three million in annual recurring revenue with evidence that acquisition can be repeated rather than recounted.

The arithmetic follows. If you raise a seed and plan twelve months of runway to reach a Series A that on average arrives at twenty, you have planned a crunch and scheduled it for a moment when you will also be trying to sell. Plan for twenty-four to thirty months, which usually means either raising more or spending less, and spending less is the option you control.

Two clocks, running at different speeds

Here is the modeling error I see most often, and I made it myself: one burn number.

An AI company has two kinds of burn and they behave nothing alike.

Headcount burn. Salaries, contractors, the tools people need. You control it completely. It moves in steps, because hires are discrete events. It is sticky in the downward direction, both practically and morally. It is almost perfectly forecastable, which is why finance people love it.

Compute burn. Inference and everything metered per request. Your customers control it. It moves continuously, it is only partly forecastable, and its defining property is that it rises when things go well.

Blend those into one line and you get a number that is wrong in both directions. It understates your risk in a good month and overstates it in a quiet one.

The scenario worth modeling explicitly is the uncomfortable one. Growth accelerates. Usage deepens. Compute burn rises immediately, because inference is consumed the moment the work happens. Revenue lags, because you invoice monthly or quarterly and business customers pay on their own schedule, which is thirty to sixty days after that. For a period of weeks, sometimes months, your runway is getting shorter precisely because you are winning. Nobody who learned software finance in the subscription era expects that shape, and it has caught out companies whose only mistake was growing quickly.

The number investors will ask for

Growth rate on its own has stopped being a sufficient answer, because anybody can buy growth. The question now is what the growth cost.

Burn multiple is net cash burned divided by net new annual recurring revenue over the same period. If you burned two million and added two million of net new recurring revenue, your burn multiple is one. Below one is excellent and rare. One to one and a half is a good business. Above two needs an explanation, and above three needs a plan rather than an explanation.

It is a better question than growth because it is difficult to flatter. It captures pricing, retention, sales efficiency and cost of goods in a single ratio, and it penalizes exactly the behaviors that look like traction and are not.

Calculate it quarterly. Show it before you are asked, because a founder who volunteers their burn multiple is telling the room something about themselves independent of the number.

The question almost nobody calculates

If you never raise another round, and your current growth rate and current cost trajectory continue, do you reach profitability before the money runs out?

That is the whole question. It has a yes or no answer and it takes an hour to compute properly. In my experience most founders have never done it, and a meaningful number of them assume the answer is no when it is actually yes with two decisions changed.

What makes it valuable is not the answer. It is what it does to the conversation. A founder who knows they can survive without raising is negotiating. A founder who does not know is hoping. Investors can tell the difference within about ten minutes, and it is the single largest determinant of terms that founders treat as unchangeable.

Three ways this goes wrong

You model revenue as booked rather than collected. Signed is not invoiced, invoiced is not paid, and the gap between them is where small companies die with healthy-looking pipelines. Model cash, on the date it arrives, with realistic terms.

You cut compute instead of headcount, or headcount instead of compute, without checking which is which. Under pressure teams reach for the lever that is easiest emotionally rather than the one that is largest. Compute burn is frequently reducible by thirty percent or more through routing and context discipline, without anyone losing a job. That work should be done before any conversation about people, and it usually is not.

You update the model quarterly. Cost per query moves weekly, prices move, and one shipped feature can change your unit economics overnight. A runway model refreshed four times a year is a historical document. Monthly is the minimum, and the person who maintains it should be a founder, not somebody a founder asks.

Headcount burn you control. Compute burn your customers control. Model them separately.

Tomorrow: the pull request nobody reviewed, and the ninety-day reckoning that follows it.

In Practice: Building an AI Company | Lesson 9: The Free Tier Is a Loan You Are Making

The usage dashboard is a bar chart with one bar per account, sorted descending. The first time I looked at it properly, the tallest bar by a considerable distance belonged to somebody who had never paid us anything.

They were not abusing it. They were using it exactly as intended, enthusiastically, every working day, for a job they genuinely had. They were, in the language of the old playbook, a fantastic signal. They were also, in the language of the invoice from Lesson 7, our third largest cost.

This is the part of the AI business model that catches people who learned software economics before 2023, and it catches them somewhere around month five.

The free user used to be free

In classic software a free user costs a row in a database, some storage, and a fraction of a shared server. The marginal cost is genuinely close to zero, which is why the freemium model became near universal. You could carry a hundred thousand free accounts on the same infrastructure that served your paying ones and barely notice.

That is no longer true, and the change is not subtle. A free user in an AI product consumes inference on every request, at the same rate as a paying one, and there is no version of the architecture where that goes away. A heavy free account can cost you more per month than a median paying account generates.

The consequence is that a free tier is no longer a marketing decision with a rounding error attached. It is a spending program, and it belongs in the same conversation as any other spending program: what is it buying, how much is it costing, and how would we know if it stopped working.

What free is actually buying

There are three legitimate answers and you should be able to name yours in one sentence.

Distribution. Free users tell other people. This works when the product is visible in use, when the user has an audience, or when the output carries your name somewhere. It works poorly for internal back-office tools, which is most business software.

Data. Usage improves the product in a way that benefits paying customers. This is real, and it is the strongest argument, but it only holds if you are actually capturing and using the signal. A free tier that generates logs nobody reads is not buying data, it is generating logs.

Qualification. Self-service reduces the cost of finding buyers. The free tier does the first three sales calls for you, and the people who arrive at your pricing page have already decided. This is the most common real answer and the easiest to measure.

If none of those three describes what you are doing, you have a free tier because everybody has a free tier. That is a preference, not a strategy, and now it has a monthly cost attached to it.

Designing a loan rather than a gift

Five design choices do most of the work.

  • Route free traffic to the cheaper model tier. The routing discussion from Lesson 7 applies with more force here. A quality difference that would be unacceptable in a paid product is entirely acceptable at a price of zero, and it can cut the cost of your free tier by most of it. Be honest about it in your documentation rather than hiding it.
  • Cap by unit of value, not by time. Fourteen days of unlimited access rewards the person evaluating you least seriously, because the serious evaluator needs three weeks to get their data in order. Ten documents, or fifty runs, or one project is better: it survives a slow start, it maps to something the user understands, and it costs you a bounded amount.
  • Rate limit, and show the limit. A visible counter is not a hostile act. It converts better than a silent ceiling because it tells the user what the paid version is for.
  • Require something. An email address costs a user nothing. A work email costs slightly more. Connecting a real data source costs real commitment and filters hard. Ask for the largest thing you can justify at the point where the user is about to receive value.
  • Measure it as a cohort and be willing to end it. Cost per free user per month, conversion rate by cohort, and time to convert. If a cohort has not converted in ninety days, it probably will not, and continuing to serve it is a decision you should be making on purpose.

The trial alternative

Freemium is not the only option and it is frequently the wrong one for business software with real serving costs.

A time-boxed trial of the full product, with payment details taken up front, does most of what freemium does at a fraction of the cost and with a far better conversion rate. Taking the card is not a trick, it is a filter, and the people it filters out were unlikely to buy.

The reverse trial is worth knowing about too: full functionality for a short period, then automatic downgrade to a genuinely limited free tier rather than a wall. The user has felt the good version and lost it, which is a much stronger motivator than never having had it, and your ongoing cost sits at the limited tier rather than the full one.

Three ways this goes wrong

You build a free tier that is good enough. The most expensive version of this mistake is not the compute bill, it is that you have shipped a competitor to yourself, staffed it, and given it away. If a meaningful share of your free users are getting the job done without paying, the tier is not a funnel, it is the product. Find the line where value becomes real and put the wall exactly there.

You underestimate abuse. A product that turns requests into money spent is an attractive target. Keys embedded in client-side code get scraped. Accounts get created in bulk. Somebody discovers your endpoint is a cheaper route to a frontier model than paying for one directly. This is not hypothetical, it is a normal Tuesday, and the defenses are ordinary: server-side keys, per-account and per-address rate limits, anomaly alerting on cost rather than on traffic, and a hard spend ceiling with somebody’s phone number attached.

You never calculate the cost per free user. It is one division and almost nobody does it. Total free-tier inference spend divided by monthly active free accounts. Put that number next to your conversion rate and your average contract value and you can answer, in about thirty seconds, whether the program is an investment or a habit.

We kept our free tier, narrowed it substantially, and routed it to a cheaper model. Conversion went up. That was not the outcome I expected and it is the one most teams report, which suggests the generous version was never doing the work we imagined.

Every free user is a loan you make in compute and hope to repay in conversion. Know the interest rate.

Tomorrow: the runway model, and the two clocks that run at different speeds.

In Practice: Building an AI Company | Lesson 8: Price the Outcome, Floor the Cost

The pricing page has three columns and the middle one is highlighted, because every pricing page has three columns and the middle one is highlighted. It took longer to agree than the architecture did, and unlike the architecture it could not be refactored quietly on a Thursday.

Pricing is the most consequential product decision most founders treat as a marketing decision. It determines which customers you attract, which ones you can afford to keep, what your salespeople argue about, and whether growth improves your margin or destroys it.

In a category where serving a customer costs real money per request, it also determines whether your most enthusiastic user is an asset or a liability.

Per seat is dying, and the reason is arithmetic

Seat-based pricing fell from around twenty-one percent of software companies to about fifteen percent inside twelve months. That is a fast move for something as sticky as a pricing model, and the cause is not fashion.

If your product means one person can now do the work that used to take ten, then per-seat pricing asks the customer to pay you in proportion to the number of people who did not get more productive. Your revenue falls as your value rises. You have built a machine that reduces the size of your own invoice.

Buyers worked this out quickly, and the more successful your deployment the faster they work it out. It is a difficult position to argue your way out of at renewal, because the customer is right.

Per seat is not dead. It remains sensible where the product augments a person who still does the job, where usage per person is roughly uniform, and where the buyer’s mental model is headcount. It is a poor fit for anything that completes work autonomously.

The four models and where each one wins

Per seat. Predictable for both sides, easy to forecast, easy to sell. Breaks when one seat can do ten seats of work, and breaks badly when consumption varies by an order of magnitude between users on the same plan.

Per unit of consumption. Charging by tokens, calls or compute. Natural for infrastructure and developer products where the buyer is technical and understands what they are consuming. In an application sold to a business buyer it creates billing anxiety, which is a real commercial problem: a finance team that cannot forecast your invoice will cap usage, and capped usage is capped value.

Per outcome. Charging when something measurable happens. Per resolved support conversation is now an established pattern at prices well under a dollar, and it has spread from support into sales and back-office work. The alignment is genuinely elegant: the vendor is paid when the thing works. The difficulty is definitional. What counts as resolved? Who decides when the customer disputes it? What happens when the model does eighty percent of the work and a human finishes it? Every one of those becomes a contract clause and eventually a support ticket.

Hybrid. A base subscription with an included allowance, plus overage above it. This is now the default. Adoption rose from roughly twenty-seven percent to forty-one percent in a year, and by some counts more than nine in ten AI software companies use some blended model with a consumption component in it.

The reason hybrid won is not that it is elegant. It is that pure models each fail in one direction. Pure subscription exposes you to the heavy user. Pure consumption exposes the customer to an unforecastable bill. Pure outcome exposes you to definitional argument and leaves money on the table with high-frequency users. Hybrid fails in none of those directions completely.

The rule

There is one principle underneath all of this and it is short enough to write on the wall.

The unit you charge for should be the unit that costs you money.

When those two diverge, the customer who loves your product most is the one damaging you most, and you find out at exactly the moment you would like to be celebrating. A flat rate plan with an unbounded heavy user is the clearest version. That account uses the product forty times more than the median, costs you real money on every request, renews without hesitation, gives you a testimonial, and quietly consumes the margin from six other accounts.

You do not have to charge per token. You do have to make sure that when consumption goes up substantially, revenue goes up too.

Designing the hybrid

Four components, and each one has a job.

  • The base. It covers your floor: the cost of existing, the cost of operating, and the cost of serving the included allowance. If the base does not cover the allowance at full consumption, you have priced a loss and made it recurring.
  • The included allowance. Size it so that the large majority of customers never exceed it. That matters psychologically more than financially. A customer who never sees an overage experiences your product as a predictable subscription, which is what their finance team wants, while the meter is still there for the ones who need it.
  • The overage rate. Priced with real margin, not at cost. Overage is not a penalty and it is not a favor. It is the part of the model that keeps you solvent at the top of the distribution.
  • A cap or an alert. Nobody should ever receive a surprise invoice from you. A notification at eighty percent of allowance and a hard ceiling that requires a decision to lift buys you more goodwill than any discount.

Three ways this goes wrong

You price against the trial month rather than year three. This is the mistake buyers make and vendors mirror. The cheapest model at low volume is frequently the most expensive at scale, and the reverse. What matters is the shape of the cost curve at projected volume: linear, sub-linear or step function. Model your own pricing at ten times current usage before you publish it, because you will live with the structure much longer than the numbers.

You adopt outcome pricing without being able to define the outcome. Outcome pricing is the most aligned model and the most operationally demanding. Before you commit, write the definition, write the dispute process, and write what happens in the partial case. If you cannot write those three paragraphs clearly today, you are not ready to sell it, and you will spend the first year arguing about invoices instead of selling.

You set it once. Willingness to pay moves. It rises as a category matures, as your product improves, and as the buyer’s alternatives get worse or better. Companies that grow well revisit pricing at least twice a year, deliberately, with data. Companies that set a price in month four and defend it for three years are leaving a great deal on the table and usually discover it during a competitive loss.

Charge for the thing that costs you money, or your best customer becomes your worst one.

Tomorrow: the usage dashboard, and the loan you are making every time somebody signs up for free.

In Practice: Building an AI Company | Lesson 7: The Cost Model Is the Business Model

The first model provider invoice arrived itemized by day. Thirty-one rows, most of them small, four of them not. The four large ones corresponded to days when we had been testing, which is to say to days when nobody was using the product at all.

What struck me was not the amount, which was trivial. It was the shape. This was a bill that moved with activity, and every previous software bill I had ever seen was a bill that moved with the calendar.

That difference is the whole subject. It is the reason an AI company is not a software company with a model attached, and it is the line item that quietly decides what business you are in.

Inference is cost of goods sold

Classic software runs at eighty to ninety percent gross margin because it is built once and the cost of serving the next customer rounds to nothing. Hosting is largely fixed and spreads across a growing base. That single fact produced two decades of business models, valuation multiples and hiring plans.

An AI product cannot do that trick. Every request spends real money. The marginal cost of the next customer is not near zero, it is a function of how much they use you.

The numbers now in circulation are consistent across sources and worth committing to memory. AI product gross margins are landing in the fifty to sixty percent band rather than the eighty to ninety percent band. One widely cited 2026 survey of several hundred software executives put the average AI product gross margin at 52 percent, up from 41 percent in 2024. Model inference accounted for somewhere around twenty to twenty-three percent of total AI product cost.

Read that last figure carefully, because the direction is the surprising part. Inference share rises as products mature. It does not fall. The intuition from classic software, where unit cost decays as you scale, is exactly inverted. You succeed, usage deepens, cost of goods grows as a proportion of the total.

A founder who assumes the old curve will build a plan in which margin improves automatically with scale. It does not. It improves with engineering, or it does not improve.

Cost per query, and why it is a weekly number

There is one metric that should be on a wall. Total inference spend for the period divided by requests served in the period.

Track it weekly, not monthly, because monthly is slow enough that a bad change ships, propagates and becomes normal before anyone notices. Early-stage products commonly sit anywhere from a fraction of a cent for a lightweight completion to fifteen cents or more for a multi-step agent run that calls tools and reasons across several turns. The spread across that range is enormous and it is entirely determined by choices you make.

When the number rises, there are only two explanations and both are worth knowing about immediately. Either you shipped a more capable feature without optimizing the path underneath it, or the amount of context being consumed per task has grown, usually because somebody added history, retrieved documents or examples to a prompt and it helped.

Neither is wrong. Both are decisions, and a decision you did not know you made is the one that shows up as a margin problem two quarters later.

The levers, in order of size

Routing. This is the largest lever by a wide margin and the one most teams reach for last. Within a single provider’s lineup the price spread between the cheapest capable tier and the most expensive frontier tier commonly runs five times or more per token. Most requests in most products do not need the top tier. Classify the incoming request, send the routine majority to the cheap tier, reserve the expensive one for the cases that genuinely need it, and measure quality on both. Teams that do this well route the large majority of traffic to inexpensive models and see no user-visible degradation.

Caching. Repeated context, whether system instructions, retrieved documents or conversation prefixes, does not need to be paid for at full rate every time. The mechanisms differ by provider and the savings on a workload with stable context are large enough to change the shape of a P&L.

Context discipline. The cheapest token is the one you never send. There is a common pattern where a retrieval step returns twenty documents, all twenty go into the prompt because it is easier, and three of them were relevant. Retrieval quality is a cost lever disguised as a quality lever.

Batching and asynchronous processing. Where the user is not waiting, work that runs on a delayed queue is materially cheaper. A surprising proportion of what teams build synchronously does not need to be.

Self-hosting. Last, deliberately. There is a crossover point where running your own inference beats paying per token, and it exists, but it arrives at volumes higher than most early companies reach and it brings an operational burden that a team of six should think hard about accepting. Do the arithmetic before the conversation, not during it.

Model three futures, not one

Unit prices for inference have fallen dramatically. Capability that cost roughly twenty dollars per million tokens from a frontier model in late 2022 was available at a small fraction of that by early 2026. Hardware rental rates for the relevant accelerators fell substantially across 2025 as well.

It is tempting to plan on that continuing. Do not plan on it exclusively. Build three scenarios.

In the first, prices keep falling at roughly ten percent a quarter and today’s fifty percent gross margin drifts toward seventy without you changing your pricing. That is the optimistic case and it is genuinely plausible.

In the second, prices hold flat. Your margin is whatever your engineering makes it.

In the third, your cost per unit of work rises even as token prices fall, because you shipped more capable features and the number of tokens consumed per completed task grew. This has been the actual experience of a lot of teams. Token consumption per task has risen by one to two orders of magnitude since late 2023 as products moved from single completions to multi-step reasoning with tool use.

If the business only works in the first scenario, that is important information and it is available today rather than in eighteen months.

Three ways this goes wrong

You put inference in operating expenses. It sits in cost of goods sold. Booked as an operating cost it disappears into a line with the design tool subscription, your reported gross margin becomes fiction, and you will not discover the error until somebody in diligence asks a question you cannot answer.

You measure the average and ignore the distribution. Usage in these products is heavily skewed. A small number of users generate a large share of consumption. An average cost per customer that looks healthy can conceal a handful of accounts that are individually unprofitable, and on a flat rate plan those are the accounts most likely to renew enthusiastically.

You optimize before you measure. The instinct to cut costs early is good and it is frequently spent in the wrong place. Instrument first: cost per request, per customer, per feature. Then optimize the thing that is actually large, which in my experience is almost never the thing anyone guessed.

In AI you do not discover your margin at year end. You design it at the start, or somebody else designs it for you.

Monday: the pricing page, and the rule that stops your best customer becoming your worst one.