Skip to content
Guide title card reading "Build your own AIOS" over a glowing open doorway standing in sunlit storm clouds

← Free GuidesPart 3 of 4, the AIOS series →

What an AIOS Is Made Of: The Three Parts

August 12, 2026

The AIOS series, part 3 of 4. Also in this series: The Operating System Built for You · 48 Jobs You Can Hand to an AI Operating System · How to Build an AIOS

Everyone is building agents. Almost nobody is building the thing agents are supposed to work inside.

That's the gap, and it's why so many impressive agent setups never turn into anything. You wire up a crew, you give them tools, and six weeks later you have a beautifully organized machine that has never once talked to a customer. The agents work on the agents.

An AIOS is the other thing. An AI operating system is the software layer that does your business's recurring work — publishes the content, sends the outreach, watches the ad spend, answers the phone — and keeps doing it whether or not you're awake. Agents are the workers. The AIOS is the company they work for.

I run mine on one box I own, for a fixed monthly ceiling I set myself. This is how I'd start it again from nothing.

One note on the name

If you go searching, you'll find "AIOS" used for something else: an academic project building an operating system kernel for language agents — schedulers, context managers, the layer between an agent and the hardware. Good work, different problem.

This is not that. This is the operating system of a small business: the thing that owns the recurring work and keeps doing it. Same word, and the distinction is worth holding onto, because most of what's written about agent infrastructure is aimed at people building platforms for other developers. You're building a company that runs itself.

The test that separates an AIOS from a toy

Before you write a line: name the job.

Not a capability. A job — something you do by hand, on a schedule, that makes or saves money. Publishing this week's content. Sending the day's outreach. Following up with every lead who didn't answer. Checking what the ads did yesterday.

If you can't name it in one sentence with a dollar attached, you're not building an operating system. You're building a demo.

Name what finished looks like in the same breath, and write both down. Published means indexed, not drafted. Sent means delivered, not queued. Those definitions become the first two lines of the file the whole system reads, and you'll be glad you wrote them while you were calm rather than at nine on a Monday.

My first one was content. I was recording ideas into my phone and then losing two hours a day turning them into something publishable, which meant on busy weeks I published nothing. So that became job one: a ninety-second voice memo goes in, a finished article comes out on ground I own, and it gets cut down and posted to every channel. My hands touch the memo. Nothing else.

One job. End to end. Finishing without me is the whole requirement — a system that gets it 80% there and hands you the rest isn't an operating system, it's a to-do list that writes itself.

Build one loop, not a platform

The instinct is to design the whole thing. Resist it, because you'll spend three weeks on plumbing and never ship the part that pays.

Take your one job and follow it all the way from trigger to receipt. What starts it? What does it need to know? What does it produce? Where does the finished thing go, and how do you find out it happened?

Then build exactly that path and nothing else.

The second job costs you a fraction of the first, because everything under it is already there. That's what the next section is about. Build the parts once, in service of one real job, and the rest of the company snaps on top.

The three parts of an AIOS

If you remember nothing else from this page, remember three words.

The machine. The parts that do the work. The workshop. Where you and your agents change it without breaking it. The leash. What it can never do without you, and how you know it actually ran.

Most of what's written about this covers the machine in detail and the other two barely at all. That's backwards for a business owner, because the machine is the part you'd expect to need. The workshop and the leash are the reason you can walk away from it, and walking away was the whole point.

mermaid
graph TD
  subgraph M[THE MACHINE — does the work]
    D[One door in] --> BR[One brain] --> LG[One ledger] --> ME[One memory] --> RB[One rulebook]
  end
  subgraph W[THE WORKSHOP — where it changes]
    RE[One repo: agents open pull requests, tests block the merge]
  end
  subgraph L[THE LEASH — why you can leave]
    PR[One proof: did anything reach a customer?]
    GA[One gate: what it may never do alone]
  end
  RB --> RE --> PR --> GA --> J[Your one job, running without you]

---

What's inside each one

The machine — five pieces

One door in. Every request enters through a single authenticated endpoint that logs what passes through it. Not a webhook here and a cron job there. When something misfires at 3am, you have one place to look.

One brain. Every call to a model goes through one function. Not "mostly." Every one. That choke point is where you choose the model, cache the prompt and count the spend. The moment one module calls out on its own, your cost numbers are fiction.

One ledger. Money is metered before the call and recorded after, against a key that makes the record impossible to write twice. A timeout does not mean nothing happened — it means you don't know. Ask before you resubmit.

One memory. Its own database, on its own box. If you can't rebuild the whole thing from a clone, an env file and one command, you don't own it — you're renting it from your own past self.

One rulebook. Every number the system judges your work against, in a plain file you can edit rather than buried in code. Then the habit that makes the file worth having: when you disagree with the machine, you edit the file. You don't argue with the machine.

The workshop — one piece, and it's the one nobody writes about

One repo. Your repository is not storage. It's the room where you and your agents work together.

The rule: no agent edits the running system. Agents open pull requests. The work arrives as a diff you can read on your phone. Another agent reviews it adversarially. Tests run automatically and block the merge if they fail. You approve, it merges, the box pulls it down.

That last part matters as much as the rest. The server is a place the truth gets copied to, never the place it lives. Fix something directly on the live box at midnight and the next deploy will quietly delete it — or that fix will be the only copy that ever existed.

And your tests live here, because tests are what let you approve a change you didn't fully read. The rule I hold: a test doesn't count until you deliberately break the code and watch that test fail. A test that has never failed has never proven anything.

The leash — two pieces, and they're why you can leave

One proof. My content system once ran four days without a single error and published nothing. Every step worked. Zero alerts. Nothing reached a customer.

The automation almost never breaks. What breaks is the step before it, the one where a person has to be present for anything to start. Nobody labels that box "Dave opens his laptop," so when Dave is out, nothing runs and nothing tells you. Silence is not success. Silence is just silence.

So measure the outcome, not the machine. Don't count jobs that ran, count things that reached a customer — and treat a zero as an alarm even when nothing failed.

One gate. Your AIOS can spend your money and publish under your name, so decide what it may never do without you. The line isn't "important," it's reversible. Anything it can undo, let it do freely. Anything it can't take back — money leaving, a message to a customer, something published in your name, anything deleted — needs your hand on it.

Make it structural rather than a rule anyone has to remember. In mine, the gates are three columns nobody but me can write. The agents prepare everything up to that line, perfectly, and stop. They aren't trusted to hold back. They're unable.

That's what makes leaving possible. Not discipline. Wiring.

---

Now go build it

That's the architecture. The actual build — the prompts to hand your coding agent for each of the eight pieces, where to run it, and a two-week order to do it in — is its own page, because it's a different kind of reading.

The AIOS Build: prompts, hosting, and a two-week order

Two rules that keep the bill sane

The thing everyone gets scared about is cost, and it's a solved problem if you're deliberate on day one.

Deterministic work never touches a model. Uploading a file, writing a row, posting to a channel, calling an API that returns exactly what you asked for — none of that needs intelligence, and routing it through a model is how people end up with a hundred-dollar day doing work a script does for nothing. Reserve the model for judgment. Use plain code for everything else.

Cheap by default, expensive on purpose. Set the small fast model as the default for every task, then promote the specific jobs that visibly need more. Most of what an operating system does all day is classification, extraction and formatting — work the cheap model does perfectly well. My whole system runs under a hard ceiling I set in the vendor console, with a second softer ceiling the system reads out of the rulebook and refuses the call before it's made rather than after. That second number belongs in the file, not the code — it's the one you'll want to move the first month you get busy.

Two ceilings, one default. That's the entire cost strategy.

Where people go wrong

Failure modes, all of which I've either hit or watched somebody hit.

Agents that work on agents. The most common one by far. If your system's output is a better version of your system, you've built a hobby with a build pipeline. Point it at revenue or turn it off.

Every module calling the model directly. It feels harmless the first time. Then you have eleven call sites, no idea what any of them costs, and no way to change models without touching all of them.

No idempotency. Retries are unavoidable. Double charges are optional.

Editing the server directly. You'll fix one thing on the live box at midnight, forget, and lose it in the next deploy — or worse, lose the only copy that ever existed.

Judgement written into code. Every limit you hard-code is a decision your future self has to re-open a pull request to change, and by then nobody remembers making it. If a number is a matter of taste, it goes in the rulebook. If it's a matter of contract, it goes in the code.

Trusting green. The one that cost me four days. Everything reporting success is not the same as anything happening. Count outcomes, not runs.

Tests that have never failed. A suite that has only ever been green is decoration. Break something on purpose today and find out which of your tests notice.

What it buys you

The obvious answer is time, and that's the small one.

The real one is that the work stops depending on you being available. Not on you being fast — I was always fast. On you being present. Every business I built before this needed me physically there for the last inch of everything, which meant a sick day cost the company a day and a vacation cost it a week.

An AIOS removes that inch. The job runs, finishes and reports, and the only thing that reaches you is the decision that genuinely needed a person — the rest were settled in the rulebook weeks ago, by you, in a calmer mood. When that's true for one job, it's a relief. When it's true for the four or five jobs that run your business, the desk stops being the place where work happens.

That's the whole point of it. Not agents. Not automation for its own sake. A business that keeps working while you're in the water.

Start with one job. Build the machine, then the workshop, then the leash. Then add the second job, and watch how much less it costs than the first.

Then the advanced series

Once the three parts are yours, each one gets a guide of its own in the next series — read in this order.

The org chart. Who does what, and the one rule that lets you leave. Draw the chart first.

The division wall. One shared file every agent reads before it starts and writes to when work ships. How twenty agents stay out of each other's way.

The shared brain. Stop paying your AI to re-read everything and forget it by morning. Give it a memory instead.

The watcher. The step that quietly stops and takes the whole chain with it. Wake an agent only when it matters.

The deploy. Reviewing and shipping from anywhere, without a laptop. How that's wired.

---

The AIOS series

  1. The Operating System Built for You — why nothing you already use was built for the person doing the work
  2. 48 Jobs You Can Hand to an AI Operating System — every job it can take off you, grouped by what each one costs
  3. What an AIOS Is Made Of — the machine, the workshop and the leash ← you are here
  4. How to Build an AIOS — nine prompts, where to run it, and a two-week order

Once yours is running, the next series takes it further — the org chart, twenty agents staying out of each other's way, the shared brain, the watcher, and shipping to the box that runs your company from wherever you happen to be: Run Your Company From Your Beach Chair.

Which part of your week should a machine be doing?

Six questions, about a minute. At the end you get a straight read on what I'd build inside your company — even if the honest answer is “nothing yet”.