What is an AI chief of staff, and what does it actually do?
It is standing work: jobs that run on a schedule inside the AI tools you already pay for, carrying your written standards, reaching the systems you point them at, and stopping where a person has to decide. It has five parts: the jobs, the standards they carry, the connectors, the gates, and a record of what went wrong. The phrase also gets attached to inbox software and to a human hire, and neither of those is this.
Questions as you read? Ask, or take the context files with you.
Where should we send it?
Sent ✓ Download it now →
What has already run by the time you open your laptop
On a Monday morning, before anyone in our firm opens a laptop, part of the week has already happened.
The running document was refreshed on Sunday evening with where every project stands and what moved. Alongside it sits a short list of proposed changes to how the system itself works, none of them applied. The week’s operations briefing has been assembled from the calendar, the inbox, the pipeline and the numbers, and written as a file rather than a notification. The outbound job has selected and staged today’s candidates and sent nothing to any of them.
Four things are waiting. Each one has a place where it stopped on purpose.
That is our own system at bosio.digital, described plainly. It is not a product anyone sold us, and no part of it required an engineer. It is a folder of written standards, a set of scheduled jobs, connections into the systems where our real data lives, and a handful of hard stops.
People search for this under the phrase AI chief of staff, which is a reasonable thing to call it and a phrase that also gets attached to two other products entirely. Those get their own section below. This article is mostly about the thing itself: what it is made of, how it gets built, and what it refuses to do.
Those three numbers say the whole thing in advance. Work can run without you. The evidence that it hands back hours is weak and you should treat it that way. And whatever you build has to cost you less attention than it returns, which is an architecture question before it is a tooling question.
The five context files.
Fill-in templates for what your AI should know, and the page that says where each one loads.
On its way. Check your inbox, or download it now →
What an AI chief of staff actually is
Start with what separates it from the assistant you already have.
An AI assistant answers when you ask. Everything good about it begins with you opening a window and typing, and everything it produces is shaped by whatever context you can remember to give it in that moment. It is genuinely useful and it is entirely dependent on your presence.
Standing work is different in three specific ways. It runs while you are elsewhere, because the trigger is something other than you. It carries your standards, because those standards are written down somewhere it reads every time instead of living in your head. And it stops, because you decided in advance which steps a person has to take.
An AI chief of staff is not something you buy. It is your own standards, written down once, attached to work that runs on a schedule and knows where to stop.
Everything below is those three properties broken into parts you can actually build.
The AI Briefing
Tuesdays. 500+ leaders. No hype, just what works.
The jobs: what runs, when, and what each one produces
A job is a piece of work with a trigger, a scope and an output you can read.
Ours run on rhythms that match the work. A running document refreshes in the early evening and a session log writes at night, so the state of the firm survives the day it happened in. Monday brings an operations briefing and a separate pull on keyword and traffic movement. Tuesday brings the client intelligence run. Friday closes the week with a wrap. Sunday evening runs a pass over the system itself. A debt refresh fires monthly on the second, and a tax check fires quarterly.
Eighteen of these are registered and fourteen are switched on. The four that are off are off deliberately, which matters more than the count.
Two design rules do most of the work here. The first is that a job produces an artifact, not an alert. A file you can open, disagree with and edit is worth more than a summary that pings you, because the file is evidence and the ping is noise. The second is that a job that quietly produces nothing is worse than a job that fails loudly, which is a lesson we paid for and will come back to.
Choosing what becomes a job is the part founders get wrong. The candidates are not the tasks you hate most. They are the tasks that recur on a rhythm you can name, where you already know what good looks like, and whose output you could check in five minutes. Everything else is a project pretending to be a job.
The standards: what good looks like, written in your words
This is the layer that makes the output yours, and it is the layer nobody ships in a box.
A standard is a written answer to a question your team currently answers by asking you. What does a finished proposal look like here. What do we never say to a client. Which numbers may appear in an external document and which may not. Who signs off on what. What counts as done, as opposed to what counts as submitted.
Most of that has never been written down anywhere, which is no failure of discipline. It is what tribal knowledge is: the standard lives in the person, gets applied consistently, and gets explained only when someone new gets it wrong.
Written standards do two things at once. They let work happen to your bar while you are not in the room, which is the actual value on offer here. And they make disagreement visible, because the moment a standard is written, someone will read it and say that is not quite how we do it, which is a conversation worth having and one that never happens on its own.
The fastest way to write one is by counter-example. Describe two pieces of output you would reject and say exactly why each fails. Rejection is specific in a way praise never manages, and most people can produce it in ten minutes for work they have been doing for years.
One rule holds across all of them. The standard is yours. The AI may draft a rule from your past work and propose it, and a person decides whether it becomes one. Inferred habits make poor standards, because a system that learns from what you did will faithfully reproduce your worst week.
The connectors: what it can reach, and what it must not
A job is only as useful as what it can see, and only as safe as what it cannot.
Connectors are the wiring into the places your real work lives: the CRM, the accounting system, the analytics, the calendar, the document store, the repository. Ours reach several of those, some through the platform’s own integrations and some through small local connections we maintain.
The interesting half of this layer is the prohibition. We hold access to two separate CRM portals, one of which belongs to a client. A report once came within a step of being assembled from the client’s portal instead of our own, which would have been a confidentiality breach and not a data error. The fix was not a reminder. A preflight check now asserts the portal identity before any pull and exits if it is wrong, and every output records which portal it read.
That is the general shape of good wiring. Give each job the smallest set of systems that lets it do its work, write down what it must never read, and make the boundary a check that runs rather than an instruction someone remembers.
The gates: the places it stops and waits for you
Here is the part that separates a system you can trust from a demo, and the part almost nothing in this category has.
The outbound job selects and stages candidates every weekday morning and then stops. It never enrolls anyone and never sends anything. Two of the touches in the sequence are set to manual inside the vendor’s platform, so the stop is enforced by software and not by our good intentions.
The weekly pass over the system reads the week’s decisions, failures and lessons, and writes proposals. It does not apply them. Every real change to how the system works goes through a person who can say no.
Publishing anything to our live site requires an explicit go. Drafting, rendering and checking all happen automatically, and nothing crosses into production on a timer.
The billing routine is switched off. It still works, and that is exactly the problem: it writes into an accounting system we no longer treat as the source of record. Left running, it would produce correct-looking entries in a ledger nobody reads, so it stays blocked until it is rebuilt.
A hook forces an approval prompt on any tool that can send something outside the firm. That one exists because a message once went out without an approval, which is usually the only reason anybody thinks to build it.
The value is not in the fourteen jobs that run. It is in the handful of places the system stops and waits for a person.
Gates are where the judgment boundary becomes physical. The AI gathers, compares, drafts from your written rules and lays out options with the evidence against each. It does not deliver its verdict on anything consequential, and it does not decide. That line is the whole argument of judgment happens in the pause, which examines the mechanism properly: people accept a confident recommendation most readily when they have not yet formed a view of their own.
The gates are also where Humans First stops being a slogan and becomes a design constraint. A system built to absorb every decision eventually produces an owner who has stopped practicing the one skill the firm actually sells.
The record: how it improves instead of repeating itself
The last part is a log of what went wrong, and it is the part that turns a set of automations into something that gets better.
A nightly backup failed silently every night for two and a half months. The process had lost a disk permission, the mirror sat frozen at an old state, and the per-run notifications said nothing useful the entire time. The fix was not the permission. The fix was that a job now has to prove it produced something.
A scheduled job fired exactly on time and wrote no file while every indicator stayed green. That class of failure now has a name in our own vocabulary, ran but produced no output, and a script that checks for it.
A fabricated statistic reached a published article and had to be corrected. Every number in a draft now needs its source in the same paragraph, and a gate lists the ones that do not have it before anything ships.
A rules file grew past twelve hundred lines while sitting outside every path the system loads. It was written to for months and read never. Rules now live where the work reads them, or they are enforced by code, because a rule nobody loads is a wish.
What makes that a feature instead of an embarrassment is the pattern. One error is an incident. The same shape of error three times is a design flaw with an address, and you cannot see the shape unless someone wrote the first two down.
How you set it up, from nothing to working
The sequence matters more than the tooling, and it is shorter than people expect.
Two things about that sequence tend to surprise people.
The standard comes before the automation, always. A job wired up before anyone wrote down what good means will produce plausible output at speed, and plausible output at speed is the most expensive failure mode available, because it takes weeks to notice and it erodes the thing clients are actually paying for.
And the gate goes in before the job is trusted, not after it disappoints you. Retrofitting a stop into a system that has been sending for a month is a policy conversation. Building it in on day one is a two-line decision nobody argues with.
What the week looks like once it runs is unremarkable, which is the point. Things arrive as files at the times you chose. A few of them wait for a yes, and clearing those takes a short pass in the morning. One or two will be wrong in an instructive way, and those go in the record. Nobody builds this for a dramatic day. They build it so the work carries their standard while they are elsewhere.
Where the platforms are already enough
Before the case for building anything, the case against it, in its strongest form.
A large share of what people want from an AI chief of staff already ships with the subscription you pay for. Claude has carried memory across sessions since Anthropic released it to Team and Enterprise on September 11, 2025 and extended it to Pro and Max on October 23, 2025, with each Project keeping its own separate memory, summaries you can read and edit, and an incognito mode that writes nothing.
Anthropic also ships a packaged library of ready-to-run business workflows wired into accounting, payments, CRM and document tools, which we covered in what the Claude for small business workflows do and where they stop.
If you live in Microsoft, Work IQ, announced at Ignite on November 18, 2025, is the layer that lets Copilot learn your style, your preferences, your habits and your workflows, and it is exposed to custom agents through Copilot Studio. If you live in Google, Workspace announced in February 2026 that Gmail search would return AI overviews and that its drafting feature would pull from your past mail, chats and files to write in your own voice.
For a lot of people, that is the whole answer. A consultant who tells you otherwise before asking what you already pay for is selling.
The five context files.
Fill-in templates for what your AI should know, and the page that says where each one loads.
On its way. Check your inbox, or download it now →
Where those features stop
They stop in the same five places, and none of it is a knock on the products.
Memory is inferred from behavior, and it is personal. It learns what you did, not what you decided or why you turned down the alternative. It holds no standards: not your review bar, not the work your firm refuses, not who signs off on what.
It also does not run on Monday morning without you, because every one of these features waits for a person to open a window and start typing. It has no gate, so nothing stops the work at the step where a human should decide. And it keeps no record of what kind of wrong it keeps being, which is the difference between catching an error and understanding why that error recurs.
Memory remembers you. Search finds what is written. Neither one knows how your firm works.
Those five gaps are the specification. Everything in the sections above is what it looks like to close them.
Two other things get sold under the same name
The phrase gets attached to two other things, and both camps are serious, so they deserve a straight answer.
One camp means a human hire. Jeffrey Bussgang, a partner at the venture firm Flybridge and a senior lecturer at Harvard Business School, argued in October 2025 for hiring an AI chief of staff reporting to the chief executive, covering operations automation, decision support, cross-functional alignment, workforce development and experimentation. His case rests on AI being a company-wide rewiring, and for the companies he writes for, it is.
The human role predates all of it. Dan Ciampa made the case for a chief of staff in Harvard Business Review in 2020, as a way of making an executive’s time, information and decision processes work better. The reason the idea keeps resurfacing is a real constraint: when Michael Porter and Nitin Nohria studied how chief executives spend their time, it took tracking leaders around the clock for thirteen weeks to reconstruct the calendar of a single week.
The other camp means software. Cortex sells an internal developer portal to engineering leaders and calls the agent on top of it an AI chief of staff. Workboard’s entry on the Workday marketplace is an objectives and key results product with an agent wrapper. Chore, a firm that sells human bookkeeping and back-office staff, publishes an explainer defining the term as software. The leading listicle for the category is published by one of the vendors in it, which ranks itself first.
That category is real and some of it is good. These products triage what arrived, protect the calendar, extract commitments from a meeting, draft the reply and hand you a brief before the day starts. Listed prices in the category’s own comparison run from about $19 a month to a little under $200 a month per person, depending on how much scheduling and workflow gets bundled in, per the vendor listicle published in September 2026.
Precision on names is worth something here. Quill, which raised a $6.5 million seed round announced in February 2026 led by Basis Set Ventures, does not use the search phrase at all. Its own term is a chief of AI staff agent, and the product is local-first meeting transcription with optional end-to-end encrypted sync and connections into Notion, Linear and Airtable.
Michael Daugherty, Quill’s chief executive, described the problem in the funding release as tools that neither talk to each other nor remember how a person actually works. A vendor conceding the coordination gap is worth more than the listicle.
Buy the software when two things are true: you are the only person who has to agree, and the help you want is triage, recall, capture and drafting. It works on the day you install it, and building your own version of inbox triage is a poor use of the one asset you cannot buy back.
A different question hides underneath this one, which is whether AI needs an owner with a title. That has its own answer in our piece on whether to hire a chief AI officer, and the difference matters: that article is about who inside the organization owns AI, and this one is about what you personally run.
The objections worth more than the pitch
Four of them, in descending order of force.
The first is that the measured effect of this entire class of tool is close to zero. Anders Humlum and Emilie Vestergaard linked Danish adoption surveys to administrative payroll records and found precise nulls on earnings and recorded hours two years after chatbot adoption, with bounds ruling out anything larger than 2%, holding even among intensive users, early adopters and workplaces that invested in rollout. Adoption changed what tasks people did without changing what they earned or how long they worked.
That finding deserves to sit in the room. The honest reading is that it measures conversational assistants inside employment, not scheduled work with gates in a firm the user owns, and that the mechanism it exposes is real anyway: time freed inside a job usually gets absorbed by the job. If your case for building rests on getting hours back, the best available evidence says be skeptical of yourself.
The second objection is that the real problem is delegation and not tooling. Gallup’s 2015 analysis of 143 chief executives from the Inc. 500 found roughly one in four scored high on delegator talent, and that the high delegators posted a three-year growth rate 112 percentage points above the low delegators. The objection lands: an owner who builds a machine to absorb work they should have handed to a person has automated the limit instead of removing it.
So be precise about which work is which. Anything that needs your standard, your judgment or your signature cannot be handed to a person or to software, and writing it down is the only way it scales at all. Anything that needs a pair of hands and a clear brief is a hiring problem wearing a technology costume. The test is whether you could write the instruction. If you can write it, someone can do it.
The third objection is that more AI means more to watch, and we published the evidence for it ourselves. Our reporting on AI brain fry covers the BCG Henderson Institute study of 1,488 workers from March 2026, which found that 14% of AI users experience cognitive overload from supervising AI, and that measured productivity peaks at around three simultaneous AI tools. Eighteen scheduled jobs against a three-tool ceiling looks like a contradiction.
It is not, and the reason is the mechanism the study identified. What costs attention is oversight of concurrent streams, not the count of jobs in a registry. Fourteen jobs that report into one place on a fixed schedule are one surface, not fourteen. The gates matter here too: they concentrate attention into a small number of decisions worth making, instead of spreading it thin across everything the system touched.
The fourth is the skeptic’s version, and it deserves its strongest form. All of this is a folder of markdown files, some scheduled prompts and a consulting invoice, and the platforms will absorb the difference within two years. The first half is true and we said it plainly two sections ago. The second half is the bet: the plumbing gets commoditized, the authored part does not, because no vendor can decide what your firm should refuse.
Build for work that meets your standard while you are not in the room. Do not build for hours returned.
Where outside help fits, and the test to apply to us
We sell this, so read the next three paragraphs with that in mind and apply the test at the end to us as readily as to anyone else.
Three things make it hard to build from inside. The first is that writing down your standards means stating things you have never had to say out loud, and most owners discover mid-sentence that their standard is not yet a standard.
The second is arbitration: the moment a second person has an opinion about what good looks like, somebody has to decide whose version becomes the rule, and that is a political act inside your own firm. The third is the loop, because the person with the least time is the one who would have to build the thing that returns it.
That is the argument for the CEO AI program, which is your own system built with you instead of handed to you. If the question turns out to be organizational instead of personal, which happens more often than owners expect, who owns your AI adoption is the piece that sorts it.
Apply one test to us or to anyone else. Ask what you own when the engagement ends. Written standards in your own accounts, gates your team can change without a phone call, and a habit of reviewing what went wrong are ownership. A system only the consultant can operate is a rental, and judgment is the one thing you cannot afford to rent.
Write One Job And The Place It Has To Stop
Pick the recurring piece of work you would hand over first if you trusted the hand. Paste this into your AI assistant. It interviews you until the job has a written standard and a stated gate, and returns both as a file you can run next week.
Context: I am writing the specification for one recurring job I want my AI to run inside my firm. The job is [DESCRIBE IT]. My firm sells [WHAT YOU SELL]. Interview me one question at a time. Push back when my answer is generic, and do not write the standard for me. Draft only what I have said.
Step 1. The trigger: Ask when this job should run: a day and a time, or an event. If I say “when I think of it”, record that the job is not yet a job.
Step 2. The inputs: Ask exactly which files, systems and past examples it should read, and which it must never read. Record the confidentiality limits verbatim.
Step 3. The standard: Interview me until I can describe what good output looks like, with two examples of unacceptable output and why each one fails. Most firms have never written any of this down, so keep asking.
Step 4. The gate: Ask what this job must never do without me: send, publish, pay, commit, promise. Write those as explicit stops, and name who approves each one.
Step 5. The failure record: Ask how I will know when the output was wrong, and where that gets written down so the pattern is visible in a month.
Output: A one-page job specification with the trigger, the inputs, the standard in my own words, the explicit stops, and an UNDECIDED list of everything I could not answer, kept separate at the end.
Whatever lands on the UNDECIDED list is a call your firm has never actually made, which is exactly why no subscription could have made it for you. See where you stand →
Start with one job
The version of this that fails is the ambitious one: a plan for a system that runs the whole business, written on a Sunday and abandoned by Thursday.
The version that works starts with one job, one written standard and one explicit stop, built inside the AI plan you already pay for. It takes an afternoon. What you learn in that afternoon is not how to configure anything. It is how much of what you consider obvious about good work has never been written down anywhere, which is also the reason nobody could sell you this finished.
The system does not become valuable when it does a lot. It becomes valuable when what it does is unmistakably yours, and when it stops in the places you decided it should stop.
Sources
- You Need to Hire an AI Chief of Staff. Now. (Jeffrey Bussgang, Flybridge and Harvard Business School, October 2025): the case for the human hire, and the scope he puts on the role
- The Case for a Chief of Staff (Dan Ciampa, Harvard Business Review, May 2020): the reference case for the human role, published before the AI framing existed
- How CEOs Manage Time (Michael E. Porter and Nitin Nohria, Harvard Business Review, July 2018): large-company chief executives tracked around the clock for thirteen weeks; cited here for the study design
- Quill launches sovereign Chief of AI Staff agent and raises $6.5M (PR Newswire, February 2026): the seed round, the local-first architecture, and the category’s own preferred term
- AI Chief of Staff (Cortex): an internal developer portal for engineering organizations sold under the phrase
- 8 Best AI Chief of Staff Tools in 2026 (get-alfred.ai, September 2026): the category listicle, published by a vendor that ranks itself first; source for the listed subscription prices
- AI Chief of Staff (Chore): an operations-outsourcing firm defining the term as software
- Claude memory (Anthropic): memory released to Team and Enterprise on September 11, 2025 and to Pro and Max on October 23, 2025, with separate memory per project
- Claude for Small Business (Anthropic, May 2026): the packaged workflow library and its integrations
- Microsoft Ignite 2025: Copilot and agents built to power the frontier firm (Microsoft, November 2025): Work IQ as the personalization layer under Copilot
- More personalized and proactive assistance in Gmail (Google Workspace, February 2026): AI overviews in Gmail search and drafting from your own past mail and files
- Large Language Models, Small Labor Market Effects (Anders Humlum and Emilie Vestergaard, Becker Friedman Institute, 2025): precise null effects on earnings and hours, ruling out anything above 2%
- Delegating: A Huge Management Challenge for Entrepreneurs (Gallup, Sangeeta Badal and Bryant Ott, 2015): 143 Inc. 500 chief executives, delegator talent and three-year growth
- AI Brain Fry: What It Is, Why 14% of Your Team Has It (bosio.digital): the BCG Henderson Institute study of 1,488 workers and the three-tool threshold
- Judgment Happens in the Pause (bosio.digital): why executives accept a confident recommendation before forming their own view
Frequently Asked Questions
What is an AI chief of staff?
It is standing work: jobs that run on a schedule inside the AI tools you already pay for, carrying your written standards, connected to the systems where your real data lives, and stopping where a person has to decide. It has five parts: the jobs, the standards they carry, the connectors, the gates and a record of what went wrong. The same phrase also gets used for inbox software and for a human hire, which is why the search results disagree with each other.
How is that different from the AI assistant I already use?
An assistant answers when you ask, shaped by whatever context you remember to give it in that moment. Standing work runs on a trigger that is not you, applies standards that are written down instead of recalled, and stops where you decided in advance that a person must approve. The assistant depends on your presence. The point of the second thing is that the work meets your bar while you are not in the room.
What should the first job be?
Work that recurs on a rhythm you can name, where you already know what good looks like, and whose output you could check in five minutes. Those three filters rule out almost everything that feels urgent and leave you with something genuinely repeatable. Write the standard before you wire anything up, run it by hand twice against that standard, and only then give it a day and a time.
Is built-in AI memory enough on its own?
If you are the only one who has to agree, often yes. Claude has carried memory across sessions since September 2025, with separate memory per project, and Microsoft and Google both ship personalization layers inside their own suites. Those features stop in five places: memory is inferred from behavior, it learns what you did and not what you decided, it holds no standards or review bar, it waits for you to start it, and it has no gate where work stops for a human decision.
What should an AI system never decide?
Anything consequential. A workable boundary is that the AI gathers, compares, drafts from written rules and lays out the options with the evidence against each, while a named person forms their own read, decides and signs off. In practice that means an outbound job that stages messages and never sends, a change process that proposes and never applies itself, and a publishing step that waits for an explicit go.
How is this different from hiring a chief AI officer?
A chief AI officer is a senior executive accountable for a company’s AI strategy, governance and returns, which is a question about who inside the organization owns AI. The question here is what you personally run. They are separable, and most owners resolve the personal system long before the org chart question becomes real.
Does the research show AI assistants actually save time?
The strongest study available says be careful. Humlum and Vestergaard linked Danish adoption surveys to administrative payroll records and found precise null effects on earnings and recorded hours two years after workplace chatbot adoption, with bounds ruling out effects larger than 2%, holding even for intensive users and workplaces that invested in rollout. Task composition changed and hours did not. Build for work getting done to your standard when you are not in the room, not for hours returned.
Will running many AI jobs overload me?
It can, and the mechanism is specific. A BCG Henderson Institute study of 1,488 workers in March 2026 found that 14% of AI users experience cognitive overload from supervising AI, and that measured productivity peaks at around three simultaneous AI tools. What costs attention is concurrent oversight, not the number of jobs in a registry. Work that runs on a schedule and reports into one place is a single surface, and explicit gates concentrate attention into the few decisions worth making.



