What makes AI training for employees actually work?
AI training works when it is a program rather than a course. The curriculum is the easy part and the part every vendor sells.
The program is everything around it: role-specific tracks so a finance manager and a sales lead are not sitting through the same generic session, each track attached to a workflow that person already runs weekly, a named owner accountable for it, reinforcement past the first week, and a measure of success that is behavior rather than course completion.
KPMG’s global study with the University of Melbourne, covering more than 48,000 people across 47 countries, found only 47% of employees have received any AI training at all. Most companies are not choosing between good and bad training. They are starting from nothing, which is an advantage if you build the program before you buy the content.
Want this made concrete for your company? Ask me — or pick one:
Start here
You have approved the budget. Someone has now asked you what the training actually is, and the honest answer is that you are not sure yet.
That is the normal position, and it is a better one than it feels like. The market is happy to answer for you: there are course catalogues, licensed libraries, half-day workshops, and platforms that will report completion rates back to you in a dashboard. All of it is buyable this afternoon.
The problem is that none of it is the decision. The decision is what happens around the content, and that part nobody sells because it cannot be packaged.
So this piece is the program, not the curriculum. Who owns it, which roles get what, how each track attaches to work someone already does, and how you know it changed anything. Buy the courses afterwards if you still need them, and you may need fewer than you think.
That first number is the one worth sitting with. Fewer than half of employees have had any AI training whatsoever, which means the common assumption behind most of these conversations is wrong. You are usually not fixing a training program that underperformed. You are building the first one, from nothing, which is a genuinely easier problem and a rare chance to get the structure right before habits form.
What most AI training gets wrong
One section on this, because the failure is well documented and it is not what this article is for. We have written separately about why your team isn’t using AI despite the rollout, and that piece covers the adoption gap in full.
The compressed version is this. The default program is a single generic session, delivered once, to everyone, disconnected from anyone’s actual job, measured by who attended. Each of those five choices is defensible on its own and together they produce a room full of people who understand what AI is and change nothing on Monday.
The fix is not better content. The content is mostly fine and getting cheaper every quarter. The fix is structural, and the rest of this is that structure.
The AI Briefing
Tuesdays. 500+ leaders. No hype, just what works.
The program is not the curriculum
Here is the distinction the whole article rests on, and it is worth being blunt about because it is what the market obscures.
The curriculum is what people learn. Prompting technique, tool mechanics, what a model can and cannot do, where the risks are. This is a solved and commoditized problem. Good material exists at every price point including free, and the quality difference between the expensive option and the adequate one is much smaller than the price difference.
The program is everything that determines whether the learning survives contact with a working week. Who it is for, in what order, attached to which task, owned by whom, reinforced how, and judged by what evidence.
Vendors sell the first and imply the second. That is not dishonest, it is just the shape of their business: a platform cannot know which workflow your operations lead runs on Tuesdays, and cannot make anyone accountable inside your company. Those are the two things that decide the outcome.
This is the whole of what we mean by Humans First, applied to training. The curriculum treats people as recipients of content. The program treats them as the thing being designed around: their actual work, their existing standards, the manager who will or will not reinforce it, and the fear nobody says out loud about what this technology means for their job. A program that works with how people actually adopt a new way of working beats one that works against it, and the difference is not motivational, it is structural.
The curriculum is a purchase. The program is a design decision. Companies keep buying the first and hoping it produces the second.
The practical consequence is a sequencing rule: design the program first, then buy the smallest amount of content that fills it. Done in that order, you often find you need one or two targeted modules rather than a library license, because most of what your people need is specific to how your company works and no catalogue contains it.
The five things a program needs
These are the parts no catalogue supplies. A program missing any one of them tends to fail in a predictable way, which is noted under each.
Two notes before the detail.
Start with one function, not the company. The instinct is a company-wide launch because it feels fair and it is easier to announce. It is also the reason most programs are shallow: designing five tracks in parallel means designing all of them badly. Pick one, build a real track, and let the second function copy a structure that already works.
Which one to pick is a real decision, and three criteria settle it. Repeatability, because a team whose work changes shape every week has nothing stable to attach training to. A clear owner, meaning someone whose authority over that team’s standards is already accepted, since you are not going to fix an ownership vacuum with a workshop. And a manager who wants it, which sounds soft and is the most predictive of the three: a skeptical manager can end a program without ever opposing it, simply by not reinforcing it.
What should not drive the choice is which team is loudest about AI, or which one is most visible to the board. Enthusiasm makes the first session pleasant and has almost no relationship to whether the behavior is still there in six weeks.
The content comes last. If you find yourself evaluating platforms before you have named the workflows, you are shopping rather than designing. No vendor can tell you which task to attach to.
And the owner’s job is a cadence, not a launch. This is the part that gets assigned as a title and never as a calendar, so it is worth stating concretely. The owner runs the second and third touches rather than delegating them. They hold a monthly thirty minutes with the managers of trained teams, asking one question: is the task being done the new way, and if not, what is in the way. They keep a running list of what people got wrong, which becomes the next cohort’s material and is the single highest-return artifact the program produces. And they are the person who decides when a track is finished and the next function starts, which stops the program becoming permanent scaffolding.
That is a few hours a month, not a role. But it is a few hours that have to be somebody’s, and naming that person is the difference between a program and an event. If nobody will take it, you have learned something more useful than any training would have taught you.
Role-specific tracks: what each function actually needs
“Role-specific” is easy to agree with and easy to fake. Faking it looks like one deck with a different first slide per department. Doing it means each track starts from the work.
The method is the same for every function and takes about an hour per track. Name the two or three tasks that person repeats most often. Pick the one that is highest volume and most inconsistent in quality. Design the training around doing that specific task with AI, end to end, using the company’s real material. Everything else is optional.
Finance and operations. The recurring work is reconciliation, variance explanation, and recurring reporting. The highest-value training is not prompting technique, it is teaching the standard the output must meet and how to check it. This group is usually the fastest to adopt and the most under-served, because training budgets tend to flow to customer-facing teams.
Sales. Call preparation, follow-up, proposal drafting, and CRM hygiene. The trap is training on generic outreach writing, which produces the exact homogenized copy your buyers are already ignoring. Train on preparation and synthesis rather than on generation.
Marketing. The most common starting point and the one most likely to go wrong, because the work is generative and the quality bar is subjective. This track needs the most time spent on what “good” looks like in your voice, and the least time on tooling.
Delivery, service, and support. Drafting responses, summarizing histories, escalation judgment. The critical content here is the boundary: what may be sent without review, and what must never be. That belongs in the training, not only in a policy document.
Everyone, and said out loud once. Before the first track runs, somebody senior has to answer the question people are already asking privately, which is what this means for their job. You do not need a perfect answer and you cannot fake one. What you need is for the question to have been acknowledged rather than left to circulate, because an unanswered version of it sits underneath every session and quietly caps how much anyone invests in learning the thing.
Leadership. A distinct track, and the one most often skipped. Executives do not need tool mechanics. They need enough fluency to tell a real result from a demo, to ask a useful question about an AI-produced number, and to make the decisions the rest of the program depends on. A program whose sponsor cannot evaluate its output is a program that will be judged on enthusiasm.
The material that makes any of these tracks work is your own. Real examples, your standards, your boundaries. That is the same asset described in our context engineering guide, and if you have written those files, most of your training content already exists.
Not sure where you stand?
Take the 90-second AI readiness read — five dimensions, a scored result, and a clear next step.
Take the readiness read →How to measure it
Completion rates are the default measure because they are the easy one to collect and the one platforms report automatically. They tell you almost nothing, because attending is not the behavior you wanted.
Three measures, and the first is the only one that is genuinely hard.
Is the task being done the new way, unprompted? Pick the workflow each track was attached to and check, four to six weeks later, whether people do it the new way when nobody is watching. This is observational rather than analytical: ask the manager, look at the actual outputs, sit in on the work. It is the measure that matters and the one no dashboard supplies.
Do this openly rather than quietly. People can tell the difference between being measured and being watched, and a program that starts feeling like surveillance loses the honesty you need from it. Say what you are looking at and why, and the reporting gets better rather than more defensive.
Depth, not headcount. How many people use AI for real work weekly, as opposed to how many have logins or have completed a module. A license is not adoption and a certificate is not a habit. This is the number to report upward, because it is honest and it moves slowly enough to be believable.
Quality of what comes out. Is the work better, or only faster? Faster and worse is a real outcome and a common one, and it shows up as rework rather than as a training metric. If your reviewers are spending more time correcting than they saved, the training taught production without teaching the standard.
Set all three before the program starts. Measures chosen afterwards are chosen to make the program look good, and everyone involved knows it.
Then act on what you find, because the measurement is not a report card. Each of the three findings has a different fix, and confusing them wastes a cycle. If the task is not being done the new way, the attachment was wrong: you trained on something adjacent to the real work rather than on the work. If usage is shallow but real, the training landed and the reinforcement did not, so add touches rather than content. And if output got faster but not better, the standard was never taught, which usually means it was never written down clearly enough to teach.
That third case is the most common and the most fixable, and it is worth checking before you conclude anything about the people. A team producing fast mediocre work has usually been given a tool and no definition of good. Our wider guide to making AI work for your teams covers the adoption side of this in more depth.
What changes with size
The five parts do not change. What changes is how many people have to agree on them, which is the same axis that governs every other build decision.
On your own, or up to about twenty-five people. You are the trainer, and that is not a compromise. The curriculum is a shared document with your real examples in it, the track is whatever you do most, and reinforcement is you correcting output and writing the correction down. At this size formal training is usually the wrong instrument entirely: what works is doing the work alongside people and capturing what you fix.
Roughly twenty-five to a hundred. Tracks split by function and the named owner becomes essential rather than nice. This is the size where the program either gets a person or gets forgotten, and where the most common failure is assigning it to whoever has the most enthusiasm rather than the most authority. Enthusiasm launches a program. Authority is what protects the time six weeks later.
Past a hundred. It becomes a program with a cadence, a budget line, and real measurement. New starters need an on-ramp, because a one-time company-wide push is invisible to everyone hired after it. This is also the size where buying content genuinely makes sense, since the per-person cost of building everything internally stops being worth it.
The five parts are constant. What scales is not the training, it is the number of people who must agree on what “good” looks like before the training can teach it.
What to buy, and what you should not outsource
An honest split, and we sell one side of it.
Worth buying. Foundational tool literacy, general prompting technique, and anything about how the models work. This is commoditized, it is well made, and building it yourself is a waste of a good month. Certification-style content for teams who want it also belongs here.
Do not outsource. Anything that encodes your standards. The definition of “good” for your deliverables, your boundaries on what may go out without review, your real before-and-after examples. An outside trainer cannot write these because they do not know them, and a program built on generic standards teaches people to produce generic work faster.
Where an outside partner genuinely helps is a narrower band than most firms will admit, and it is mostly the parts that are awkward from inside. Designing the track structure when nobody internally has run one before. Facilitating the sessions where the standard gets argued out, which goes better when the person holding the pen has no stake in whose version wins. And sustaining the cadence through the first two quarters, which is where almost every internally-run program quietly stops.
That is what our AI change management and role-by-role training engagement is, and it is deliberately scoped to hand the program back. Apply the same test here you would apply anywhere: ask what you own when it ends. If you finish with tracks your own managers can run, written standards, and a measurement habit, that is capability transfer. If you finish needing us to run the next cohort, you have bought a subscription to your own training. That is also the standard to apply when choosing any AI consulting partner.
One more thing that is not outsourceable: the decision about who owns this. If nobody owns your AI adoption, a training program will not create that ownership. It will surface its absence about six weeks in, when the reinforcement does not happen and nobody is accountable for the fact.
Your first program, in 30 days
A single track, one function, start to finish. This is deliberately small.
Week one: pick and attach. Choose the function with the most repeatable work and the clearest owner. Name the two or three tasks that team repeats weekly, and select one. Write down what a finished, good version of that task looks like, in enough detail that two reviewers would agree. That document is now the core of your curriculum and you wrote it in an afternoon.
Week two: build the track. Design ninety minutes around doing that one task with AI, end to end, using real company material rather than examples. Include the boundary, meaning what may go out without review and what may not. Buy a short foundational module if the group needs tool basics first, and keep it separate from your track rather than blended into it.
Week three: run it, then run it again three days later. The second session is where the program is won and it is the step that gets cut. It is short, it is about what people actually tried, and it is where the real questions surface, because nobody knows what they do not understand until they have attempted the work.
Three days is deliberate. Long enough that everyone has had a real attempt at the task, short enough that the attempt is still fresh and the ones who struggled have not yet quietly reverted. Run it as thirty minutes with one question: what did you try, and where did it not go the way you expected? Then have the person who solved it show the group, rather than the trainer. Peer demonstration outperforms instruction here for a straightforward reason, which is that a colleague solving your exact problem with your company’s material is proof, and a trainer doing it is a demo.
Whatever comes up in that half hour is your second cohort’s curriculum. Write it down as it happens.
Week four: watch and write down. Look at the actual outputs. Ask the manager whether the task is being done the new way when nobody is watching. Collect what people got wrong and add it to the material, which is what makes the second cohort cheaper and better than the first.
Then pick the second function and hand them a structure that already works, rather than designing another one from scratch. If you want the wider architecture this fits into, the AI operating system piece covers the five stages a company builds, of which this is a piece of the second one.
Design One Real Training Track This Week
Paste this into your AI assistant. It interviews you about one team’s actual work and returns a training track built around it, not a generic curriculum.
Context: I am designing an AI training track for one function in my company, not a company-wide program. We are a [INDUSTRY] company with [NUMBER] people. The function is [TEAM]. Interview me one step at a time, and push back if my answers are generic. I want to finish with something I could run in two weeks.
Step 1. Find the work: Ask what tasks this team repeats weekly. Then ask which is highest volume and which produces the most inconsistent quality depending on who does it. Pick one. Do not accept “reporting” or “communications” as an answer, make me name the specific recurring task.
Step 2. Define good: Interview me until you can write the standard a finished version of that task must meet, specific enough that two different reviewers applying it would reach the same verdict. Push back on adjectives. Ask for a real example of a good one and a bad one.
Step 3. Find the boundary: Ask what may go out without review, what needs a named approver, and what must never leave the company. Mark anything I have not decided as UNDECIDED rather than guessing.
Step 4. Build the ninety minutes: Design one session that walks this team through doing that task with AI end to end, using our own material. Tell me exactly what I need to prepare beforehand.
Step 5. Design the second touch: A short follow-up three days later, built around what people will have struggled with. Tell me what to ask them.
Output: A one-page track plan (the task, the standard, the boundary, the session outline, the follow-up), a list of what I must prepare, my UNDECIDED list, and the single behavioral measure I should check at week four.
The UNDECIDED list is usually the useful part. Those are the standards your company has never actually settled, which means training cannot teach them yet and no vendor could have supplied them. See where you stand →
Sources
- Trust, attitudes and use of artificial intelligence: A global study 2025 · KPMG with the University of Melbourne, April 2025 (more than 48,000 people across 47 countries, fielded November 2024 to January 2025, led by Professor Nicole Gillespie and Dr Steve Lockey; only 47% of employees say they have received AI training, and only 40% say their workplace has a policy or guidance on generative AI use)
- You Rolled Out AI. Your Team Still Isn’t Using It. Here’s the Real Reason. · bosio.digital, 2026 (the adoption gap in full: deployment against adoption, the forgetting curve, and why the standard program does not change behavior)
- Making AI Work for Your Teams · bosio.digital, 2026 (the practical adoption guide this program sits inside)
- Who Owns Your AI Adoption? The Org Chart Just Answered Wrong. · bosio.digital, July 12, 2026 (why a training program cannot create ownership that does not exist)
- How to Give Your AI the Five Things It Can’t Work Out on Its Own · bosio.digital, August 10, 2026 (the written standards that become most of your training material)
- Four Things Get Called an AI Operating System. Only One Runs a Company. · bosio.digital, August 18, 2026 (the five-stage build sequence this program is part of)
- Beyond the Big 4: A Mid-Market Leader’s Guide · bosio.digital, 2026 (choosing a partner by fit, and the ownership-transfer test)
Frequently Asked Questions
What is AI training for employees?
AI training for employees is structured teaching that helps staff use AI tools in their actual work. In practice it has two halves that are usually confused: the curriculum, meaning tool mechanics, prompting technique and risk awareness, which is widely available and largely commoditized; and the program, meaning who is trained in what order, which workflow each track attaches to, who owns it, and how success is measured. The curriculum can be bought. The program has to be designed around how your company works.
How long should an AI training program take?
A single role-specific track takes about four weeks from selection to first measurement: one week to pick the workflow and write the standard, one to build the session, one to run it plus a short follow-up, and one to observe whether behavior changed. Company-wide programs take longer, but designing five tracks in parallel is how they get shallow. Build one properly, then let other functions copy a structure that already works.
What should AI training actually cover?
Start with one task the team already repeats weekly, and teach that task end to end using your own material rather than sample data. Cover the standard a finished version must meet, and the boundary, meaning what may go out without review and what may never leave the company. Tool basics and general prompting technique are worth buying separately if the group needs them, but they are the foundation rather than the program.
How do you measure whether AI training worked?
Not by completion rates, which measure attendance rather than behavior. Check three things instead: whether the specific task the track was attached to is now being done the new way without prompting, four to six weeks later; adoption depth, meaning how many people use AI for real work weekly rather than how many hold licenses; and whether output quality improved rather than only speed. Agree all three before the program starts, because measures picked afterwards get picked to flatter it.
Should AI training be different for each department?
Yes, and this is the difference between a program that changes behavior and one that is politely attended. A finance manager and a sales lead repeat different work and need different things, and both know it within ten minutes of a generic session. Role-specific means each track starts from that function’s actual recurring work, not one shared deck with a different opening slide.
Do we need to hire an outside firm for AI training?
Not for the foundational content, which is commoditized and well made at every price point. Outside help earns its place in three narrower places: designing the track structure when nobody internally has done it, facilitating the sessions where your standards get argued out, which is easier for someone with no stake in whose version wins, and sustaining the cadence through the first two quarters. Whatever you buy, judge it on what you own at the end.
Who should own AI training internally?
One named person with protected time, not a committee and not a department in the abstract. Between twenty-five and a hundred people this is the single most load-bearing decision in the whole program, and the common error is assigning it to whoever is most enthusiastic rather than whoever has the authority to protect the time. Enthusiasm launches a program; authority is what keeps it alive at week six.
Why doesn't AI training usually change behavior?
Because the standard shape is a single generic session, delivered once, to everyone, disconnected from anyone’s real job, and measured by attendance. Each of those choices is individually reasonable and together they produce understanding without changed behavior. The fix is structural rather than editorial: role-specific tracks attached to real workflows, a named owner, reinforcement past the first week, and success defined as the task being done the new way unprompted.
Turn scattered AI into a system your company runs on.
CompanyOS is the AI operating system your whole company runs on — governed accounts, real adoption, and visibility you own.
See CompanyOS → Not sure where to start? Take the 90-second readiness read →


