The Content Firehose: Why AI-Made Courses Aren’t Improving Learning Outcomes

Gen AI
The Content Firehose: Why AI-Made Courses Aren’t Improving Learning Outcomes
Ankush Mathur

Written by

Ankush Mathur

Updated on

August 24, 2026

Read time

13 mins read

Quick Answer: AI made course production nearly free, and production was never the bottleneck. Learning outcomes depend on cognitive load management, instructional design, effortful retrieval, assessment aligned to job behavior, and timely feedback. Generated content improves none of these by default, and often makes several worse.

71% of L&D professionals are now exploring, experimenting with, or integrating AI into their work, yet the teams shipping ten times the content are rarely seeing any movement in the metrics that matter: behavior change, retention, and business results. The fix is not better prompts. It is moving AI from the author’s chair to the assistant’s chair, inside an LMS designed around outcomes instead of output.

Somewhere in your organization right now, a course is being generated. Not written. Generated. The topic went in, twelve modules came out, and the library counter ticked up again. Your team has never produced this much training this fast in its history.

So here is the uncomfortable question this article exists to answer: if content production is up 10x, why are the numbers that actually matter (behavior on the job, skill retention, performance outcomes) exactly where they were before the AI budget line appeared?

This is not an anti-AI argument. LinkedIn’s Workplace Learning Report, surveying 937 L&D and HR professionals, found 71% already exploring, experimenting with, or integrating AI into their work, and they are right to. The problem is a specific misuse pattern, and it rhymes with what happened everywhere else AI landed: MIT’s NANDA initiative found that 95% of enterprise GenAI pilots produced zero measurable P&L impact, not because the technology failed, but because it was pointed at the wrong part of the problem. In learning, the wrong part is content volume. Content was never the constraint.

Content Is Not Learning

The core confusion is a category error: treating content as if it were learning. Content is an input. Learning is a change in what a person can do, measured after the course is over, ideally weeks after.

The largest dataset on this predates AI entirely. Harvard and MIT’s four-year study of 290 online courses and 4.5 million participants found that in a typical course, about 7,900 people access content and roughly 500 earn a certificate. Access to high-quality content, produced by two of the best universities on earth, moved almost nobody to completion, let alone mastery. If elite human-made content could not carry learning on its own, generated content certainly will not.

AI changed the economics of production, and in doing so it broke the last natural constraint on volume. When a course cost $10,000 and six weeks to build, someone asked whether it was needed. When it costs an afternoon, nobody asks, and the library fills with material nobody requested, sequenced by nobody, owned by nobody. The dashboard says the program is thriving. We wrote about why that dashboard lies in the vanity metric trap; AI content is the accelerant version of the same fire.

The Cognitive Overload Problem

There is a harder version of this claim, and it comes from cognitive science: past a point, more content does not just fail to help. It actively hurts.

Cognitive load theory, developed by John Sweller, splits the demand on a learner’s working memory into three parts: intrinsic load (the difficulty of the concept itself), germane load (the useful effort of building understanding), and extraneous load (everything else: padding, repetition, decorative detail). Working memory is fixed and small. Every word of filler competes with the concept for the same limited capacity.

One working memory, three kinds of load: intrinsic load from the concept itself, germane load from useful effort, and extraneous load from padding and recaps, which AI-generated content inflates past working memory capacity

Now consider what large language models produce by default: fluent, expansive, generously padded prose, with an introduction, a recap, and three examples where one would do. Generated content inflates precisely the load that teaches nothing. Cutting words is an instructional act, and AI does the opposite unless a skilled human forces it to. An AI-padded forty-minute course can genuinely be harder to learn from than a disciplined fifteen-minute one covering the same concept.

Missing Instructional Design

Ask an instructional designer what they actually do and very little of the answer is “write content.” The craft is in the decisions around the content: what to include and, more importantly, what to cut, in what order concepts build, where practice goes, what a learner must demonstrate before moving on, and what evidence would prove the whole thing worked.

That discipline has a name, backward design, popularized by Wiggins and McTighe: define the outcome first, define the evidence second, and only then build the material. AI generation runs this exactly backward. It starts from a topic and produces material, with the outcome undefined and the evidence never specified. The result is not a curriculum. It is a pile of plausible modules wearing a curriculum’s clothes.

This is why “we generated a course on X” and “we built training for X” are different claims. The first is a statement about content existing. The second, if it is honest, is a statement about a designed path to a defined outcome. AI can accelerate the second. It can only ever fake the first into looking like it.

Engagement vs Information: The Effort Is the Point

The most counterintuitive finding in learning science is that smooth is bad. In a foundational experiment published in Psychological Science, Roediger and Karpicke had students either repeatedly restudy material or take recall tests on it. A week later, the tested group retained substantially more. The restudy group had something else instead: higher confidence. Passive review felt like learning while producing less of it.

Generated content is the smoothest content ever made. It reads effortlessly, summarizes perfectly, and asks nothing of the reader. That polish is exactly the property that suppresses the effortful retrieval learning depends on. A learner who glides through six AI modules feels informed, scores the course highly, and retains a fraction of what a rougher, retrieval-heavy experience would have left behind. The feeling of learning and the fact of it have quietly diverged, and every standard LMS metric measures the feeling.

The Assessment Gap

AI will happily generate a quiz for every module, instantly, and this is where the misuse gets expensive: the quiz it generates tests recall of its own wording. Five questions asking learners to recognize sentences they read four minutes ago is not assessment. It is proofreading with a scoreboard.

Real assessment starts from the job, not the text: what should this person be able to do differently, and what evidence would show it? That means scenario judgments, worked problems, and application tasks, written from the competency, then checked against the content. The distance between those two kinds of question is the distance between a certificate and a capability. Here is the pattern across the whole stack:

What AI produces by default What learning actually requires The fix
More content, faster Less content, better sequenced Curate ruthlessly; cap course length
Fluent explanations Effortful retrieval Replace recap pages with recall checks
Instant quizzes on its own text Assessment aligned to job behavior Write items from tasks, not paragraphs
Personalized tone Personalized path Branch on demonstrated mastery
Answers at any hour Feedback at the moment of error Build timed feedback loops into the flow

Delayed Feedback Kills the Loop

Even a good assessment fails if its result arrives late. Feedback works when it lands close to the attempt, while the learner’s reasoning is still available to correct. A score emailed three days later corrects nothing; the mistake has already consolidated.

The deeper timing problem is the forgetting curve, first mapped by Ebbinghaus and replicated in a controlled 2015 study by Murre and Dros: retention collapses fastest in the first days after learning, then levels off. A course with no retrieval events after completion is a course whose content is mostly gone within a week, however good it was. The counter is spaced retrieval: short recall checks at growing intervals, each one interrupting the decay and resetting it on a shallower slope.

The forgetting curve after a course ends: recall collapses within days on content alone, while spaced recall checks at day 2, day 7 and day 30 keep retention substantially higher

Notice what this implies about where the engineering effort belongs. The highest-leverage feature in a learning platform is not a bigger library. It is a scheduling engine that brings the right three questions back to the right learner on the right day. That is a build decision, and it is exactly the kind of decision that gets skipped when the roadmap is busy generating more modules.

The Personalization Myth

“AI-personalized learning” is doing a lot of unearned work in vendor decks right now. In most implementations it means the same content, reworded: friendlier tone for one learner, more formal for another, the learner’s name in the introduction. The path (what comes next, what can be skipped, what must be repeated) is identical for everyone.

Real personalization branches on demonstrated mastery. A learner who proves competence on a diagnostic skips ahead; a learner who fails a retrieval check gets a different explanation and another attempt, not the same paragraph again in a warmer voice. That requires a competency model, tagged content, and assessment infrastructure, none of which a text generator provides. We called the surface version “personalization theater” in the vanity metric trap, and AI has made the theater dramatically cheaper to stage without making the real thing any more common.

Shipping more courses and moving fewer numbers?

Techuz builds learning platforms around the parts AI can’t generate: competency models, spaced retrieval engines, behavior-aligned assessment, and analytics that measure mastery instead of clicks.

Talk to our LMS team

Where AI Actually Helps: The Assistant’s Chair, Not the Author’s

None of this argues for banning the tools. It argues for seating them correctly. The dividing line is simple: does a human own the instructional decisions, or just the publish button?

AI as the author versus AI as the assistant: risky uses include fully generated courses, auto-quizzes testing their own wording and reworded personalization, while effective uses include expert-revised first drafts, scenario variations at scale, human-approved practice items and tutoring inside guardrails

In the assistant’s chair, AI is genuinely transformative for L&D. It drafts, and an expert cuts. It produces twenty scenario variations from one validated template, which is exactly the kind of task where generation shines because a human defined the pattern. It generates candidate assessment items that an instructional designer approves, rejects, or sharpens. It answers learner questions on demand, inside guardrails, grounded in approved material rather than its own imagination. Each of these is the pattern that works everywhere AI works: it removes the drafting bottleneck while a human keeps the judgment, the same division of labor we mapped in where generative AI actually improves efficiency.

What all the effective uses share is that none of them increase published volume by default. They increase the quality and speed of a process a human still owns.

The Outcome-Driven LMS Strategy

For a corporate L&D head, the practical question is what to build and buy around, and the answer is an LMS architected backward from outcomes rather than forward from content. Concretely, that means five things.

First, a competency model as the spine: every course, assessment item, and retrieval check tagged to a defined capability, because personalization and skills reporting are impossible without it. Second, assessment designed from job tasks, with generated items allowed in only through human review. Third, a spaced retrieval engine that schedules recall checks after course completion, because that is where retention is actually won. Fourth, feedback loops timed to the moment of error, not the end of the module. Fifth, analytics built on Kirkpatrick’s four levels, reporting behavior change and business results, not reactions and completions; if the dashboard cannot show level three, the program cannot prove it works.

This is the difference between buying a content library with a quiz feature and commissioning custom LMS development around how learning actually happens. It is also where AI fits best of all: a generative AI development company can wire generation into that architecture as a governed assistant (drafting, varying, tutoring within guardrails) instead of an ungoverned author. The technology is the same. The seat it occupies changes everything.

The L&D Head’s Reality Checklist

Before the next quarter of AI-accelerated production, five questions worth answering honestly:

  • Can we name the specific on-the-job behavior each of our last ten courses was built to change?
  • Does any assessment item test application, or only recognition of the course’s own wording?
  • Does anything bring content back to a learner after the completion date, even once?
  • Does our “personalization” change the path, or only the phrasing?
  • If the board asked for evidence at Kirkpatrick level three, what would we show them?

If most of these are uncomfortable, the constraint on your learning outcomes is not content volume, and the next AI-generated course will not move it.

Build learning that survives day 30

As an LMS development company, Techuz builds outcome-driven platforms with competency models, retrieval engines, and behavior-level analytics, with AI seated where it helps and governed where it doesn’t.

Start a conversation

FAQs

Why isn’t our AI-generated training improving performance metrics?

Because content volume was probably never your constraint. Outcomes depend on instructional design, effortful retrieval, behavior-aligned assessment, and timely feedback, and generated content improves none of these by default. Diagnose which of those layers is missing before generating anything else.

Is AI-generated course content bad for learning?

Not inherently. Unreviewed generated content tends to be padded, which raises extraneous cognitive load, and its smoothness suppresses the effortful retrieval that drives retention. The same technology used as a governed assistant, drafting material an expert revises and generating practice items a human approves, genuinely accelerates good design.

What should we measure instead of course completions?

Completion to mastery ratio, retention at 30 and 90 days via spaced recall checks, and behavior change on the job, which is Kirkpatrick level three. If your platform can only report reactions and completions, it is measuring activity, not learning.

What does real AI personalization in an LMS look like?

Branching on demonstrated mastery: learners who prove competence skip ahead, and learners who fail a recall check get a different explanation and another attempt. That requires a competency model and tagged content. Rewording the same material in a friendlier tone is personalization theater.

Where does AI genuinely help an L&D team right now?

Drafting first versions an expert cuts down, generating scenario variations from validated templates, producing candidate assessment items for human review, and answering learner questions inside grounded guardrails. The common thread is that a human keeps the instructional decisions. If you want that wired into your platform properly, an experienced LMS development company can build the governance in from the start.

Sources

Looking for timeline and cost estimates for your app?

Contact us Edit Logo Edit Logo
Ankush Mathur

Ankush Mathur

Ankush Mathur leads technology at Techuz as CTO & Technical Project Manager, where he helps startups and enterprises architect and scale their software. He's spent his career moving from hands-on development to technology leadership, and enjoys writing about engineering practices, AI, and the decisions behind building solid products.