Type a few details into ChatGPT and it drafts you a six-month marathon plan in under a minute. That’s the whole appeal, and it’s also the problem. A January 2026 study tested eight AI chatbots against real marathon training science and found the plans look right on the screen and miss half of what actually keeps you healthy and on pace.
Researchers ran Claude, ChatGPT, Gemini, and DeepSeek through the same test: write a 6-month marathon plan for a beginner, an intermediate, and an advanced runner. Every model nailed one thing. More than 80% of the prescribed volume landed at easy, low-intensity effort, matching the 80/20 rule coaches have used for decades. Then the wheels came off.
Several models skipped the actual weekly mileage numbers a runner needs to execute the plan. Some blurred intermediate and advanced into a single tier. Pacing guidance got shakier, not steadier, for the runners who needed it most.
A plan that looks right on the screen can still fail on the road.
What the 2026 Study on AI-Generated Marathon Plans Actually Found
The research ran in the British Medical Bulletin in early 2026. A team led by Montaruli and colleagues tested eight AI systems: two Claude models, three ChatGPT models, two Gemini models, and DeepSeek R1. Each one wrote a 6-month plan for three runner levels, producing 24 separate plans in total. Researchers then checked every plan against the peer-reviewed marathon training literature.
Here’s the scorecard.
| Plan component | What the study found |
|---|---|
| Intensity distribution | Got it right — over 80% low intensity across all 8 models, matching the 80/20 principle |
| Weekly mileage numbers | Missing or vague in several models |
| Beginner vs. intermediate vs. advanced | Blurred into one tier in a subset of the 8 systems |
| Pacing guidance | Inconsistent, and worse for advanced runners specifically |
| Taper phase | Present in shape, but not timed to actual accumulated load |
A separate 2024 study scored ChatGPT’s exercise advice against official ACSM guidelines for comprehensiveness. The gap was wide enough to get cited across secondary coverage at roughly 41% of expected guideline content present. Translation: the chatbot said true things. It didn’t say most of the things.
Why “Good Enough” Isn’t Good Enough: The Injury-Risk Math
Here’s why the missing mileage numbers matter so much. Sports scientists track a number called acute:chronic workload ratio, or ACWR. It compares your training load this week against your average load over the past month, calculated from your logged training stress score (more on how that’s built in our TSS vs. TRIMP breakdown). In plain terms, it flags when you’re ramping up faster than your body can absorb.
A systematic review of 27 studies found the safest zone sits between 0.80 and 1.30. Push your ratio to 1.50 or higher and injury odds jump 2.3 to 3 times. Push past 2.0 and the odds ratio climbs to 4.00.
A chatbot that doesn’t know your last four weeks of training can’t calculate this number. It can’t even see it.
No training history, no ACWR, no real safety check.
Where the “Just Ask ChatGPT” Myth Came From
The myth isn’t that people love ChatGPT. It’s the assumption that everyone else is already using it to build their marathon plan. A 2024 pilot study in Innsbruck checked that assumption directly.
Researchers surveyed 119 recreational athletes. About 54.6% used a structured training plan of some kind. Of those, only 25% (16 people) used an AI-generated one. Most runners with a real plan still got it from a coach, a template, or a running club.
There’s a second wrinkle. Six experienced coaches blind-reviewed a batch of plans and correctly spotted the AI-written one 4 times out of 6.
The plan reads well. The details give it away.
The Correct Model: A Plan Is a Guess Until It Meets a Feedback Loop
A single chatbot prompt works like cruise control. It locks in a speed and holds it, no matter what’s happening on the road ahead. A real training plan needs something closer to adaptive cruise control, a system that reads the traffic (your fatigue, your missed runs, your last hard week) and adjusts in real time.
This is also where the classic “10% rule” falls apart. Most static plans, chatbot-written or not, quietly assume you shouldn’t raise weekly mileage more than 10% at a time. It sounds cautious. It isn’t well supported.
A Danish study tracked 60 novice runners by GPS for 10 weeks. Thirteen got injured. The 47 who stayed healthy actually increased mileage by 22.1% per week on average, more than double the “10% rule.” Injured runners increased by over 30%. The line between safe and risky isn’t a flat percentage. It depends on the runner, the week, and what came before it.
A fixed rule can’t replace watching what your body is actually absorbing.
The Reproducibility Problem: Same Prompt, Different Plan
Ask a chatbot the same question twice and you’d expect the same answer, especially with settings turned to their most consistent mode. A 2026 study tested exactly that.
Researchers ran three models, GPT-4.1, Gemini 2.5 Flash, and Claude Sonnet 4.6, through 6 exercise-prescription scenarios, 20 times each. That’s 360 total outputs. Then they measured how closely each model’s 20 answers matched one another.
| Model | Semantic similarity (20 runs) | What it means |
|---|---|---|
| GPT-4.1 | 0.955 | Nearly always a fresh, unique answer (100% unique outputs) |
| Gemini 2.5 Flash | 0.950 | High similarity, but heavy repetition — only 27.5% unique outputs |
| Claude Sonnet 4.6 | 0.903 | Most variation run to run of the three |
Ask Gemini the same question twice and you have roughly a 1-in-4 chance of a genuinely new answer. The researchers’ own conclusion is blunt: which model you pick “directly affects exercise prescription consistency” and should be treated as a real decision, not a technical footnote.
The model you happen to open changes the plan you get.
One Runner’s Experiment: Three Prompts, Three Different Plans
Take a runner I’ll call Marcus. He’s 34, chasing a sub-3:30 marathon after finishing his first one in 3:45. He asks ChatGPT for a 16-week plan, likes it, then asks Claude the same question to double-check. The two plans don’t match. ChatGPT’s long run peaks at 20 miles in week 12. Claude’s peaks at 22 miles in week 10, two weeks earlier.
Marcus splits the difference and stitches together his own hybrid. Neither plan adjusted for the 9 days he’d missed in February with the flu. Nobody warned him. By week 8, chasing the higher of the two numbers, his mileage jumps 34% across two weeks. His knee starts complaining. He loses 12 training days to rest.
Two experts giving two answers isn’t expert advice. It’s a guess, twice.
How AthleteOS Builds Your Plan From Your Training Data, Not a Prompt
AthleteOS doesn’t start from a single text prompt. Every plan runs off three numbers that update after each workout you log: your fitness score, your fatigue score, and your form score, the same fitness-fatigue-form framework covered in our guide to CTL, ATL, and TSB.
Where the British Medical Bulletin study found models blur intermediate and advanced into one generic plan, AthleteOS reads your actual training history, not a level you picked from a dropdown. Where the reproducibility study found the same model gives two different answers to the same question, your plan doesn’t regenerate from scratch every time you open the app. It adjusts from the sessions sitting in your workout calendar.
Miss a week with the flu, like Marcus did? Your fatigue score and fitness score both move, and the AI coach adjusts your ramp instead of pretending nothing happened. AthleteOS also tracks your drift ratio on every long run, so a fading aerobic base gets caught before race day, not after.
Sign up for AthleteOS and build a marathon plan from your actual training data, not a one-time guess.