Training

Should you trust your watch's FTP and threshold estimates?

Your device's threshold numbers are a real model fitted to your real data — and an extrapolation from efforts you may never have done. Here is how to tell which one you are looking at.

10 min read

Yes as a starting point, no as a fact. The FTP or threshold number your watch shows you is a genuine model fitted to your own data, and for an athlete who has never run a test it beats every alternative — including the number you would have guessed. What it is not is a measurement. It is an inference drawn from the efforts you happened to record, and its accuracy depends almost entirely on whether those efforts came anywhere near your limit. That single point explains nearly everything about when the estimate is good, when it drifts, and what to do about it.

How a watch arrives at a threshold number

The specifics are proprietary and they differ between platforms, but the ingredients are not mysterious. Garmin, Polar, COROS, Wahoo, Apple, and the training apps that sit on top of them are all solving the same problem with some combination of the same three approaches.

It watches the relationship between effort and output. Across many sessions your device builds a picture of what heart rate you produce at a given pace or power. That relationship is close to linear through the easy and moderate range, and it bends at the point where you start accumulating lactate faster than you can clear it. Fit the line, find the bend, and you have a threshold estimate.

It tracks your best efforts at each duration. Every time you push hard for five minutes, twelve minutes, twenty minutes, the device records the best power or pace you held for that long. Those points describe a curve, and on a well-populated curve threshold sits in a fairly predictable place. This is the most direct method available — and it works well exactly to the extent that you have supplied hard efforts.

It reads finer heart-rate signals. Some platforms analyze the beat-to-beat pattern rather than just average heart rate. As you cross from moderate into hard, the variation between beats changes in a way that is detectable before you consciously feel the shift. Estimating threshold from that transition is genuinely clever, and it is why a watch can sometimes revise a threshold from a session you never intended as a test.

With very little data to work from, most platforms fall back on population averages — age, weight, sex, self-reported activity level — until you give them something better. That is the right thing to do, and it is also the reason a brand-new user's first estimate should be treated as a placeholder.

Why this is a good starting point, not a bad one

It is worth saying plainly, because "your watch is wrong" is the lazier and more popular take. For an athlete in their first season, an on-device estimate is a real improvement over the alternatives, which are guessing, copying a training partner, or training with no zones at all.

The estimate is anchored to your data rather than someone else's. It updates for free, without you having to schedule twenty minutes of misery. And for the purpose most people need it for — separating easy from moderate from hard — being within a few percent is plenty, because the gap between those efforts is much wider than the error.

The trouble starts when the number gets used for things that need more precision: prescribing interval paces, pricing training load, and tracking fitness over months.

The limits are in the data, not the vendor

Every platform inherits the same constraints, because they are all working from the same raw material — the efforts you actually recorded.

A model can only see what you gave it. Estimating your threshold from training in which you never approached your limit is like estimating your one-rep max from sets of ten. It can be done, the answer is not silly, and it carries all the uncertainty of a long extrapolation. Athletes who train mostly easy — which is most athletes, correctly — hand their device very little data near the interesting part of the curve.

Your training mix skews each sport differently. If your running is almost all easy running, the heart-rate-to-pace line has to be extended a long way past any point you actually visited, and a small error in the slope becomes a large error at the far end. Meanwhile your bike number, if you do structured intervals with a power meter, may be sitting on solid ground. Same watch, same week, two very different levels of confidence in two numbers presented identically.

Pool swimming is usually invisible to threshold logic. Threshold estimation grew up around running and cycling, where GPS and power give a clean measure of output. In a pool there is no GPS, stroke and turn detection are approximate, and there is nothing to price effort against. Most devices simply do not produce a swim threshold, and where a number does exist it usually came from something you typed in rather than something the watch measured. Swim threshold has to come from a test you do on purpose — critical swim speed is that test.

An automatic bump follows your best day, not your normal day. Think about what the session behind a "new FTP" notification usually looks like: cool morning, rested legs, decent fuel, a tailwind on the way out, maybe a group to sit in. The device sees the output, not the circumstances. Any estimator that responds to your best recent effort is, by construction, fitting the top of your distribution rather than the middle of it. That is the mechanism behind a very common experience — a celebratory notification after one great session, followed by six weeks of intervals that feel just out of reach.

Heart-rate-based estimates inherit heart rate's bad days. Heat, dehydration, caffeine, illness, poor sleep, cardiac drift late in a long session, and wrist optical sensors losing the signal during hard intervals or in cold weather all distort the effort-to-output relationship the model is trying to fit. A chest strap removes one of those problems and none of the others.

None of this is a knock on any particular company. It is what happens when you ask an honest model to infer a ceiling from data collected below it.

What a 5–10% threshold error actually costs

A few percent sounds harmless. It is not, because the error propagates into everything downstream and the load model squares part of it.

It shifts every zone. Suppose your true threshold power is 250 W and the estimate says 275 W. Your sweet-spot session at 90% now asks for 247 W — that is 99% of your actual threshold. You have been handed a threshold session with the word "sweet spot" on the label, and told to do it three times a week. The run version is just as sharp: a true threshold pace of 4:12/km recorded as 4:00/km is twelve seconds per kilometer, which over 5 × 1 km is the difference between a repeatable session and one that falls apart in the fourth rep.

It misprices every session's load. Training Stress Score scales with the square of your intensity relative to threshold, so a threshold error gets roughly doubled on its way into the load number.

It quietly corrupts fitness and fatigue tracking. CTL and ATL are running averages of those daily scores, so a mispriced session is a mispriced day is a mispriced season.

If your threshold isYour zones areYour load scores areWhat you experience
5–10% too highToo hardUnder-scored by roughly 10–17%Sessions consistently feel harder than they are labeled, while your fitness line under-reports what you are absorbing
5–10% too lowToo easyOver-scored by roughly 10–20%Everything feels comfortable, the fitness line climbs nicely, race day disappoints

The second row is more common than athletes expect, and it is self-inflicted rather than the device's fault: a threshold set at the start of a base period stops being true the moment the base period works. If any of that arithmetic is unfamiliar, training load explained covers how TSS, CTL, and ATL fit together and where the model stops telling the truth.

Treat the estimate as a hypothesis, then confirm it

The fix is not to distrust your watch. It is to stop treating an inference as a result. Every so often, go and produce the maximal data the model has been missing.

SportTestWhat it gives youHow often
Bike20-minute maximal time trial; FTP is about 95% of your average power (a ramp test or 2 × 8 min are reasonable alternatives)FTP, and threshold heart rate from the same effortEvery 6–10 weeks in a build block
RunA hard 5 km or 10 km race, or a solo 30-minute time trial — take the average pace and heart rate over the final 20 minutesThreshold pace and threshold heart rateTwo or three times a season
SwimCritical swim speed: time trial 400 m and 200 m with full recovery between; CSS in m/s is 200 divided by the difference in secondsCSS, and pace per 100 m for every swim zoneEvery 8–12 weeks

Test conditions matter more than the protocol you pick. Same course or trainer, similar weather, similar time of day, similar fueling, and ideally the same equipment. A test you cannot repeat under comparable conditions gives you a number you cannot compare to the last one.

Race results are the best threshold data you will ever get

A race is a maximal effort you were going to do anyway, with a clock, a course, and other people to stop you pacing it politely. That makes race data better ground truth than any test you can talk yourself out of halfway through.

An hour-long bike leg with a power meter is very nearly the definition of functional threshold power. A 10 km road race gives you a run threshold pace with real confidence — for most trained runners, threshold pace sits somewhere between 10 km and half-marathon race pace, closer to the half the more running you do. Even a hard local parkrun, run honestly, is more informative about your ceiling than a month of comfortable training.

This is also why the sensible way for software to handle thresholds is to seed them from your history and then ask. PaceBeats sets your starting thresholds from your imported training data and proposes an update when it sees a race-quality effort that contradicts the number on file — as a suggestion you confirm, rather than silently rewriting your zones and every load number behind them.

How to actually use the number

A short set of working rules, in rough order of how much time they save.

Do not act on a single jump. If your threshold estimate moves several percent after one session, wait. If it was real, it will happen again within a few weeks. If it was a tailwind, it will not.

Sanity-check against the hour. Threshold is roughly what you could hold for about an hour when it hurts the whole way. If you could not hold your stated threshold pace or power for an hour on a good day, the number is too high, whatever the watch says.

Sanity-check downward too. If your easy sessions feel like work and you cannot speak in full sentences during them, your threshold is probably set too high and dragging every zone boundary up with it. Heart rate, pace, and power zones covers how the boundaries relate.

Test each sport separately. There is no conversion. Being a strong cyclist tells you nothing useful about your run threshold, and neither tells you anything about your CSS.

Update deliberately, not continuously. Changing your threshold every time the estimate wobbles makes your own training history unreadable, because the same session gets a different price each month. Change it when you have evidence, write down the date and what the evidence was, and leave it alone until the next one.

The estimate is not the problem. Treating an extrapolation as a measurement is the problem — and the cure costs you one uncomfortable session a couple of times a season.

Questions athletes ask

Why did my watch suddenly raise my FTP?

Almost always because you produced an unusually good effort and the estimator responded to it. Threshold models are fitted to your best recent data, so a cool morning, fresh legs, a tailwind, or a fast group ride can all move the number without your fitness having changed. It may well be real, but a single session is not enough to know. Wait a few weeks and see whether you can produce something similar again before you rebuild your training zones around it.

Is a lab test worth it?

For most age-group athletes, no. A lab test gives you a more precise threshold and extra data such as VO2 max and, in some protocols, substrate use — but the precision gain over a well-executed field test is smaller than the day-to-day variation in how you feel, and the number goes stale in a couple of months just the same. A lab test earns its cost if you have a specific physiological question, if you are chasing performance at a level where small margins decide outcomes, or if you simply cannot pace a solo field test honestly. Otherwise the money is better spent on a repeatable field test protocol you will actually run.

How often should I test my thresholds?

Every 6 to 10 weeks during a build block is a reasonable default for the bike, two or three times a season for the run, and every 8 to 12 weeks for swim critical swim speed. Test more often when fitness is changing quickly — a beginner's first season, or a return from a long break — and less often when you are holding steady. Race results count as tests and are better data than any protocol, so a season with regular racing needs fewer standalone test sessions.

My watch's threshold pace feels too fast — what do I do?

Trust your body first and confirm with a test. Threshold pace should be sustainable for roughly an hour on a good day, so if your threshold intervals fall apart in the third or fourth rep, the number is likely too high. Run a 30-minute solo time trial or a hard 5 km to 10 km race and set threshold from that result rather than from the estimate. In the meantime, run your threshold sessions by feel or by heart rate — an effort you can hold for the whole set is more useful than an effort that matches a number.

Does the FTP estimate work without a power meter?

Partly. Without power, a bike threshold estimate has to be inferred from heart rate and speed, which means it inherits every condition that affects those — wind, terrain, drafting, heat, and cardiac drift. That makes it usable as a rough anchor and unreliable for prescribing interval targets. If you train mostly indoors on a smart trainer, the trainer's power data solves this. Outdoors without power, heart-rate-based zones are the more honest tool, and a periodic field test on a consistent course is how you keep them anchored.

Next step

Turn the theory into next week's training.

Start free