What Oura, WHOOP, Apple Watch and Recovery Scores Can—and Can’t—Tell You
Wearable technology has become so deeply embedded in fitness culture that simply waking up is no longer enough. Now we wake up and immediately ask a device how we slept. How much REM did I get? What was my HRV? How low was my resting heart rate? What is my recovery score? Am I “ready” to train today?
In some ways, this is remarkable. A small ring or watch can collect information throughout the night that would have been impractical for the average person to monitor not long ago. Wearable technology has become one of the dominant forces in the fitness industry, and I understand the appeal because I was an enthusiastic user myself.
For years, I wore an Oura Ring and used the information to explore my sleep, recovery and readiness. It helped me notice some genuinely interesting patterns. I could see what happened after a recovery week, what changed after a vacation in the sun, and even what appeared to happen to my REM sleep after evening yoga.
That information had value. The problem came when there was a danger of allowing the device to move from being one source of information to becoming the final authority. Those are two very different things.
THE BEST USE OF A WEARABLE: FIND PATTERNS
The strongest case for a fitness tracker is not that it can tell you exactly how healthy or recovered you are at 7:03 tomorrow morning. Its real value is that repeated measurements can help reveal patterns over time.
I saw this firsthand in Vacation Readiness Reset, after returning from a week-long vacation in the Dominican Republic. My Oura readiness scores jumped to 100 and 98 after being considerably more sluggish before the trip. I had spent the week outdoors, relaxing, getting sunlight and temporarily stepping away from my usual training and work demands. Whatever combination of factors was responsible, the device gave me another piece of information suggesting that something about the change in environment and routine had improved recovery.
I noticed something similar during what I call The Deset Week. By periodically reducing training volume while maintaining meaningful intensity, workouts become shorter and overall recovery demands decrease. During these lower-volume periods, I repeatedly noticed some of my best sleep and readiness scores.
I documented another interesting observation in Evening Yoga Improves REM Sleep. On two occasions, an evening yoga session was followed by unusually long periods of REM sleep according to my Oura Ring, along with feeling particularly refreshed the following morning. I was careful even then to treat it as a limited n=1 experiment, not proof that evening yoga had suddenly become a clinical treatment for REM sleep.
That is how I think wearable data is most useful: observe what happens, record it, look for repeating patterns, change one thing when possible, and see whether the result appears again. The danger begins when “That’s interesting” quietly turns into “My device says this, therefore it must be true.”
THEN MY READINESS SCORE GOT IT COMPLETELY WRONG
I wrote about one particularly revealing example in Better Than an Oura Ring, and it ultimately captured the limitation better than anything else.
My Oura readiness score was 94. By the number, I should have been primed for a terrific workout. I started with bent-knee deadlifts and planned to work from 305 to 335 pounds for five sets of five, but the first work set felt like I was trying to lift a tank. After two repetitions, I stopped.
Years of training experience were telling me something the algorithm was not: I was not ready to train heavy that day. Instead of forcing the session because a ring had given me permission, I packed it up, went for a walk and returned the following morning.
This time the weights moved easily and the workout went exactly as planned. The funny part was that my readiness score that morning was 93—essentially identical to the 94 I had received on the terrible day.
One point separated the two scores. Actual performance separated the two days.
That experience permanently changed the way I looked at readiness technology.
A READINESS SCORE IS NOT READINESS
Oura, WHOOP, Garmin, Apple and similar devices can collect or estimate useful variables such as heart rate, resting heart rate, activity, movement, temperature trends and heart-rate variability. Their software then takes selected pieces of that information and converts them into something far more attractive to the consumer: a single score.
Recovery. Readiness. Sleep. Stress. Whatever label is placed on it, the important point is that the number is an algorithmic interpretation of measurements and estimates. It is not a direct measurement of whether 315 pounds will feel heavy in your hands that afternoon.
There is a major difference between saying, “Your resting heart rate was seven beats higher than your recent baseline” and saying, “You are only 61% recovered today.” The first describes a physiological variable. The second interprets what several variables supposedly mean for you.
That interpretation can be useful. It simply needs context.
YOUR WATCH DOESN’T ACTUALLY WATCH YOU SLEEP
Sleep tracking is perhaps the easiest place to understand the limitation.
A wearable is not looking inside your brain and directly identifying every transition between wakefulness, light sleep, deep sleep and REM. It combines movement, heart-rate signals and other sensor data and then applies an algorithm to estimate what was happening during the night.
The clinical standard for detailed sleep staging is polysomnography, which measures brain activity along with several other physiological signals. Consumer trackers have improved substantially, and some perform surprisingly well for certain measurements, but estimates of individual sleep stages are still not equivalent to a laboratory sleep study.
That does not make the information worthless. It simply changes how it should be interpreted. If your wearable repeatedly shows that your sleep duration deteriorates after late alcohol consumption, improves when you maintain a consistent bedtime, or changes dramatically during periods of excessive training stress, that pattern may be quite useful even if the device cannot perfectly determine whether you spent exactly 83 minutes in REM sleep.
This is the distinction that matters throughout the entire article:
Wearables can be excellent for trends without being infallible on individual readings.
HRV IS USEFUL—UNTIL WE TURN IT INTO A CRYSTAL BALL
Heart-rate variability, or HRV, deserves more respect than many of the simplified scores built from it. HRV reflects variation in the intervals between heartbeats and can provide useful information about autonomic nervous system activity, particularly when measurements are collected consistently under similar conditions.
The problem is rarely HRV itself. The problem is what we ask HRV to tell us.
A single low reading can be influenced by hard training, illness, alcohol, poor sleep, psychological stress, travel and numerous other factors. For that reason, your relationship to your own baseline and longer-term trend generally matters far more than whether today’s number looks impressive compared with someone else’s.
Even when HRV is used to guide training, it is not magic. It is another piece of information that may help refine a decision, not a crystal ball capable of telling you exactly how you will perform several hours later.
THE SCORE SHOULD START A QUESTION, NOT END THE CONVERSATION
Suppose you wake up and your recovery score is unusually poor. That is useful information, but the score should trigger investigation rather than immediately dictate behaviour.
How did you sleep? Did you train particularly hard yesterday? Are you coming down with something? Did you drink alcohol? Are you dealing with unusual stress? Did you travel? Is your resting heart rate elevated? Is HRV outside its normal range? How do you actually feel?
For someone preparing to train, there is another question that may matter even more: what happens when you start moving?
The opposite situation deserves the same scrutiny. A beautiful green readiness score does not give you permission to ignore unusual fatigue, pain, poor coordination or a warm-up weight that suddenly feels twice as heavy as normal.
Use the score to ask better questions. Don’t use it to stop asking questions.
YOUR BODY CAN DISAGREE WITH THE DEVICE
This was one of the biggest reasons I eventually retired my Oura Ring, a decision I discussed in Why I Retired My Oura Ring. I increasingly noticed a disconnect between how the device said I should feel and how I actually felt.
Some mornings followed mediocre scores even though I felt energetic and performed exceptionally well. Other mornings looked terrific on the dashboard while I felt flat. The natural temptation is to assume that the technology must know something you do not, because numbers look objective and feelings appear subjective.
But subjective information is not automatically inferior information. Your perception of fatigue, soreness, mood, motivation and physical readiness contains information too. Training research has repeatedly shown that subjective monitoring can be highly responsive to changes in workload and recovery.
That does not mean your feelings are always right either. It means the most useful approach is not device versus intuition. It is device plus intuition, history and performance.
WHEN TRACKING STARTS HURTING THE THING YOU’RE TRACKING
Sleep introduces an even stranger problem: you can become so concerned about achieving perfect sleep that the concern itself makes sleeping harder.
Researchers have used the term orthosomnia to describe the perfectionistic pursuit of ideal sleep driven by sleep-tracker data. The concept resonated with my own experience because there comes a point where waking up and immediately checking whether a machine approved of your night can become counterproductive.
You may wake up feeling perfectly fine. Then you look at the app and see a score of 62. Suddenly, you do not feel quite as fine.
That may sound psychological because it is. Feedback shapes expectations, and expectations influence how we interpret our own condition. Research using deliberately false sleep feedback has demonstrated that being told you slept poorly can influence how sleepy, fatigued and alert you report feeling later, even when the feedback itself is not real.
Think about the implication: a number can change how you interpret your own day—even when the number is wrong.
That is a very different issue from whether the ring accurately measured your resting heart rate.
ORTHOSOMNIA DOESN’T MEAN WEARABLES ARE EVIL
It would be equally foolish to overcorrect and declare that wearables are harmful.
Not everyone who owns an Oura Ring or Apple Watch becomes anxious about sleep. Some people look at the information, find a few useful patterns and move on with their day. Research has also suggested that wearable sleep feedback can sometimes be helpful, particularly when it is paired with appropriate guidance about what the information does and does not mean.
That last part is crucial: data plus context can help. Data without context can confuse.
The problem is not necessarily the tracker itself. The problem begins when the tracker becomes an oracle whose verdict is allowed to overrule every other piece of information.
DON’T LET A BAD SCORE GIVE YOU A BAD DAY
If you wake up feeling rested and energetic, an imperfect sleep score should not automatically convince you that something is wrong. Conversely, if the app awards you a spectacular score but you are exhausted, unusually sore, irritable and moving poorly, the number should not persuade you to ignore those signals.
The goal of health technology should be to increase self-awareness. If using it eventually causes you to distrust every internal signal unless an algorithm validates it, something has gone backwards.
A good wearable should help you understand yourself better. In an ideal world, the lessons you learn from it should eventually make you less dependent on the device, not more.
THE JOURNAL MAY TELL YOU WHAT THE DASHBOARD CANNOT
As I discussed in Journaling Like Leonardo da Vinci, I have kept journals for years. My approach came partly from reading How to Think Like Leonardo da Vinci. Rather than treating a journal exclusively as an emotional diary, I use it as an information repository. When I learn something useful, notice a pattern or develop an idea, I write it down and revisit it later.
The same philosophy works beautifully for health and training.
You could record how rested you actually feel, training performance, muscle soreness, stress, travel, alcohol intake, unusual aches or pains, recovery methods, changes in training volume and anything else that seems relevant. Then compare those observations with the wearable data.
Over time, the combination can reveal things a dashboard by itself may never explain. Perhaps alcohol consistently destroys your HRV. Maybe late training does not affect your sleep at all. Perhaps reducing training volume every third or fourth week reliably improves both subjective recovery and tracker data. Evening yoga may repeatedly coincide with better sleep, or a sunny vacation may produce a dramatic recovery reset.
The point is not to collect more information for the sake of collecting it. It is to understand your own patterns rather than outsourcing interpretation of your life to an algorithm.
YOUR TRAINING LOG MAY BE THE BEST READINESS TRACKER YOU OWN
Then we come back to performance.
If today’s planned squat weight normally moves quickly and suddenly feels brutally heavy during the warm-up, that matters. If your usual chin-up load feels unusually effortless, that matters too. Poor coordination, unfamiliar joint discomfort, changes in bar speed and the general feel of a familiar movement all provide immediate information about your current state.
A ring worn during sleep cannot completely capture that.
This is where years of keeping a training log become incredibly valuable. You gradually build a personal database of what different levels of readiness actually feel like and, more importantly, how they translate into performance.
My deadlift experience with a readiness score of 94 remains the perfect example. The device said go; the bar said no.
The bar was right.
THE THREE-LAYER READINESS SYSTEM
LAYER 1 — THE DEVICE: DATA
Use the wearable for things it can do reasonably well: sleep duration trends, resting heart rate, HRV, activity, temperature trends and other repeated measurements.
LAYER 2 — THE JOURNAL: CONTEXT
Add what the device cannot know completely: how you feel, psychological stress, muscle soreness, travel, nutrition, lifestyle changes, previous training and anything unusual about the day.
LAYER 3 — PERFORMANCE: REALITY
Finally, look at what actually happens when you perform: warm-up quality, strength, speed, coordination, pain, technique and actual output.

WHEN A WEARABLE SHOULD GET YOUR ATTENTION
None of this means ignoring meaningful physiological information.
If your resting heart rate suddenly changes substantially and remains altered, that deserves attention. If a device repeatedly identifies an unusual heart rhythm, unexplained oxygen abnormality or another concerning pattern—particularly when accompanied by symptoms—that information should not be dismissed simply because consumer wearables have limitations.
It should also not be self-diagnosed from an app.
A wearable may occasionally identify a signal worth investigating. When the issue moves from general fitness and recovery into possible medical territory, the appropriate next step is discussion with a qualified healthcare professional.
Your fitness tracker can provide useful data, but it is still not your doctor.
THE BEST WAY TO USE WEARABLE TECHNOLOGY
If you enjoy Oura, WHOOP, Apple Watch, Garmin or another tracker, I would not tell you to throw it away. I would tell you to keep it in perspective.
Look at weeks and months rather than one morning. Establish your own baseline instead of obsessing over comparisons with other people. Experiment with behaviours and see whether patterns repeat. Compare the device with what you record in your journal and what actually happens during training.
Pay particular attention when several independent signals agree. Poor subjective sleep, elevated resting heart rate, depressed HRV, unusual fatigue and terrible warm-ups together tell a much more compelling story than one isolated recovery score.
And periodically ask yourself a simple but important question:
Is this device helping me make better decisions—or is it simply giving me more things to worry about?
If the answer eventually becomes the latter, taking it off for a while may be more informative than checking another score.
THE BOTTOM LINE
Wearables are extraordinary tools. They can collect information we once had no practical way to gather at home, encourage activity, expose patterns, support behavioural change and make people more curious about sleep, recovery and health.
But more data does not automatically produce more wisdom.
Sleep stages are estimates. HRV needs context. Readiness scores are interpretations. Algorithms can disagree with your experience, and the feedback itself can sometimes change how you perceive your own condition.
That is why my view of wearable technology has evolved. I no longer think the question is whether trackers are good or bad. The better question is whether we are using them for the job they actually do well.
Let the wearable collect the data. Let the journal provide the context. Let your actual performance provide the reality check. And when the issue crosses from fitness monitoring into diagnosis, let an appropriate healthcare professional take over.
Your fitness tracker can tell you a lot.
It just shouldn’t get the final word.
SELECTED REFERENCES
- Performance of Consumer Wrist-Worn Sleep Tracking Devices Compared to Polysomnography: A Meta-Analysis
- A Performance Validation of Six Commercial Wrist-Worn Wearable Sleep-Tracking Devices for Sleep Stage Scoring Compared to Polysomnography
- Performance of Three Consumer Sleep-Tracking Devices Compared With Actigraphy and Polysomnography
- Validation of Nocturnal Resting Heart Rate and Heart Rate Variability in Consumer Wearables
- Orthosomnia: Are Some Patients Taking the Quantified Self Too Far?
- Prevalence of Orthosomnia in a General Population Sample
- Sham Sleep Feedback Delivered via Actigraphy Biases Daytime Symptom Reports in People With Insomnia
- Monitoring the Athlete Training Response: Subjective Self-Reported Measures Trump Commonly Used Objective Measures


