Readiness & training
Readiness
A 0–100 read on how recovered you are this morning, blended from your overnight physiology against your own personal baselines, not a population average.
How it’s built
Starts at 100. HRV against your personal band adds or subtracts the most; resting heart rate above baseline subtracts; sleep is penalised below about 7 h and above about 8.5 h; raised respiration and wrist temperature apply small illness penalties.
Range
80+ Prime · 60–79 Good · 40–59 Moderate · below 40 Low.
HRV against your own baseline; illness panel validated by COVI-GAPP —
PMID 35728900.
Today’s call (HRV gate)
Turns your recovery into a concrete training decision: proceed, ease, swap to easy, or rest.
How it’s built
HRV against your personal band is the primary signal. Form (TSB) and the illness panel can only downgrade a hard day toward rest; they never upgrade it. Easy days are never gated. With no HRV this morning, or no personal band yet, the call is unknown and the app says nothing rather than guess, unless an illness signal or very low form is real evidence on its own.
Range
Only hard days are gated; the underlying HRV, TSB and illness numbers are always shown.
Training status
One verdict on where your training is heading, with every input shown rather than hidden.
How it’s built
Combines your CTL (fitness) slope, ACWR (acute:chronic load) and Foster monotony. A strain or spike signal overrides a “productive” read (priority-ordered).
Range
Detraining · Maintaining · Productive · Overreaching · Strained · Mixed.
Banister fitness-fatigue model; Foster monotony and strain —
PMID 9662690. ACWR is a contested soft guardrail, and framed as one in the app.
Cumulative load
The slow-moving “are you quietly digging a hole?” read that daily readiness can’t capture.
How it’s built
An honest reserve from training load (TSB), HRV trend, resting-heart-rate drift and sleep debt, with each component shown. It refuses to fake a minute-by-minute battery.
Range
70+ Fresh · 50–69 Balanced · 30–49 Accumulating · below 30 Depleted.
Session analysis
What one workout actually did: how its minutes spread across your heart-rate zones, how much load it carried and, when your watch measured it, how fast your heart rate came down afterwards.
How it’s built
Every heart-rate sample recorded during the session is placed in your personal zones (from your LTHR when set, else your max HR); a sample stands for the time until the next one, capped at 60 s so a dropout can’t smear one reading across minutes. Session strain uses the same Banister TRIMP weighting and 0–21 scale as Day Strain, so the two compare directly. The intensity label is descriptive: easy (80% or more of minutes in Z1–2), threshold (35% or more in Z3), hard (20% or more in Z4–5) or mixed. It does not estimate a “training effect”. One-minute recovery is shown only when your watch measured it; it is never computed from an unstructured stop.
Range
Well-trained endurance athletes do about 80% of sessions easy and 20% genuinely hard. A one-minute heart-rate drop of 12 bpm or less was the abnormal cut-off in a large exercise-test cohort; wrist readings after an unstructured stop run lower, so compare against your own sessions.
Intensity distribution: Seiler 2010 —
PMID 20861519; heart-rate recovery and mortality: Cole 1999 —
PMID 10536127; Banister TRIMP training-impulse framework.
Sleep
Sleep score
One number for last night, built only from the parts of sleep the research supports.
How it’s built
Three components, weighted and renormalised over whatever was measured: regularity (the Sleep Regularity Index below) 40%, duration 35%, efficiency 25%. Duration scores 100 inside 7–9 h (7–8 h from age 65), losing 30 points per hour short and 12 per hour over, floored at 40. Efficiency scores 100 at 90% or better, 80 at 85%, and 0 at 65% or below. Regularity follows the UK Biobank distribution: an SRI of 87 scores 100, the median of 81 scores 78, and 60 scores 0. A score needs at least two components; duration alone is not a score, and a missing component is shown as missing, never filled in.
Range
85+ Excellent · 70–84 Good · 55–69 Fair · below 55 Poor.
Deliberately left out
Sleep-stage percentages: consensus supports continuity measures (efficiency, wakings, time to fall asleep) as quality indicators, with little agreement on architecture. Wake after sleep onset isn’t scored separately, because with only asleep and awake data it is the same two numbers as efficiency. A night whose source reports no awake time at all isn’t treated as flawless. Long nights are penalised more gently than short ones: the higher mortality of habitual long sleep is confounded by illness, and a single long night is usually catch-up.
Sleep regularity (SRI)
How consistent your sleep timing is from one day to the next. In a UK Biobank study of 60,977 people, it predicted mortality more strongly than sleep duration did.
How it’s built
The Sleep Regularity Index: the probability that you’re in the same state, asleep or awake, at the same minute on two consecutive days, scaled 0–100. Spline computes it minute by minute from the sleep your devices recorded, as the study did: waking in the night counts as awake, and naps count as sleep. Only nights exactly one calendar day apart are paired, so a missing night never compares Monday with Wednesday. It uses the latest 13 pairs within 28 days and needs at least 5.
Range
85+ Excellent · 75–84 Good · 65–74 Fair · below 65 work on consistency. In the study, the median was 81; 85 sat near the 70th percentile and 65 near the 10th.
Polar’s data carries only bed and wake times plus total interruptions, so an SRI computed from a Polar device reads somewhat high.
Definition: Phillips 2017. Mortality and population distribution: Windred 2024 —
PMID 37738616.
Staging confidence & sleep debt
Two honesty layers most apps skip.
Confidence
Spline borrows sleep staging from your device rather than owning it, so it flags low-confidence nights (phone-only detection, very short or fragmented sleep) instead of presenting stage percentages as fact. Nothing in the sleep score depends on staging, so the score stays valid on those nights.
Debt
Gross 7-day shortfall against a 7.5 h/night need. It reports gross shortfall rather than netting long nights one-for-one: a weekend lie-in only partially repays a weekday deficit.
One long night doesn’t repay chronic restriction: Banks 2010 —
PMID 20815182.
Tonight’s plan
When to turn the lights out tonight so you still wake at your usual time with the sleep you need, worked out from your own recent nights rather than a generic eight hours.
How it’s built
Your wake time stays where it usually is: the median of your last 14 nights, rounded to five minutes and never moved, because a lie-in trades the bigger lever (regularity) for the smaller one. From it the plan counts back the sleep you need: 7.5 h, plus up to an hour while you are repaying debt, spread over several nights, divided by your own sleep efficiency, plus 15 minutes to fall asleep. Wind-down starts 45 minutes before lights-out and caffeine stops 6 hours before it. One night never moves your bedtime more than 45 minutes earlier than usual, and the plan needs five recent nights before it says anything. The same plan appears on the widgets and your watch in the evening.
Range
A planning rule built on the evidence below, not a measurement; individual need spans roughly 7 to 9 hours.
Adults need 7 h or more: Watson 2015 —
PMID 26039963; regular timing outranks duration for mortality: Windred 2024 —
PMID 37738616; one long night does not repay chronic restriction: Banks 2010 —
PMID 20815182; sleep is hardest to start in the hours before your usual bedtime: Lavie 1986 —
PMID 2420557; caffeine 6 h before bed: Drake 2013 —
PMID 24235903; normal time to fall asleep: Ohayon 2017 —
PMID 28346153.
Alcohol & caffeine
The two logged habits with the largest measured effect on overnight recovery, shown beside your sleep numbers and never folded into readiness.
How it’s built
Alcohol: drinks logged in Apple Health from noon yesterday to 6:00 today, at 14 g of pure alcohol each, divided by your body weight and placed in a dose band. Without a body weight, no band is claimed, because the bands are per kilogram. Caffeine: milligrams logged within six hours of your usual bedtime (the median of your last 14 nights). 200 mg or more is flagged; smaller amounts are only marked “watch”, because smaller doses weren’t tested.
Range
Alcohol up to 0.25 g/kg low, 0.25–0.75 moderate, above 0.75 high. Overnight recovery fell by 9.3%, 24.0% and 39.2% across those bands in 4,098 within-person comparisons.
Why not in readiness
Readiness stays measured physiology. Mixing a self-report into a measured score is the black-box move Spline exists to avoid.
Alcohol dose and overnight recovery: Pietiälä 2018 —
PMID 29549064; caffeine 0, 3 and 6 hours before bed: Drake 2013 —
PMID 24235903.
Everyday signals
Today’s lead signal
The one signal furthest outside your own normal range this morning, named first. Or nothing, when nothing is.
How it’s built
Candidate signals are ranked by how far each sits outside your personal range. The card cannot be shown without the measurement that triggered it: the reason is a required part of it, not optional text. It stays silent unless a usable HRV band and a readiness score exist, because “nothing is out of range” would otherwise only mean “nothing was measured”. There is deliberately no all-clear card: on a normal day, absence is the signal.
Check-ins
A single, specific question when a signal persists, so the app learns context it can’t measure.
How it’s built
Asked only when a signal has persisted for three or more consecutive days, never after a one-day spike. Your answer becomes a journal tag on that day, which the pattern engine can later test. One question at a time, strongest first; answering or dismissing hides it for 14 days. It all happens on your iPhone.
Habit patterns Pro
Which of your habits and journal tags travel with better or worse recovery, found automatically and reported with the caution they deserve.
How it’s built
Every habit and tag is compared against the next morning’s HRV, resting heart rate and sleep, never the same morning. A pattern is reported only if the effect is at least small-to-moderate (|Hedges’ g| of 0.3 or more), its bootstrap confidence interval excludes zero, each side has at least 10 nights, and it survives Benjamini–Hochberg false-discovery correction across every test run. Findings are worded as associations, never causes, with the option to test one properly as an N-of-1 experiment.
False-discovery rate: Benjamini & Hochberg 1995; Hedges’ g effect size with bootstrap confidence intervals.
Stress history
How your daytime stress has trended, described rather than scored.
How it’s built
Daytime stress is derived from each day’s continuous heart rate and stored once a day from then on. The history reports the days above your own median, and whether the newest third of the window differs from the oldest third (thirds, not halves, so a fortnight-long shift isn’t averaged away). With fewer than 14 days it says nothing.
There is deliberately no “cumulative stress” score. A composite like that would be exactly the kind of undisclosed number Spline exists to avoid.
Body, longevity & bloodwork
Biological age (wearable) Pro
A transparent wearable estimate of how old your body “reads”, against your calendar age.
How it’s built
Each factor (VO₂max, resting heart rate, HRV, sleep, sleep consistency, body fat and activity) adds or subtracts years against a sex-specific healthy anchor. The factors sum to your offset, and each shows its own ± years. It needs at least three measured markers; a marker with no data is left out rather than counted as zero.
Range
An estimate with wide individual uncertainty, labelled as one. The dial’s scale is honest: your calendar age sits at the top, with a ±15-year window either side.
Composite of validated markers; VO₂max and mortality: Mandsager 2018 —
PMID 30646252.
Blood biological age (PhenoAge) Pro
A mortality-linked biological age computed from nine standard chemistry and blood-count markers.
How it’s built
The Levine PhenoAge formula: a weighted combination of albumin, creatinine, glucose, CRP, lymphocyte %, MCV, RDW, alkaline phosphatase, white-cell count and age, mapped through a Gompertz mortality model. All unit conversions are pinned by tests.
Range
Lower than your chronological age is favourable. A risk surrogate with wide individual uncertainty, not a diagnosis.
Marker change significance (RCV) Pro
Whether a lab result actually moved, or just wobbled within assay and biological noise.
How it’s built
Reference Change Value from each marker’s within-subject and analytical variation (asymmetric for right-skewed markers like CRP and triglycerides). Readings in different units or assays are refused rather than wrongly compared. Retest cadence follows the same biological-variation data.
EFLM / Westgard biological-variation estimates; conservative, for “is this real?”, not clinical decisions.
What actually moves each marker
When a biomarker is out of range, the app lists the highest-leverage interventions, each with an explicit evidence grade and a PubMed link, and an honest flag when the strong option is a prescription only a doctor can start.
Evidence grades
Strong · Moderate · Emerging · Clinical
It never poses as medical advice: “Clinical” means bring the numbers to your doctor. Lp(a), for example, is framed as genetically fixed: the actionable response is to drive every other atherogenic particle low, not to chase a number lifestyle barely changes.
Waist-to-height ratio
Waist circumference divided by height: a better single predictor of cardiometabolic risk than BMI or waist alone, and it doesn’t misread muscular builds the way BMI does.
How it’s built
Your waist from Apple Health or a manual entry (which is written back to Health), divided by your height.
Range
Below 0.5 healthy · 0.5–0.6 increased · 0.6 and above high.
Heart & blood pressure
Blood pressure (home cuff)
A blood-pressure cuff at home must be judged differently from one at the clinic: on a multi-day average, against lower home thresholds. Most apps quietly apply the office 140/90 to home readings, which is wrong.
How it’s built
Averages your recent home readings (a 7-day morning and evening average is the clinically valid number; a single reading is not a diagnosis). Category uses the ESC 2024 European scheme on home values (Non-elevated below 120/70, Elevated 120–134/70–84, Hypertension 135/85 and above), assigned by the higher of systolic and diastolic. ACC/AHA 2017 (US, home threshold 130/80) is shown as a comparison.
Also derived
Mean arterial pressure; pulse pressure (systolic minus diastolic), flagged above about 60 mmHg as a large-artery-stiffness and independent cardiovascular-risk marker; and short-term variability.
Range
Home target below 135/85. Risk is continuous: each 20/10 mmHg lower roughly halves vascular mortality, down to 115/75 mmHg.
ECG rhythm screening
When an Apple Watch supplies a single-lead ECG, the app shows its rhythm result, honestly framed.
What it is, and isn’t
A single-lead flag (sinus rhythm, possible atrial fibrillation, or inconclusive) is a screening signal, never a diagnosis. A possible-AFib result means: have a clinician confirm it with a medical 12-lead ECG; never start or stop medication based on the device. Inconclusive results are shown as such, not hidden.
The Withings BPM Core’s on-device ECG is not shared to Apple Health (only blood pressure and heart rate are), and apps can’t read non-Apple-Watch ECG, so the ECG readout appears only if you also record on an Apple Watch.
Confirmation required: ESC 2024 AF guideline,
PMID 39210723; asymptomatic-screening evidence insufficient: USPSTF 2022,
PMID 35076659.
Food
Food log numbers
How a typed meal becomes calories without the AI inventing a number.
How it’s built
Apple’s on-device model, or a built-in parser where Apple Intelligence isn’t available, only reads food names and amounts from what you wrote. Nutrition comes from food databases: Germany’s BLS, France’s CIQUAL, Denmark’s Frida, the USDA’s FNDDS (foods as eaten, with measured portion weights), and Open Food Facts for packaged products. Every portion shows where its weight came from: your own earlier correction, the food’s own portion, the label serving, a similar USDA food, or a density-aware estimate. Your corrections are remembered on your iPhone, where you can review or forget them.
The one exception
A nutrition-label photo is the only place a number comes from an image, so it is cross-checked against the 4/4/9 kcal Atwater rule before you log it.
BLS 4.0 (Max Rubner-Institut, CC BY 4.0); Table CIQUAL 2025 (ANSES, Etalab Open Licence 2.0); Frida 5.5 (DTU National Food Institute, CC BY 4.0); FNDDS 2021–2023 (USDA Agricultural Research Service, public domain); Open Food Facts.
Ask Spline
How answers are built Pro
Why the coach can’t make up a number.
How it’s built
Your question is routed to a topic (readiness, sleep, training, HRV, a data lookup and so on), first by whole-word keywords in English and German, and only if that finds nothing, by Apple’s on-device model. The model may choose a topic; it never writes the answer. The app then composes the answer from the same values your cards show and the engines’ own verdicts (the gate, tonight’s plan, the strain target), so the coach and the cards cannot disagree. Missing data is stated as missing. A data question such as “when was my HRV lowest last month?” becomes a query over the same data the charts draw, and the interpreted query is shown above the answer so a misread question is visible.
Suggestions
Levers are drawn from a fixed, cited list and phrased as things to try one at a time, not prescriptions.
Honesty notes that ship in the app. Some claims were corrected during a PubMed audit and are framed accordingly: HDL is U-shaped (high isn’t automatically better); ApoB and other atherogenic particles are monotonic, never shown as a U-curve; sodium uses the 2,300 mg chronic-disease-risk limit, not an invented optimum. Where the evidence is weak (cold exposure, cycle-based periodization) the app says so rather than overselling it.