Smartwatch Calories — Why Two Watches Disagreed by 80% on One Ride

Smartwatch Calories — Why Two Watches Disagreed by 80% on One Ride

Search terms: calories wrong · Fitbit vs Samsung calories · Pixel Watch calorie accuracy · Samsung Health workout calories · active vs total calories · gross vs net calories · MET formula · Keytel · energy expenditure accuracy · cardiac drift · heat inflated heart rate · EPOC

Two watches, one wrist-pair, one ride. Every sensor agreed. The calorie numbers differed by 80%. This page is the method for working out which number to believe — and the answer for why the honest verdict was neither, yet.


The case

5.56 mi bike ride on paved suburban streets, Phoenix, 105 °F, 15 Aug 2026.

Pixel Watch 4 (Fitbit) Galaxy Watch Ultra (Samsung Health) Δ
Distance 5.56 mi 5.56 mi 0%
Duration 27m 48s 27m 28s 1.2%
Avg speed 12.00 mph 12.1 mph 0.8%
Avg heart rate 147 bpm 148 bpm 0.7%
Elevation gain 120 ft 148 ft 21%
Calories 487 271 80%

Every sensor agrees. Only the model disagrees.

Distance, duration, speed and heart rate are measured. Calories is the only figure on that screen that is computed from a stored profile plus an undisclosed equation. That is why it is the only one that diverged, and it is why you cannot fix it by trusting one watch’s hardware over the other’s.


The three candidate causes, ranked

1. Profile data — explains 0% or ~97%. Nothing in between.

Check this first, because it is binary and it costs two minutes.

A stale body weight deflates every workout, because calorie models scale close to linearly with mass. The tell is arithmetic:

8.0 METs x 3.5 x 70 kg / 200 x 27.6 min  =  270.5 kcal      Samsung reported 271

70 kg is the textbook MET reference weight — the default baked into the standard formula. The ratio matched too: 122/70 = 1.75 against an observed calorie ratio of 487/271 = 1.80.

In this case the hypothesis was REFUTED — Samsung Health held 267 lb, correct to within daily fluctuation. Check it anyway on any new discrepancy; it is the cheapest possible explanation and when it is the cause, everything downstream is moot.

Verify by computing reported_kcal x (true_kg / stored_kg) and seeing whether it lands on the other device’s number.

2. Algorithm — explains most of the remainder

  • Fitbit/Pixel is heart-rate driven. Google’s own help text says heart rate “is also included, especially to estimate calories burned during exercise.”
  • Likely: Samsung does not use heart rate for cycling — only for walking and running, falling back to a MET/profile table otherwise. Flagged Likely: the primary sources are Samsung community threads that return HTTP 403 and could not be read firsthand.

Solve backwards for what Samsung was actually crediting, at the correct mass:

271 kcal / 27.6 min / (3.5 x 121 kg / 200)  =  4.63 METs

4.6 METs is “bicycling under 10 mph, leisure.” The ride was 12 mph at 148 bpm. A table-driven engine graded a hard effort as a gentle one, because it never looked at the heart rate.

Neither vendor publishes its equation. Anyone quoting a specific Fitbit or Samsung calorie formula is reconstructing it.

3. Gross vs net — the intuitive explanation, and the smallest

This is the one everybody reaches for first: “one includes my metabolism and the other doesn’t.” It is real, and it is nearly worthless.

Resting metabolism over a 27.6-minute ride, for a 121 kg 49-year-old male at 5'10":

Mifflin-St Jeor  ->  2,082 kcal/day  ->  1.45 kcal/min  ->  40 kcal for the ride

40 kcal is 18.6% of a 216-kcal gap — and that is the ceiling, reached only if one device reports gross and the other net.

Likely: both are already gross. Fitbit’s API defines its calorie series as “inclusive of BMR” and keeps a separate MarginalCalories field for the active-only concept; Samsung maps exercise calories to Health Connect’s TotalCaloriesBurnedRecord, which Android defines as including basal energy. If both are gross, this explains zero.

Personal constant: BMR ≈ 1.45 kcal/min → 43 kcal per 30 minutes. A ~50/30min rule of thumb over-subtracts by about 7 kcal per half hour — not worth correcting.

Beware the MET shortcut for this: 1 MET = 3.5 ml O2/kg/min is calibrated on a ~70 kg reference person and assumes every kilogram is metabolically active. At 121 kg with high fat mass it overstates resting metabolism by roughly 46% (2.12 vs 1.45 kcal/min). Use Mifflin-St Jeor.


Why the physics could not arbitrate

Three independent-looking estimates were built blind — MET table, heart-rate equation (Keytel), and a mechanical power model (rolling resistance + aero + climbing ÷ gross efficiency).

MET table (naive)      ~473 kcal
Keytel, male           ~476 kcal     <- female coefficients give ~222; sex matters enormously here
Cycling power model    ~280 kcal

The Pixel is closer if and only if the true gross cost exceeded (487+271)/2 = 379 kcal. Centrally the models land 325–391 kcal — straddling 379. No arbitration is possible at that resolution.

Two traps worth remembering:

  1. The MET and Keytel estimates are not independent corroboration. Both are calibrated on normative populations, both are blind to mechanical work, and both run high the same way for a heavy body. Two methods agreeing is meaningless when they share the bias.
  2. The physics model’s honest envelope is 193–615 kcal across plausible tyre, drag, efficiency and gradient constants — a 3.2x range, against a device disagreement of only 1.80x. A model whose uncertainty exceeds the thing it is measuring cannot referee.

The two tests that actually settle it

Modelling was exhausted. These are measurements.

Test A — gross vs net. 15 minutes, free.

Turn OFF auto-pause on both watches first. Both platforms auto-pause cycling when stationary; if it triggers, the workout accrues nothing and both devices read as false “active-only.”

Sit completely still. Start Outdoor cycling on both — outdoor, so the speed-based engine sees a genuine zero, and so the activity type matches the ride under investigation. Stop after 15 minutes.

Reading Verdict
~21–22 kcal (15 x 1.45) reports gross — BMR included
~0–4 kcal active-only

Test B — which engine is calibrated. 20 minutes, indoors.

Needs a displayed wattage, not an owned power meter — a smart bike, spin bike, BikeErg or Peloton all qualify. Ride steady, wearing both watches. Mechanical work becomes known rather than modelled:

gross kcal  ~=  watts x minutes x 0.068 to 0.080

(from P x t / GE / 4184, gross efficiency 0.18–0.21)

Steady watts 20 min → true gross
100 W 137–159 kcal
125 W 171–199 kcal
150 W 205–239 kcal
175 W 239–279 kcal

Whichever watch lands nearest that band has the better engine. Because Test A already established which readings include BMR, the comparison is like-for-like — the two tests compose.

Use the bike’s watts, not the bike’s calories. A smart bike measures mechanical work directly, which genuinely beats a wrist — but converting watts to calories still requires assuming an efficiency, and consumer equipment tends to assume a generous one. Record its calorie figure anyway; comparing it to the band above tests whether the bike is optimistic.

Run Test B on a fresh morning, not after a hard session. Residual heat load and fluid deficit keep heart rate elevated for hours, which inflates the HR-driven watch against a known wattage — contaminating the exact thing the test exists to isolate.

Test C — isolating the heat. Autumn.

Repeat the identical route at the identical average speed on a 70–75 °F morning. Same mass, same distance, marginally more mechanical work (cool air is denser). If the HR-driven number drops sharply while distance, duration and speed are unchanged, the original was heat-inflated. In Phoenix, a genuine 70–75 °F dawn realistically means October.


Heat: heart rate stops meaning work

Likely: heart rate rises roughly 1 bpm per °C of air temperature above thermoneutral at a fixed workload — about 15 bpm on a 105 °F afternoon that buys no extra mechanical work. The mechanism is skin blood flow and falling stroke volume, not extra oxygen consumed.

The cleanest demonstration: cyclists held at constant heart rate in 35 °C consumed 24% less oxygen and produced 37% less power by 45 minutes. Identical HR, less metabolism.

Consequences that matter for reading the screen:

  • 85% of max heart rate on a hot afternoon does not mean 85% of maximum metabolic work.
  • “17 of 28 minutes in peak zone” is a heart-rate distribution, not a work distribution.
  • An HR-driven engine reads that drift as calories. A table-driven engine ignores it entirely. On a hot day these two designs must diverge, and in exactly this direction.

Elevated resting HR afterwards is mostly not EPOC. At 40 minutes post-effort, residual core temperature and fluid loss dominate; EPOC’s contribution after a ~28-minute effort is small. Working backwards through Keytel, closing the whole 216-kcal gap by heat correction alone would require a true working heart rate of 98 bpm — not credible. Heat explains roughly a third of the discrepancy, not all of it.

The useful recovery metric is next morning’s resting HR, not anything on the workout screen. Back to baseline = strain handled. Still elevated = heat load and fluid deficit haven’t cleared.


How accurate these numbers can ever be

Both are estimates with a wide band. Neither is a measurement.

  • Fuller et al. 2020 (systematic review, 305 energy-expenditure comparisons): only 9.2% fell within ±3% of laboratory reference, and no brand managed it more than 13% of the time.
  • A 2022 Fitbit meta-analysis reports limits of agreement of −5.32 to +5.70 kcal/min — wider than the average burn rate of most exercise.
  • Best available Samsung validation (n=148, treadmill, indirect calorimetry): group bias near zero but per-session MAPE ~10–12%, Bland-Altman limits −62 to +66 kcal on a 214 kcal workout — about ±29% for a single session. Caveats: running not cycling; sample averaged 65.5 kg; and it was Samsung-funded.
  • No peer-reviewed energy-expenditure validation of any Pixel Watch generation exists. Anyone quoting a Pixel Watch calorie-accuracy percentage is fabricating it.

Realistic single-session band: ±20–30%, and wider at high body mass, because error grows with body-fat percentage and every published validation sample averaged ~65 kg.

For a record shared with a coach or doctor

  • Pick one device and stay on it. Switching mid-record inserts an 80% step change that looks like a physiological event and isn’t.
  • Label the column with the device name and the word estimated.
  • Read the trend across weeks, not the level of any one session. A number consistently 15% high still tracks direction correctly, which is what matters.
  • Distance, duration, speed and heart rate agreed within ~1% across both watches. Those are the trustworthy columns. Calories are not a clinical quantity.

The transferable lesson

The same shape keeps appearing across this KB: a number that reports confidently on something it does not measure.

Here it was two watches showing a calorie field with the authority of the distance field beside it, when one is sensed and the other is invented. The diagnostic move that worked was not more modelling — it was asking which fields are measured, which are computed, and what stored value feeds the computation.

And the discipline that mattered: the leading hypothesis was elegant, arithmetically striking (270.5 vs 271), and wrong. It was killed by a two-minute check of a stored profile field. State the falsification condition before running the test, then actually run it.