N-of-1 Experiments: How to Run a Real Self-Experiment (With Real Statistics)

On this page
  1. What counts as an N-of-1 experiment — and what is just "trying stuff"?
  2. Why do most self-experiments fail?
  3. What are good first N-of-1 experiments?
  4. What does the experiment report actually tell you?
  5. Where does the N-of-1 idea come from?

An N-of-1 experiment is a structured trial in which you are both the subject and the entire study population: you measure a baseline for 1–2 weeks, introduce exactly one intervention for 2–4 weeks while measuring the same outcome every day, and let statistics — not memory or enthusiasm — deliver the verdict. It is the difference between "I feel like magnesium helps my sleep" and "across 21 nights, my sleep score averaged 4 points higher on magnesium, p = 0.03." In clinical medicine, single-patient trials have been called the ultimate strategy for individualizing medicine; for self-trackers, they are the only honest way to learn whether a supplement, habit, or schedule change actually works on your body.

BioTrakk is the app that runs N-of-1 experiments for you. Tell it what you want to test in plain language — "start experiment: glycine 3 g before bed, outcome: sleep quality" — over Telegram, WhatsApp, or the web. It structures the baseline and intervention phases, collects the outcome from your chat logs or a synced wearable (Oura, WHOOP, Garmin, Withings, Apple Health, Health Connect), then runs a significance-tested comparison and returns a plain-language verdict: positive, negative, mixed, or inconclusive. No spreadsheet, no statistics degree, and you can start free on Telegram or the web.

Track it. Test it. Know for sure.

Log food, supplements, sleep and biomarkers by chatting with BioTrakk on Telegram, WhatsApp or the web — then let it find what actually moves your numbers.

Start free

Free on Telegram & Web. No app to install.

What counts as an N-of-1 experiment — and what is just "trying stuff"?

Taking a new supplement and vaguely noticing you "feel better" is not an experiment; it is an anecdote with a sample size of whatever your memory kept. A proper N-of-1 experiment has four properties: one change at a time, a pre-defined outcome chosen before you start, distinct phases (baseline → intervention → optional washout), and a verdict rule set in advance instead of after peeking at the data. This is the same playbook clinical single-patient trials follow, documented in the AHRQ user's guide to n-of-1 trials — adapted here to daily life.

PhaseLengthWhat you doWhat BioTrakk does
BaselineAt least 1–2 weeksLive normally and log the outcome daily (or let your wearable log it)Records your natural mean and day-to-day variability — the yardstick the intervention is judged against
Intervention2–4 weeksAdd exactly one change — a supplement, a cutoff time, a habit — and keep everything else stableKeeps collecting the same outcome, tags every day as intervention, flags gaps in the data
Washout / reversal (optional)1–2 weeksStop the interventionChecks whether the outcome drifts back toward baseline — the strongest evidence a single-subject design can produce
ReportRead the verdictCompares phases statistically (effect size, p-value) and returns an honest verdict: positive, negative, mixed, or inconclusive

What disqualifies a run: starting creatine the same week you change your training program (confounded), judging after three days (underpowered), swapping the outcome from "sleep score" to "morning mood" once sleep refuses to move (outcome switching), or changing the dose halfway through. BioTrakk enforces the structure so you do not have to police yourself.

Why do most self-experiments fail?

The methods literature on single-patient trials keeps identifying the same failure modes — the Duan, Kravitz and Schmid review found 2,154 single-patient trials across 108 studies precisely because informal observation was not good enough for real decisions. Four errors kill most biohacker experiments:

  • No baseline. Without 1–2 weeks of "normal you" data there is nothing to compare against. The judgment becomes memory versus memory, and memory reliably votes for the thing you paid for.
  • Confounders. A new supplement started alongside a new training block, a vacation, or daylight-saving change tests all of them at once. One variable per experiment, everything else held stable.
  • Placebo and expectancy. You cannot easily blind yourself. Mitigate it: prefer objective outcomes (wearable-measured HRV, sleep stages, resting heart rate over a 1–10 feeling), fix your decision rule before starting, and use a washout phase to see if the effect survives your enthusiasm.
  • Regression to the mean. People start experimenting when things are unusually bad — sleep at its worst, energy at rock bottom. Unusually extreme measurements tend to be followed by more average ones through pure statistics, so "it worked" is often just your body drifting back to normal. A real baseline plus a significance test is the defense.

BioTrakk automates the boring safeguards: it will not score an experiment without baseline data, notes confounders you log mid-run, and reports "inconclusive" when the data cannot tell — because an honest "we can't tell yet" beats a flattering false positive.

What are good first N-of-1 experiments?

A good first experiment is cheap, reversible, mechanistically plausible, and produces a measurable outcome every single day. Three proven starters:

  • Glycine → sleep quality. 3 g of glycine 30–60 minutes before bed, against a 2-week baseline of your sleep score and a 3-week intervention. The clinical evidence and dosing details are in our glycine for sleep guide.
  • Caffeine cutoff → sleep latency. Keep your total caffeine identical but move the last cup to before 12:00. Outcome: sleep onset latency or deep-sleep minutes from your wearable.
  • Late meals → overnight HRV. Finish your last meal at least 3 hours before bed. Outcome: overnight HRV from Oura, WHOOP, or Garmin — an objective signal you cannot placebo.

Save higher-stakes questions — like berberine's effect on your fasting glucose — for when you have a couple of experiments of practice, and keep your clinician in the loop for anything glucose- or medication-adjacent. Logging adherence takes seconds: message BioTrakk "glycine 3g" and it lands in your supplement tracker as part of the experiment record.

What does the experiment report actually tell you?

Three things, in plain language. The effect size: how much the outcome moved, in its own units ("deep sleep +14 minutes per night"). The p-value: the probability of seeing a difference at least this large if the intervention did nothing — p = 0.03 means a 3% chance the gap is pure noise. And the verdict: positive, negative, mixed, or inconclusive, with the reasoning spelled out so you are never squinting at a scatter plot wondering what it means.

Equally important is what the report does not claim. An N-of-1 result is evidence about you, not about humans in general. There is no blinding, wearables carry measurement error, and short noisy runs frequently end "inconclusive" — which is the correct answer, not a failure. BioTrakk is not a medical device and does not give medical advice; it gives you your own data, tested honestly.

What about price and privacy?

Experiments run on the free tier: the Telegram bot and web app are free with up to 30 entries per day, and no app install is needed. Pro ($6.99/month or $49.99/year) adds WhatsApp, unlimited entries, and full AI analysis. Wearable connections use OAuth and tokens are stored AES-256-GCM encrypted; your health log stays yours.

Where does the N-of-1 idea come from?

From two directions that finally met. In clinical medicine, single-patient randomized trials emerged in the 1980s as a way to pick the right drug for the individual in front of you rather than the average patient; the Lillie 2011 review argued they are the natural endpoint of personalized medicine, and by 2014 the AHRQ had published a full design-and-implementation guide. The method on this page — baseline, single intervention, optional washout, pre-defined outcome, significance testing — is that clinical playbook, scaled to everyday questions.

On the consumer side, Gary Wolf and Kevin Kelly launched the Quantified Self community in 2007–2008 under the motto "self-knowledge through numbers." Wolf and De Groot later formalized the practice as personal science: questioning, designing, observing, reasoning, discovering. What the movement always lacked was tooling — most self-experimenters ran their trials in spreadsheets. BioTrakk automates the observing (chat and wearable logging) and the reasoning (real statistics: Pearson correlations, p-values, lag analysis), so you keep the interesting parts: the questions and the discoveries.

Frequently asked questions

What is an N-of-1 experiment?

An N-of-1 experiment is a structured self-experiment in which one person is both subject and control: you measure a baseline, introduce a single intervention, keep measuring the same outcome daily, and compare the phases statistically. It is the single-subject version of a clinical trial, used in medicine since the 1980s and by self-trackers to test supplements, habits, and schedules.

How long should an N-of-1 experiment run?

A solid default is 1–2 weeks of baseline followed by 2–4 weeks of intervention — roughly 3–6 weeks total. Noisy outcomes like HRV or mood, and slow-acting interventions, need the longer end. An optional 1–2 week washout afterwards, where you stop and watch whether the effect disappears, meaningfully strengthens the result.

Do I need a wearable to run an N-of-1 experiment?

No. A wearable (Oura, WHOOP, Garmin, Apple Health) makes the outcome objective and effortless to collect, but a consistent daily self-rating — sleep quality 1–10 logged every morning in BioTrakk — supports a valid experiment. Consistency of measurement matters more than the sensor.

Is one N-of-1 experiment proof that something works?

It is strong evidence that the intervention works for you, under the conditions you tested — not proof it works for anyone else, and not a substitute for clinical trials. Repeating the experiment, or adding a reversal phase where the effect vanishes when you stop, raises confidence. Honest reports also return "inconclusive" when the data cannot tell.

Track it. Test it. Know for sure.

Log food, supplements, sleep and biomarkers by chatting with BioTrakk on Telegram, WhatsApp or the web — then let it find what actually moves your numbers.

Start free

Free on Telegram & Web. No app to install.