Blog weather-sensitivity · symptom-tracking · self-tracking

How to keep a symptom diary that actually finds your triggers

Memory is a bad judge of weather triggers. Here's how to log symptoms so statistics — not vibes — can tell you what's actually a pattern in your own data.

Every article in this series so far has landed on the same honest, slightly unsatisfying conclusion. Pressure and headache: a real but modest link, tangled up with belief running ahead of the data (weather sensitivity: what the science says). Joints and rain: genuinely conflicting studies, split by research design as much as by anything else. Geomagnetic storms: a small population-level signal sitting on evidence too thin to say what it means for any one person. Heat, humidity, and fatigue: the one link solid enough to call basic physics, undercut by how wildly the personal threshold varies. Read all four and one thread runs through every single one — the population-level answer is never "yes" or "no," it's "weakly, for some people, and it depends." That's not a dodge. It's the actual, honest shape of this science, and it means the only useful next question is not "does weather affect people" but "does it affect me, and how would I actually know." This article is about answering that second question properly — with a diary built to be checked by statistics, not by memory. A note before we start: this is informational material, not medical advice. If a symptom is new, severe, or frightening, that's a conversation for a doctor, not an app or a spreadsheet.

Why your memory is the wrong judge

Start with the most direct experiment on exactly this question. In 1996, Donald Redelmeier and Amos Tversky — a physician and the psychologist who, with Daniel Kahneman, pioneered the study of cognitive bias — tracked 18 arthritis patients for more than a year, recording their pain alongside the real local weather that each patient believed affected them. When they ran the numbers, there was no statistically significant association between any patient's pain and their own implicated weather condition — not for any single patient in the group (Redelmeier & Tversky, 1996). That result alone might just mean weather truly doesn't matter to joints, which lines up with plenty of what this series has already covered. But the same paper ran a second, cleaner experiment: it showed 97 college students purely random, computer-generated sequences of two variables with no built-in relationship at all, and asked them to judge whether the sequences were correlated. People reliably reported seeing correlations in the noise — patterns that, by construction, were not there to find. The authors' conclusion was that the belief in weather-pain links likely survives not because the data supports it, but because the human mind is built to detect patterns whether or not they exist — what researchers call illusory correlation.

That tendency has a well-documented partner: confirmation bias, the habit of noticing and remembering the days that fit a story you already hold — a stormy day with a bad headache — while quietly discounting the stormy day that passed without incident, and never even registering the clear day that also happened to be bad. Add in plain recall bias — the fact that reconstructing "how bad was my headache three weeks ago" is a much less accurate act than rating it right now — and you get a memory system that is, structurally, a poor instrument for this specific question. A study comparing 181 children's retrospective headache questionnaires against the same children's real-time four-week diaries found exactly this pattern in practice: compared with the diary, the retrospective questionnaire systematically overestimated both headache intensity and headache duration, with the size of the error tracking the child's age, headache severity, and depression scores (van den Brink, Bandell-Hoekstra & Abu-Saad, 2001). None of this means people are careless or dishonest about their own symptoms. It means memory was never built to function as a measuring instrument, and asking it to do statistics-grade work is asking it to do a job it was never designed for.

What a diary needs to get right

The fix researchers reach for has a name: ecological momentary assessment, or EMA — recording a symptom close to the moment it happens, in the setting it happens in, rather than reconstructing it later from memory. The method was formalized in a foundational 1994 paper by Arthur Stone and Saul Shiffman, who argued that assessing symptoms, mood, and behavior in real time and in a person's actual daily environment sidesteps the recall problems that undermine retrospective reporting (Stone & Shiffman, 1994). Applied to a personal weather-and-symptom log, that principle translates into a short list of concrete habits, not a vague instruction to "pay attention":

  1. Log close to the moment, not at the end of the week. A same-day entry is a measurement; a Sunday-night reconstruction of the whole week is already drifting back toward memory and its biases.
  2. Log the misses, not just the hits. A calm day with normal pressure and no headache is exactly as important a data point as a stormy day with a bad one — leaving it out is confirmation bias with extra steps, just wearing the costume of a diary.
  3. Keep the categories consistent. Rating a symptom on the same 0–10 scale, or checking the same short symptom list, every single day is what makes day 3 comparable to day 60. A diary that changes its own units halfway through can't be analyzed at all.
  4. Let the weather attach itself automatically where possible. The less a log depends on you separately remembering to check and record the day's pressure or humidity, the less room recall bias has to creep back in through a side door.
  5. Keep going past the point where it feels like you already know the answer. This is the hardest rule and the one most people break first, for a reason the next section makes precise.

Why it takes statistics, not a glance, to know

Say, hypothetically, that after a month of daily logging your data shows a correlation of r = 0.34 between low-pressure days and how bad your headache felt that day. That's a real-looking number — the kind of tidy correlation that would make for a satisfying "aha." But with only 30 days of data behind it, a correlation of that size works out to a p-value of roughly 0.07 by the standard statistical test for correlations — which does not clear the conventional significance bar of 0.05. In plain terms: with that little data, a correlation that size is still quite plausibly a coincidence, even though it looks meaningful on the screen. Calling it a finding at that point would be committing exactly the error Redelmeier and Tversky's college students committed with pure noise — seeing a pattern before there's enough evidence to distinguish it from chance.

Here's the part that makes logging longer worth the effort: keep recording, and suppose that same r = 0.34 relationship holds up once you reach 90 days — about three months of consistent entries. Run the identical statistical test on the larger sample and the same-sized correlation now works out to a p-value of roughly 0.001 — comfortably past the significance threshold, because a correlation of a given size becomes much harder to explain by chance once it's measured across three times as many days. Nothing about the underlying relationship changed between day 30 and day 90; only the amount of evidence behind it did. That's the honest, unglamorous reason "give it time" isn't a platitude — statistical power to detect a real effect (or correctly rule out a fake one) climbs with the number of observations, and a month of enthusiastic tracking is usually not enough of them.

The multiple-comparisons trap — and why FDR matters

There's a second trap that opens up the moment you start checking more than one weather variable at a time, which is exactly what most people do — pressure, temperature, humidity, wind, and more, all tested against a symptom in the same sitting. Test eight different weather factors against your symptoms using the ordinary 0.05 significance cutoff on each one individually, and pure chance alone means that roughly one in twenty of those tests will look "significant" even if none of the underlying relationships are real — and across eight simultaneous tests, that adds up to a meaningfully elevated chance that at least one false alarm turns up looking like a discovery. This is the multiple-comparisons problem, and it's the statistical version of the confirmation-bias trap from earlier in this article: the more questions you ask of the same data at once, the more chances chance itself gets to hand you a false positive that looks exactly like a real one.

The standard fix is a false discovery rate, or FDR, correction — a method introduced by Yoav Benjamini and Yosef Hochberg in 1995 specifically for this situation, where many hypotheses are tested at once and a plain, uncorrected significance cutoff would let too many false positives through (Benjamini & Hochberg, 1995). Rather than demanding an impossibly strict bar for every single test — which tends to bury real effects along with the fake ones — FDR correction adjusts the threshold based on how many comparisons were run, so that among whatever associations end up flagged as "significant," the expected share that are actually false stays controlled and known, rather than left to guesswork. Applied to a personal weather log, this is the difference between "pressure showed up as significant, so it must be my trigger" and "pressure held up even after accounting for the seven other factors I also checked" — the second claim is the one worth trusting, and it's the one that needs correction to make honestly.

Illustrative example, not real data: the same r=0.34 correlation is not significant after 30 days (p≈0.07) but is after 90 days (p≈0.001); and when eight weather factors are tested at once, only associations that survive false-discovery-rate (FDR) correction count as real.SIGNIFICANCE × DAYS · SAME_r=0.34p=0.05n=30·p≈.07n=90·p≈.001Not significant — coincidence still plausibleSignificant — unlikely to be chanceDAYS_LOGGED →8_FACTORS_TESTEDUNCORRECTED_p<.05AFTER_FDROnly associations that survive FDR correction countFIG.06 · SIGNIFICANCE + FDR · ILLUSTRATIVE_EXAMPLE · NOT_A_MEASUREMENT

What this actually buys you: a sample size of one, done properly

This is the logic behind a research design called the n-of-1 trial: rather than asking whether a treatment or trigger works for people in general, you take one person's own repeated, consistent observations as the dataset, with that same person's other days serving as the comparison — no separate control group required, because everyone is their own control across time (Duan, Kravitz & Schmid, 2013). That's exactly what a properly kept symptom-and-weather log is: not a substitute for population research, but a legitimate, individually meaningful dataset in its own right, provided it's built the way the sections above describe — logged close to the moment, including the ordinary days, over enough time, checked with a method that accounts for testing more than one thing at once.

This is also, plainly, why MeteoHealth exists as an app rather than a spreadsheet template. It logs your symptoms and mood in two taps, pairs each entry with local weather and the Apple Health metrics you already collect, and runs the actual statistics on-device — Pearson correlations, ANOVA, and false-discovery-rate correction across the weather variables it checks — so that what surfaces as a pattern in your own data has already been through the same discipline described in this article, rather than a glance at a chart. It does not diagnose or predict anything; it shows you your own numbers, with the uncertainty left honestly attached, and it's built entirely on-device so the personal health data behind those numbers never has to leave your phone. Everything else in this series — the weak population averages, the disagreement between studies, the gap between belief and data — is the reason that approach is necessary in the first place. A population study can tell you what tends to be true on average, across thousands of people it will never individually know. It cannot tell you whether the ache in your knee this Tuesday tracks the pressure drop outside. Only your own record, read by statistics rather than memory, can do that.

Bottom line

Across this whole cluster, the honest answer to "does weather affect health" has been some version of "weakly, unevenly, and it depends on the person" — never a clean yes, never a clean no. That's not an unsatisfying place to land; it's the correct place to land, and it points directly at what to do next. Memory is a poor judge of these questions, reliably finding patterns in randomness and remembering the days that confirm a story while forgetting the days that don't. A diary fixes that only if it's logged close to the moment, includes the ordinary days along with the notable ones, and runs long enough to give real statistics — not a glance at a chart — enough evidence to tell a genuine pattern from noise, correctly accounting for the fact that checking several weather factors at once makes false alarms more likely unless the analysis is built to handle it. Do that, and you're not guessing anymore, and you're not trusting a memory built to mislead you here. You're running a legitimate, single-person study on the one dataset that was always going to matter most: your own.

Sources [1..5]
  1. On the belief that arthritis pain is related to the weather Redelmeier & Tversky, Proceedings of the National Academy of Sciences, 1996.
  2. The occurrence of recall bias in pediatric headache: a comparison of questionnaire and diary data van den Brink, Bandell-Hoekstra & Abu-Saad, Headache, 2001.
  3. Ecological Momentary Assessment (EMA) in Behavioral Medicine Stone & Shiffman, Annals of Behavioral Medicine, 1994.
  4. Single-patient (n-of-1) trials: a pragmatic clinical decision methodology for patient-centered comparative effectiveness research Duan, Kravitz & Schmid, Journal of Clinical Epidemiology, 2013.
  5. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing Benjamini & Hochberg, Journal of the Royal Statistical Society, Series B, 1995.