How to read your own symptom odds: 7 out of 9 versus 2 out of 20
A single percentage about tomorrow doesn't exist. Here's how to read your own weather-symptom counts — like 7 out of 9 — against the ordinary days you need beside them.
Open a symptom-and-weather log after a few weeks of use and the question in your head is almost always the same one: will tomorrow be bad? That specific question — a forecast, a number attached to a day that hasn't happened yet — is not one your own data, or anyone else's, can actually answer. What your data can answer, once there's enough of it, is narrower and more useful: on days that looked like this one before, how often did things go badly for you, compared with how often they go badly on an ordinary day? That answer looks less like a headline and more like a pair of fractions — bad on 7 of the last 9 days after a sharp pressure drop, bad on 2 of the last 20 ordinary days — and learning to read that pair correctly is most of the actual skill behind using any personal health log well. A note before we start: this is informational material, not medical advice. MeteoHealth is a wellness tracker, not a diagnostic tool; if a symptom is new, severe, or getting worse, that's a conversation for a doctor, not this article.
Why the population verdict isn't your verdict
One of the more careful attempts to settle the weather-and-symptoms question at a population level is worth sitting with, because the result runs against what most people tracking their own symptoms expect to hear. Zebenholzer and colleagues followed 238 people with migraine who lived within 25 km of a single weather station in Vienna, had each keep a daily diary for 90 days, and checked those diaries against eleven meteorological measurements plus seventeen categorized weather situations recorded for the same days (Zebenholzer et al., 2011). Several signals that looked promising in the first pass — a ridge of high pressure, lower wind speed, a shift in sunshine duration — did not hold their statistical significance in the final analysis. The researchers' own conclusion was blunt: weather is considerably overrated as a trigger for headaches, well behind more mundane factors like stress and, for many patients, menstruation. That's an inconvenient line for an app built around weather and symptoms to quote, and it's quoted here anyway, because it's exactly why the rest of this article matters. If weather were a strong, reliable trigger across a whole population, a personal record wouldn't need careful reading — it would just confirm the average. It doesn't confirm that, and a second study sharpens the mismatch further. Among 77 people with migraine whom Prince and colleagues tested against measured weather conditions, roughly half turned out to be objectively sensitive to at least one weather factor — but more patients believed weather was their trigger than the measurements supported (Prince et al., 2004). Belief runs ahead of data in both directions here: some people certain weather affects them aren't shown that in the numbers, and a verdict of "overrated, on average" doesn't rule out that it's real for one particular person inside that "half." Averages and beliefs are both the wrong instrument for the individual question, for a specific reason. Fisher, Medaglia, and Jeronimus, working across six repeated-measures studies, found that the variance around the expected value was two to four times larger within individuals than within groups (Fisher, Medaglia & Jeronimus, 2018). Their broader argument: a group-level finding transfers cleanly to an individual only for what statisticians call an ergodic process — one where the pattern across people and the pattern within one person over time share the same shape — and psychological processes like symptom fluctuation are typically not ergodic. A population study, however well conducted, was never built to answer "what happens to me." A personal record isn't a lesser substitute for that answer; it's closer to the only instrument capable of giving it. For the fuller population picture behind that "overrated" verdict, weather sensitivity: what the science says and barometric pressure and headache go deeper into the mechanism debate — this piece picks up where both leave off.
Natural frequencies: why "7 out of 9" beats "78%"
Even once someone accepts that their own record is the right place to look, the way that record gets reported can undo the whole exercise. Gigerenzer and colleagues, reviewing how badly health statistics get communicated across medicine, wrote plainly that many doctors, patients, journalists, and politicians alike do not understand what health statistics mean or draw wrong conclusions without noticing — collective statistical illiteracy, the widespread inability to understand the meaning of numbers (Gigerenzer et al., 2007). Their proposed fix is a specific, almost old-fashioned format: natural frequencies, the way humans encoded information before mathematical probabilities were invented in the mid-17th century — simple counts that are not normalized with respect to base rates, unlike relative frequencies and conditional probabilities, and for that reason far easier for a brain to digest. Their standard illustration is a mammography scenario: with a cancer prevalence of 1%, a test sensitivity of 90%, and a false-positive rate of 9%, most people asked what a positive result means badly overestimate the odds of cancer. Laid out as natural frequencies against 1,000 women, the picture clears up fast. Ten of the 1,000 actually have cancer, and about 9 of those 10 test positive. Of the 990 without cancer, roughly 89 also test positive, purely as false alarms. Add the two groups and about 98 women get a positive result in total — only 9 of them, roughly 1 in 10, actually have cancer. Those four natural frequencies add up to the real total of 1,000, whereas the equivalent conditional probabilities do not add up to 100% — each pair instead gets separately normalized against its own base rate, which is exactly the step that hides what's going on. A figure like "a positive result means about a 90% chance of cancer" sounds precise and isn't checkable by anyone reading it. "9 of the 98 who tested positive actually have cancer" is arithmetic anyone can redo. The same logic carries straight over to a personal weather log. "Bad on 7 of the last 9 days with a sharp pressure drop" is a natural frequency: two whole numbers, nine real days that were actually lived through, arithmetic you can verify against your own diary. "A 78% chance of a bad day when pressure drops" is the same fact wearing a costume that hides the 9 — and it invites you to stop asking what's behind it.
The number that means nothing without its twin
Even a clean natural frequency like "7 out of 9" is an unfinished sentence on its own, and this is where most personal tracking, inside or outside an app, quietly goes wrong. Seven bad days out of nine pressure-drop days sounds dramatic by itself. Whether it says anything about pressure depends entirely on a number that has to stand right next to it: how often the same person has a bad day when nothing in particular happened — an ordinary day, logged the same way, with no notable pressure change. Suppose the ordinary-day count comes out to 2 bad days out of 20. Put the two fractions side by side: roughly 78% of pressure-drop days were bad (7 of 9), against 10% of ordinary days (2 of 20). That's a real, sizeable gap, and it's the gap — not the raw 7-out-of-9 figure alone — doing the actual work of suggesting a pattern. Now change only the ordinary-day count and watch the conclusion invert. Suppose ordinary days instead come out to 15 bad days out of 20 — someone going through a rough stretch generally, for reasons that may have nothing to do with weather. The pressure-drop rate is still 7 of 9, still about 78%. But the ordinary-day rate is now 75% — 15 of 20 — close enough that there's essentially no daylight between the two groups. The same headline figure, "bad on 7 of 9 pressure-drop days," describes two completely different situations depending on what the other days looked like: in the first case it points at pressure, in the second it points at almost nothing, because this person is simply having more bad days than good ones right now, weather aside. A bare frequency reported for one condition is not yet a finding. It becomes one only once it's set against the matching frequency for the days without that condition, counted the same way on both sides. A sense of what a "pressure-drop day" can even mean in this kind of accounting shows up in a Japanese diary study: Okuma and colleagues had 34 people with migraine keep daily records and found attacks clustering on days when atmospheric pressure sagged to roughly 1003–1007 hPa, 6 to 10 hPa below the standard 1013 hPa (Okuma et al., 2015). The exact drop that counts varies by person and season, but the shape of the comparison — drop days against ordinary days, counted the same way — is the part that generalizes.
Why the test itself becomes a choice at this scale
There's a further wrinkle in moving from "these two fractions look different" to "different enough that it's probably not chance." At the sample sizes an early personal log actually has — a handful of pressure-drop days against a few dozen ordinary ones — the statistical test used to compare two counts isn't a formality any off-the-shelf method handles the same way. A review by Campbell of chi-squared and Fisher–Irwin tests on small two-by-two tables found that the best-performing option overall was a corrected version of chi-squared — the so-called "N−1" test — whenever the smallest expected count in the table is at least 1, falling back to the Fisher–Irwin exact test by Irwin's rule below that (Campbell, 2007). That's worth pausing on: older, more familiar guidance called for a minimum expected count of at least 5 in every cell before trusting an ordinary chi-squared test — a bar a log with 9 pressure-drop days, and a bad-day cell of two or three entries, doesn't clear for a while. None of this means an ordinary chi-squared test is secretly wrong, or that the exact test is the single correct choice everywhere; it means that at 9 days, or 20, or 30, which test gets used stops being background detail and becomes a decision with real consequences for whether a gap between two fractions gets called meaningful. Leaning toward the more conservative option at this scale tends to under-claim a pattern rather than over-claim one — the safer direction to be wrong in. It's also one more reason a lone percentage, offered without the underlying counts, deserves the same scrutiny as any other unchecked arithmetic claim.
The days that don't get logged don't count
There's a failure mode no statistical test can fix after the fact, because it happens earlier, at the point of logging. If a symptom log only gets an entry on days that already feel notable — a bad headache, a rough night, a day worth complaining about — the ordinary-day comparison from the previous section doesn't exist by construction, not because ordinary days were rare, but because nobody wrote them down. Whatever frequency comes out of a log built that way will look worse than reality, systematically, because the denominator is missing its largest, most boring category. This is a narrower slice of a broader logging discipline — what to log, how close to the moment, how to keep memory from quietly editing the record — covered in full in how to keep a symptom diary that actually finds your triggers. The short version here: an ordinary day, logged plainly, is exactly as valuable an entry as a bad one, because it's the only thing that turns "7 bad days" into a fraction that means anything.
When the answer is "no pattern" — and why that's still an answer
Once enough ordinary days are on the record, one of two things happens, and only one of them makes a satisfying headline. Sometimes the pressure-drop rate and the ordinary-day rate really do separate — closer to the first arithmetic example above than the second — and that's a real, personally useful result. Just as often, weeks of careful logging land closer to the second example: the two rates sit close enough together that there's no meaningful gap between them, which in plain language means pressure drops don't appear to be doing much for this particular person. That outcome lines up with what Zebenholzer and colleagues found at the population level, where several signals that looked real in the first pass didn't survive the final analysis, and it deserves to be taken exactly as seriously as the opposite result. A gap that closes under scrutiny isn't a failed experiment — it's the same process that would have flagged a real gap, correctly reporting that this particular one isn't there. Reading "no pattern" as a disappointment rather than as information is the same mindset that lets illusory correlations survive in the first place: it keeps looking for the pattern it expected instead of accepting the one the counts actually show.
What actually shows up in the app
MeteoHealth's part in this is deliberately narrow. It logs symptoms next to the weather and Apple Health data for the same day, and once there's enough history, it shows a personal frequency — bad days after a pressure drop — sitting right next to the matching frequency for ordinary days, as counts, not a lone percentage. While the log is still short, it says how many more entries would make that comparison worth reading, instead of offering a number early and letting you assume it already means something. The arithmetic runs on-device, on data that doesn't need to leave the phone to be useful. None of this amounts to a diagnosis or a forecast: the app doesn't diagnose a condition, doesn't predict tomorrow, and doesn't tell you what caused a bad day. It shows what your own days looked like, side by side, and leaves the reading of that comparison where it belongs — with two numbers and arithmetic anyone can redo. Weather may be linked to how a specific person feels, or it may not be. The only way to find out which is true for one particular body is to look at that body's own days, counted properly, against the ordinary ones they're being compared to.
- Helping Doctors and Patients Make Sense of Health Statistics — Gigerenzer et al., Psychological Science in the Public Interest, 2007.
- Lack of group-to-individual generalizability is a threat to human subjects research — Fisher, Medaglia & Jeronimus, Proceedings of the National Academy of Sciences, 2018.
- Migraine and weather: A prospective diary-based analysis — Zebenholzer et al., Cephalalgia, 2011.
- Chi-squared and Fisher–Irwin tests of two-by-two tables with small sample recommendations — Campbell, Statistics in Medicine, 2007.
- The effect of weather on headache — Prince et al., Headache, 2004.
- Examination of fluctuations in atmospheric pressure related to migraine — Okuma et al., SpringerPlus, 2015.