How we rate evidence.
Every guide and brief on Restadian carries one of six evidence ratings. This page explains what those ratings mean, how we arrive at one, and what would make us change it. If a rating on a page looks wrong to you, this is the document to argue with.
Six levels, defined before the page is written.
The rating describes the state of the research on a question, not our enthusiasm for the answer. A well-evidenced boring finding rates higher than an exciting thin one. Each level also fixes the verb the page above it is allowed to use, which is the part that keeps the rating honest: confident prose under a weak rating is a defect, and so is hedged prose under a strong one.
Strong
A well-described mechanism supported by repeated findings across independent groups, measured in people, with results that point the same direction. New work is unlikely to reverse it.
Stated plainly, with the size of the effect where one is known.
Reasonable to act on. Individual response still varies, so the timing and the amount are yours to adjust.
Moderate
A real effect, consistent enough to be useful. How large it is, how long it lasts, or which people it was measured in still leaves real room to move.
Written as "appears to" or "is associated with", never as a promise.
Worth trying as an experiment you can reverse. Watch what happens over a fortnight, not a night.
Limited
Thin. Small studies, short follow-up, a single group, or a stand-in measurement rather than the outcome you actually care about. Plausible and under-tested.
Written as "early work suggests", with the size of that work named.
Interesting, not directive. We report it so you know it exists, not so you rearrange your evening around it.
Mixed
Credible work disagrees, and the disagreement has not resolved. Different methods, populations, or definitions produce different answers.
Written as "studies disagree", followed by who disagrees and why.
Treat any confident claim in either direction with suspicion, including ours. We explain the disagreement rather than pick a side.
Insufficient
There is research, and it does not answer the question either way. Often because the thing being asked about cannot yet be separated from something it travels with.
Written as "there is not enough evidence to say", and then what would settle it.
Nothing here to act on. Where a mechanism points a direction we say so, and label it as mechanism rather than as outcome.
Not studied adequately
Nobody has properly looked, at least not in people. Animal work, laboratory models, and confident popular repetition all sit here until a human study exists.
Written as "this has not been studied in people".
We describe what is not known rather than filling the gap. A firm recommendation on this from anyone is not coming from evidence.
How a rating is decided
What counts as a source
We prefer, in roughly this order: systematic reviews and meta-analyses; randomised or controlled experimental work; well-conducted observational studies with a plausible mechanism; and consensus statements from professional sleep and circadian bodies.
We link directly to the source rather than to coverage of it, so you can check what it actually says. Where a source sits behind a paywall we say so and link the abstract.
We do not cite press releases, single-outlet news write-ups, or a product manufacturer describing its own product as evidence for that product.
How we judge quality
Beyond study design, we look at who was studied and whether the finding is likely to hold outside that group. Sleep research is heavily weighted towards young adults, university populations, and laboratory conditions with controlled light. A result from that setting is real, but it may be smaller or differently shaped in a household with children, a night shift, or a northern winter.
We look at effect size, not just whether a result reached significance. An effect that is reliably detectable but small enough to disappear inside normal night-to-night variation is not something we will tell you to rearrange your evening around.
We look for replication. One study is a finding. Several independent groups pointing the same way is a basis for advice.
What we do with disagreement
Where credible work disagrees, the page is rated mixed and the disagreement is described. We do not average two opposing findings into a confident middle, and we do not quietly cite only the side we find more convincing.
Sometimes the disagreement is definitional — two studies measuring "sleep quality" in incompatible ways will reach incompatible conclusions. Where that is what is going on, we say so, because it changes what the argument is actually about.
Why the bottom of the scale has three levels
"Studies disagree", "the research cannot settle it", and "nobody has looked" are three different situations, and a reader needs to be able to tell them apart. A scale that offers one word for all three turns every uncertain answer into the same shrug, which is how careful health writing becomes useless.
Mixed means good studies point in different directions and the argument is live. Insufficient means the work exists but does not answer the question — often because the thing being asked about cannot be separated from something it travels with. Not studied adequately means there is no human evidence at all, however much confident advice is circulating.
The practical difference is what you should do next. A mixed rating means the answer depends on details we will try to name. An insufficient rating means waiting is reasonable. A not-studied-adequately rating means anyone telling you what to do is not getting it from research.
Conflicts of interest
Restadian does not currently accept sponsored content, affiliate revenue, or paid product placement. If that changes, it will be disclosed on this page and on every page it affects.
When a cited study was funded by a party with a commercial interest in its outcome, we note that alongside the citation. Industry funding does not invalidate a study, but it is information a reader is entitled to have.
We do not rate a product category we sell into, because we do not sell into any.
What changes a rating
A rating moves up when independent replication arrives, when a mechanism that was inferred gets measured directly, or when a finding is confirmed in a population closer to the one reading the page.
A rating moves down when a replication fails, when a review finds the original effect was smaller than reported, or when a claim turns out to rest on fewer independent sources than it first appeared to.
When a rating changes, the guide records the change and the reasoning. Substantive changes are also published on the corrections page.
Review cadence
Every guide carries a review date and a next-review date, both visible on the page. The interval depends on how quickly the underlying evidence moves: fast-moving or disputed topics are scheduled more frequently than settled mechanisms.
A review is not a proofread. It means the sources have been re-checked, newer work has been looked for, and the rating has been reconsidered against the definitions above. If nothing has changed, the review date still updates — that itself is useful information, because it tells you the page was checked rather than forgotten.
A page that passes its next-review date without being reviewed is flagged internally. We would rather mark a page as overdue than let it quietly age.
Think a rating is wrong?
That is a reasonable thing to think, and we would rather hear it. Point us at what we missed.