Methodology
How a claim gets a grade, which sources are allowed to appear as a citation, and what the grade does and does not mean.
The grading scale
Grading follows HEALM, applied per claim. The letter fixes how the claim may be phrased, and the phrasing is not interchangeable: a C-grade claim is never written in A-grade language.
Several studies of good quality agree on the effect, its direction and roughly its size.
The studies point the same way, with less agreement between them or fewer of them.
The early work is promising and not yet settled. Often a small number of trials, or a body of observational work.
The studies disagree. Grade D earns no medal and is never the basis of a recommendation, but it is shown by default: it is graded evidence, and hiding it would curate by leaving it out rather than by saying what it is. Narrow the evidence floor to A–C, A–B or A only whenever you want less.
A grade is confidence in the evidence for one specific claim, not a rating of the activity. A simple habit whose benefit has been replicated many times can grade A, while a demanding practice supported by one small trial grades C. The grade says how sure the research is, not how much the activity is worth doing.
Two levels of grade, both shown
Each entry carries its own grade, judged across its whole source set. That is the grade on the badge and the one the evidence filter runs on. Each cited study also carries its own grade, shown beside it in the citation list.
The two can differ, and it matters that they can: a strong individual study can sit under a more cautious overall claim, because one good result is not consistency. Both are on the surface so that difference is visible rather than averaged away.
Which sources count
Peer-reviewed systematic reviews and meta-analyses are preferred, because the question being graded is whether a finding holds across studies rather than whether it appeared once. Primary trials are cited where they are the best available evidence for a claim, or where they are the landmark result a review is built on.
None of these is shown as graded evidence:
- Guidelines, position statements and consensus documents. A guideline is a recommendation derived from evidence, not the evidence. Citing one would move the audit trail one step further from the data.
- Protocols and registrations, including INPLASY and PROSPERO records. A protocol is a plan for a review, published before it has results.
- Preprints, which have been posted by their authors but not reviewed.
- Conference abstracts and proceedings, which are summaries rather than reviewed papers.
Records of these kinds are not deleted, from the reviewed pool or from an entry’s sources, because the number of records screened is part of an honest count of the evidence base and because a reader who can see what a source is can judge it. What they lose is the medal: wherever one appears, in a focus area’s full study listing or among an entry’s key sources, it carries a neutral label and its resolving DOI instead of a grade.
- Protocol / registration. A registered plan for a review or trial, not its results.
- Preprint. Posted by its authors before peer review.
- Conference abstract. A conference summary rather than a peer-reviewed paper.
- Peer-review report. A referee's report on a paper, not the paper itself.
- Guideline / position stand. A recommendation derived from evidence, which is a step further from the data than the evidence itself.
- Book chapter. A chapter in a scholarly book rather than a journal paper.
- Type not recorded. Our record does not say what kind of publication this is, so no grade is shown. That is a gap in our data, not a judgement about the research.
The last of those is the honest floor of this screen. A record’s kind is derived from its DOI, its venue and its title, and where those do not settle the question the record is marked unconfirmed rather than assumed to be a finished paper. That is deliberately the cautious direction: a handful of genuine papers carry the label because their record is too thin to verify, which costs a medal, where the opposite error would put a grade on something that never earned one.
A small number of sources attached to entries before this screen existed are of these kinds, and a few remain. They are labelled wherever they appear rather than quietly removed, because withdrawing a source would change what an entry’s claim rests on, and that is a judgement for a human reviewer and not a display rule.
A focus area’s listing says how many of its records were screened this way, because that count is itself part of an honest picture of the evidence base.
The four citation gates
Every candidate citation passes four gates in order. Each gate can only remove a source, never promote one, so a source that fails at any point is not shown whatever the earlier gates said.
- 1. Statistical relevance
The candidate source has to be about the claim, measured. Relevance is scored on the reported outcome and population rather than on title similarity, so a paper that mentions the topic in passing does not qualify as support for it.
- 2. Reading-based judgment
What survives the first gate is read. The question is whether the paper actually reports the effect in the direction and population the claim describes, and what it says about size and certainty. A source that turns out to qualify or contradict the claim on reading is removed here, and the claim is reworded or regraded.
- 3. Research-library cross-check
Every surviving source is checked against the wider research library for the focus area: does it sit inside the reviewed evidence base, is it the version of record, and is it already represented by a stronger source covering the same finding. Duplicates and superseded versions are dropped here.
- 4. Venue screen
The last gate is where the work was published. Peer-reviewed venues pass. Registered protocols, registrations, preprints and conference abstracts do not, and neither do guidelines and consensus statements, however authoritative, because a guideline is a recommendation rather than the evidence for one.
Filtering, not trimming
The library is comprehensive and curation happens at the moment you read it. The evidence filter has three settings, A only, A to B, and A to C, with A to C as the default. Nothing weaker is deleted from the library; it is filtered out of the view, and widening the filter brings it back. An institution can set the floor for everyone in it.
The same principle applies to how many results you see. Every entry that clears your filters is listed. Nothing is capped to a chosen few.
Doses are the evidence's, durations are yours
Each practice carries the dose the evidence was gathered at. You can schedule any length you like, and the app labels where that lands: at or above the evidence-based dose, or a partial benefit below it. A shorter session is never presented as producing the same outcome as the full one.
Where the research states no minute count, none is invented. The scheduled length is described as a calendar slot and the dose is quoted in the research’s own words.
What the app writes, and what it does not
The library is curated and human-verified. It is a fixed, versioned set of entries, claims, doses, grades and citations. Assembling a week means selecting from that set and explaining the selection using what those entries already say.
Nothing in the library is generated: no invented entries, claims, citations, doses or grades, and no advice written beyond what the reviewed content supports. If it is on the screen as evidence, it is in the library, and the DOI resolves.
Every source in an entry’s trail says how it was read: full text, abstract only, or screened only. An entry with any abstract-only source says so, with the count, so partially verified content is never presented as fully verified. The seven non-physical dimensions carry a separate beta label: their review set is version 0.9-beta until the Round-2 systematic reviews. To see this in practice, open any graded entry and follow it to the studies behind it.