Herbora
← All posts

How we grade evidence, and why four labels beat a star rating

The hardest design decision in Herbora was not the camera or the offline packs. It was what to put next to a sentence like “used traditionally for digestive complaints”.

Put nothing there and you have published folklore as fact. Put a five-star rating there and you have invented a precision nobody has. We landed on four labels, and the reasoning is worth writing down.

The four labels

Strong evidence Moderate evidence Limited evidence Traditional use

Strong means multiple well-conducted human trials, or a systematic review that pools them, point the same way. This is a high bar and relatively few plant–use pairs clear it.

Moderate means there is human evidence, but less of it, or it is smaller, or results are mixed enough that a careful reader would hedge.

Limited means the research exists but does not carry much weight on its own: small or uncontrolled studies, or laboratory and animal work that has not been tested in people. This label does a lot of work, because a great deal of exciting-sounding plant research is cell studies that have never left the bench.

Traditional use means exactly what it says: this use is documented in a healing tradition, and we are not claiming a clinical result for it.

Why traditional use is a finding, not a failure

It would be easy to read the fourth label as the bottom of a ladder — the shelf where plants sit until science rescues them. We do not think of it that way.

“No trials have been run” and “trials were run and found nothing” are completely different statements. Most plants are in the first category, and a rating scale that blurs them is lying to you.

Absence of evidence is not evidence of absence, and it is also not evidence of benefit. A plant with centuries of documented use and no modern study is genuinely interesting — as ethnobotany, as a research lead, as cultural knowledge worth recording. It is simply not a proven treatment, and the label says so plainly rather than implying a verdict that has never been reached.

Why not a score out of five?

A numeric score invites arithmetic that the underlying data cannot support. If chamomile scores 3.5 and ginger scores 4.0, a reader will conclude ginger is “better” — better for what, in whom, at what preparation and dose? The number implies a comparison that was never made.

Four named categories are coarser, and that coarseness is the point. They describe the shape of the evidence rather than pretending to measure its magnitude.

Grades attach to a use, not to a plant

This matters more than it sounds. A plant is not “strong evidence”. A specific use of a specific preparation may be. The same plant can carry a moderate grade for one use and traditional-use only for three others, and in the app they sit side by side on the plant page rather than being averaged into a headline.

Averaging would produce the single most misleading number we could show you.

Every label carries its sources

A grade with no citation is just our opinion in a coloured pill. Each rating links to the references behind it, so you can check whether the review we leaned on actually says what we think it says. If you find a case where it does not, the plant page has a contribute button and we would genuinely like to hear about it.


Herbora is for education only. Nothing here is medical advice, and an evidence grade is not a recommendation to take anything. Talk to a clinician about your health, particularly if you are pregnant, breastfeeding or taking medication.