GRADE certainty and Summary of Findings table

Rate certainty one outcome at a time across the five downgrade and three upgrade domains, and export the Summary of Findings table. The rating follows the published GRADE rules, so randomised evidence starts at high and observational evidence at low. Nothing is uploaded and nothing is stored.

Your outcomes

GRADE is rated per outcome, so one review can carry high certainty for one outcome and very low for another.

Certainty is easier when the evidence is already in one place

In a Verflux project the risk-of-bias ratings, the heterogeneity statistics and the bias tests are already computed, so the GRADE domains start from your own data instead of your memory, and the Summary of Findings table fills itself in.

Start a free project See what else it does

Cite this tool

If this tool helped with a published review, a citation is the only payment we ask for.

View on Google Scholar
Rehman, N. U., Saif Ullah, K., & Tufail, U. (2026). Verflux: A browser-based platform for end-to-end systematic reviews and meta-analysis (Version 1.0) [Software]. https://verflux.com
@software{verflux2026, author = {Rehman, Naeem Ur and Saif Ullah, Khuram and Tufail, Usman}, title = {Verflux: A browser-based platform for end-to-end systematic reviews and meta-analysis}, year = {2026}, version = {1.0}, url = {https://verflux.com} }

How a GRADE rating is built

Certainty is about the body of evidence for one outcome, not about a single study and not about the review as a whole.

randomised trials start at high certainty
observational studies start at low certainty
each serious concern −1 level, very serious −2
upgrade +1 or +2, for observational evidence only

The five reasons to downgrade

Risk of bias
The studies contributing to this outcome carry design or conduct problems. This is where your RoB 2 or ROBINS-I judgements arrive, which is why the two are usually done together.
Inconsistency
The results disagree more than chance explains: wide I², little overlap between confidence intervals, or effects that point in opposite directions without a reason.
Indirectness
The evidence answers a nearby question rather than yours. A different population, a surrogate outcome, an indirect comparison between two treatments never compared head to head.
Imprecision
The confidence interval is wide enough to include both a worthwhile effect and no effect, or the total sample falls short of what the question needs.
Publication bias
There is reason to think the missing studies are not missing at random: an asymmetric funnel plot, a literature made of small industry-funded trials, no registered protocol.

The three reasons to upgrade

These apply to observational evidence, where the starting point is low.

  1. Large effect. A risk ratio around 2 or 0.5 that survives adjustment, or around 5 or 0.2 for two levels.
  2. Dose-response gradient. More exposure, more effect, in a pattern confounding would struggle to produce.
  3. Plausible confounding working against the effect. The biases you can identify would shrink the observed effect, so the true effect is probably larger.

What reviewers query most

  1. One certainty for the whole review. GRADE is per outcome, and a table with a single row usually means the assessment was done once rather than per outcome.
  2. Downgrading twice for the same problem. Small studies that are also imprecise is one concern reported in two columns unless the reasons genuinely differ.
  3. Upgrading randomised evidence. It already starts at high; the upgrade domains exist for observational evidence.
  4. No stated reason. Every step down needs a sentence a reader can disagree with, which is what the reasons column in the table is for.

Questions

What are the five GRADE downgrade domains?

Risk of bias, inconsistency, indirectness, imprecision and publication bias. Each can lower certainty by one level, or by two when the concern is very serious.

When can certainty be upgraded?

Only for observational evidence, and only for a large effect, a dose-response gradient, or plausible confounding that would reduce rather than create the observed effect.

Why do randomised trials start at high certainty?

Randomisation balances known and unknown confounders, so the design itself removes the usual reason to doubt a comparison. Observational studies start at low because confounding remains plausible.

Is GRADE rated per outcome or per review?

Per outcome. The same review can carry high certainty for one outcome and very low for another, which is why the table has one row per outcome.

Is it free?

Yes, with no sign-up. It runs entirely in your browser, so nothing you type is uploaded or stored.

More free tools

No account needed for any of them, and nothing you type is uploaded.

PRISMA 2020 flow diagram Build the flow diagram from your screening counts, with the arithmetic checked. Risk of bias figures Traffic-light and summary figures for RoB 2, ROBINS-I, Newcastle-Ottawa, QUADAS-2 and AXIS.