How we test
Anything carrying our Field-tested badge was used by a person on this team, for a stated length of time, before we wrote a word about it. This page explains the method and, more usefully, its limits.
Most publishers describe testing in adjectives. We would rather describe it in numbers, because an adjective cannot be checked. This page sets out what happens before a guide is published: how long something is used, what is recorded while it is used, who reviews the result, and what the exercise is incapable of proving. It applies to every guide on this site, whether or not that guide ends up carrying a badge.
Our starting position is that first-hand use and published research answer different questions. Research tells you what happens on average across many people. Using something yourself tells you what it is like to live with, which no study measures. A guide is strongest when it says which of the two a given claim rests on, so that is what we label.
The three evidence labels
Every guide falls into one of three categories, and the category is visible at the top of the page rather than buried here. We separate them because merging them is how readers get misled.
| Label | What it means | What it does not mean |
|---|---|---|
| Field-tested | Someone here used it, for a stated duration, and recorded what happened. The duration is printed on the article. | That the result would replicate in a controlled trial. |
| Research-backed | The claims rest on peer-reviewed studies or official guidance, cited inline and listed in full at the end. | That we personally verified the outcome. |
| Neither | Explanatory or practical writing where the honest basis is craft knowledge, and we say so in the article. | That it was tested or that it carries research behind it. |
A guide can carry both of the first two labels. The two badges answer different questions and are counted separately: the research badge reports how many sources a guide cites and how many of those come from government bodies or peer-reviewed databases.
What the Field-tested badge requires
The badge is not awarded for having an opinion about something. Four conditions have to be met before it appears, and if any one of them fails, the guide publishes without it.
- A stated duration. Not "extensively" or "for months", but a number printed on the article. If we cannot say how long, we did not test it.
- A defined protocol, written before we start. What we will do, how often, and what we will record. Deciding what counts as success after seeing the result is the most common way honest people fool themselves.
- A record kept during use, not from memory. Sessions, loads, wear, failures and dates, written down as they happen.
- A reported failure condition. What would have made us call it useless. If nothing could have, the test proved nothing.
How long we use something
Minimum durations are set by category, and they are minimums rather than targets. They exist because the useful information about most things arrives late: shoes are comfortable on day one and shaped by week six, and training programmes feel productive long before they demonstrably are.
| Category | Minimum before we write | Why that long |
|---|---|---|
| Training programmes | One full cycle, minimum four weeks | Early strength gains are largely neural, so anything shorter measures novelty rather than adaptation. |
| Clothing and footwear | Six weeks of regular wear | Fit, wash behaviour and wear points only separate good from bad over time. |
| Tools and gear | Until it fails or clearly will not | Durability is the whole question, and it cannot be sampled. |
| Grooming products | Four weeks of daily use | Skin and hair respond slowly, and the first week is not representative. |
| Habits and routines | Eight weeks | Anything can be sustained for a fortnight, which is why a fortnight proves nothing. |
What we borrow from published standards
We are a publication rather than a laboratory, and we do not pretend otherwise. What we can do is take the parts of established research practice that survive translation to a small publisher, and apply them consistently.
- Pre-registration of the protocol. Trial registries exist because deciding the outcome measure in advance prevents a result being reshaped to fit. The CONSORT reporting guidance formalises this for clinical trials. We keep the small version: the protocol is written before the first session.
- Reporting dropouts. Where a programme is run with a group, we report how many people started, how many finished, and why the others stopped. A result quoted only from the people who completed it is the most flattering number available and the least honest one.
- Separating the reviewer from the author. The person who checks a guide is never the person who wrote it.
- Deferring to systematic reviews over single studies. Where a body of evidence exists, we cite the pooled analysis rather than the most quotable individual trial. Bodies such as Cochrane (official source) exist because single studies mislead more often than they inform.
Where public guidance overrides us
On anything touching health, our own experience is subordinate to official guidance, and we say so in the guide rather than quietly picking whichever supports the more interesting claim. Where the two disagree, the guidance wins and the disagreement is stated.
The baselines we defer to include the Physical Activity Guidelines for Americans (official source), the NHS physical activity guidelines, and the evidence summaries published by the National Institutes of Health (official source). Links to official sources are marked with a badge throughout the site so you can tell a government or peer-reviewed source from a blog at a glance.
Who reviews a guide
Anything that touches physical risk, mental health or money is reviewed before publication by someone with a relevant qualification, named at the top of the article. The reviewer's job is not to improve the prose. It is to find the claim that is wrong, the caveat that is missing and the recommendation that is unsafe for some readers.
A reviewer can block publication. Where a reviewer and a writer disagree and the disagreement cannot be resolved by evidence, the guide states that the question is contested rather than picking a side for the sake of a clean answer. The full workflow is described in our editorial process.
What our testing cannot tell you
This is the section most testing pages leave out, which is precisely why it is here. Our method has real limits and knowing them is what lets you weigh what we publish.
- Sample sizes are small. A programme run with a handful of beginners is an observation, not a trial. We report the number every time so you can weight it accordingly.
- There is no control group. When people improve during a four-week programme, part of that improvement would have happened from any structured training.
- Nobody is blinded. We know which product we are using and what we expect, and expectation influences perception. This is why our subjective impressions are reported as impressions.
- Durability testing ends when we publish. A jacket that survives six weeks may not survive two winters. Where we later learn it did not, we update the guide and date the change.
- One body is not every body. Fit, tolerance and response vary enormously, and a guide written from one person's use cannot account for your injuries, your build or your history.
When we get it wrong
We update the guide, change the recommendation and date the revision, in public. Every article carries an update history listing what changed and when, so a reader arriving a year later can see how the advice moved and why. We do not quietly delete advice that turned out to be wrong, because a record of having been wrong is more useful to you than a clean archive is to us.
Where a recommendation changes because the product changed rather than because we were mistaken, we say that too. The two are different and conflating them is a way of never admitting the second one.
Money, and why it cannot reach this
We buy what we test. Where a company sends something unsolicited, we return it, donate it, or name it in the guide. No brand sees a guide before it publishes, and affiliate links never decide what we recommend or what we say about it. The complete set of commercial rules is on our advertising policy page.
If a guide reads as though something here was not followed, tell us. That is a serious complaint and it is treated as one.
Last reviewed: 11 August 2026. This page covers what the field-tested badge means and what it does not.