9 min read · 1,877 words
A brand tracking study that hasn’t moved in two years is usually not evidence that your marketing failed. It is evidence that the study was designed badly. Most B2B trackers use a saturated metric, a sample that contains people who will never buy, and a base too small to detect real movement.
Some years ago I inherited a tracker in an industrial equipment business, the kind that sells to contractors, fleet owners and plant engineers across a handful of Asian and Middle Eastern markets. Two waves a year. Roughly 300 interviews a market.
For two years the top-line number moved a point or two in either direction and then came back. I remember presenting that chart in a regional review and realising, halfway through my own slide, that I could not honestly tell the room whether the campaign we had just spent a year on had done anything. The tracker was not lying. It was not built to answer the question I was being asked, and nobody, including me, had checked whether it could.
I have rebuilt that measurement twice since. What follows is what I got wrong, and the arithmetic I wish someone had put in front of me the first time. If you run brand measurement inside a product-led manufacturing business, some of this will be familiar.
Why does a brand tracking study go flat?
Three design decisions, usually taken separately and by different people, combine into a tracker that cannot move.
The first is the metric. Prompted awareness for an established brand in a defined industrial category is typically already high, and a number sitting at 80% has nowhere left to travel.
The second is the sample. B2B trackers are commonly fielded to “business decision makers”, a definition broad enough to include hundreds of people whose employer will never buy the category at any price.
The third is cadence and base. Two waves a year of 300 interviews gives one comparison per year, on a base where anything under eight percentage points is indistinguishable from nothing.
Underneath all three sits a structural fact about business buying. Research by the LinkedIn B2B Institute with the Ehrenberg-Bass Institute found roughly 95% of potential business buyers are out of market at any moment, because purchase cycles are long: about 80% of companies change banking services once every five years, and 75% buy computers once every four years. You are surveying a mostly dormant population about a decision most of them will not take this year.
Prompted recall saturates, and then it stops telling you anything
Prompted awareness asks whether someone recognises your name from a list. In a category with six credible suppliers and a trade press that has covered all of them for twenty years, almost everyone recognises almost everyone. The metric hits its ceiling and stays.
Jenni Romaniuk of the Ehrenberg-Bass Institute is blunt about the older measures. Speaking to Contagious in 2023 around the publication of Better Brand Health, she said top-of-mind awareness “has all the bad biases”, disadvantaging non-users and smaller brands, precisely the groups a growing brand needs to understand. Those measures date from the 1960s, before memory science reached marketing. Prompted awareness does one narrow job: it establishes that non-buyers know your brand belongs in the category.
That is worth knowing once. It is not worth 24 months of tracking budget. Ehrenberg-Bass research reported in Marketing Week in 2022 put it more sharply: “There are still people measuring mental availability with top-of-mind awareness. They should stop.”
If you sell TMT bars, cement or heavy vehicles in India, your prompted recall among specifying engineers was high before you commissioned the tracker. Tata, JSW, Ashok Leyland and Blue Dart do not have an awareness problem. They have a buying-situation association problem, which prompted recall was never designed to detect.
In B2B, sample composition matters more than sample size
This is the part I got most wrong. I spent a year arguing for a bigger sample when the real problem was who was in it.
A consumer tracker can define its universe loosely because almost everyone buys shampoo. A B2B category may have a few thousand genuine buying organisations in a country, and inside each one the person who specifies is rarely the person who signs. McKinsey’s 2026 B2B Pulse survey of nearly 4,000 decision makers across 13 countries found buyers move through an average of ten channels per purchase. That is not one respondent. That is a committee with different memories of your brand.
The How B2B Brands Grow work from Ehrenberg-Bass and the B2B Institute, based on a survey of 1,200 B2B buyers in the US and UK, found active brand rejection running at around 10% in banking and insurance. Even for well-known brands, far more potential customers were unaware than had rejected them. Your problem is rarely that buyers dislike you. It is that the right buyers do not think of you at the right moment.
So the sample frame is the whole ball game. Screen on category buying role, not job seniority. Quota by buying stage, not company size band. And accept that 400 correctly screened specifiers, contractors and procurement heads will tell you more than 1,500 generic “business professionals” ever will. This is the same base-quality trap that makes NPS so unreliable in B2B.
The margin of error arithmetic nobody runs
Here is the calculation that should open every tracker debrief and almost never does.
The margin of error on a single survey percentage at 95% confidence is 1.96 multiplied by the square root of p(1-p)/n, where p is the proportion and n the sample size. Qualtrics sets out that formula and the z-score of 1.96 for 95% confidence, with a worked example: 52% on a sample of 1,000 carries a margin of about plus or minus 3.1 points.
Now the part that gets forgotten. Comparing two waves means comparing two independent estimates, each with its own error. Pew Research Center puts the consequence plainly: the margin of error on a difference is about twice the margin on a single number. A poll with a plus or minus 3-point margin needs a 6-point gap before the gap is real.
Apply that to a typical B2B tracker, at the worst case of p = 50%.
| Interviews per wave | Margin of error on one number | Smallest wave-on-wave change that is real | Is a 3-point move detectable? |
|---|---|---|---|
| 200 | ±6.9 points | ±9.8 points | No |
| 300 | ±5.7 points | ±8.0 points | No |
| 500 | ±4.4 points | ±6.2 points | No |
| 1,000 | ±3.1 points | ±4.4 points | No |
| 2,100 | ±2.1 points | ±3.0 points | Just barely |
Read the last row again. To call a three-point movement in prompted awareness statistically real, you need roughly 2,100 correctly screened interviews in every wave. That is about seven times what a standard B2B tracker fields, and in most industrial categories there are not 2,100 qualified buyers in the country to interview twice a year.
There is a second, nastier consequence. Twelve metrics across four waves produces around 48 comparisons a year. At 95% confidence, roughly one in twenty will look significant purely by chance. So an unmoving tracker still hands you two or three exciting false positives a year, and someone builds a plan on one of them.
If your brand tracker hasn’t moved in two years, you are probably not measuring brand at all — you are measuring the margin of error.
How often should a B2B brand tracker run?
Continuously, in small batches, reported as a rolling average rather than as discrete waves. This one change does more than any other single fix.
Eight quarterly waves of 300 gives 2,400 interviews over two years. Compared wave against wave, that data is nearly useless. Pooled into a rolling four-quarter base and modelled as a trend line, the same 2,400 interviews show a real slope with a genuine confidence interval. You did not need a bigger budget. You needed a different reporting frame.
Between waves, use a metric with no sampling error at all. Les Binet presented share of search to the IPA’s EffWorks Global conference in 2020, showing that share of search correlates with market share across automotive, energy and mobile handsets, with a lead time of up to a year in cars. Around 60% of search volume, he found, comes from long-term advertising effects and 40% from short-term. In a narrow industrial category the query set takes work to build, including generic product terms and local-language variants, but it costs nothing and updates weekly.
What should replace prompted awareness?
Three measures, in order of how much work they take to stand up.
Category entry point association. Instead of asking whether people know you, ask which brands come to mind for each buying situation: the plant is expanding, the current supplier missed a delivery window, a tender requires certified material. Ehrenberg-Bass research reported in Marketing Week found each additional entry point a customer links to a brand lowered defection probability by around 5% in the US insurance sector studied, and recommended a “long shortlist” of five to eight.
Mental market share. Mental penetration multiplied by network size. It behaves like a share metric, so it moves when you take associations from a competitor, and it tracks closely with sales market share. Unlike awareness, it has room to move for a decade.
Share of search. Weekly, free, and unusually good at flagging trouble before revenue does.
None of this argues against brand investment. It argues for measuring it properly, alongside the harder conversation about how much a B2B company should spend on brand. Binet and Field’s B2B analysis, reported by The Drum in 2019, recommended 46% brand building and 54% sales activation, against their 60:40 consumer benchmark. That 46% needs a measurement system capable of showing whether it worked.
What to fix before your next wave
Six things, in the order I would do them.
- Calculate your own detectable difference. Take your base size and run 1.96 times the square root of 0.5 divided by n. If the answer exceeds any movement you have ever reported, stop before commissioning another wave.
- Audit the screener, not the questionnaire. Ask what share of last wave’s respondents work at organisations that could plausibly buy from you. If your agency cannot answer within a day, that is the finding.
- Move to continuous fieldwork with rolling four-period bases. Same annual spend, far more usable data.
- Replace the awareness headline with five to eight category entry points drawn from real buying situations, not from your positioning deck.
- Stand up share of search this month. The underlying logic sits in How B2B Brands Grow, published in November 2023 by Jenni Romaniuk, Byron Sharp, John Dawes and Sahar Faghidno.
- Put the confidence interval on every chart. Error bars change a leadership review more than any slide you will ever write.
My honest view, having sat on both sides of this: most B2B brand trackers are bought as insurance, not as instrumentation. They exist so the marketing function can prove it measures something, which is a different objective from finding out what is true, and it produces a very different research design. The cheapest move this quarter is not commissioning a better tracker. It is opening the last one and calculating what it was ever capable of detecting.
So before you approve the next wave: what is the smallest change your tracker could actually see, and has your brand ever moved that far in a year?
-
Previous Post
Crisis Communication Plan: The First 60 Minutes