Mental Models Library · Prototype, provisional scores
The Evidence Line
178 marketing ideas, sorted by how much evidence actually backs them. The higher up, the better the proof. The bigger the bubble, the more people believe it.
Of the 24 best-known ideas in marketing, 20 sit below the line.
George Box wrote that all models are wrong and some are useful. The sentence everyone forgets comes next: how wrong does a model have to be before it stops being useful?
It depends on the job. A framework that organises a conversation costs almost nothing when it is wrong, because it makes no prediction. A model that allocates money on a prediction, such as retention being cheaper than acquisition, costs a great deal when it is wrong, and the cost is invisible because the counterfactual never appears on the dashboard.
Of the 31 concepts taught in the two best-selling marketing textbooks, 27 sit below the line. The pattern holds the whole way down: the better known an idea is, the worse the evidence behind it. Of the 56 ideas a marketer could use without explaining, 75% fall below the line, median score 17. Of the 93 that need explaining, 28% fall below, median score 65.
Fame and evidence are not merely unrelated here. They run in opposite directions.
This is a map of which beliefs you can safely make lists with, and which ones you should stop making bets with.
| Model | Type | Score | Design /35 | Repl. /30 | Indep. /15 | Effect /20 | Direction | Believed /6 | Claim scored | What the evidence says | Source | Status |
|---|
Why this matters
All models are wrong. The question is which ones you are betting on.
In December 1799 George Washington woke with a sore throat. His physicians drained close to 40% of his blood in twelve hours and he died that night. They were not quacks. They were the best doctors in America, practising a method that had survived two thousand years because the theory behind it explained every outcome. If the patient lived, the bleeding worked. If he died, it came too late. A model that can explain any result predicts none of them. Most marketing frameworks are built the same way.
George Box wrote that all models are wrong and some are useful, and the sentence everyone forgets comes next: the practical question is how wrong a model has to be before it stops being useful. The answer depends on the job. A framework that organises a conversation, such as SWOT or the 4Ps, costs almost nothing when it is wrong because it makes no prediction. A model that allocates money on a prediction, such as retention being cheaper than acquisition, targeting beating reach, or three exposures being needed before an ad works, costs a great deal when it is wrong, and the cost is invisible because the counterfactual never appears on the dashboard. So this chart is not a list of things to stop believing. It is a map of which beliefs you can safely make lists with and which ones you should stop making bets with.
What other fields learned when they grew up
Prediction is the only referee. Ptolemy's model put the Earth at the centre and predicted eclipses for 1,400 years. Copernicus was more elegant and, at first, no more accurate. Kepler and Newton won because they predicted better, and they were adopted when the stakes rose: calendars tolerated error, ocean navigation did not. The purchase funnel and last-click attribution were good enough when digital was a tenth of the budget. At 60% of budgets, with acquisition costs up 222% since 2015, the epicycles are showing.
Counting came before understanding. Pierre Louis had no germ theory. In the 1830s he counted pneumonia patients who were bled and those who were not, and the bled ones died more. Medicine started improving outcomes decades before it could explain why. Andrew Ehrenberg did the same in 1959: he counted purchase panels and found Double Jeopardy long before memory science offered a mechanism. Read outcomes before activity. Base rate, then regression to the mean, then counterfactual. Most claims that a campaign worked do not survive step two.
Correct findings get rejected when they insult the profession. Semmelweis cut maternal mortality from 18% to 2% with handwashing and was driven out of Vienna, because the finding implied doctors were killing patients and came with no mechanism. Ehrenberg's laws sat in the journals from the 1960s and reached practitioners in 2010 with How Brands Grow. Same shape: they implied loyalty programmes, tight targeting and differentiation were mostly wasted money. Our own survey of 149 Canadian marketers found 73% agreed it costs five times more to win a customer than to keep one, and agreement rose with experience. Expect the lag, and expect it to feel personal.
Evidence without a home gets lost. James Lind ran a controlled trial on scurvy in 1747. The Navy adopted citrus in 1795. A century later it switched to a lime juice that had lost its vitamin C, nobody knew why the cure had worked, and scurvy returned on polar expeditions. Marketing forgot too: Lodish's 389 split-cable experiments in 1995 showed that half of TV campaigns produced no sales lift and that the successful ones doubled their effect over the following two years. A decade of digital dashboards erased both findings and Binet and Field had to recover them from the IPA databank. This library exists so that nobody has to rediscover adstock a third time.
Incentives keep wrong models alive. Everyone knew stress caused ulcers, and an antacid industry rested on it, until Barry Marshall drank a flask of H. pylori. Ask who profits from a belief. Ad-tech profits from targeting accuracy (third-party age data is right 24% of the time). Agencies profit from rebrands (Tropicana lost 20% of sales in two months). An ETF was built on a satisfaction index whose creator ran the study; four independent replications found nothing. Independence is one of the four criteria in the rubric for this reason.
Wrong is about domain, not truth. Newtonian mechanics is wrong and it got us to the Moon. The engineering question is never whether a model is true but whether it is accurate inside the domain you are using it in. The Dirichlet needs a different switching parameter in subscription markets. Binet and Field's 60/40 is an average that runs from 50/50 to 80/20 and comes from award entries. Both are excellent inside their domain and misleading at the edges. State the boundary conditions or do not use the model.
What changes
Thomas Kuhn's most useful idea is not revolution but normal science: once a field has a paradigm with a good predictive record, the productive work is disciplined problem-solving inside it. Marketing has one. Ehrenberg-Bass for how buyers behave; Binet, Field and the econometric literature for how advertising pays back; Simon and McKinsey for price. The work is not to find a new framework. It is to operate inside the ones that predict, and to know their edges.
- Write the prediction down first. Effect, metric and horizon, before the money moves. If you cannot, it is not a strategy.
- Count before you explain. Base rate, regression to the mean, counterfactual, in that order.
- State the domain. One line per model: where it holds, where it breaks.
- Ask who profits from the belief. If it is the person selling the tool, the claim needs a holdout test.
- Swap, do not just delete. Every refuted belief on this chart has an evidenced counterpart in the same family. Hover a bubble and its relatives light up.
- Teach the next cohort. Resistance is strongest when a finding implies the listener has wasted money. Calibrate rather than accuse; the Quick Diagnosis is built for that.
Evidence did not make medicine certain. It moved the odds: childbirth mortality from roughly one in six to one in ten thousand. Marketing's odds are measurable too. Binet and Davis found the industry became 4% more efficient and 11% less profitable by optimising the wrong number. Better odds compound over a career. The strategist's job does not disappear. It changes from being the source of truth to being the person who knows which map to use where.
How the score is built
What gets scored: the model's central claim as practitioners state it (shown as "Claim" on every bubble), not whether the tool is useful. A thinking tool can be worth using and still score near zero; it just can't be cited as proof. Where a belief and its evidenced opposite both circulate, they get two bubbles (Ads Create Demand and Ads Are a Weak Force; Frequency of Three and Reach Over Frequency).
Score = (Design + Replication + Independence + Effect) × Direction. Each criterion is awarded exactly one of its listed levels (the ladders below), never a value in between. The four parts grade the evidence base out of 100; the multiplier applies what the best counter-evidence says. The evidence line sits at 50, which the rubric makes mean one thing: at least one credible independent study, showing a real effect, that nothing contradicts.
Design best available evidence type · up to 35
Replication breadth and independence of repeats · up to 30
Independence who paid, and is the method public · up to 15
Effect size and consistency · up to 20
Direction what the counter-evidence says · multiplier
Comparators
The grey column is not marketing. It holds ten familiar claims from physics, medicine, economics and psychology scored on the identical rubric, spread from gravity (100) to left brain / right brain (12), so the marketing scores have something everyone recognises to stand next to. Smoking and lung cancer scores 93 without a single randomised trial, which is worth remembering when someone dismisses the Ehrenberg-Bass panels as "only observational".
Bands
Strong 85 and over · Good 70 to 84 · Promising 50 to 69 · Thin 35 to 49 · Slight 20 to 34 · None or contradicted under 20.
Worked example, Cost of Dull: design 28 (IPA databank analysis), replication 8 (one study), independence 5 (System1 co-authored and sells the emotion score), effect 14, supported × 1.0 = 55. One good study, sponsor-adjacent, sitting just above the line. That is what the line is for.
Bubble size
How widely the model is believed or used, 1 to 5. Where Marc's CMA member surveys measured agreement with a belief directly, the tooltip shows the measured figure; those bubbles are the first to move from judgement to data. 5 in the introductory textbooks and in routine boardroom or agency use · 4 common in practitioner discourse · 3 known within a specialism (effectiveness, media, research) · 2 known mainly to researchers · 1 obscure outside its originating team. This is the softest number on the page; a later version can replace it with measured proxies (search volume, LinkedIn mentions, textbook index presence), the way the original Snake Oil chart used Google hits.
Status and disputes
Verified means the work and its headline numbers were checked against a primary or institutional source; Partial means the work is confirmed but a quoted figure is not yet; sponsored studies are named as such in the source line. Related claims are grouped into families (media weight and scheduling, creative, pricing and promotion, and so on): hover any bubble and its relatives take a dark ring. Related is not identical, which is why a family can hold a supported claim and a contradicted one side by side. Scores are provisional until reviewed. To dispute one, argue the criterion: "you gave X a 10 for independence and it should be 5" moves the bubble by the arithmetic, not by opinion.