Why You Think Like a Bayesian But Were Taught Like a Frequentist

We're wired to reason like Bayesians but trained as frequentists. Here's why that gap exists and what it means for how we understand probability.

Why You Think Like a Bayesian But Were Taught Like a Frequentist

The Schoolroom Distortion

Think back to your first high school statistics class. The teacher walks to the blackboard, draws a coin, and asks a seemingly simple question: “What is the probability of flipping heads?” You and everyone else in the room answer instantly: “Fifty percent.” If pressed for an explanation, you’d probably say that if you were to flip that coin an infinite number of times, half of those flips would land on heads.

This is Frequentism, and it is the standard operating process of our educational system. It defines probability through the cold, objective lens of repeatable data. In the frequentist world, probability is understood through what would happen if we could repeat the same experiment over and over again. It is a neat, comforting framework designed for a world made of rolling dice, shuffled decks of cards, and endless time.

But there is a catch. The moment you step out of the classroom, the laboratory walls crumble. Real life rarely offers us the luxury of infinite trials. You cannot marry someone a thousand times to measure the probability of a happy marriage, nor can a company launch the exact same product campaign on Instagram a million times to test the engagement gain. In the messy arena of human existence, the frequentist definition of probability starts to feel less like a tool and more like a straitjacket.

The Reality: Built to Be Bayesian

And at this very point is where the grand illusion of our education becomes apparent: while we were trained to think like frequentists on paper, we seem to be naturally inclined to reason in Bayesian-like ways.

In the real world, probability isn’t about counting repetitions in the infinite future; it is about quantifying uncertainty in the present. This is the core of Bayesian statistics. Instead of demanding endless data before making a judgment, a Bayesian starts with a Prior — an initial belief based on past experience or intuition. Then, as new evidence arrives, they update that belief to arrive at a Posterior probability.

We don’t need a math degree to do this. The human brain often behaves like a Bayesian prediction engine. Our ancestors on the savannah didn’t have the luxury of waiting for a rustling bush to move ninety-nine times to calculate a p-value before running from a predator. They had a powerful prior: “Rustling bush equals danger.” They saw a tiny bit of new data — maybe a flicker of yellow fur — instantly updated their probability, and survived. It is not hard to see why a fast, Bayesian-like way of updating beliefs could have been useful for survival.

Of course, none of this means we are flawless statisticians. Decades of work by Kahneman and Tversky showed something paradoxical: when asked to reason about probabilities explicitly, with numbers on paper, humans are notoriously bad. We routinely ignore base rates and get textbook Bayesian problems wrong. But this is exactly the point. We run the Bayesian engine beautifully when we don’t think about it; we just can’t read its dashboard. Our intuition is the posterior; our conscious math is the bug.

The Supermarket and the Waiting Game

To see this evolutionary machinery in action, we don’t need to fight off tigers. We just need to look at how we navigate ordinary, adult life.

Imagine walking into a premium supermarket while traveling in a foreign country. You spot a high-end chocolate bar on a shelf, but there is no price tag. A strict frequentist approach would have a harder time here: you have zero observations for this exact item in this exact store. Theoretically the price could be two euros or two hundred, and both are equally unknown.

But you don’t panic, because your brain instantly deploys a prior. You know roughly what chocolate costs in general. You register the elegant lighting, the wooden shelving, the fact that everything else here is somehow imported. Before you have even touched the wrapper, you are carrying a probability distribution in your head, centred somewhere around four euros with a tail that would not be shocked by seven.

Notice what just happened. Nobody handed you data on this product. You built a prior out of context — the store, the country, the packaging. This is not sloppy thinking; it is the most information-efficient move available to you.

Then the cashier says “That will be 14 euros.” Your prediction error spikes as the number sits far out in the tail of what you were expecting. And by the time you walk out the door, you have quietly revised something bigger than the price of one chocolate bar: your whole sense of what “premium” costs in this country. The next unlabelled shelf you encounter, you will guess higher — or maybe you will walk away entirely.

If pricing chocolate feels too trivial, consider the higher stakes of a first date. The date goes well, the conversation flows, you laugh at the same jokes, you float home. You believe there is roughly an 80% chance they want a second date. Before falling asleep, you send a text: “I had a wonderful time tonight.”

Now the evidence starts arriving. A reply in two minutes nudges your belief up towards 95% and you sleep like a baby. But two hours pass. Then five. The next morning your screen is still blank, and this silence is deeply improbable under your original hypothesis. Without ever opening a statistics textbook, your brain grinds through a brutal round of updating. Eighty percent becomes fifty, and fifty becomes thirty. Your prior stays exactly where it was, of course, because a prior is what you believed before. What is quietly dying overnight is your posterior.

A frequentist would frame the question differently. Instead of asking “what is the probability that this person likes me?”, they would ask what would happen to the response rate across repeated, comparable situations. But you don’t have a thousand attempts. You have one.

The Mathematical Straitjacket

If the Bayesian approach is so natural, so deeply embedded in our cognitive wiring, why does the educational system still force-feed us frequentism? Why did most of the 20th century treat Bayes as a curiosity rather than a tool? The answer is a mix of fierce philosophical warfare and one dirty mathematical secret.

In the early 1900s, the architects of modern statistics — Ronald Fisher, and later Jerzy Neyman and Egon Pearson — were on a mission to make science objective. When they looked at Bayes’ theorem, they recoiled at one specific term: the prior. The idea of a scientist injecting personal belief into a mathematical equation felt like heresy. Fisher was deeply skeptical of it. Where do you even get a prior from, and how do you defend it to a referee or in a conference?

It is a fair question, and it deserves a fair answer. Notice that your prior about the chocolate bar didn’t come from nowhere. It came from every price you had ever seen, filtered through the lighting of that particular store. A prior isn’t a wish but compressed experience, written down honestly instead of smuggled in through the back door. The frequentists make assumptions too — they just don’t have to declare them loudly.

Fisher wasn’t persuaded. To rescue science from subjectivity, that generation built a different framework: p-values, t-tests, confidence intervals. The ambition was to make statistical conclusions depend on the observed data and a clearly specified sampling procedure, rather than on a scientist’s prior beliefs about the unknown parameters.

Bayes never quite disappeared, to be fair. Harold Jeffreys was arguing publicly with Fisher throughout the 1930s, and Alan Turing was using Bayesian reasoning at Bletchley Park to break Enigma — work that stayed classified for decades. But these were exceptions, and the reason Bayes stayed at the margins had less to do with philosophy than most people assume.

Even a Bayesian statistician who had won the philosophical argument would immediately hit a wall, and the wall was made of arithmetic. Bayes’ theorem contains one term that, for realistic models, was often impossible to compute. That single term is what kept the entire approach out of practical reach, and it is worth looking at it directly.

Anatomy of a Nightmare

Let’s look at the monster Fisher refused to fight. Bayes’ theorem itself is deceptively simple:

$$P(\theta \mid D) = \frac{P(D \mid \theta), P(\theta)}{\int P(D \mid \theta), P(\theta), d\theta}$$

Here, θ (theta) is everything you don’t know: all the parameters of your model, bundled together. D is your data. The left side, P(θ|D), is what you want: the posterior — your updated belief after seeing the data. The numerator is friendly. P(D|θ) is the likelihood (how well a specific guess of θ explains the data), and P(θ) is your prior. For any single guess of θ, both are easy to compute.

The trouble lives entirely in the denominator. That integral has a name — the marginal likelihood — and it is commonly written in the compact form P(D). A reasonable question is: “why is it an integral at all?” or “why isn’t it simply a number you can look up?”

The intuition is this: P(D) asks “How probable is my data, full stop?” Not “how probable is my data if the price effect variable is 0.4,” but how probable it is overall, without committing to any particular version of θ. To answer that, you must average the likelihood across every possible value θ could take, weighted by how plausible each value is under your prior. In models with even a modest number of parameters, that integral spans a space so vast that no analytic solution exists and no computer can evaluate it by brute force — and this computational barrier, more than any philosophical objection, is what kept Bayesian inference on the sidelines for most of the twentieth century.