Experimentation Theatre: The A/B Test That Taught You Nothing
Or: Why Running 200 Experiments a Quarter Doesn’t Mean You’re Learning
It was 4:20 on a Wednesday when Somchai walked into the standup room carrying two sheets of printed paper — actual paper, warm from the printer on level 7. The air conditioning was fighting a losing battle against twelve bodies and a projector that had been running since noon. He didn’t say anything at first, just pinned the pages to the whiteboard next to six weeks’ worth of experiment results that had been slowly draining the optimism from the room. Everyone’s laptops stayed open but the typing stopped.
That silence is what real experimentation sounds like. Not the confident click of “ship variant B,” but the uncomfortable pause before someone says, “I think we’ve been looking at this wrong.”
As the Nobel laureate Richard Feynman once warned, “The first principle is that you must not fool yourself — and you are the easiest person to fool.” He was talking about scientific rigour, but he might as well have been reviewing the experiment culture at most software companies. We’ve built entire platforms for running A/B tests, hired data scientists to interpret them, and created dashboards that glow with statistical significance. And yet, most engineering organisations have adopted the ritual of experimentation without the discipline of it.
We run A/B tests the way we run standups — because someone told us to, not because we’re genuinely learning. The gap isn’t tooling. It’s mindset. And in that gap lives a family of sins that statisticians have a name for: p-hacking — the art of manipulating how and when you analyse experiments until you get a statistically significant result, even if the effect isn’t real. You’re not falsifying data. You’re just peeking, slicing, and tweaking your way into a low p-value. And most teams don’t even realise they’re doing it.
Here at Agoda, we run experiments at serious scale. We’ve been doing it long enough to know that the difference between a team that experiments and a team that learns is the same difference between a chef who owns a thermometer and one who actually checks it. And we’ve made enough mistakes along the way to have opinions about where things go wrong.
The Confirmation Experiment
Let’s start with the most common sin: the experiment you’ve already decided to ship.
You know the pattern. Someone has a feature they’re invested in. Maybe they spent two sprints building it. Maybe their OKR depends on it landing. So they “run an A/B test” — but the success criteria were never defined upfront, the failure scenario was never discussed, and if variant B loses, well, let’s try a different segment. Or a different metric. Or let’s just run it one more time.
As the statistician W. Edwards Deming put it, “Without data you’re just another person with an opinion.” Fair enough. But running a test you’ve already decided the outcome of doesn’t give you data — it gives you a permission slip.
Let me tell you about a product grid experiment that nearly broke this pattern in the worst way before it broke it in the best.
We were testing a new layout for our product listing grid — the page where users see hotels, prices, and make their first comparison. The team built variant B, ran it for two weeks. It lost. Not ambiguously — it lost. So they ran it again. Lost again. Six weeks in, the pressure to kill the experiment was real. The two-copy tax was brutal: every other team making changes to anything near that grid had to maintain both codepaths, effectively doing their work twice. Merge conflicts in unrelated PRs became a daily friction. The experiment’s duration was being decided not by what we were learning, but by engineering convenience.
And then Somchai walked in with those printed pages.
He’d pulled competitor screenshots. He’d segmented the data by market — not just “did it win or lose” but who was it winning or losing with, and why. The insight was deceptively simple: in the new grid layout, the quantity dropdown had been placed in the same cell as the price. For most users, this was a minor inconvenience. But for Chinese and Indonesian users — markets with high price sensitivity where users obsessively compare prices across options — it made the grid nearly unusable for the one thing they cared about most. The price column had become noisy, harder to scan, harder to compare. Moving the quantity selector into its own column fixed it. That was the whole thing. A dropdown in the wrong cell, invisible in the aggregate data, obvious once you knew where to look.
That’s a story about real experimentation. The team didn’t just read the p-value and ship. They didn’t read the p-value and kill it either. They dug. They had a hypothesis about why. And that “why” led to a targeted solution that nobody would have found by reading a dashboard.
The contrast here matters: the confirmation experiment asks “did we win?” The real experiment asks “what did we learn?”
The Slot Machine
Here’s a conversation that happens in every data-driven organisation, and it happened to us too: someone presents an experiment result. It lost. The Product Owner pauses, then asks: “Can we run it again?”
That question sounds reasonable. Maybe the timing was off. Maybe there was a seasonal effect. Maybe the sample wasn’t quite right. So you rerun it. And sometimes — just often enough to be dangerous — it wins.
Statisticians have a name for this: optional stopping — deciding when to end an experiment based on the results you’re seeing rather than a predetermined sample size or duration. Its close cousin is peeking: checking results daily and stopping the moment the dashboard turns green. Both feel rational in the moment. Both massively inflate your false positive rate.
The economist Charles Goodhart observed that “when a measure becomes a target, it ceases to be a good measure.” We’d made statistical significance the target, and smart people were finding ways to hit it.
We knew the odds because we’d measured them. We used to run two-week experiments, which captured enough traffic to measure effects confidently. But to understand our actual false positive rate, we ran something unusual: roughly 200 “test experiments” that ran continuously, where variant A and variant B were identical. No change at all. Just determining whether the system would flag a winner when there was nothing to find.
About 5% of them showed a statistically significant result. One in twenty. Which means every rerun of a losing experiment is another pull on the slot machine — a 1 in 20 chance it “wins” purely by accident.
And we started seeing exactly that behaviour. Certain Product Owners would get a loss and immediately queue up a rerun. Not maliciously — they believed in their feature, and the platform made it easy. One more spin. The system didn’t prevent it, didn’t flag it, didn’t even track it. If you wanted to rerun an experiment three times, nobody would know unless they went looking.
We had quiet conversations with the repeat offenders, but we also recognised this was a system problem, not a people problem. You can’t give people a slot machine and then blame them for pulling the lever.
So we designed an evaluation run process. After a win, a new automated run would start — normally only a week, but dynamically adjusted based on data volume and signal strength. If the second run showed similar results, the original win was validated. If it didn’t, we knew the first result was noise dressed up as signal. The rerun wasn’t optional, wasn’t manually triggered, and wasn’t something a PO could skip because they were confident in the result.
The lesson: if your experimentation platform doesn’t account for the humans using it, you’ve built a casino, not a laboratory.
The Partial Experiment
Now for the anti-pattern that should terrify you, because most teams don’t even know they have it.
You’ve carefully designed your A/B test. Frontend shows variant B to 50% of users. Clean split, proper bucketing, everything by the book. Except three microservices deep in your backend, the experiment variant context has been lost. The OpenTelemetry baggage that was supposed to carry the variant assignment through the request path didn’t make it all the way. Half your backend is running variant A logic while the frontend displays variant B.
You’re not running an experiment. You’re running theatre.
This is experimentation at the infrastructure level, and it’s where the gap between “we A/B test” and “we A/B test correctly” becomes a chasm. The variant propagation problem — ensuring that every service in your request path knows which experiment arm a user is in — is genuinely hard in a microservices architecture. Context gets lost at async boundaries. Services that were written before the experiment framework existed don’t know to look for variant assignments. Middleware that was never designed to forward baggage headers silently strips them.
The result is contaminated data. Your “clean” 50/50 split is actually a messy distribution of partially-applied changes, and your metrics are measuring an experiment that doesn’t exist in the form you think it does. You ship variant B believing it won, and the effect evaporates in production — because in production, all services run B, and it turns out the “win” was caused by the weird interaction between B on the frontend and A on three backend services.
If you’re running experiments across a service-oriented architecture, variant propagation isn’t a nice-to-have — it’s infrastructure that requires real investment. Not from one team. Not from your “experimentation platform” squad working in isolation. From the entire engineering organisation. Every service owner needs to understand how variants flow through the system. Every team building a new microservice needs to propagate experiment context as naturally as they propagate authentication tokens. Experimentation can’t be something that lives in a platform team’s backlog — it needs to be in your engineers’ veins, as fundamental to how they build software as logging or error handling. Because if only the teams running the experiment care about variant consistency, and the thirty services between the frontend and the database don’t, you’re not running experiments. You’re generating random numbers and calling them insights.
The Two-Copy Tax
This one is rarely discussed but universally experienced. Every running experiment doubles a portion of your codebase. Variant A and variant B are both alive, both need to compile, both need to pass tests, and both need to be understood by every engineer who touches adjacent code.
The product grid experiment I mentioned earlier? Six weeks of maintaining dual codepaths. Every team working near that grid had to implement their changes twice — once for each variant. Merge conflicts became a double tax. Code reviews took longer because reviewers had to understand both paths. The cognitive load wasn’t just on the system owner; it radiated outward.
This creates a perverse pressure: the longer an experiment runs, the more expensive it becomes for the entire organisation, not just the team running it. So experiments get cut short — not because the team has learned enough, but because the engineering cost of continuing has become politically untenable.
As the management theorist Peter Drucker observed, “What gets measured gets managed.” But what doesn’t get measured gets ignored — and nobody measures the cost of an experiment on the teams that didn’t run it. The merge conflicts, the duplicated effort, the slowed velocity of adjacent work. It’s invisible in your experiment platform and very visible in your engineers’ frustration.
The solution isn’t to stop running experiments. It’s to factor their organisational cost into the decision about how long they should run, and to build infrastructure that minimises the blast radius of concurrent experiments. Feature flags that cleanly isolate variant logic. Experiment platforms that make it easy to remove the losing path. Engineering culture that treats experiment cleanup as seriously as experiment launch.
The Metric Orphan
The final anti-pattern is the most philosophically damaging: the experiment with no hypothesis.
“Let’s A/B test it and see what happens.” You’ve heard this. You may have said it. It feels data-driven. It feels scientific. It is neither.
There’s a worse variant of this that researchers call HARKing — Hypothesising After Results are Known. That’s when you run a test, see that conversion didn’t move but time-on-site went up, and retroactively declare that time-on-site was what you were testing all along. It makes random noise look like insight. The original hypothesis quietly disappears, replaced by whatever metric happened to turn green.
Related is the multiple comparisons problem: if you test enough metrics, segment by enough dimensions — mobile users, iOS users in Thailand, users who clicked at least twice on a Wednesday — something will eventually light up by pure chance. Test five metrics and you’ve got roughly a 23% chance that at least one shows significance even when there’s no real effect. That’s not a rounding error. That’s a coin flip dressed up as science.
A test without a hypothesis is observation without theory. You’ll get a result — variant B moved metric X by Y% — but you won’t know why, you won’t know if it’ll hold, and you won’t know what to do next. There’s no pre-registered belief about what should change, by how much, and for what reason.
The philosopher Karl Popper argued that science advances not by proving theories right, but by designing tests that could prove them wrong. An experiment without a hypothesis can’t be falsified because there’s nothing to falsify. It’s just… watching numbers move.
The discipline looks like this: “We believe that displaying prices in local currency for Southeast Asian markets will increase click-through rate by at least 3%, because our qualitative research shows users in these markets spend disproportionate time on price comparison. We’ll measure CTR on the product detail page over two weeks. If CTR doesn’t increase by at least 2%, we’ll consider this disproven.”
That’s not more work. That’s the same work, but with a reason attached. And when the experiment concludes — win or lose — you’ve actually learned something transferable. The win tells you your model of user behaviour was right. The loss tells you it was wrong, which is arguably more valuable because it forces you to update your understanding.
Naming the Disease
If you’ve read this far and thought “we do some of this,” you’re not alone. These patterns are so common in the industry that researchers have formal names for all of them. It’s worth knowing the vocabulary, because naming a problem is the first step toward recognising it in the wild.
P-hacking is the umbrella: manipulating your analysis — consciously or not — until you land on a significant result. It includes optional stopping (ending the test when you like what you see), peeking (checking daily and reacting to interim results), HARKing (rewriting your hypothesis after seeing the data), and the multiple comparisons problem (testing enough metrics or segments that something will be significant by chance alone).
None of these require bad intentions. That’s what makes them dangerous. A Product Owner who reruns a losing test isn’t trying to game the system — they believe in their feature. An analyst who segments by market after seeing flat results isn’t being dishonest — they’re being curious. A team that celebrates a significant secondary metric isn’t lying — they’re looking for the win in a sea of ambiguity.
The fix isn’t vigilance. It’s structure. Pre-registered hypotheses. Fixed sample sizes determined by power calculations before launch. Primary metrics that drive decisions, with secondary metrics explicitly labelled as exploratory. Automated evaluation runs that validate wins before they’re taken.
The point isn’t to make experimentation harder. It’s to make the results trustworthy enough to actually act on.
What Good Actually Looks Like
Let me take you back to Somchai and his printed charts.
What made that moment work wasn’t the data — everyone had the data. It was the thinking behind the data. He didn’t look at the aggregate metric and declare victory or defeat. He segmented by market because he had a theory about price sensitivity. He pulled competitor screenshots because he wanted to understand how users actually used the grid — scanning prices, comparing options — not just whether they clicked. And when he found it — a dropdown sharing a cell with the price, cluttering the one column price-sensitive users needed to be clean — it wasn’t a breakthrough of data science. It was a breakthrough of curiosity. He printed the charts because he wanted people to look at them together, not glance at a Slack link between meetings.
That’s the difference between experimentation culture and experimentation theatre. Theatre is the infrastructure, the dashboards, the weekly experiment review. Culture is the engineer who spends a Friday afternoon asking “but why did it lose in Indonesia?” when everyone else has moved on.
The uncomfortable question every team should ask themselves: of your last ten experiments, how many could you have predicted the outcome of before running them? If the answer is most of them, you’re not experimenting — you’re confirming. And if you couldn’t predict any of them, you might not understand your product well enough to be running experiments at all.
As the NFL coach Bill Walsh believed, the score takes care of itself — but only if the process is disciplined. The same is true for experimentation. The learning takes care of itself when your team has the discipline to sit with an uncomfortable result and ask what it actually means. Skip the discipline, and you’re just keeping score on a game you’re not really playing.
The Bottom Line
We’ve reached an interesting inflection point in our industry. Nearly every serious engineering organisation has adopted experimentation. The infrastructure exists. The cultural permission exists. We’ve won the argument that “data beats opinions.”
But we’ve replaced one set of problems with another. Instead of building the wrong thing because we didn’t test, we’re now building the marginally-better thing because we tested badly. We’ve moved from no data to noisy data, from gut decisions to cargo-cult statistics, from “we don’t experiment” to “we experiment but don’t learn.”
The fix isn’t more experiments. It’s better ones. Pre-registered hypotheses that can actually be falsified. Infrastructure that ensures end-to-end variant consistency. Evaluation runs that catch false positives before they become shipped features. Teams that celebrate the experiment that was killed as much as the one that was shipped — because killing a bad idea early is the highest-ROI outcome an experiment can deliver.
Your experimentation platform is probably fine. Your experimentation culture might not be. And the difference between the two is the difference between a team that runs two hundred experiments a quarter and learns nothing, and a team that runs twenty and changes everything.
Now, if you’ll excuse me, I need to go review an experiment that’s been running for four weeks. The Product Owner assures me that this time the variant will win — they’ve already updated the roadmap slide. I’m sure it’ll be fine.