Beer & Servers Don't Mix

Your Engineers Don’t Trust Your Data (And They’re Probably Right)

Or: Why Your Analytics Pipeline Is the Most Expensive Lie in Your Stack

It was 2:47 on a Wednesday when Somchai forwarded the screenshot. Not to the data team’s channel — to ours. A spreadsheet, clearly hand-built, with booking numbers that didn’t match anything on our official Grafana dashboard. The vinyl floor tiles ticked faintly under the aircon as six engineers leaned back from their monitors in near-unison, the way people do when they recognise something they’ve been privately thinking for months. Nobody said “the data is wrong.” Somchai just asked, very quietly, “So which number do we use for the review tomorrow?”

That question — which number do we use — is the most expensive question in your engineering organisation, and you’re paying for it whether you hear it asked aloud or not.

Shadow IT never begins as rebellion. It begins as embarrassment. Your engineers aren’t building shadow dashboards to undermine you. They’re doing it because the official data is wrong, late, incomprehensible, or — worst of all — unverifiable. And they still need to ship. When you’re running ten scrum teams across a microservices architecture, the “single source of truth” is often a polite fiction. The truth is scattered across event streams, database replicas, cached aggregations, and that one Grafana dashboard someone built at 2am during an incident that somehow became the canonical source for a business metric.

As Sherlock Holmes put it, “It is a capital mistake to theorize before one has data.” But there’s a darker corollary for engineering orgs: it’s an equally capital mistake to have data that nobody trusts, because then every team theorises from their own private dataset instead.

The Four Failure Modes of Data Trust

Engineers don’t distrust data on principle. They distrust it for specific, diagnosable reasons. In my experience leading over a hundred engineers, these failures fall into four categories — and the fourth is the one most orgs completely miss.

Wrong — “The number doesn’t match reality.” The pipeline has bugs. The ETL silently drops records. A join condition went stale three sprints ago and nobody noticed because the output still looked plausible. Engineers spot this first because they’re closest to the systems producing the data. They see the API return a 200 for a booking, then check the dashboard the next morning and it’s not there. They file a ticket. Nothing happens. They file another one. Still nothing. So they write a query. Then a script. Then a “temporary” dashboard. And now you have shadow infrastructure.

The insidious part: wrong data doesn’t announce itself. It just sits there, looking authoritative, while decisions get made on top of it.

Late — “By the time it’s official, it’s archaeology.” The dashboard refreshes daily but decisions happen hourly. During an incident, your engineer has already queried the database directly, correlated it with logs, and pushed a fix before the “official” metric even registers that something happened. The official number becomes a historical curiosity — useful for quarterly reviews, useless for the work.

Latency kills trust because it teaches engineers that the official path is the slow path. And once they’ve built the fast path, they never come back.

Incomprehensible — “Nobody agrees on what this means.” The metric exists but the definition is a political negotiation. Is “active user” someone who logged in? Someone who performed a search? Someone whose session lasted more than thirty seconds? When three teams each have a different understanding of the same metric name, they’re not disagreeing about data — they’re disagreeing about reality. And no dashboard can resolve that.

This is especially toxic in microservices architectures where different services own different slices of the user journey. Each team’s local definition makes perfect sense in isolation. The contradiction only appears when you try to stitch them together — which is exactly what your analytics pipeline is supposed to do.

Black Box — “Show me the workings.” This is the one most BI teams underestimate, and it’s the biggest trust killer I’ve seen. You get a round number at the end of a pipeline and no way to validate it. No intermediate steps. No ability to trace a single record from source to summary. Just a number on a dashboard, presented with the authority of scripture and the transparency of a magic trick.

Think about what you’re asking engineers to do: accept a number they can’t verify, that measures their team’s performance, that influences resourcing decisions. These are people who write unit tests because they don’t trust their own code to be correct. And you’re asking them to trust a pipeline they can’t inspect?

This is where data democratisation stops being a buzzword and starts being an architectural decision. At Agoda, we took a specific approach: we try to democratise data by making every decision traceable back to super-flat tables in a data lake. One giant table with hundreds of columns for every booking made on the platform, with all the properties you’d want to query on. The same for search. So when you see a report, you can go look up the source data yourself, in all its detail. You can run the same query. You can check the numbers. You can verify.

The flat table approach isn’t just a technical choice — it’s a trust architecture. It says: we’re not hiding anything. The workings are right there. Go check. When engineers can validate the numbers themselves, shadow dashboards become unnecessary. You don’t need to build your own system when the official one is transparent enough to interrogate.

The opposite — a deeply nested pipeline of transformations, aggregations, and business logic buried in dbt models three layers deep — is a trust vacuum. It doesn’t matter how correct the output is if nobody can verify it. Correctness you can’t demonstrate is indistinguishable from wrongness.

Trust vs. Buy-In — The Measurement Problem Nobody Wants to Talk About

There’s a distinction that matters here, and it goes beyond data quality: the difference between trusting a metric and being bought in to it.

DORA metrics are the perfect case study. Deployment frequency, lead time for changes, change failure rate, mean time to recovery — these are well-defined, well-researched, and broadly accepted in the industry. The data can be perfectly accurate. The pipeline can be fully transparent. And your engineers can still reject the whole thing.

At Agoda, we weren’t bought into DORA. So we created our own.

Why? Because trust in the number is necessary but not sufficient. You also need buy-in to the premise. If a team doesn’t believe that deployment frequency is a meaningful measure of their effectiveness — maybe because they’re doing large, complex migrations, or because their domain requires careful, infrequent releases — then showing them an accurate, transparent dashboard just makes them resent you more precisely.

This is where three related concepts become critically important, because they explain the mechanics of how well-intentioned measurement goes wrong.

Goodhart’s Law — “When a measure becomes a target.” Charles Goodhart’s original 1975 observation was that any observed statistical regularity tends to collapse once pressure is placed upon it for control purposes. Marilyn Strathern later distilled it beautifully: when a measure becomes a target, it ceases to be a good measure.

In engineering, this manifests as engineers breaking tasks into smaller pieces to game completion metrics, or writing low-value tests to meet percentage requirements rather than ensuring real coverage. You see it with DORA too: teams can inflate deployment frequency by splitting releases into trivially small increments. The number goes up. The actual delivery capability doesn’t change.

The moment your engineers start optimising for the metric rather than the thing the metric was supposed to represent, you haven’t improved performance — you’ve just made your dashboard lie in a more sophisticated way.

The McNamara Fallacy — “If it can’t be measured, it doesn’t exist.” Robert McNamara attempted to reduce the Vietnam War to a mathematical model, using enemy body counts as a measure of military success. Daniel Yankelovich described the four-step descent: measure what’s easy, disregard what can’t be measured, presume the unmeasurable isn’t important, then conclude it doesn’t exist.

An InfoWorld piece applied this directly to our industry, arguing that the software industry is increasingly in danger of falling victim to this fallacy, as measuring the development process has become easier and easier with modern tools. We can now observe deployment frequency, PR review time, and cycle time with ease. But the temptation is to stop there — to treat what’s measurable as the complete picture.

What you can’t easily measure: whether the team has psychological safety to raise concerns. Whether the architecture is accumulating hidden coupling. Whether the senior engineer is quietly burning out. Whether that “high velocity” is actually reckless speed. The McNamara Fallacy says: because you can’t put these in a dashboard, they slowly stop existing in your decision-making.

As the economist John Kenneth Galbraith observed, “The conventional view serves to protect us from the painful job of thinking.” A dashboard full of green metrics is exactly this kind of conventional view — comfortable, authoritative, and potentially masking the things that actually matter.

Surrogation — “The metric IS the strategy.” This is the least known of the three and probably the most dangerous. Coined by Choi, Hecht, and Tayler in management accounting research, surrogation is the tendency for managers to lose sight of the strategic constructs that measures are intended to represent, acting as though the measures are the constructs themselves.

The Wells Fargo scandal is the extreme case — management inadvertently replaced their “build long-term relationships” strategy with their “cross-selling” metric, resulting in a massive account fraud scandal. But the critical finding from the research is even more sobering: this happens even without incentive compensation. Simply providing a measure to managers can trigger the substitution. It’s a cognitive bias, not just a perverse incentive problem.

You don’t need to tie bonuses to DORA metrics for surrogation to kick in. Just putting the dashboard on the wall is enough.

This is why buy-in matters as much as accuracy. If your engineers understand why a metric exists, what strategic question it’s answering, and what its known limitations are, they can use it as a lens without mistaking it for reality. If they just see a number and a target, surrogation is inevitable.

Metric Quality → Trust → Performance (Yes, There’s Research)

There’s empirical evidence that this isn’t just a feelings problem. A 2024 study of 152 middle managers in healthcare found that metric quality — specifically accuracy, sensitivity, and verifiability — moderates the relationship between performance management and interpersonal trust, which is subsequently linked with unit performance.

Read that chain carefully: metric quality → interpersonal trust → performance. Bad metrics don’t just give you wrong answers. They actively erode trust between people. And that erosion of trust is what actually hurts performance.

This maps directly to engineering orgs. When your DORA dashboard shows a team’s deployment frequency tanked, but the team knows it’s because the pipeline was broken for a week — not because they stopped shipping — two things break simultaneously: trust in the measurement system, and trust in leadership’s judgement for using it. If leaders make resourcing decisions based on metrics the team knows are misleading, the team doesn’t just distrust the data. They distrust the leaders.

And here’s the feedback loop that makes it vicious: once engineers distrust the official metrics, they stop investing in data quality. Why bother fixing the pipeline if leadership is going to misinterpret the output anyway? The data gets worse. Trust drops further. More shadow systems appear. The “official” metrics become a performative exercise — everyone goes through the motions in the quarterly review, then goes back to their private spreadsheet for actual decisions.

Jerry Muller captured this dynamic in The Tyranny of Metrics: metric fixation is the seemingly irresistible pressure to measure performance, publicize it, and reward it, often in the face of evidence that this just doesn’t work very well. The pressure to measure isn’t the problem. The pressure to measure without earning trust in the measurement is.

Building Trust in Data — What Actually Works

Trust in data is earned the same way trust between people is earned: through transparency, consistency, and demonstrated good faith over time. Here’s what that looks like in practice.

Kill the black box — make lineage visible. If engineers can’t trace a number back to its source query, they won’t trust it. This is the single highest-leverage thing you can do. The flat-table approach I described earlier — where every report can be validated against the same source tables that anyone can query — removes the most common objection: “I don’t know where this number comes from.”

This doesn’t mean every engineer will go check the source data. Most won’t. But knowing they can changes their relationship with the number entirely. It’s the difference between “trust me” and “check me.” One builds dependence, the other builds confidence.

Separate learning metrics from judgement metrics. A team cannot explore honestly if the same metric is used to punish them. This is perhaps the most important design principle for engineering metrics. If deployment frequency is a tool for the team to understand their own delivery patterns, it’s useful. If deployment frequency is what determines whether the team gets headcount next quarter, it’s a target — and per Goodhart, it immediately stops being a good measure.

Be explicit about which metrics are for learning — the team owns them, uses them for retrospectives, experiments with improving them — and which are for judgement. Mixing them is how you get gaming.

Use contradictory metrics. Roger Martin’s advice on fighting surrogation: use multiple metrics that create dynamic tension. If you measure deployment frequency, also measure change failure rate. If you measure velocity, also measure defect escape rate. You can’t game all of them simultaneously without actually improving the underlying capability.

The DORA framework already does this to some extent — the four metrics are designed to balance each other. But orgs often cherry-pick the one or two that are easiest to improve and ignore the rest, which defeats the purpose entirely.

Co-create metrics with the teams they measure. The surrogation research found something actionable: involving managers in the selection of a strategy reduces their tendency to surrogate. If teams participate in choosing what gets measured and how, they’re far more likely to understand the metric as a proxy rather than mistaking it for the thing itself.

This is the buy-in problem solved at the source. Don’t hand teams a dashboard and ask them to believe in it. Bring them into the room when you’re deciding what to measure. Let them argue about definitions. Let them point out the failure modes. A metric that survives that scrutiny is one people will actually use honestly.

Earn trust incrementally — fix one number. Don’t try to overhaul the entire data platform at once. Pick one metric that everyone knows is broken. Fix it. Make it accurate, fast, well-defined, and transparent. Then do the next one.

Trust compounds. So does distrust. Every broken metric you fix is a deposit in the trust account. Every metric you leave broken — especially after someone reported the issue — is a withdrawal. The rate at which you fix known data problems is itself a signal of whether you’re serious about data quality, and your engineers are absolutely watching.

When someone builds a shadow dashboard, promote it. Don’t punish it. Don’t treat it as a governance violation. Treat it as the most honest piece of user feedback you’ll ever get about your data platform. That shadow dashboard is telling you exactly three things: what data your engineers actually need, where the official system fails to provide it, and how much latency and accuracy matter for real decisions.

Ask the engineer who built it to help fix the official system. Give them credit. Make the shadow dashboard unnecessary by making the official one good enough that nobody needs to route around it.

The Bottom Line

You cannot mandate trust in data any more than you can mandate trust between people. You can mandate that teams use the official dashboard in their quarterly reviews, but you can’t mandate that they believe it. And the gap between usage and belief is where shadow systems live.

The path forward isn’t better dashboards. It’s not more metrics. It’s not a fancier data platform. It’s the hard, slow work of making your data transparent, your metrics co-owned, and your response to data quality issues fast enough that people believe you when you say “we care about getting this right.”

Campbell’s Law warns us: the more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort the social processes it’s intended to monitor. The antidote isn’t to stop measuring. It’s to measure with humility, transparency, and an honest acknowledgement that the map is not the territory.

Your engineers already know this. The question is whether your data culture reflects it.

Now, if you’ll excuse me, I need to go check why the number in last week’s review deck doesn’t match the number I’m seeing in the source table. I’m sure it’s fine. It’s always fine.