Three Days of Silence: What a Dead Tracking Script Taught Me About Watching, Not Just Monitoring
A tracking script had been dead for three days before anyone noticed - not because anything broke, but because nothing looked like it had.
A few weeks ago I learned that a piece of tracking on one of my sites had been quietly dead for almost three days. Not broken in any way a visitor would notice. The pages loaded. People clicked around, filled out forms, bought what they came to buy. But somewhere in a routine deploy, one snippet had replaced another, and the wire that reported all of that activity back to me had gone dark. Nobody saw an error. There was nothing to see. That's the part that stayed with me.
I didn't catch it because anything told me to look. There was no alert, no dashboard turning red, no ticket dropping into a queue. I caught it because I went into the reports on an ordinary afternoon, expecting to see the usual shape of the week, and the shape wasn't there. Three days of nothing, sitting underneath three days of perfectly normal-looking traffic. If I hadn't opened that specific report on that specific day, it could have stayed dark for another month. Maybe longer - I genuinely don't know, and that's not a comfortable thing to type out loud.
It would be tidy to call this a tooling failure. Wrong snippet, one deploy, an easy fix once you finally see it. But the actual gap wasn't in the code. It was in the three days nobody was checking, because checking was never automated in the first place. Nothing "missed" this. There was no system that was supposed to catch it and didn't. What caught it was a habit - someone who still goes and looks, on purpose, without being told to, because they've learned not to trust that things are fine just because nothing looks broken.
It's tempting to file this under annoying, no real harm done. Nothing crashed. No revenue was lost at the point of sale. But tracking data isn't just a record of what happened - it's the thing every downstream decision leans on, and a silent gap doesn't announce itself to any of those decisions either. Evaluate a campaign during those three days and it looks underwater, because half its performance never got recorded. Read an experiment during that window and the read comes out wrong in a way nobody thinks to question, because "no data" doesn't look like a data problem. It looks like a real, disappointing outcome. A dashboard can't tell the difference between a channel that stopped converting and a channel that stopped being measured. It reports both the same way: a number that went quiet.
Not the missing rows, in other words. The decisions built on top of a floor that wasn't there.
This failure mode doesn't look like a failure while it's happening. An outage announces itself. A 500 error announces itself. A checkout that stops working announces itself loudly, usually within minutes, because someone's revenue depends on it and someone else is watching that number like a hawk. A dead tracking pixel produces exactly the same output as a healthy site having a quiet week: fewer numbers, no errors, nothing red. The absence of data looks identical to the absence of activity. That's not a bug in any one analytics tool. That's just what silence looks like from the outside.
The instinct, especially lately, is to solve this with more automation - more dashboards, more alerting rules, maybe a model flagging anomalies before a human ever has to notice. All of that helps, genuinely; I build a living out of instrumentation and I'm not arguing against it. But automation is only as good as the failure modes someone thought to describe to it ahead of time. An alert can tell you a number dropped below a threshold. It's much worse at telling you a number that should exist simply stopped existing, especially if it fades out slowly, or lives in a corner of the funnel nobody built a threshold for. Run five monitoring tools in parallel and you can still have a blind spot shaped exactly like the thing nobody thought to watch for.
There's a quieter shift happening underneath all this, too. The more dashboards a team has, the more everyone assumes somebody else is watching them. Observability turns into a checkbox instead of a practice - we have the tooling, therefore we're covered - and the actual watching, the unglamorous act of opening a report and asking does this look right to me, starts to feel redundant. Optional. Something you'd only bother with if you didn't trust the tools. Except the tools were never built to distrust themselves. That job was always going to belong to a person.
I didn't walk away from this wanting more dashboards. I walked away wanting a better relationship with the ones I already had.
The biggest change was small, almost embarrassingly low-tech: a standing block on my calendar, once a week, with no other agenda than sitting with the numbers. Not building a report. Not answering a stakeholder's question. Just looking at the shape of things - traffic, conversions, the handful of metrics that actually matter - and asking whether that shape matches what I'd expect. It sounds trivial, because it is. It's also the exact habit that caught this failure in the first place, and the first thing that gets dropped when a week gets busy, which is precisely why it had to become scheduled and non-negotiable instead of something I trusted myself to just remember. (I didn't, for the record. There was a stretch in July where it slipped for three weeks running. Make of that what you will.)
I also stopped trusting any single source of truth for the metrics that matter most. If one tracking implementation goes dark, a second one, built differently and deployed through a completely different mechanism, should visibly disagree with it. That disagreement is the signal. One tool reporting zero looks like nothing happened. Two tools that suddenly stop agreeing with each other looks like exactly what it is: something broke. Redundancy isn't just about surviving an outage in one vendor's infrastructure - it's about giving silence something to contrast against.
The harder one to operationalize, because it's a discipline and not a system, is treating "everything looks fine" as a claim that needs evidence rather than a default state. Dashboards answer the questions someone thought to ask them. They don't answer whether there's a question they should be asking that they're not. Only a person walking through the data with some skepticism catches that kind of gap. It isn't a threshold problem. It's a does-this-match-my-mental-model-of-the-business problem, and mental models don't live in monitoring configs.
There's a version of this story that ends with "and that's why you need better tooling," and that version isn't wrong exactly. Better tooling would have shortened three days to three hours. It would still have needed someone to build the right check, anticipate the right failure, and eventually, still, look. Tooling moves the floor. It doesn't remove the need for someone standing on it.
Underneath the SQL and the dashboards and the tag managers, that's the actual job. Not deploy hygiene, though I've tightened that too. Somebody still has to open the report and decide if it looks right. That part doesn't automate.
Better tooling narrows the gap. It doesn't replace the habit of still checking. See what that habit looks like applied elsewhere in our case studies.
Get a Monitoring Check